How RecoKit Recommends from Day 0, Without Historical Data
By Grégory Le Goff ·
Every online store reaches the same point at some stage: the catalogue is ready, the website is live, but the "You may also like" section is either empty or filled with best-sellers that have little to do with the product the visitor is looking at.
This is not a configuration problem. It is a structural issue known as cold start, and it affects recommendation engines of all kinds.
The good news is that cold start is not a new problem. Over the years, several approaches have been developed to deal with it. They each solve part of the problem, at a different level. Understanding these approaches helps explain what a recommendation engine can realistically deliver from the very first day.
The Starting Point: No Clicks, No Patterns
Most recommendation engines have historically relied on collaborative filtering. Users are compared based on their behaviour, such as clicks, add-to-cart events and purchases. The system then recommends products that users with similar behaviour have interacted with.
This approach is powerful, but it depends heavily on behavioural data. If there are no interactions, there is nothing to compare:
users → interactions → similar users → recommendations
A variation of this approach appeared later, mainly to improve scalability: item-to-item filtering, popularised by Amazon in 2003 [1].
Instead of comparing users, the system looks at purchase co-occurrences between products. This leads to the familiar "customers who bought X also bought Y" approach.
The calculation is done at product level, but the underlying principle remains the same: it is a statistic based on past purchases. A product that has just been added to the catalogue, with no purchases yet, has no purchase co-occurrences to rely on.
Item-to-item therefore does not really escape behavioural data. It simply changes the level at which the statistics are calculated.
The same problem appears with a new store, a new product or a newly launched seasonal collection. The key ingredient, namely interactions, is missing.
The recommendation engine may remain largely silent for days or even weeks while enough data is collected. This is the classic cold-start problem, which has been studied in recommender-system research for many years, both for new users and new products [2].
There is another, older and simpler approach: move up a level and reason in terms of product families rather than individual products.
Generic rules can be defined once, such as "cameras → memory cards" or "drills → drill bits". These rules often come from market basket analysis, a data-mining technique popularised by the Apriori algorithm [3].
The advantage is obvious. A whole product family can immediately benefit from rules that already exist for its category, even if a particular product has no history of its own.
The downside is that the rule is generic. It does not necessarily check whether the products are actually compatible. Two pool pumps may belong to the same product family while having very different flow rates, for example.
Generic recommendations can also become repetitive. To introduce some variation, systems sometimes add a degree of randomness to the recommendations. Diversity and serendipity are themselves well-known topics in recommender-system research [4].
Building Block 1: Embedding-Based Similarity
The first approach that really moves away from behavioural data is to stop looking at what customers do and start looking at what the product is.
The product page itself becomes the source of information, independently of clicks and purchases.
The title, description, category and product specifications, and sometimes the image, can be transformed into an embedding: a numerical representation that captures the semantic meaning of the product.
This technique has become widely used in natural language processing, with models such as Sentence-BERT [5].
Products that are close to each other in this semantic space are likely to be similar in the real world. A "200 m waterproof diving watch" and a "100 m waterproof watch", for example, will naturally end up close to each other without requiring a single customer interaction.
This is the principle behind content-based recommendation:
products → semantic vectors → nearest neighbours → recommendations
The benefit is immediate. Recommendations can be generated as soon as the catalogue is ingested, without waiting for any purchase or click history.
This has become a basic building block of modern recommendation systems, from large e-commerce platforms to smaller solutions.
But there is an important limitation.
Similarity is not complementarity.
An embedding-based approach can easily identify that one drill is similar to another drill. It does not automatically know that a drill needs drill bits.
That relationship is not primarily about semantic similarity between two product descriptions. It is about how the products are used together. A drill and a drill bit may be very different products from a semantic point of view, while being directly related in practice.
Building Block 2: Transferring Complementary Products by Analogy
Before turning to external knowledge, there is another useful technique that can make better use of what is already available in the catalogue.
Suppose product A has a known and validated complementary product C. Now suppose a new product A' is very similar to A.
C can then become a candidate complementary product for A'.
In other words, known relationships can be transferred through the embedding space instead of being learned from scratch for every new product:
A has complementary product C + A' is similar to A → C becomes a candidate for A'
This is a relatively simple way of using existing knowledge to deal with new products.
However, there is an important catch.
Semantic similarity does not necessarily mean physical compatibility.
Two pool pumps can have almost identical product descriptions while having different flow rates. A filter that works with one may be completely unsuitable for the other.
This means that analogy cannot be used safely without a compatibility check.
Before recommending C for A', important product characteristics need to be checked. Depending on the product category, this could include dimensions, flow rate, capacity, voltage, connectors or other technical specifications.
Without this additional layer, the system can inherit both good and bad analogies.
Building Block 3: Using the General Knowledge of LLMs
This is where a more recent development becomes interesting.
Large language models, or LLMs, have learned much more than the meaning of words. During their training, they have also acquired a large amount of general knowledge about the world: how objects are used, which products are normally used together, and what constraints or precautions are associated with different uses.
Recent research suggests that this knowledge can be used to produce competitive recommendations in cold-start situations, even when there is no interaction history [6]. This has led to a growing body of research around the use of LLMs for cold-start recommendation [7].
The important point is that this knowledge does not have to be learned again from the behaviour of a particular website.
An LLM can already know that a drill is generally used with drill bits, a drill stand or safety glasses. An experienced salesperson in a hardware store would know the same thing without having to analyse the purchase history of that particular store.
This knowledge can be used to build several things:
- Purchase-intent taxonomies, such as drilling and fixing, protection or precision measurement, connected through business relationships rather than learned purely from statistics.
- Complementarity relationships that can be established before the first cross-purchase has taken place on the website.
- Semantic validation of proposed product pairs, helping to reject combinations that may look similar in vector space but make little sense in practice.
This is an important difference from the two previous approaches.
The question is no longer simply "what looks like this product?" or "what can we infer from similar products in the catalogue?"
It becomes:
"What logically goes with this product in the real world, even if there is no comparable product or purchase history in the catalogue yet?"
What This Means in Practice
By combining these three building blocks, a recommendation engine can cover the different stages of the cold-start problem:
-
Day 0: embeddings provide a reliable similarity layer as soon as the catalogue is ingested.
-
Still on Day 0: known relationships from similar products can be transferred by analogy, provided that compatibility is checked. When the catalogue already contains similar and validated cases, this is often a very efficient way to generate recommendations.
-
Still on Day 0, with a broader range of possibilities: the general knowledge of an LLM can add complementarity and purchase-intent information even when no useful analogy exists. There is still no need to wait for historical data.
-
Over time: real behavioural signals become available. They can then be used to refine and correct the initial recommendations produced by these three layers.
The result is a hybrid system. The recommendation engine does not have to start from scratch once behavioural data becomes available.
In Summary
Cold start is not simply a binary problem where a store either has data or does not.
It is a problem that can be addressed layer by layer.
Embedding-based similarity has already made it possible to recommend products without waiting weeks for enough traffic to accumulate. Analogy-based transfer makes it possible to reuse relationships that are already known in the catalogue, provided that compatibility is checked. LLMs take this one step further by bringing in general knowledge about how products are used and which products naturally go together.
The result is a recommendation system that can do more than find products that look similar. It can start identifying products that actually go together, from the very first day of a catalogue's existence.
The next question is how to implement this in practice. That is the subject of a dedicated article, based on our own implementation at RecoKit.
Sources
- G. Linden, B. Smith, J. York, Amazon.com Recommendations: Item-to-Item Collaborative Filtering, IEEE Internet Computing, 2003.
- J. Gope, S. K. Jain, A survey on solving cold start problem in recommender systems, IEEE ICCCA, 2017.
- R. Agrawal, R. Srikant, Fast Algorithms for Mining Association Rules in Large Databases, VLDB, 1994.
- M. Kaminskas, D. Bridge, Diversity, Serendipity, Novelty, and Coverage: A Survey and Empirical Analysis of Beyond-Accuracy Objectives in Recommender Systems, ACM TiiS, 2016.
- N. Reimers, I. Gurevych, Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks, EMNLP-IJCNLP, 2019.
- S. Sanner et al., Large Language Models are Competitive Near Cold-start Recommenders for Language- and Item-based Preferences, arXiv:2307.14225, 2023.
- W. Zhang et al., Cold-Start Recommendation towards the Era of Large Language Models (LLMs): A Comprehensive Survey and Roadmap, arXiv:2501.01945, 2025.