Similar Products

Similar products finds products that are alike based on how you describe them, not on what shoppers did. For a given seed product, the output is the products whose titles and descriptions read most similarly. A leather weekend bag surfaces other leather bags and travel duffels, because their catalog text says similar things.

In the algorithm picker where you create a recipe, this algorithm is called Similar items.

Use it on product detail pages ("You may also like"), on out-of-stock pages (close substitutes for the thing the visitor wanted), or anywhere a content-based "more like this" makes sense.

How it works

The other co-occurrence algorithms (viewed together, bought together) learn from shopper behavior, so they need traffic and order history before they can say anything. Similar products takes a different route:

  1. It reads your Live Catalog. For every product, the engine takes the title and description stored in your Live Catalog.
  2. It turns each product's text into an embedding. An embedding is a numerical fingerprint of what the text means: products described in similar language end up with similar fingerprints, even when they share no exact words.
  3. It builds the feed ahead of time. During the run, the engine compares every product's fingerprint against all the others, keeps the closest matches, and writes the results into the recipe's output dataset: a normal versioned A vs B dataset, served from the edge like any other.

Nothing is computed while a visitor waits. The heavy lifting happens during the run; serving is the same fast key lookup every dataset gets.

Works from day one: the cold-start answer

Because Similar products needs only your Live Catalog, it produces full results with zero shopper traffic. A brand-new store, a brand-new product, a market you just launched in: all covered from the first run. This makes it a natural fallback step for behavior-based recipes: viewed-together for products with history, similar-products for everything else.

Requirements

One thing: your project's Live Catalog, the always-fresh copy of your products A vs B keeps (see Set up your product catalog). It is the same catalog every recipe uses; you do not link a separate dataset. Where other algorithms use the catalog for enrichment, Similar products requires it: the catalog text is the data source.

The engine reads these fields from each product:

FieldRequired?Used for
titleAt least one of title / descriptionThe text that gets embedded
descriptionAt least one of title / descriptionThe text that gets embedded: richer descriptions give better matches
categoryOptionalThe sameCategoryOnly parameter
availabilityOptionalThe excludeOutOfStock filter, applied live at serve time

A product with no title and no description cannot be matched. It is skipped and counted in the run stats: it will not appear as a seed or as a recommendation.

Better descriptions, better matches

The engine matches on meaning, so the quality of your catalog text is the quality of your recommendations. A one-word title gives the engine little to go on; a sentence or two of real description gives it a lot.

Parameters

ParameterTypeDefaultDescription
topKnumber (1–50)12How many similar products to include per seed product.
sameCategoryOnlybooleanfalseWhen true, a product is only matched against products in the same catalog category. Use this when cross-category matches (a "leather" wallet next to a "leather" sofa) would look wrong.
excludeOutOfStockbooleanoffOff when you create a recipe, so out-of-stock products can still be recommended until you turn this on. Once on, the check runs live at serve time against each product's Live Catalog availability: a product is dropped when its availability is out_of_stock or removed, so a product that sells out stops appearing within about a minute, with no re-run.

The order-based parameters (windowDays, topN, minSupport, perCategory) do not apply to Similar products and are ignored if set.

When it runs

Like every recipe, an enabled Similar products recipe runs nightly at 03:00 UTC. Each nightly run picks up any Live Catalog changes since the last run automatically (see "What you see during a run" below).

On top of the nightly schedule, enabling the recipe kicks off a run immediately. Flipping it on re-embeds and matches right away, so you never sit on a stale (or empty) feed. A Live Catalog change on its own does not trigger an immediate run: it's picked up on the next nightly run instead.

The manual Run now button works as usual.

What you see during a run

A Similar products run has two stages, and the recipe page shows which one is in progress:

  1. "Embedding catalog…": the engine is turning catalog text into fingerprints. This stage is skipped entirely when the Live Catalog has not changed since the last run (same catalog epoch), which is why nightly runs on a stable catalog finish quickly.
  2. "Building similar-products feed…": the engine is matching every product against the rest and writing the output rows.

When the run completes, the output dataset gets a new live version atomically: exactly like every other recipe. A failed run never overwrites the previous version.

Run stats

After a run, the recipe page shows these Similar-products-specific stats alongside the standard ones:

Stat shown on screenMeaning
Products embeddedHow many products were turned into embeddings in this run.
Skipped (no text)Products with no title and no description: they were left out entirely. A non-zero number here means some catalog rows need text added.
Seeds with no matchesSeed products for which no recommendations survived the filters: for example, everything similar was filtered out by sameCategoryOnly or excludeOutOfStock. These seeds get no output row, so lookups for them fall through to the recipe's fallback chain.
Dropped (no catalog match)Items excluded because their SKU was not found in the Live Catalog. These are not written to the output. In the API response this field is called droppedNoCatalog.

The embeddedCatalogEpoch field in the stored run stats records which Live Catalog snapshot the embeddings were built from: when this matches the current catalog epoch, the next run skips the embedding stage.

Reading the output

The output is a standard FEED dataset, keyed by the seed product's SKU, with the same row shape as viewed-together and bought-together: { sku, score } per item, enriched with live-catalog product details. Price, compare-at price, availability, and stock resolve live at serve time (price in minor units with a currency code). Everything on the Reference page about keys, row shape, and reading applies.

Was this helpful?