Ecommerce Recommendation Quality: Eligibility, Judgement Sets, Regression Tests, and Measurement
Recommendation quality is not a click-through rate. It is the integrity of the path from shopper context to eligible candidates, a defensible relationship, useful ordering, a visible module, and interpretable behaviour.
A shelf can earn clicks while showing duplicates, incompatible products, or a misleading heading. It can also be excellent and receive few clicks because it appears when the shopper does not need another choice. Quality review must inspect the answer before using outcomes to decide whether the answer helped.
Quality contract
Six layers must remain true at the same time
Separate the layers so a reviewer can locate the first broken contract. Changing ranking will not fix a missing storefront block. Changing the heading will not fix incompatible candidates.
| Layer | Question | Failure | Evidence |
|---|---|---|---|
| Context | What product, basket, page, market, and shopper job is the shelf about? | The system evaluates the wrong anchor or a stale cart. | Recorded context and placement |
| Eligibility | Which candidates are allowed to appear at all? | Hidden, unavailable, duplicate, current, or incompatible products enter the set. | Exclusion reason and eligible count |
| Relationship | Why is each candidate relevant to this context? | Popularity, title overlap, or category membership masquerades as a useful relationship. | Named strategy and relationship source |
| Ordering | Why should one eligible candidate appear before another? | The first cards are technically valid but weak, repetitive, or commercially unhelpful. | Rank order and useful tradeoffs |
| Presentation | Can the shopper understand and act on the shelf? | Wrong heading, hidden module, mobile overflow, unclear price or unavailable action. | Visible cards and interaction state |
| Measurement | Can exposure and downstream actions be interpreted? | Returned products count as views, missing events become zero engagement, or attribution is treated as causality. | View, click, cart, order, empty, and error states |
Offline measures
Use several measures because recommendation quality has several ways to fail
One score cannot represent eligibility, retrieval, ordering, diversity, and coverage without hiding the reason a shelf changed. Keep hard violations visible and choose softer measures from the decision the strategy promises.
The formulas below are evaluation definitions, not industry benchmarks. Each store still chooses its K, its judgement set, and the boundary that makes a result acceptable.
| Measure | Question | Definition | Boundary |
|---|---|---|---|
| Eligible-set pass rate | How often does every visible candidate pass the non-negotiable rules? | Reviewed shelves with zero eligibility failures ÷ reviewed shelves | A high rate says nothing about whether the candidates are useful. |
| Precision at K | Among the first K candidates, how many satisfy the reviewed recommendation job? | Relevant candidates in the first K ÷ K | Requires a clear relevance judgement and does not measure missing good candidates. |
| Expected-candidate coverage | How often does the set contain products reviewers said should be considered? | Expected candidates retrieved ÷ expected candidates recorded | The judgement set is incomplete by design and should not become an exhaustive catalogue truth. |
| Prohibited-candidate rate | How often do known-wrong products enter the visible set? | Prohibited visible candidates ÷ visible reviewed candidates | A zero rate is necessary for sensitive relationships, but it does not prove ranking quality. |
| Set diversity | Does the shelf provide distinct tradeoffs rather than repeated families? | Store-defined coverage of useful difference across the visible set | More diversity is not always better. Compatibility and coherence remain constraints. |
| Empty-placement rate | How often does a configured placement have no eligible visible result? | Empty placements ÷ attempted placements | An empty result can be correct. Segment by strategy, placement, and catalogue context. |
Hard and soft rules
Do not average a hard failure into a relevance score
A product that cannot be purchased, must not be shown, is incompatible, or is already the anchor should fail eligibility. It should not survive because it has a strong title match, high sales, or a favourable relationship score.
Once candidates pass hard rules, soft ranking can balance closeness, usefulness, diversity, price, style, popularity, recency, or merchant priorities. This split prevents an optimisation from trading away a non-negotiable constraint.
Hard gate
Allowed or not allowed. Failure removes the candidate.
Examples
Visibility, availability, market, compatibility, identity, current basket
Soft ordering
Better or worse among eligible candidates.
Examples
Relationship strength, diversity, price proximity, style, recency, performance
Reviewer calibration
Reviewer disagreement is evidence that the relationship is underspecified
A merchandiser may value visual coherence, a support specialist may recognise a compatibility risk, and a storefront owner may see that the card cannot explain the difference. Their disagreement should not be flattened into an average score before the rule is understood.
Calibrate on a small set, discuss the reason behind each label, then refine the relationship contract. The goal is not universal agreement on taste. It is consistent handling of the boundaries the store considers important.
- 1
Write the relationship first
Define the shopper job, hard constraints, allowed tradeoffs, and visible comparison unit before showing candidates to reviewers.
- 2
Review independently
Have category and storefront reviewers score the same anchors without seeing each other’s labels.
- 3
Resolve disagreement by rule
Turn repeated disagreement into a clearer compatibility, diversity, identity, or heading rule rather than averaging incompatible opinions.
- 4
Preserve the reason
Save why a product is expected, acceptable, weak, or prohibited so future catalogue changes can be judged against the same contract.
- 5
Revisit stale judgements
Recheck fixtures after assortment, market, compatibility, product-family, or strategy changes.
Human review
Use a rubric that can reject a shelf before traffic reaches it
Review several named anchors with a consistent rubric. Eligibility is binary. The remaining dimensions describe the usefulness of the set, not just each card in isolation. Keep reviewer notes so disagreements reveal where the recommendation contract is unclear.
| Dimension | Scale | Acceptance meaning |
|---|---|---|
| Eligibility | Pass / fail | No candidate violates visibility, availability, identity, market, or compatibility rules |
| Relationship | 0 to 3 | 0 unrelated; 1 weak association; 2 credible; 3 clearly supports the named shopper job |
| Diversity | 0 to 2 | 0 duplicates; 1 limited choice; 2 useful tradeoffs without incoherence |
| Explainability | 0 to 2 | 0 heading conflicts; 1 plausible; 2 shopper can understand why the row belongs |
| Actionability | 0 to 2 | 0 cannot buy; 1 friction remains; 2 required price, variant, stock, and action are clear |
A useful approval rule
Every card must pass eligibility. No card may score zero for relationship or explainability. The set must provide at least two genuinely different choices unless fewer eligible candidates exist. Store-specific risk can make this rule stricter.
Release gates
A recommendation change is a release across data, answer, interface, and evidence
Changing a strategy can expose a different dependency in the catalogue and a different risk in the storefront. Release approval should therefore prove more than “the cards loaded.”
| Gate | Evidence | Decision |
|---|---|---|
| Catalogue gate | Required context and relationship fields are present for the intended categories. | Block when the missing field can produce unsafe or misleading products. |
| Offline quality gate | Judgement and regression sets meet store-defined eligibility and relevance conditions. | Block on prohibited candidates; review softer score movement. |
| Render gate | The shelf is visible, labelled honestly, responsive, and connected to the intended product action. | Block on inaccessible or misleading presentation. |
| Event gate | View, empty, error, product action, and order evidence can be distinguished. | Do not interpret engagement when exposure or joins are incomplete. |
| Experiment gate | The causal question, treatment, primary metric, guardrails, and stop rule are written. | Do not run a broad test that changes candidates, placement, cards, and search together. |
| Rollback gate | The prior strategy or placement configuration can be restored and its evidence remains identifiable. | Do not publish a high-risk change without a known recovery path. |
The broader search QA guide explains how to maintain fixtures, severity, evidence, and independent release gates across discovery surfaces.
Regression coverage
Protect catalogue edges, not only the best-looking examples
A regression set should include common, sparse, sensitive, unavailable, multi-context, and market-specific fixtures. Save the catalogue version, strategy, placement, context, expected positives, prohibited candidates, and visible order.
Popular anchor
Changes to the highest-traffic product family
Expected
Known credible candidates remain near the top
Prohibited
Anchor and duplicate family do not reappear
Sparse-data product
Unsafe fallbacks and dishonest headings
Expected
A labelled fallback or honest empty state
Prohibited
Unrelated popularity row presented as similarity
Compatibility-sensitive product
Relationship evidence overriding hard fit
Expected
Only verified compatible candidates
Prohibited
Shared-number or shared-title false positives
Low-stock or unavailable anchor
Recovery quality and candidate eligibility
Expected
Purchasable alternatives when the job calls for them
Prohibited
Unavailable candidates or a second copy of the anchor
Multi-item cart
Context collision and repeated basket items
Expected
One remaining useful need is represented
Prohibited
Products or covered needs already in the cart
Market-restricted product
Context leakage across availability boundaries
Expected
Candidates purchasable in the active market
Prohibited
Products visible only in another market
Failure states
Missing, empty, wrong, unseen, and unmeasured are different incidents
| Observed state | Possible boundary | First useful response |
|---|---|---|
| No shelf | Disabled placement, missing storefront block, render failure | Confirm configuration and runtime before judging recommendation quality |
| Empty shelf | No eligible candidates, missing context, sparse relationship evidence | Preserve empty as evidence; inspect candidate and exclusion counts |
| Wrong products | Bad relationship, unsafe fallback, stale catalogue, missing eligibility rule | Save anchor, returned IDs, context, strategy, and catalogue version |
| Right products, weak interaction | Poor placement, heading, cards, density, or low shopper need | Review visibility and action path before changing candidates |
| Clicks without trusted orders | Broken lineage, reporting window, checkout identity, or genuinely exploratory use | Validate joins before interpreting revenue |
| Attributed orders without proven lift | Observed sequence is being treated as causal effect | Use a controlled experiment for the causal question |
Behaviour and causality
Measure whether the shelf was seen before interpreting what happened next
Returned candidates are not impressions. A module can be below the fold, hidden on mobile, fail to render, or be replaced before it becomes visible. Record visible exposure, product action, direct add, cart state, checkout, verified order, empty response, and error as different facts.
Attribution connects an observed recommendation touchpoint to a later order. It helps explain the path and reconcile revenue. It does not prove the order would not have happened without the shelf. Use a controlled experiment when the decision depends on incremental effect, and keep product-page conversion, checkout completion, empty-module rate, and quality regressions as guardrails.
The recommendation analytics guide provides the complete event and attribution contract.
Operating loop
Move from incident evidence to the smallest responsible change
- 1
Name the job and surface
State the anchor or cart context, intended strategy, placement, heading, and expected shopper action.
- 2
Reproduce the visible state
Save the actual cards, order, device width, market, stock state, and whether the module became visible.
- 3
Find the first failed layer
Check context, eligibility, relationship, ordering, presentation, then event integrity.
- 4
Change one contract
Repair data, exclusion, strategy, placement, heading, density, or event handling without bundling unrelated changes.
- 5
Run judgement and regression sets
Recheck expected and prohibited candidates across high-risk catalogue fixtures.
- 6
Observe and decide
Review visible exposure and guardrails, then use an experiment only when the causal decision warrants it.
ParticleSearch fit
ParticleSearch preserves recommendation state as evidence, including when nothing appears
ParticleSearch recommendation placements report visible recommendation views separately from empty outcomes and errors. Product actions retain the recommendation strategy, placement, and available relationship evidence. The current product and contextual products are excluded from the visible shelf.
Merchants can keep a placement theme-controlled or manage its strategy, heading, product count, and shown state in the dashboard. Recommendation experiments compare named strategies, while revenue reporting only presents Shopify-verified order totals when a verified recommendation touchpoint can be connected to the order path.
This does not make every recommendation correct automatically. It gives the team enough state to distinguish no candidate, no view, failed render, weak interaction, and verified downstream action, then change the placement or strategy without treating every problem as “bad recommendations.”
Apply the framework to a recommendation job
Use the similar-products guide for alternatives and the cross-sell guide for complements and cart context.