Skip to article
Product Recommendations 2026-07-2934 min read

Ecommerce Recommendation Quality: Eligibility, Judgement Sets, Regression Tests, and Measurement

Recommendation quality is not a click-through rate. It is the integrity of the path from shopper context to eligible candidates, a defensible relationship, useful ordering, a visible module, and interpretable behaviour.

A shelf can earn clicks while showing duplicates, incompatible products, or a misleading heading. It can also be excellent and receive few clicks because it appears when the shopper does not need another choice. Quality review must inspect the answer before using outcomes to decide whether the answer helped.

Quality contract

Six layers must remain true at the same time

Separate the layers so a reviewer can locate the first broken contract. Changing ranking will not fix a missing storefront block. Changing the heading will not fix incompatible candidates.

LayerQuestionFailureEvidence
ContextWhat product, basket, page, market, and shopper job is the shelf about?The system evaluates the wrong anchor or a stale cart.Recorded context and placement
EligibilityWhich candidates are allowed to appear at all?Hidden, unavailable, duplicate, current, or incompatible products enter the set.Exclusion reason and eligible count
RelationshipWhy is each candidate relevant to this context?Popularity, title overlap, or category membership masquerades as a useful relationship.Named strategy and relationship source
OrderingWhy should one eligible candidate appear before another?The first cards are technically valid but weak, repetitive, or commercially unhelpful.Rank order and useful tradeoffs
PresentationCan the shopper understand and act on the shelf?Wrong heading, hidden module, mobile overflow, unclear price or unavailable action.Visible cards and interaction state
MeasurementCan exposure and downstream actions be interpreted?Returned products count as views, missing events become zero engagement, or attribution is treated as causality.View, click, cart, order, empty, and error states

Offline measures

Use several measures because recommendation quality has several ways to fail

One score cannot represent eligibility, retrieval, ordering, diversity, and coverage without hiding the reason a shelf changed. Keep hard violations visible and choose softer measures from the decision the strategy promises.

The formulas below are evaluation definitions, not industry benchmarks. Each store still chooses its K, its judgement set, and the boundary that makes a result acceptable.

MeasureQuestionDefinitionBoundary
Eligible-set pass rateHow often does every visible candidate pass the non-negotiable rules?Reviewed shelves with zero eligibility failures ÷ reviewed shelvesA high rate says nothing about whether the candidates are useful.
Precision at KAmong the first K candidates, how many satisfy the reviewed recommendation job?Relevant candidates in the first K ÷ KRequires a clear relevance judgement and does not measure missing good candidates.
Expected-candidate coverageHow often does the set contain products reviewers said should be considered?Expected candidates retrieved ÷ expected candidates recordedThe judgement set is incomplete by design and should not become an exhaustive catalogue truth.
Prohibited-candidate rateHow often do known-wrong products enter the visible set?Prohibited visible candidates ÷ visible reviewed candidatesA zero rate is necessary for sensitive relationships, but it does not prove ranking quality.
Set diversityDoes the shelf provide distinct tradeoffs rather than repeated families?Store-defined coverage of useful difference across the visible setMore diversity is not always better. Compatibility and coherence remain constraints.
Empty-placement rateHow often does a configured placement have no eligible visible result?Empty placements ÷ attempted placementsAn empty result can be correct. Segment by strategy, placement, and catalogue context.

Hard and soft rules

Do not average a hard failure into a relevance score

A product that cannot be purchased, must not be shown, is incompatible, or is already the anchor should fail eligibility. It should not survive because it has a strong title match, high sales, or a favourable relationship score.

Once candidates pass hard rules, soft ranking can balance closeness, usefulness, diversity, price, style, popularity, recency, or merchant priorities. This split prevents an optimisation from trading away a non-negotiable constraint.

Hard gate

Allowed or not allowed. Failure removes the candidate.

Examples

Visibility, availability, market, compatibility, identity, current basket

Soft ordering

Better or worse among eligible candidates.

Examples

Relationship strength, diversity, price proximity, style, recency, performance

Reviewer calibration

Reviewer disagreement is evidence that the relationship is underspecified

A merchandiser may value visual coherence, a support specialist may recognise a compatibility risk, and a storefront owner may see that the card cannot explain the difference. Their disagreement should not be flattened into an average score before the rule is understood.

Calibrate on a small set, discuss the reason behind each label, then refine the relationship contract. The goal is not universal agreement on taste. It is consistent handling of the boundaries the store considers important.

  1. 1

    Write the relationship first

    Define the shopper job, hard constraints, allowed tradeoffs, and visible comparison unit before showing candidates to reviewers.

  2. 2

    Review independently

    Have category and storefront reviewers score the same anchors without seeing each other’s labels.

  3. 3

    Resolve disagreement by rule

    Turn repeated disagreement into a clearer compatibility, diversity, identity, or heading rule rather than averaging incompatible opinions.

  4. 4

    Preserve the reason

    Save why a product is expected, acceptable, weak, or prohibited so future catalogue changes can be judged against the same contract.

  5. 5

    Revisit stale judgements

    Recheck fixtures after assortment, market, compatibility, product-family, or strategy changes.

Human review

Use a rubric that can reject a shelf before traffic reaches it

Review several named anchors with a consistent rubric. Eligibility is binary. The remaining dimensions describe the usefulness of the set, not just each card in isolation. Keep reviewer notes so disagreements reveal where the recommendation contract is unclear.

DimensionScaleAcceptance meaning
EligibilityPass / failNo candidate violates visibility, availability, identity, market, or compatibility rules
Relationship0 to 30 unrelated; 1 weak association; 2 credible; 3 clearly supports the named shopper job
Diversity0 to 20 duplicates; 1 limited choice; 2 useful tradeoffs without incoherence
Explainability0 to 20 heading conflicts; 1 plausible; 2 shopper can understand why the row belongs
Actionability0 to 20 cannot buy; 1 friction remains; 2 required price, variant, stock, and action are clear

A useful approval rule

Every card must pass eligibility. No card may score zero for relationship or explainability. The set must provide at least two genuinely different choices unless fewer eligible candidates exist. Store-specific risk can make this rule stricter.

Release gates

A recommendation change is a release across data, answer, interface, and evidence

Changing a strategy can expose a different dependency in the catalogue and a different risk in the storefront. Release approval should therefore prove more than “the cards loaded.”

GateEvidenceDecision
Catalogue gateRequired context and relationship fields are present for the intended categories.Block when the missing field can produce unsafe or misleading products.
Offline quality gateJudgement and regression sets meet store-defined eligibility and relevance conditions.Block on prohibited candidates; review softer score movement.
Render gateThe shelf is visible, labelled honestly, responsive, and connected to the intended product action.Block on inaccessible or misleading presentation.
Event gateView, empty, error, product action, and order evidence can be distinguished.Do not interpret engagement when exposure or joins are incomplete.
Experiment gateThe causal question, treatment, primary metric, guardrails, and stop rule are written.Do not run a broad test that changes candidates, placement, cards, and search together.
Rollback gateThe prior strategy or placement configuration can be restored and its evidence remains identifiable.Do not publish a high-risk change without a known recovery path.

The broader search QA guide explains how to maintain fixtures, severity, evidence, and independent release gates across discovery surfaces.

Regression coverage

Protect catalogue edges, not only the best-looking examples

A regression set should include common, sparse, sensitive, unavailable, multi-context, and market-specific fixtures. Save the catalogue version, strategy, placement, context, expected positives, prohibited candidates, and visible order.

Popular anchor

Changes to the highest-traffic product family

Expected

Known credible candidates remain near the top

Prohibited

Anchor and duplicate family do not reappear

Sparse-data product

Unsafe fallbacks and dishonest headings

Expected

A labelled fallback or honest empty state

Prohibited

Unrelated popularity row presented as similarity

Compatibility-sensitive product

Relationship evidence overriding hard fit

Expected

Only verified compatible candidates

Prohibited

Shared-number or shared-title false positives

Low-stock or unavailable anchor

Recovery quality and candidate eligibility

Expected

Purchasable alternatives when the job calls for them

Prohibited

Unavailable candidates or a second copy of the anchor

Multi-item cart

Context collision and repeated basket items

Expected

One remaining useful need is represented

Prohibited

Products or covered needs already in the cart

Market-restricted product

Context leakage across availability boundaries

Expected

Candidates purchasable in the active market

Prohibited

Products visible only in another market

Failure states

Missing, empty, wrong, unseen, and unmeasured are different incidents

Observed statePossible boundaryFirst useful response
No shelfDisabled placement, missing storefront block, render failureConfirm configuration and runtime before judging recommendation quality
Empty shelfNo eligible candidates, missing context, sparse relationship evidencePreserve empty as evidence; inspect candidate and exclusion counts
Wrong productsBad relationship, unsafe fallback, stale catalogue, missing eligibility ruleSave anchor, returned IDs, context, strategy, and catalogue version
Right products, weak interactionPoor placement, heading, cards, density, or low shopper needReview visibility and action path before changing candidates
Clicks without trusted ordersBroken lineage, reporting window, checkout identity, or genuinely exploratory useValidate joins before interpreting revenue
Attributed orders without proven liftObserved sequence is being treated as causal effectUse a controlled experiment for the causal question

Behaviour and causality

Measure whether the shelf was seen before interpreting what happened next

Returned candidates are not impressions. A module can be below the fold, hidden on mobile, fail to render, or be replaced before it becomes visible. Record visible exposure, product action, direct add, cart state, checkout, verified order, empty response, and error as different facts.

Attribution connects an observed recommendation touchpoint to a later order. It helps explain the path and reconcile revenue. It does not prove the order would not have happened without the shelf. Use a controlled experiment when the decision depends on incremental effect, and keep product-page conversion, checkout completion, empty-module rate, and quality regressions as guardrails.

The recommendation analytics guide provides the complete event and attribution contract.

Operating loop

Move from incident evidence to the smallest responsible change

  1. 1

    Name the job and surface

    State the anchor or cart context, intended strategy, placement, heading, and expected shopper action.

  2. 2

    Reproduce the visible state

    Save the actual cards, order, device width, market, stock state, and whether the module became visible.

  3. 3

    Find the first failed layer

    Check context, eligibility, relationship, ordering, presentation, then event integrity.

  4. 4

    Change one contract

    Repair data, exclusion, strategy, placement, heading, density, or event handling without bundling unrelated changes.

  5. 5

    Run judgement and regression sets

    Recheck expected and prohibited candidates across high-risk catalogue fixtures.

  6. 6

    Observe and decide

    Review visible exposure and guardrails, then use an experiment only when the causal decision warrants it.

ParticleSearch fit

ParticleSearch preserves recommendation state as evidence, including when nothing appears

ParticleSearch recommendation placements report visible recommendation views separately from empty outcomes and errors. Product actions retain the recommendation strategy, placement, and available relationship evidence. The current product and contextual products are excluded from the visible shelf.

Merchants can keep a placement theme-controlled or manage its strategy, heading, product count, and shown state in the dashboard. Recommendation experiments compare named strategies, while revenue reporting only presents Shopify-verified order totals when a verified recommendation touchpoint can be connected to the order path.

This does not make every recommendation correct automatically. It gives the team enough state to distinguish no candidate, no view, failed render, weak interaction, and verified downstream action, then change the placement or strategy without treating every problem as “bad recommendations.”

Apply the framework to a recommendation job

Use the similar-products guide for alternatives and the cross-sell guide for complements and cart context.