ParticleSearch Experiments: Test Search and Recommendations Without Guesswork
A storefront experiment is useful when two responsible choices could serve the same shopper job and the better option is uncertain. It is not a substitute for fixing a defect, cleaning product data, or deciding what the experience is supposed to accomplish.
ParticleSearch supports controlled variants across search and recommendation surfaces. This guide explains how to design a merchant-useful test, protect important behavior, and interpret click and revenue evidence without turning uncertainty into a confident story.
For example, comparing six recommendation cards with ten is a valid experiment only if both layouts are usable and the question is whether extra choice improves the shopping path. If the six-card version hides the only compatible accessory, the test mixes a quality defect with a presentation choice. Repair the defect first, then test the genuine uncertainty.
Product contract checked August 12, 2026. ParticleSearch behavior was verified against the current experiment workspace, session-assignment contract, storefront request path, recommendation event context, and revenue-attribution reporting.
The design and interpretation guidance is grounded in Microsoft Research's published experimentation patterns and the NIST statistical handbook. Store-specific thresholds and ship decisions remain merchant judgements.
The short answer
Experiment only where two responsible choices are genuinely uncertain
Use experiments for tradeoffs such as discovery versus density, one recommendation strategy versus another, or a layout choice whose outcome is unclear. Do not randomize defects, catalog truth, exact identifier behavior, or the basic ability to complete a purchase.
Good experiment question
Does a more prominent filter entry point increase useful product actions for mobile discovery queries without increasing empty intersections?
Not an experiment question
Is the mobile filter button broken or hidden behind another element? Repair the defect first, then test a deliberate design choice.
Good experiment question
Does a recommendation strategy based on cart context help shoppers complete a purchase better than a broad popularity baseline?
Not an experiment question
Should an exact SKU return the correct variant? That is an acceptance requirement, not an opinion to randomize.
Chapter 1 · The experiment's job
Isolate one decision while keeping the experience operational
A useful experiment has one question, a control, one or more variants, an eligible audience, a primary metric, and guardrails. The session remains in a stable variant so the recorded experience is coherent.
Stable assignment is not personalization. The shopper is not being profiled into a supposedly ideal experience. The experiment is creating comparable observations for a defined product decision.
Define one question
Name the surface, audience, primary metric, guardrail, and decision rule.
Assign a session
Eligible sessions receive a stable experiment variant for the measured experience.
Observe outcomes
Views, clicks, empty states, errors, and available order evidence stay connected to the variant.
Make a bounded decision
Adopt, reject, continue, or redesign based on evidence and operational context.
A worked hypothesis is more useful than “test the new thing”
Example: “For descriptive furniture searches, a meaning-first posture will increase product actions because the first page will contain more useful alternatives, while exact table-model queries remain unchanged.” That sentence names the audience, the change, the mechanism, the primary outcome, and the protected behaviour. If the result cannot be written that specifically, the test is probably still a product discussion rather than an experiment.
Chapter 2 · Protect the comparison
Random assignment is useful only when the groups differ by the intended choice
The value of an experiment is its counterfactual. Control estimates what would have happened without the change, while the variant estimates what happened with it. Random assignment helps make those groups comparable before the experience begins. It cannot rescue a test in which eligibility changes after exposure, the storefront delivers the wrong treatment, or a second release changes one path halfway through.
This is why assignment is a design choice rather than an implementation detail. Microsoft Research's pre-experiment guidance notes that the unit should match the product interaction and remain stable, and that network or shared-resource effects can violate a simple user-level comparison. For a merchant, the plain-language question is: can one shopper stay in one coherent experience, and can that shopper's outcome be observed without changing who qualifies?
Assignment unit
What is randomized, and can it stay in one experience?
ParticleSearch keeps an eligible session in one variant. That is useful for immediate search and recommendation outcomes. It is not the right design for a long-term customer question when the same shopper can return in another session, device, or channel.
Eligible population
Who could have received either choice before the treatment acted?
Define eligibility from facts available to both groups. Filtering later to people who clicked, saw a treatment-only module, or behaved differently after exposure selects on the outcome and breaks the comparison.
Treatment integrity
Did the two groups receive the intended difference and nothing else?
Confirm the selected profile, recommendation strategy, cards, requests, and events on both paths. A label in a dashboard is not proof that the storefront delivered the variant correctly.
Interference
Can one group change what the other group receives?
Shared inventory, campaign changes, manual merchandising, caches, or another experiment can alter both paths while the test runs. Record these changes because random assignment does not make the environment stand still.
A valid comparison path
Eligible before treatment
Stable random assignment
Verified experience and events
Comparable outcome
Do not condition the comparison on a treatment-only action. Compare the population assigned, then use exposure and diagnostic metrics to explain whether the treatment reached them.
An unexpected split is a stop sign, not a footnote. A planned 50/50 test will not produce exactly equal counts by chance, so eyeballing the ratio is not enough. A sample-ratio-mismatch test asks whether the difference is too large for the planned allocation and sample size. Microsoft documents allocation failures caused by assignment, redirects, missing logs, biased joins, and treatment-dependent filtering. Diagnose the cause before interpreting the outcome.
Design constraint · Define the evidence contract
Define what a decision-ready result must contain before choosing the test
The comparison-integrity chapter established that assignment alone cannot make a result trustworthy. Before choosing a test question, define the evidence the question must be able to produce: comparable exposure, a response tied to the hypothesis, protected guardrails, and a decision the merchant can take. This contract prevents a convenient dashboard metric from choosing the experiment after the fact.
A control and variant make a comparison more useful when the session assignment, exposure, event definition, and observation window are held steady. They do not remove seasonality, stock changes, campaigns, consent loss, or a broken event stream. Those conditions belong in the result record.
The practical question is not “which card has the larger percentage?” It is “did the variant answer the stated shopper job better, while the protected behaviours remained acceptable?” That is why a merchant should read exposure, response, guardrails, and decision together.
Exposure
Did eligible shoppers actually receive the control or variant?
Check assignment and views before reading a response rate. A missing or uneven exposure record makes the comparison unreliable.
Response
Did the experience change the behaviour named in the hypothesis?
Choose the closest observable outcome: a product action for discovery, a cart action for completion, or an order signal when the path supports it.
Guardrails
Did the change create a new failure while improving the primary metric?
Inspect exact queries, empty states, errors, latency, availability, and handoff. A higher click rate is not a win if the answer becomes less trustworthy.
Decision
What will the team do with this result?
Adopt, reject, continue, or redesign should be decided before the test ends. Otherwise the dashboard produces a number without an operating consequence.
ParticleSearch gives the comparison a usable operating home. The workspace keeps the hypothesis, audience, variants, schedule, metrics, and guardrails together, then leaves the result open to merchant judgement. It reduces the chance that a promising percentage gets separated from the exact queries or storefront states it was meant to protect.
With that evidence contract fixed, the next question is which product decision is uncertain enough to test. The guided questions below outrank a generic feature test because each one names the shopper job, competing choices, and guardrail the result must preserve.
Chapter 3 · Guided test questions
Start with a product question, not a dashboard feature
Meaning-first search
Does a meaning-led search posture help discovery queries without harming precise ones?
Control
Current search behavior
Variant
Meaning-first search behavior
Guardrail
Protect exact identifiers and known high-value queries.
Similar vs frequently bought together
Does the placement help shoppers compare substitutes or complete the purchase?
Control
Similar products
Variant
Frequently bought together
Guardrail
Watch relevance, empty modules, product clicks, and order evidence.
Similar vs complete the look
Does visual coordination outperform close product similarity for this placement?
Control
Similar products
Variant
Complete the look
Guardrail
Keep product eligibility and placement context comparable.
Best sellers vs cart context
Does cart-aware relevance improve the module over a broad popularity baseline?
Control
Best sellers
Variant
Cart-context recommendations
Guardrail
Check empty states, latency, click behavior, and verified order evidence.
Chapter 4 · Build the experiment brief
Define the decision before traffic enters the test
Hypothesis
State why a specific change should help a defined shopper job.
Surface
Choose search, recommendations, or all only when the same question genuinely spans both.
Variants
Use two clear variants when possible. ParticleSearch supports between two and eight.
Traffic
Choose how much eligible traffic enters the experiment. Keep a deliberate holdout.
Schedule
Set a start and end when campaigns, releases, or reporting discipline require it.
Primary metric
Choose the one behavior most directly connected to the hypothesis.
Guardrails
Protect errors, empty states, exact queries, latency, and important commerce paths.
Decision
Define what will cause adoption, rejection, continuation, or a redesigned test.
Discovery question
Primary metric: product action rate or product CTR. Guardrails: no-result rate, exact-query position, errors, and latency.
Recommendation question
Primary metric: recommendation product action or direct add. Guardrails: empty modules, duplicate products, product relevance, and order evidence.
Completion question
Primary metric: a defined cart or order outcome. Guardrails: search coverage, product handoff, availability, and the direct versus assisted attribution split.
ParticleSearch prevents overlapping running experiments on the same surface and schedule. This protects interpretation, but it does not make a weak hypothesis strong. The brief still needs a coherent shopper job and a decision the merchant can act on.
Chapter 5 · Plan the evidence
Traffic answers how fast you learn; the decision defines how much evidence you need
There is no universal “enough traffic” number. A test that tries to detect a large change in a common product action needs less evidence than a test looking for a small change in a rare order outcome. NIST's sample-size guidance makes the dependency explicit: evidence needs are set by the baseline rate, the change worth detecting, the accepted false-alarm risk, and the power to detect that change when it is real.
For merchants, the most important input is the minimum worthwhile effect. It turns a vague hope for “lift” into a business boundary. If a change would not be worth the engineering, operational, or shopper cost unless it improves the chosen outcome by a certain amount, the test must be capable of distinguishing at least that amount from noise.
Baseline
How often does the primary outcome happen today, using the same unit and denominator the experiment will use?
Minimum worthwhile effect
What is the smallest improvement large enough to justify the implementation cost and any trade-off?
Detection standard
How much risk of a false alarm and a missed worthwhile effect will the decision tolerate?
Eligible traffic
How many independent assignment units actually reach this surface and qualify for the test?
Sample adequacy
Can the available assignment units distinguish the minimum worthwhile effect from ordinary variation? If not, the honest output is a wider uncertainty range or an inconclusive result, not a universal minimum copied from another store.
Calendar coverage
Did the test include the weekly shopping rhythm and the conditions the decision must survive? More days can expose weekday, campaign, and novelty effects. They do not repair biased assignment, broken logging, or a treatment that was never delivered.
Do not stop because the dashboard turns green. Repeatedly checking an ordinary fixed-horizon test and ending it as soon as a threshold is crossed increases the chance of a false win. Microsoft's during-experiment guidance similarly calls for methods that account for repeated testing and early peeking. Either follow the pre-declared duration and analysis rule, or use a statistical method designed for sequential monitoring. Stop early for clear harm or broken data; do not improvise a success rule after seeing the result.
Chapter 6 · Read the evidence
Separate exposure, response, reliability, and commercial context
Exposure
Confirm eligible sessions and variant views before comparing response. Uneven or missing exposure can make a percentage look more precise than the underlying evidence.
Response
Use clicks and click-through rate when the hypothesis concerns discovery. Add-to-cart or order evidence may be more relevant for a commerce-completion question.
Reliability
Inspect empty responses, errors, eligibility, runtime changes, campaigns, and inventory before attributing the difference to the variant.
Commercial context
Shopify-verified attribution can connect order evidence to a variant when available. Missing revenue is not zero lift, and associated revenue still needs the attribution boundary.
Read the effect as a range, not a verdict generated by one threshold. The point estimate is the best single estimate from this sample. Its uncertainty range shows which other effects remain compatible with the evidence under the analysis assumptions. A statistically detectable difference can still be too small to matter, while a non-detectable result can still include an important benefit or harm when the sample is weak.
Keep absolute and relative differences together. Moving a rare action from one small rate to another can produce an impressive relative percentage while affecting few shopping journeys. The absolute change expresses reach; the relative change expresses scale against the baseline. Neither should be interpreted without counts and uncertainty.
1 · Can the comparison be trusted?
Inspect
Check planned versus observed assignment, missing exposure, event joins, treatment delivery, and major runtime changes.
Boundary
If allocation or instrumentation is unexplained, diagnose it before reading a winner.
2 · What changed, and by how much?
Inspect
Report the control value, variant value, absolute difference, relative difference, and uncertainty range for the primary metric.
Boundary
A small p-value does not tell you whether the effect is commercially useful.
3 · What did the improvement cost?
Inspect
Read guardrails and diagnostic metrics beside the primary outcome, using the trade-offs written in the brief.
Boundary
A product-action lift is not a win if exact identifiers, availability, latency, errors, or checkout regress.
4 · Is the effect stable enough to generalize?
Inspect
Inspect the result by date and the few pre-declared segments that can change the decision.
Boundary
A first-day spike, one campaign, or a tiny post-hoc segment is a reason to learn more, not a reason to ship.
5 · Which decision does the evidence support?
Inspect
Adopt, reject, continue, redesign, or declare the result inconclusive. Preserve the reasoning, not only the winning label.
Boundary
“Inconclusive” is correct when the data cannot separate a useful effect from a trivial or harmful one.
Worked interpretation · illustrative scenario
A higher estimate does not automatically support adoption
This scenario teaches the reasoning sequence. It is not a ParticleSearch benchmark and the observations are not claims about another store.
Question
For mobile descriptive furniture queries, compare the current search posture with meaning-first search. The promised mechanism is a more useful first page, not simply more clicks.
Read: The audience, mechanism, primary outcome, and exact-query guardrail are explicit.
Integrity
Assignment and view counts pass the planned allocation check. Query fixtures confirm that descriptive queries differ between variants and protected identifiers still reach the same products.
Read: The comparison is eligible for interpretation. If this check failed, the analysis would stop here.
Effect
The variant estimate is higher for product actions, but the uncertainty range still includes effects too small to justify changing the store-wide posture.
Read: Promising is not the same as decision-ready. The result has not yet resolved the business question.
Time and guardrails
The difference is largest on the first day and narrows later. Exact-query behaviour remains acceptable, while no-result and latency signals do not show a material regression.
Read: The guardrails are reassuring, but the time pattern could be novelty, traffic mix, or ordinary noise.
Decision
Keep the test running to its pre-declared evidence boundary, then adopt only if the sustained effect excludes changes too small to matter. Otherwise record an inconclusive result and redesign.
Read: The experiment protects the decision from an early, attractive percentage.
Treat segments as questions, not escape hatches
Pre-declare segments such as mobile, market, query class, or placement when each has a mechanism that can change the decision. If dozens of cuts are searched after the result, one will often look extreme by chance. Label that finding exploratory and confirm it in a later test.
Inspect the effect over time
A new interface can earn curiosity clicks at first or perform poorly while shoppers learn it. Campaigns, weekends, and inventory can also change the observed effect. A date view helps expose those mechanisms; it does not license choosing only the days that support the preferred story.
For the difference between direct, assisted, and unmatched order evidence, use the ParticleSearch revenue attribution guide.
Chapter 7 · When not to experiment
Some decisions need repair, judgment, or more evidence first
An obvious defect
If a result is broken, an event is missing, or a mobile layout is unusable, repair it and verify the fix.
Several unrelated changes
A new layout, new ranking posture, new cards, and new filters in one variant cannot explain which decision mattered.
A tiny or unstable audience
Insufficient traffic, short promotions, or changing inventory can make the result too noisy to support the decision.
A question with no merchant action
Do not collect experiment data if no outcome would change the product or operating decision.
A causal claim from attribution alone
Associated revenue can add context, but attribution and controlled comparison answer different questions.
Chapter 8 · Close the loop
End with a decision and a preserved record
Record the hypothesis, variants, dates, evidence, caveats, and decision. If the result is inconclusive, say what prevented the decision. If a variant wins, keep monitoring the protected searches and operational guardrails after it becomes normal storefront behavior.
For store-level search posture, read the ParticleSearch search profiles guide. For query-level controls, use the ranking, synonyms, and redirects guide.
The experiment earns a decision only when the comparison is valid, the observed effect is large enough to matter, the uncertainty is narrow enough for the choice, and the protected shopping paths remain acceptable. If any of those conditions fail, “no decision yet” is more useful than a winner chosen from an attractive chart.
The shareworthy learning is not that one variant won on one store. It is the mechanism the test clarified: which shopper, in which context, responded to which product decision, at what cost, and with what remaining uncertainty. That record helps the next merchant decision even when the answer is to keep the current experience.
Primary references
These sources establish the experimentation principles used above. ParticleSearch product behaviour is described separately and was checked against the current implementation.
- Microsoft Research: pre-experiment trustworthiness patterns Hypotheses, success metrics, power, assignment units, counterfactual logging, and engineering bias.
- Microsoft Research: during-experiment trustworthiness patterns Metric taxonomy, early peeking, guardrails, stable segments, and novelty effects.
- Microsoft Research: post-experiment trustworthiness patterns Treatment integrity, denominator changes, triggered analysis, trade-offs, reproduction, and preserved learning.
- Microsoft Research: diagnosing sample ratio mismatch Why an unexpected allocation is a data-quality symptom that must be diagnosed before reading effects.
- NIST: sample sizes required for tests of proportions How baseline proportion, detectable change, significance, and power determine evidence requirements.
- Microsoft Research: external validity of online experiments Why day-to-day movement, novelty, and future conditions limit how confidently a short test generalizes.
Acceptance judgement
A useful experiment earns a decision, not just a winner
For the worked meaning-first search example, accept the result only if allocation and exposure are trustworthy, the effect is meaningful enough for the written rule, the exact-query and commerce guardrails pass, and the time pattern is not carrying the conclusion. Otherwise keep the current experience or continue measuring.
Strongest alternative: when the question is really a catalogue, placement, or instrumentation defect, repair that layer before experimenting. A controlled comparison cannot make a broken input or missing event contract interpretable.
Final decision
Ship only when the experiment record names the mechanism, population, primary metric, guardrails, uncertainty, and next owner. Treat “inconclusive” as the correct acceptance judgement when the evidence cannot separate a useful effect from noise or harm.