ParticleSearch Experiments: Test Search and Recommendations Without Guesswork
A storefront experiment is useful when two responsible choices could serve the same shopper job and the better option is uncertain. It is not a substitute for fixing a defect, cleaning product data, or deciding what the experience is supposed to accomplish.
ParticleSearch supports controlled variants across search and recommendation surfaces. This guide explains how to design a merchant-useful test, protect important behavior, and interpret click and revenue evidence without turning uncertainty into a confident story.
Experiment behavior checked July 28, 2026. The guide was verified against the current ParticleSearch experiment workspace, assignment contract, storefront request path, recommendation event context, and revenue-attribution reporting.
The short answer
Experiment only where two responsible choices are genuinely uncertain
Use experiments for tradeoffs such as discovery versus density, one recommendation strategy versus another, or a layout choice whose outcome is unclear. Do not randomize defects, catalog truth, exact identifier behavior, or the basic ability to complete a purchase.
Good experiment question
Does a more prominent filter entry point increase useful product actions for mobile discovery queries without increasing empty intersections?
Not an experiment question
Is the mobile filter button broken or hidden behind another element? Repair the defect first, then test a deliberate design choice.
Good experiment question
Does a recommendation strategy based on cart context help shoppers complete a purchase better than a broad popularity baseline?
Not an experiment question
Should an exact SKU return the correct variant? That is an acceptance requirement, not an opinion to randomize.
Chapter 1 · The experiment's job
Isolate one decision while keeping the experience operational
A useful experiment has one question, a control, one or more variants, an eligible audience, a primary metric, and guardrails. The session remains in a stable variant so the recorded experience is coherent.
Stable assignment is not personalization. The shopper is not being profiled into a supposedly ideal experience. The experiment is creating comparable observations for a defined product decision.
Define one question
Name the surface, audience, primary metric, guardrail, and decision rule.
Assign a session
Eligible sessions receive a stable experiment variant for the measured experience.
Observe outcomes
Views, clicks, empty states, errors, and available order evidence stay connected to the variant.
Make a bounded decision
Adopt, reject, continue, or redesign based on evidence and operational context.
A worked hypothesis is more useful than “test the new thing”
Example: “For descriptive furniture searches, a meaning-first posture will increase product actions because the first page will contain more useful alternatives, while exact table-model queries remain unchanged.” That sentence names the audience, the change, the mechanism, the primary outcome, and the protected behaviour. If the result cannot be written that specifically, the test is probably still a product discussion rather than an experiment.
Chapter 2 · Guided test questions
Start with a product question, not a dashboard feature
Meaning-first search
Does a meaning-led search posture help discovery queries without harming precise ones?
Control
Current search behavior
Variant
Meaning-first search behavior
Guardrail
Protect exact identifiers and known high-value queries.
Similar vs frequently bought together
Does the placement help shoppers compare substitutes or complete the purchase?
Control
Similar products
Variant
Frequently bought together
Guardrail
Watch relevance, empty modules, product clicks, and order evidence.
Similar vs complete the look
Does visual coordination outperform close product similarity for this placement?
Control
Similar products
Variant
Complete the look
Guardrail
Keep product eligibility and placement context comparable.
Best sellers vs cart context
Does cart-aware relevance improve the module over a broad popularity baseline?
Control
Best sellers
Variant
Cart-context recommendations
Guardrail
Check empty states, latency, click behavior, and verified order evidence.
Chapter 3 · Build the experiment brief
Define the decision before traffic enters the test
Hypothesis
State why a specific change should help a defined shopper job.
Surface
Choose search, recommendations, or all only when the same question genuinely spans both.
Variants
Use two clear variants when possible. ParticleSearch supports between two and eight.
Traffic
Choose how much eligible traffic enters the experiment. Keep a deliberate holdout.
Schedule
Set a start and end when campaigns, releases, or reporting discipline require it.
Primary metric
Choose the one behavior most directly connected to the hypothesis.
Guardrails
Protect errors, empty states, exact queries, latency, and important commerce paths.
Decision
Define what will cause adoption, rejection, continuation, or a redesigned test.
Discovery question
Primary metric: product action rate or product CTR. Guardrails: no-result rate, exact-query position, errors, and latency.
Recommendation question
Primary metric: recommendation product action or direct add. Guardrails: empty modules, duplicate products, product relevance, and order evidence.
Completion question
Primary metric: a defined cart or order outcome. Guardrails: search coverage, product handoff, availability, and the direct versus assisted attribution split.
ParticleSearch prevents overlapping running experiments on the same surface and schedule. This protects interpretation, but it does not make a weak hypothesis strong. The brief still needs a coherent shopper job and a decision the merchant can act on.
Chapter 4 · Read the evidence
Separate exposure, response, reliability, and commercial context
Exposure
Confirm eligible sessions and variant views before comparing response. Uneven or missing exposure can make a percentage look more precise than the underlying evidence.
Response
Use clicks and click-through rate when the hypothesis concerns discovery. Add-to-cart or order evidence may be more relevant for a commerce-completion question.
Reliability
Inspect empty responses, errors, eligibility, runtime changes, campaigns, and inventory before attributing the difference to the variant.
Commercial context
Shopify-verified attribution can connect order evidence to a variant when available. Missing revenue is not zero lift, and associated revenue still needs the attribution boundary.
For the difference between direct, assisted, and unmatched order evidence, use the ParticleSearch revenue attribution guide.
Chapter 5 · When not to experiment
Some decisions need repair, judgment, or more evidence first
An obvious defect
If a result is broken, an event is missing, or a mobile layout is unusable, repair it and verify the fix.
Several unrelated changes
A new layout, new ranking posture, new cards, and new filters in one variant cannot explain which decision mattered.
A tiny or unstable audience
Insufficient traffic, short promotions, or changing inventory can make the result too noisy to support the decision.
A question with no merchant action
Do not collect experiment data if no outcome would change the product or operating decision.
A causal claim from attribution alone
Associated revenue can add context, but attribution and controlled comparison answer different questions.
Chapter 6 · Close the loop
End with a decision and a preserved record
Record the hypothesis, variants, dates, evidence, caveats, and decision. If the result is inconclusive, say what prevented the decision. If a variant wins, keep monitoring the protected searches and operational guardrails after it becomes normal storefront behavior.
For store-level search posture, read the ParticleSearch search profiles guide. For query-level controls, use the ranking, synonyms, and redirects guide.