Key takeaways
- Start with a decision and causal hypothesis, not a backlog of button colors.
- Change one coherent experience or proposition with enough contrast to matter while protecting assignment and instrumentation.
- Use valid completion, lead quality, accessibility, performance, complaints, and downstream outcomes as appropriate guardrails.
- Predefine sample assumptions and stopping; report uncertainty and implement only after checking operational consequences.
01
Write the decision before the variant
State what the team will change if the treatment improves, harms, or does not resolve the primary outcome. A hypothesis should connect a customer problem to a mechanism: showing price range and scope assumptions will reduce unsuitable inquiries without reducing qualified opportunities.
Choose a primary outcome close enough to measure and meaningful enough to act on. Click-through is often too early; closed revenue may be too rare or delayed. Valid form completion, scheduled consultation, or qualified opportunity can be useful when definitions and follow-up are stable.
02
Create a coherent treatment
Change the information, proof, offer, interaction, or sequence implied by the hypothesis. A treatment can include several coordinated page elements when they form one experience. Document the exact difference and prevent unrelated releases from drifting into only one arm.
Verify assignment, persistence across sessions where appropriate, analytics, consent, performance, accessibility, and downstream identifiers before launch. Internal staff and bots should not become experimental customers.
Swipe to compare every column
| Decision | Primary outcome | Guardrails |
|---|---|---|
| Publish pricing range | Qualified inquiry rate | Overall opportunity volume, complaints |
| Shorten lead form | Valid completion | Spam, qualification, sales workload |
| Add evidence-led case study | Consultation progression | Page speed, comprehension, credibility feedback |
| Change navigation path | Task completion | Wrong turns, exits, accessibility |
03
Respect uncertainty and delayed outcomes
Estimate baseline, minimum detectable effect, sample, and run time before launch. Account for day-of-week and the business conversion cycle. Avoid repeated peeking and stopping at a convenient moment. Segment analysis should be planned or treated as exploratory, not as a machine for finding a favorable subgroup.
Report effect estimate and interval, allocation, exclusions, dates, sample, guardrails, and known confounders. “No clear difference” is useful when it prevents a costly redesign based on noise.
04
Close the loop after the result
If the treatment is adopted, verify production matches the tested experience and monitor the downstream metrics that mature later. Archive the hypothesis, screenshots, code or configuration, result, and decision so the next team does not rerun the same ambiguous test.
A healthy program runs fewer tests than a random-ideas factory and learns more from each one because the decision was real before the dashboard appeared.
Primary sources and further reading
Use the source material to validate details against your own context and current platform configuration.
This field note follows the XenGrowth editorial policy: primary sources where available, visible limitations, material review dates, and no invented first-hand experience.
Stay with the problem



