Key takeaways
- Write the decision, hypothesis, unit, intervention, outcome, guardrails, and stop rule before launch.
- Preserve assignment, exclusions, changes, incidents, and analysis—not a polished result alone.
- Record uncertainty and alternative explanations beside the conclusion.
- Make negative and inconclusive results searchable so the organization does not repeat them unknowingly.
01
The archive starts before the treatment begins
A result cannot be interpreted without knowing what the team planned to change, whom it expected to affect, and what decision the evidence was meant to support. Write the record before launch, then lock or version the original hypothesis and analysis plan.
NIST describes experimental design as a detailed plan laid out in advance so the data can support valid and objective conclusions. GTM work may not always permit a clean randomized test, but it still benefits from explicit objectives, factors, responses, and assumptions.
Swipe to compare every column
| Record | Before launch | After launch |
|---|---|---|
| Decision | What choice will the evidence inform? | Decision made, owner, and date |
| Design | Unit, assignment, exposure, duration, sample | Deviations and contamination |
| Measurement | Primary outcome, guardrails, maturity window | Estimate, interval, missingness |
| Context | Audience, market, channel, offer, concurrent changes | Incidents and plausible alternatives |
02
Use the strongest design the constraint allows
Randomization protects against systematic differences when it is feasible and ethical. Blocking can account for important nuisance factors. NIST’s guidance is clear that every experiment has nuisance factors and summarizes the principle as blocking what can be controlled and randomizing what cannot.
When randomization is not possible, state the comparison method and its likely bias. A before-and-after campaign test can be affected by seasonality, competitor activity, sales capacity, and measurement changes. Naming those limits is part of the result, not an apology attached later.
03
Keep the messy run history
Record eligibility queries or versions, assignment seed where relevant, launch time, sample exclusions, tracking checks, delivery, spend, creative or offer versions, pauses, incidents, and deviations. Link the immutable data snapshot and code or query used for analysis when possible.
Do not silently change the primary outcome after seeing early data. If the team explores another metric, label it exploratory. Record data maturity and late-arriving outcomes so a fast read is not confused with the final one.
04
Archive the decision, including “we still do not know”
Summarize the estimate, uncertainty, guardrails, segment differences that were planned, operational cost, and practical significance. “Statistically detectable” is not the same as commercially worthwhile, and “not detected” is not proof of no effect.
Tag the record by audience, channel, offer, market, funnel stage, and decision. Link follow-up tests and later contradictions. Review the archive before approving new work. The purpose is institutional memory: a future team should understand what happened without asking the person who built the spreadsheet.
Primary sources and further reading
Use the source material to validate details against your own context and current platform configuration.
- NIST: What is experimental design?
- NIST: Steps of design of experiments
- NIST: Randomized block designs
This field note follows the XenGrowth editorial policy: primary sources where available, visible limitations, material review dates, and no invented first-hand experience.
Stay with the problem



