Key takeaways
- Define the business decision and the counterfactual before choosing an experiment design.
- Use randomized holdouts when feasible; observational attribution answers a different question.
- Check power, contamination, interference, and implementation fidelity before reading the lift estimate.
- Report uncertainty and absolute outcomes alongside a percentage lift.
01
Begin with the decision that could change
“Did marketing work?” is too broad for an experiment. A usable question names the intervention, eligible population, exposure period, outcome window, and decision that follows. For example: should paid search in these regions keep its present budget next quarter? That wording forces the team to decide what “without the ads” means.
Attribution distributes credit across observed interactions. Incrementality estimates the difference between an outcome under treatment and the outcome that would have occurred without it. The second quantity is a counterfactual; it cannot be recovered merely by choosing a more elaborate attribution model.
Swipe to compare every column
| Design choice | Question to settle | Failure it prevents |
|---|---|---|
| Unit | Person, account, store, region, or time block? | Pretending dependent observations are independent |
| Treatment | What exactly changes, and by how much? | Testing a bundle nobody can repeat |
| Outcome | Revenue, qualified pipeline, or another stable event? | Optimizing a convenient proxy |
| Window | When can the effect reasonably appear? | Stopping before delayed outcomes mature |
02
Protect the contrast between treatment and control
Randomization is valuable because it makes the treatment assignment independent of potential outcomes in expectation. It does not repair a broken launch. Budget leakage, overlapping campaigns, sales territories that cross geo boundaries, and customer movement can contaminate the contrast.
Keep an implementation log: intended spend, delivered spend, audience exclusions, outages, creative changes, promotions, and material competitor events. A technically sophisticated estimate built on an undocumented treatment is hard to interpret and harder to repeat.
03
Ask whether the test can detect a useful difference
Power analysis is not a ceremonial calculation performed after the regions have been chosen. Estimate baseline volume, variation, feasible holdout size, expected treatment strength, and the smallest effect that would change the decision. When the business cannot supply enough units or time, narrow the question or accept that the test may remain inconclusive.
A non-significant result does not prove zero effect. A statistically detectable lift can still be commercially trivial. Put the estimate, interval, absolute outcome, media cost, and operational caveats on the same page so neither statistical nor commercial importance disappears.
04
Turn the result into a bounded operating decision
State where the finding applies: the tested markets, spend range, audience, offer, creative, and period. Extrapolating far beyond that support is a new assumption, not part of the result. Record follow-up conditions that would trigger another test.
Incrementality testing is most useful as a repeated calibration practice. It can challenge platform reporting, inform model priors, and show where another dollar has evidence behind it. It cannot issue a permanent certificate that a channel always works.
Primary sources and further reading
Use the source material to validate details against your own context and current platform configuration.
- Google Research: Near Impressions for Observational Causal Ad Impact
- Google Research: Incremental Clicks—The Impact of Search Advertising
- Google Meridian: About MMM as a causal inference methodology
This field note follows the XenGrowth editorial policy: primary sources where available, visible limitations, material review dates, and no invented first-hand experience.
Stay with the problem



