A working reference for designing, running, and reading experiments in a live product — the checks that catch a bad test before it costs you a decision.
A hypothesis names three things: the change, the mechanism, and the direction of the expected effect. "Showing the estimated delivery time on the checkout button increases conversion, because customers are more confident the order will actually arrive" is testable. "Make checkout better" is not — there's no version of the result that could contradict it.
Write the hypothesis down before writing any code. It's the only artifact that later tells you whether a result should have surprised you, or whether you're just describing noise in language that sounds like insight.
A metric chosen after looking at the data isn't a hypothesis test — it's a story fitted to a result.
One primary metric decides the outcome. Everything else is a guardrail: a metric that isn't allowed to get worse, even if the primary metric improves. A promo-banner test might optimize checkout conversion while guarding against a drop in average order value or a rise in cancellation rate.
Pick the guardrails before launch. Picking them after the fact turns every side effect into a justification for whatever result you already wanted.
Primary: checkout conversion rate. Guardrails: average order value, cancellation rate, support tickets per order.
Three numbers decide how long a test needs to run: the minimum detectable effect (MDE) — the smallest lift worth acting on — the significance level, and the statistical power. Setting the MDE too small chases a lift you could never afford to notice; setting it too large means a real, useful improvement slips through as "not significant."
Where a pre-experiment covariate is available — last week's spend, historical order frequency — CUPED (Controlled-experiment Using Pre-Experiment Data) removes the variance that covariate explains before testing, which shrinks the sample size needed for the same MDE without touching the randomization itself.
Running the calculation on the metric's mean when the metric is a rate, or vice versa, is the most common sizing mistake — the variance term changes shape depending on which one you're testing.
Randomize at the unit that matches the hypothesis. Order-level randomization for a change a customer will notice across sessions lets the same person land in both arms and contaminates the comparison — randomize by customer instead.
Check for sample ratio mismatch (SRM) within the first day: compare the observed split against the intended one with a chi-square test. A flagged SRM invalidates the test regardless of what the results later show, because it means the two groups were never comparable to begin with — the imbalance is usually a bug in the assignment logic, not bad luck.
In a two-sided marketplace, a demand-side change can bleed into supply and back into the control group — a promo that pulls orders toward one courier pool changes what every courier sees, control included.
Statistical significance and practical significance answer different questions. A p-value tells you how surprising the result would be if there were truly no effect. It says nothing about whether the effect, if real, is big enough to justify shipping. Report both the p-value and the confidence interval on the lift — the interval is what tells a stakeholder the range of outcomes they should actually expect.
Match the test to the metric's shape. A Welch's t-test handles a roughly normal metric like conversion rate without assuming equal variance between arms; a right-skewed metric like order value, with a long tail of large baskets, is usually better served by a Mann-Whitney U test than by comparing means directly.
Decide the stopping rule before the test starts, and don't check significance daily and stop the moment it crosses the threshold — repeated peeking inflates the false-positive rate well past the nominal 5%, even though each individual check looks legitimate.
Testing five guardrail metrics at α = 0.05 gives roughly a 23% chance one of them turns up "significant" by chance alone. Correct for multiple comparisons, or treat guardrail breaches as a flag to investigate rather than a verdict.
A test resolves to one of three calls: ship, hold, or iterate. Ship when the primary metric clears the MDE with guardrails intact. Hold when the interval straddles zero — that's a real "we don't know yet," not a soft no. Iterate when a guardrail moved even though the primary metric looks good; the test found something worth understanding before it goes to everyone.
For anything sensitive to novelty — a new visual treatment, a redesigned flow — expect the early lift to fade as the initial curiosity wears off. Extend the test or rerun it on returning users only before trusting a first-week number for a retention-relevant decision.
Document the result whether or not it shipped. A null result still tells you a variant of that idea, at that effect size, isn't worth testing again — and that's exactly the kind of thing a team re-tests every eighteen months if nobody wrote it down.
The fields worth filling in before a single line of assignment code gets written — shown here with a worked example.