A/B Testing Reference

Running an A/B test that actually tells you something

A working reference for designing, running, and reading experiments in a live product — the checks that catch a bad test before it costs you a decision.

unchanged experience the thing you're testing
01

Frame the hypothesis

A hypothesis names three things: the change, the mechanism, and the direction of the expected effect. "Showing the estimated delivery time on the checkout button increases conversion, because customers are more confident the order will actually arrive" is testable. "Make checkout better" is not — there's no version of the result that could contradict it.

Write the hypothesis down before writing any code. It's the only artifact that later tells you whether a result should have surprised you, or whether you're just describing noise in language that sounds like insight.

Watch out

A metric chosen after looking at the data isn't a hypothesis test — it's a story fitted to a result.

02

Choose your metrics

One primary metric decides the outcome. Everything else is a guardrail: a metric that isn't allowed to get worse, even if the primary metric improves. A promo-banner test might optimize checkout conversion while guarding against a drop in average order value or a rise in cancellation rate.

Pick the guardrails before launch. Picking them after the fact turns every side effect into a justification for whatever result you already wanted.

Example

Primary: checkout conversion rate. Guardrails: average order value, cancellation rate, support tickets per order.

03

Size the test

Three numbers decide how long a test needs to run: the minimum detectable effect (MDE) — the smallest lift worth acting on — the significance level, and the statistical power. Setting the MDE too small chases a lift you could never afford to notice; setting it too large means a real, useful improvement slips through as "not significant."

n ≈ 2 · (z_α/2 + z_β)² · p(1 − p) / δ² for a two-arm proportion test, where δ is the MDE and p is the baseline rate. At α = 0.05 and 80% power, (z_α/2 + z_β)² ≈ 7.9.

Where a pre-experiment covariate is available — last week's spend, historical order frequency — CUPED (Controlled-experiment Using Pre-Experiment Data) removes the variance that covariate explains before testing, which shrinks the sample size needed for the same MDE without touching the randomization itself.

Watch out

Running the calculation on the metric's mean when the metric is a rate, or vice versa, is the most common sizing mistake — the variance term changes shape depending on which one you're testing.

04

Run it

Randomize at the unit that matches the hypothesis. Order-level randomization for a change a customer will notice across sessions lets the same person land in both arms and contaminates the comparison — randomize by customer instead.

Check for sample ratio mismatch (SRM) within the first day: compare the observed split against the intended one with a chi-square test. A flagged SRM invalidates the test regardless of what the results later show, because it means the two groups were never comparable to begin with — the imbalance is usually a bug in the assignment logic, not bad luck.

Watch out

In a two-sided marketplace, a demand-side change can bleed into supply and back into the control group — a promo that pulls orders toward one courier pool changes what every courier sees, control included.

05

Read the results

Statistical significance and practical significance answer different questions. A p-value tells you how surprising the result would be if there were truly no effect. It says nothing about whether the effect, if real, is big enough to justify shipping. Report both the p-value and the confidence interval on the lift — the interval is what tells a stakeholder the range of outcomes they should actually expect.

Match the test to the metric's shape. A Welch's t-test handles a roughly normal metric like conversion rate without assuming equal variance between arms; a right-skewed metric like order value, with a long tail of large baskets, is usually better served by a Mann-Whitney U test than by comparing means directly.

Decide the stopping rule before the test starts, and don't check significance daily and stop the moment it crosses the threshold — repeated peeking inflates the false-positive rate well past the nominal 5%, even though each individual check looks legitimate.

Watch out

Testing five guardrail metrics at α = 0.05 gives roughly a 23% chance one of them turns up "significant" by chance alone. Correct for multiple comparisons, or treat guardrail breaches as a flag to investigate rather than a verdict.

06

Decide and ship

A test resolves to one of three calls: ship, hold, or iterate. Ship when the primary metric clears the MDE with guardrails intact. Hold when the interval straddles zero — that's a real "we don't know yet," not a soft no. Iterate when a guardrail moved even though the primary metric looks good; the test found something worth understanding before it goes to everyone.

For anything sensitive to novelty — a new visual treatment, a redesigned flow — expect the early lift to fade as the initial curiosity wears off. Extend the test or rerun it on returning users only before trusting a first-week number for a retention-relevant decision.

Document the result whether or not it shipped. A null result still tells you a variant of that idea, at that effect size, isn't worth testing again — and that's exactly the kind of thing a team re-tests every eighteen months if nobody wrote it down.

Test design worksheet

The fields worth filling in before a single line of assignment code gets written — shown here with a worked example.

Checkout ETA badge example
Hypothesis
Showing estimated delivery time on the checkout button increases conversion, because customers gain confidence the order will arrive on time.
Primary metric
Checkout conversion rate
Guardrails
Average order value, cancellation rate, support tickets per order
MDE
+1.5 percentage points on a 24% baseline
Randomization unit
Customer ID, sticky across sessions
Sample size
≈ 38,000 customers per arm (α = 0.05, power = 0.8)
Duration
2 full weeks, covering both weekday and weekend ordering patterns
Stopping rule
No interim checks before day 14; read once, at the planned sample size

Quick reference

p-value
The probability of seeing a result this extreme, or more so, if there were truly no effect. Not the probability the hypothesis is true.
Statistical power
The probability of detecting a real effect of a given size, if it exists. Conventionally set at 80%.
MDE
Minimum detectable effect — the smallest lift the test is designed to reliably catch. Smaller MDE, larger sample.
CUPED
A variance-reduction technique that adjusts the metric using a pre-experiment covariate, shrinking the sample size needed for the same MDE.
SRM
Sample ratio mismatch — when the observed traffic split deviates from the intended one by more than chance. Usually signals a bug in assignment, not a fluke.
Type I / II error
Type I: calling a real non-effect significant (false positive). Type II: missing a real effect (false negative).
One-sided vs two-sided
One-sided tests only for a lift in one direction; two-sided also catches an unexpected drop. Default to two-sided unless a drop is truly impossible.
Guardrail metric
A metric that must not get worse, chosen before launch, regardless of what happens to the primary metric.