The method moved from direct-mail marketing into online experimentation in the early 2000s and became dominant in product development through Google, Microsoft, and Booking.com’s published work on experimentation at scale. The core mechanics: randomise users into variants, run until the required sample is reached (see sample size calculator), measure the primary metric, declare the winner if the difference is statistically significant.
Common frameworks: classical hypothesis testing (frequentist, used by Optimizely and VWO), Bayesian inference (used by Statsig and some in-house platforms), and sequential testing (allows early stopping without inflated error rates). The choice of framework affects stopping rules and interpretation.
Without a controlled test, teams attribute outcomes to the wrong causes — a seasonal uplift gets credited to a product change, a drop gets explained by a design decision that didn’t cause it. A/B testing is the mechanism that produces causal evidence instead of correlational storytelling. For live products with meaningful traffic, it’s the gold standard for validation.
When you have the traffic to detect the effect you care about (use the sample size calculator), when the change is substantial enough that stakeholders need causal evidence to decide, and when you can measure a metric that genuinely represents the thing you want to improve. Not useful for discovery, for pre-launch work, or when the metric lags months behind the exposure.
A homepage test replacing a text CTA with an icon-and-text CTA: 50/50 random split, primary metric = click-through to pricing, secondary metrics = time on page and bounce. Run until the calculated sample size is hit (often 2–6 weeks on moderate traffic). Declare a winner only if the primary metric’s confidence interval excludes zero.
The result is almost always smaller than the stakeholder prediction. Teams that run many tests find average effect sizes well under 5% — the memorable 40% wins are outliers. Build the roadmap on realistic effect assumptions, not on the biggest reported case studies.
A/B testing sits in the broader experimentation family alongside multivariate testing and bandit algorithms. It intersects with sample size calculation, statistical inference, and prioritisation (which test to run first). It’s the validation step for hypotheses generated by qualitative research.
Reviewed 2 October 2026 · Editorial standards