Method · UX Glossary

A/B testing

A/B testing (also called split testing or online controlled experimentation) is a method for comparing two versions of an interface by randomly assigning users to one or the other and measuring the difference in a primary metric. It provides causal evidence that one design beats another on a specific metric at a specific significance level.

What it is

The method moved from direct-mail marketing into online experimentation in the early 2000s and became dominant in product development through Google, Microsoft, and Booking.com’s published work on experimentation at scale. The core mechanics: randomise users into variants, run until the required sample is reached (see sample size calculator), measure the primary metric, declare the winner if the difference is statistically significant.

Common frameworks: classical hypothesis testing (frequentist, used by Optimizely and VWO), Bayesian inference (used by Statsig and some in-house platforms), and sequential testing (allows early stopping without inflated error rates). The choice of framework affects stopping rules and interpretation.

Why it matters

Without a controlled test, teams attribute outcomes to the wrong causes — a seasonal uplift gets credited to a product change, a drop gets explained by a design decision that didn’t cause it. A/B testing is the mechanism that produces causal evidence instead of correlational storytelling. For live products with meaningful traffic, it’s the gold standard for validation.

When to use it

When you have the traffic to detect the effect you care about (use the sample size calculator), when the change is substantial enough that stakeholders need causal evidence to decide, and when you can measure a metric that genuinely represents the thing you want to improve. Not useful for discovery, for pre-launch work, or when the metric lags months behind the exposure.

What it looks like in practice

A homepage test replacing a text CTA with an icon-and-text CTA: 50/50 random split, primary metric = click-through to pricing, secondary metrics = time on page and bounce. Run until the calculated sample size is hit (often 2–6 weeks on moderate traffic). Declare a winner only if the primary metric’s confidence interval excludes zero.

The result is almost always smaller than the stakeholder prediction. Teams that run many tests find average effect sizes well under 5% — the memorable 40% wins are outliers. Build the roadmap on realistic effect assumptions, not on the biggest reported case studies.

What people get wrong

Learn next

A/B testing sits in the broader experimentation family alongside multivariate testing and bandit algorithms. It intersects with sample size calculation, statistical inference, and prioritisation (which test to run first). It’s the validation step for hypotheses generated by qualitative research.

Related terms & references
JP
Associate Director, Experience Design at JD.com · Previously Head of UX at Selfridges & Co · Building UX Companion

Reviewed 2 October 2026 · Editorial standards