Skip to content
Getting Digital

A/B Testing

Also: split testing, online controlled experiment, A/B test, multivariate testing, statistical significance, experimentation

A/B testing is a controlled experiment on a live audience: visitors are assigned at random to two or more versions of a page, an email or a feature, and the versions are compared on one pre-chosen outcome so that the difference can be attributed to the change rather than to who happened to see it.

Assessment. Random assignment is what gives the method its value, and much of what is described as A/B testing lacks it: one version shown this week against another shown last week, or a result declared as soon as a chart separates. A valid test requires a fixed outcome, a sample size decided in advance and a run to completion. A business without the traffic to reach that sample size within a few weeks should improve its pages by observation and reserve testing for the questions that justify the wait.

The method comes from clinical trials and agriculture and reached the web through the large platforms, which run thousands of experiments a year on search results, feeds and checkouts. The design is simple. Choose one metric that defines success before the test starts. Split arriving visitors at random, so that time of day, device, campaign and mood are spread evenly across the versions. Show each group its version for the whole period. Compare the metric with a statistical test that says how surprising the observed difference would be if the versions were in fact equal. Everything else, the tool, the dashboard, the confidence meter, is presentation around those four steps.

Where tests go wrong is well catalogued, and Kohavi, Tang and Xu's book on online controlled experiments is the catalogue. Peeking: checking daily and stopping on the first significant reading, which turns a five-in-a-hundred false-positive rate into a near certainty of one. Underpowered tests: too few conversions to detect the size of change that is plausible, so the test ends inconclusive after a month. Metric drift: optimising sign-ups and discovering later that the winning version produced sign-ups who never bought. Novelty: a redesign wins for a fortnight because it is new, then loses. Each has a known remedy, a sample size fixed in advance, a guardrail metric, a longer run, and each is skipped under deadline pressure.

The tooling changed in 2023 when Google retired Optimize, its free testing tool, on 30 September and named AB Tasty, Optimizely and VWO as the testing tools it integrates with its analytics. The retirement moved testing from a free add-on to a decision with a price, which has had a useful side effect: teams now ask whether a test is worth running before setting one up. A change that removes a broken field needs no test; ship it and watch. A pricing page, a checkout step or a headline on the page that earns the revenue deserves one, and the test is cheap against the cost of guessing wrong.

  • One primary metric, chosen before the test, plus a guardrail that must not get worse.
  • Sample size fixed before the start, derived from the current rate and the smallest change worth detecting.
  • Random assignment, held for the whole run, with the same visitor always seeing the same version.
  • No peeking: the result is read once, at the planned end.
  • Segments read with suspicion: a win in one browser among twenty is the expected signature of chance.

In practice

  • The question: does showing delivery costs on the product page, rather than at checkout, change the purchase rate?
  • The plan: purchase rate as the metric, average order value as the guardrail, a sample size that the shop's traffic reaches in three weeks, visitors split at random by a cookie.
  • The pressure: on day four the new version leads and someone asks to ship it. The plan says day twenty-one.
  • The reading: on day twenty-one the lead has held and the order value has not fallen. The change ships, and the next question goes into the queue.

Often confused with

Conversion Rate Optimisation (CRO)
Conversion rate optimisation is the whole practice of raising the share of visitors who act; A/B testing is its measuring instrument for one change at a time. Much of the practice happens without tests, on observation and removal.
Data Analysis
Data analysis reads what happened; an A/B test arranges for something to happen in two ways at once so that the cause is known. Analysis of existing data can suggest a change; only an experiment can attribute the result to it.

Key takeaways

  • →Random assignment plus one pre-chosen metric is what separates an experiment from a comparison.
  • →Decide the sample size first and read the result once. Stopping early produces false winners.
  • →Test the changes that earn the revenue; ship the obvious fixes without a test and watch.

Related concepts

Where this concept sits in the field

FAQ

What does statistically significant mean in a test report?
That a difference as large as the one observed would be unlikely if the two versions were in fact equal, usually judged at a five-in-a-hundred threshold. It says nothing about the size of the effect or whether it matters commercially, and it is only valid if the test ran to its planned size without being stopped early.
How much traffic does a test need?
Enough conversions, not visits, to detect the smallest change worth acting on; the calculators built into every testing tool take the current rate and that change and return a number. Pages with several hundred monthly conversions can test large changes; pages with a dozen cannot test anything.
Is multivariate testing better than A/B?
It tests several elements at once and needs correspondingly more traffic to reach a conclusion for each combination. Few sites outside the largest have enough. One change at a time, run properly, answers more questions per year.

Sources

The primary text this definition rests on. Read it before relying on this one.

Last reviewed 3 October 2026 · Getting Digital