The method comes from clinical trials and agriculture and reached the web through the large platforms, which run thousands of experiments a year on search results, feeds and checkouts. The design is simple. Choose one metric that defines success before the test starts. Split arriving visitors at random, so that time of day, device, campaign and mood are spread evenly across the versions. Show each group its version for the whole period. Compare the metric with a statistical test that says how surprising the observed difference would be if the versions were in fact equal. Everything else, the tool, the dashboard, the confidence meter, is presentation around those four steps.
Where tests go wrong is well catalogued, and Kohavi, Tang and Xu's book on online controlled experiments is the catalogue. Peeking: checking daily and stopping on the first significant reading, which turns a five-in-a-hundred false-positive rate into a near certainty of one. Underpowered tests: too few conversions to detect the size of change that is plausible, so the test ends inconclusive after a month. Metric drift: optimising sign-ups and discovering later that the winning version produced sign-ups who never bought. Novelty: a redesign wins for a fortnight because it is new, then loses. Each has a known remedy, a sample size fixed in advance, a guardrail metric, a longer run, and each is skipped under deadline pressure.
The tooling changed in 2023 when Google retired Optimize, its free testing tool, on 30 September and named AB Tasty, Optimizely and VWO as the testing tools it integrates with its analytics. The retirement moved testing from a free add-on to a decision with a price, which has had a useful side effect: teams now ask whether a test is worth running before setting one up. A change that removes a broken field needs no test; ship it and watch. A pricing page, a checkout step or a headline on the page that earns the revenue deserves one, and the test is cheap against the cost of guessing wrong.
- One primary metric, chosen before the test, plus a guardrail that must not get worse.
- Sample size fixed before the start, derived from the current rate and the smallest change worth detecting.
- Random assignment, held for the whole run, with the same visitor always seeing the same version.
- No peeking: the result is read once, at the planned end.
- Segments read with suspicion: a win in one browser among twenty is the expected signature of chance.
