Skip to content
Getting Digital

Hypothesis Testing

Also: statistical significance, p-value, null hypothesis, significance test, confidence interval, statistical inference

Hypothesis testing is the statistical procedure for deciding whether an observed difference, between two groups, two periods or a sample and a claimed value, is larger than chance variation would produce if there were no real difference, by computing how probable a result at least that extreme would be under that assumption.

Assessment. The p-value answers a narrow question, how unusual the data would be if there were no effect, and most of the harm done with it comes from treating that answer as a verdict on whether the effect is real, large or important. The American Statistical Association's 2016 statement set out the limits in six principles, and an analyst who reports an effect size with an interval, states the sample size in advance and reads a p-value as one piece of evidence among several is using the method as intended. A result below a threshold is where the examination of a finding begins.

The procedure starts from a null hypothesis, the assumption that there is no difference, and asks how probable the observed data, or data more extreme, would be if that assumption held. That probability is the p-value. A small p-value means the data would be unusual under the null hypothesis; a conventional threshold, 0.05, has been used for a century to call such a result statistically significant. The threshold is a convention, not a law of nature, and the ASA's statement was issued because the convention had become, in its words, a gatekeeper for what is publishable and had encouraged practices that search for small p-values rather than for understanding.

  1. A p-value can indicate how incompatible the data are with a specified statistical model.
  2. It does not measure the probability that the hypothesis is true, or that the data arose by chance alone.
  3. Scientific conclusions and business or policy decisions should not rest only on whether a p-value passes a threshold.
  4. Proper inference requires full reporting and transparency.
  5. A p-value does not measure the size of an effect or the importance of a result.
  6. By itself, a p-value is not a good measure of the evidence for a model or hypothesis.

Those are the ASA's six principles, restated. The second and fifth are the ones most often violated in practice. A p-value of 0.03 does not mean a 97 per cent chance that the effect is real, and a p-value of 0.001 does not mean the effect is large; it may be tiny and measured on a very large sample. The remedies the statement points to are routine in good analytical work: report the effect size and a confidence interval, which states the range of effects consistent with the data; decide the sample size before collecting it, based on the smallest effect that would matter; and report every comparison that was made, since twenty comparisons at a 0.05 threshold will produce one significant result by chance.

In the applied setting the procedure is the reading step of an A/B test: two groups assigned at random, one metric chosen in advance, a test at the planned sample size that asks whether the difference between the groups exceeds what random assignment alone would produce. Kohavi, Tang and Xu's treatment of online experiments spends most of its length on what makes the test invalid, early stopping, multiple metrics, segments examined after the fact, rather than on the arithmetic, which the tools perform. The same discipline applies to any comparison an analyst draws from existing data, with the added caution that groups which were not randomly assigned differ in ways the test cannot see. The statistics and experimentation courses cover the methods; Google's Advanced Data Analytics certificate includes a statistics course that tests them.

Reporting a result

State the comparison, the sample size, the effect size with its confidence interval, the test used and the p-value, in that order. A report that gives only the p-value withholds the information needed to judge whether the result matters.

In practice

  • The claim: a new onboarding flow raises the share of users who complete setup.
  • The test: users assigned at random; a sample size fixed from the current completion rate and a minimum effect of two percentage points; the metric read once, at the end.
  • The result: completion up by 2.4 percentage points, with a 95 per cent confidence interval from 0.8 to 4.0; p-value 0.003.
  • The reading: the data would be unusual if the flow had no effect; the effect is probably between one and four points; the decision rests on whether two points justifies the engineering cost, which the p-value does not address.

Often confused with

A/B Testing
An A/B test is the experimental design, random assignment and a fixed metric; hypothesis testing is the statistical reading of its result. The design is what makes the reading valid, and the reading is one step of the design.
Data Analysis
Data analysis is the whole work of answering a question from data; hypothesis testing is one of its tools, used when the question is whether a difference is larger than chance. Much analysis is descriptive and needs no test.

Key takeaways

  • →The p-value states how unusual the data would be if there were no effect, and nothing more than that.
  • →Report the effect size and its confidence interval; set the sample size before the test; disclose every comparison.
  • →A significant result is the starting point of the examination of a finding.

Related concepts

  • The tool used when the question is whether a difference exceeds chance.

  • The statistical reading of an experiment's result.

Where this concept sits in the field

Certifications that test this

Vendor exams whose syllabus covers this concept: facts, cost and a preparation path on each page.

More courses from these categories

Courses from the categories where this concept is taught. Details, price and the provider link are on each course page.

FAQ

What does a confidence interval add to a p-value?
The range of effect sizes consistent with the data. A result can be statistically significant with an interval that runs from a negligible effect to a large one, which tells the decision-maker that the size is uncertain even though the direction is probably right.
Is 0.05 the right threshold?
It is a convention, and the ASA statement's third principle says a decision should not rest on it alone. Fields with costly false positives use stricter thresholds; an analyst reporting to a business should present the effect size and the interval and let the threshold be a secondary consideration.
Why does testing many metrics cause problems?
Each test has a chance of a false positive, and the chances accumulate. Twenty metrics examined at a 0.05 threshold produce, on average, one spurious significant result. The remedies are one primary metric decided in advance, a correction for multiple comparisons, or a separate confirmation test.

Sources

The primary text this definition rests on. Read it before relying on this one.

Last reviewed 3 October 2026 · Getting Digital