The procedure starts from a null hypothesis, the assumption that there is no difference, and asks how probable the observed data, or data more extreme, would be if that assumption held. That probability is the p-value. A small p-value means the data would be unusual under the null hypothesis; a conventional threshold, 0.05, has been used for a century to call such a result statistically significant. The threshold is a convention, not a law of nature, and the ASA's statement was issued because the convention had become, in its words, a gatekeeper for what is publishable and had encouraged practices that search for small p-values rather than for understanding.
- A p-value can indicate how incompatible the data are with a specified statistical model.
- It does not measure the probability that the hypothesis is true, or that the data arose by chance alone.
- Scientific conclusions and business or policy decisions should not rest only on whether a p-value passes a threshold.
- Proper inference requires full reporting and transparency.
- A p-value does not measure the size of an effect or the importance of a result.
- By itself, a p-value is not a good measure of the evidence for a model or hypothesis.
Those are the ASA's six principles, restated. The second and fifth are the ones most often violated in practice. A p-value of 0.03 does not mean a 97 per cent chance that the effect is real, and a p-value of 0.001 does not mean the effect is large; it may be tiny and measured on a very large sample. The remedies the statement points to are routine in good analytical work: report the effect size and a confidence interval, which states the range of effects consistent with the data; decide the sample size before collecting it, based on the smallest effect that would matter; and report every comparison that was made, since twenty comparisons at a 0.05 threshold will produce one significant result by chance.
In the applied setting the procedure is the reading step of an A/B test: two groups assigned at random, one metric chosen in advance, a test at the planned sample size that asks whether the difference between the groups exceeds what random assignment alone would produce. Kohavi, Tang and Xu's treatment of online experiments spends most of its length on what makes the test invalid, early stopping, multiple metrics, segments examined after the fact, rather than on the arithmetic, which the tools perform. The same discipline applies to any comparison an analyst draws from existing data, with the added caution that groups which were not randomly assigned differ in ways the test cannot see. The statistics and experimentation courses cover the methods; Google's Advanced Data Analytics certificate includes a statistics course that tests them.
Reporting a result
State the comparison, the sample size, the effect size with its confidence interval, the test used and the p-value, in that order. A report that gives only the p-value withholds the information needed to judge whether the result matters.
