Why It Is Called the Null Hypothesis
Null was never about nothingness. Fisher's null is the hypothesis that can be nullified — set up so the data have something to argue against. This piece follows one line: how a tea party produced the first p value, how Neyman and Pearson turned it into a decision rule, how textbooks welded the two together into something neither man designed, and where that weld cracks.
On this page
In every statistics course you are told that the null hypothesis is the one you set out to reject. That is right, and it is what I dutifully memorised.
What nobody told me is why the thing is called null.
This is not idle etymology. Read null as "empty, nothing there" and the whole logic tilts: you start thinking a test proves that nothing happened, that establishes the absence of an effect, that the null is a claim about the world. All three are wrong, and all three are common.
Fisher's own sentence runs: the null hypothesis is never proved or established, but is possibly disproved in the course of experimentation; every experiment may be said to exist only in order to give the facts a chance of disproving the null hypothesis. The weight is on the last clause. The null is the hypothesis built to take the attack — null here is closer to "to be nullified", nullifiable, rather than to "empty".
Honestly, the word has no authoritative etymology to point at. The Oxford English Dictionary dates its first use of "null hypothesis" to 1935, which is this book of Fisher's; why he reached for that particular word, he never explains anywhere I could find. What we can pin down is how he used it.
And the trouble goes further: the significance testing in today's textbooks is two incompatible logics welded together, and the weld runs straight through this word. This piece takes the weld apart.
A lady, and eight cups of tea
Rothamsted Experimental Station, England, some afternoon in the 1920s. Muriel Bristol, an algologist, says she can taste whether the milk went into the cup before the tea or after. Fisher does not argue with her about the physiology of taste. He designs an experiment: eight cups, four poured milk-first and four tea-first, presented in random order, and she must pick out the four that had milk first.
Write the null hypothesis down first, because the whole experiment rests on it:
The null hypothesis: she cannot taste the difference at all. The four cups she picks have nothing to do with which four actually had milk first — what she is doing is equivalent to choosing four cups out of eight with her eyes shut.
Notice how it is phrased. It does not say she is unable to taste anything; it is a claim put up there to be knocked down. The job of the experiment is to give the data a chance to knock it down.
Once that sentence exists, everything else can be worked out in advance: if she really is guessing, what happens? Four cups out of eight can be chosen in seventy ways, all equally likely to a guesser. Exactly one of those seventy is completely correct. So a woman with no ability whatsoever gets all four right with probability , about 0.014.
That is what a null hypothesis is for: it turns the vague sentence "she cannot taste it" into a concrete world you can enumerate and compute in. Without it, "she got four right" is just a fact, and you have no idea whether four is impressive. With it, you have something to compare against.
Try it yourself:
That is the earliest p value there is. It was counted, not estimated: list every outcome available to a guesser, count the ones at least as good as hers, divide by seventy. Listing every rearrangement and counting is what later came to be called a permutation test. That is what the null is for — it hands you a world small enough to count all the way through. Without it, all you know is that she got four right, and not whether four is impressive.
As for how many she actually got: Fisher never says. The story that she identified all eight correctly is a recollection passed along decades later; Fisher's daughter, in her biography of him, writes only that she got enough right to make her case. If that means six of eight, then — nowhere near significant by anyone's standard. The most famous experiment in the subject has no recorded result.
What a p value actually looks like
The tea experiment is small enough to count out by hand. Real studies rarely are, so they lean on sampling distributions instead. The question does not change: if the null were true, how often would data this extreme show up?
Follow that sentence one step further and you get a result most people never quite absorb. When the null is true, p values are uniform. The probability of is exactly 0.05; the probability of is exactly 0.01. In a world where nothing is happening, one study in twenty comes out significant.
Set the effect to zero and run it two thousand times:
Delta is the real gap between the two group means, measured in standard deviations: if two classes differ by 2 points on a test whose scores have a standard deviation of 10, delta is 0.2. Delta = 0 means the groups really are identical, and any difference you see is sampling noise.
5.1% of 2000 studies came out below the threshold. The effect is zero, so every one of them is a false positive — at exactly the rate you set, 0.05.
The delta on that slider is the real gap between the two group means, measured in standard deviations. The unit matters: "the two classes differ by two points" means nothing until you know how spread out the scores are. If the standard deviation is ten points, then two points is 0.2 of one, and the two distributions almost completely overlap. Delta = 0 means the groups really are identical, and every difference you see is sampling noise.
Push the slider right and the distribution starts piling up against the left wall. How fast it piles up is what power means: the chance of catching an effect that is really there. Same picture, same machinery — the null simply is not true any more.
Neyman and Pearson turn it into a decision
In 1933 Neyman and Pearson published something that looks similar and rests on entirely different foundations. Their question was: if you must choose between two hypotheses, what rule makes you wrong least often in the long run? They called the test a rule of behaviour, and meant it literally.
Picture a factory stamping metal plates for surgical instruments. Gigerenzer uses exactly this scene to explain their logic: the machine may drift out of tolerance, and inspecting costs money. A false alarm shuts down a working line for an afternoon. A miss puts a defective plate inside somebody. The two errors cost different amounts, so the manager fixes a very small rate for the first (a Type I error), tolerates a larger one for the second, and from those two numbers works out how many plates to pull off the line each day. Nobody in this story is asking how strong the evidence is. They are asking how much the rule will cost them over a year.
Out of that come paired with , an fixed before you look, and, as the price of both, the second kind of error and the power to avoid it. Saying is a statement about the rule: of all the occasions the null holds, you will call it wrong on 5% of them. It says nothing about how confident you should be right now.
The procedure has its own house rules. Fix the threshold before you see the data. Report a decision, reject or do not reject. Do not write down what p came to — in this logic it is only an intermediate quantity to compare against the line, and carries no evidential weight of its own.
Fisher hated this, and the two camps fought about it for twenty years. Textbooks settled the fight by taking both.
The weld
One dataset, three readings:
One dataset: 12 observations in the treated group, 12 in the control group. t = 2.267, df = 22, two-sided p = 0.0336. All three schools are handed the same numbers.
Report p = 0.0336, and stop.
No alternative hypothesis, no threshold agreed beforehand, no act of acceptance. The p value measures how badly the data sit with the null. The null is there to be knocked down; it never gets proved. What happens next is the investigator’s judgement: run it again, or take the result provisionally.
That third panel is the sentence almost every paper writes. It claims two things that exclude each other. "The smaller the p, the stronger the evidence" is Fisher's, and the price of it is giving up the long-run guarantee. "Significant at the 0.05 level, so we reject the null" is Neyman and Pearson's, and the price of that is fixing the threshold beforehand and never reading p as evidence. The sentence pays neither price and claims both authorities. Gigerenzer calls it the null ritual: a procedure with no author, which nobody will defend, performed daily by an entire discipline.
The 0.05 threshold has no theoretical standing either. Fisher said one in twenty was a convenient line to draw. Convenient, and easy to tabulate in 1925.
Where the weld cracks
A p value is not the probability that the null is true. It is the probability of data like these given that the null is true. Those two sentences look almost identical, and the condition runs in opposite directions.
Change the setting and it becomes obvious. Suppose a disease affects ten people in every thousand. The test is decent: of the people who have it, eight in ten test positive; of the people who do not, 5% test positive anyway. You test positive. What are the chances you have the disease?
Skip the formula and count people. Out of a thousand:
- Ten actually have it. Eight of them test positive; two are missed.
- Nine hundred and ninety do not. About fifty of them test positive anyway.
Fifty-eight people got a positive result, and only eight of them are ill. A positive test means about an 8 in 58 chance of being ill — roughly 14%. Not 80%, and certainly not 95%. (That quantity has a name: positive predictive value.)
What changed? "Eight in ten sick people test positive" reads from the disease towards the test. "How many people who test positive are sick" reads from the test back towards the disease. Reverse the direction and the answer drops from 80% to 14%. The thing that went missing in between is the base rate: how rare the disease was to begin with. The test itself knows nothing about that.
The p value is that test. tells you how rare data like this would be if the null were true — reading from the null towards the data. The question you actually care about runs the other way: given data like this, how likely is the null to be false? Answering it needs the one thing the p value has no access to — how many hypotheses in this field were true in the first place. That is the research version of a base rate.
Turn the dials yourself. Prior odds of 1:10 mean that of every eleven hypotheses someone tests in this field, about one is true, which is roughly what exploratory work and screening experiments look like:
| Per 1000 hypotheses tested | significant | not significant |
|---|---|---|
| real effect (91) | 73 | 18 |
| no effect (909) | 45 | 864 |
Of everything that comes back "significant", 62% is real. p < 0.05 never meant "95% likely to be true". That 95% is computed assuming the null holds, and the question here runs the other way.
Lower the power and the picture gets worse: in a field where most hypotheses are wrong and power is low, most of what comes back "significant" is false. Every single test was performed correctly and nobody cheated. This is the arithmetic behind Ioannidis's 2005 claim that most published research findings are false — he accused nobody of fraud; he applied the head-counting above to the scientific literature as a whole.
A small p does not mean a large effect, either. The 1988 aspirin trial is the standing example: 22,071 physicians, 104 heart attacks on aspirin against 189 on placebo, . About as small as a p value gets. And the effect size, expressed as a correlation, is roughly 0.03 — two numbers from the same table, one of which looks like proof and one of which you can barely see.
Freeze the effect at something too small to care about and turn up the sample size:
d = 0.15 means the two group means really differ by 0.15 of a standard deviation — about 1.5 points on a test whose scores have a standard deviation of 10. Nobody would care about a gap that size. It never changes across this chart; only the number of people does.
With 80 people per group, run 400 times, the typical study comes back with p = 0.257. Not significant yet. Collect a few more people and it will be.
Past roughly 500 people per group, the typical study drops below 0.05 and starts coming out "significant".
A p value measures how unlike the null the data are. Collect enough data and even a trivial difference looks nothing like the null. Sometimes is just telling you the sample was big.
And then there is the oldest trick that isn't a trick — the problem of multiple comparisons. Twenty outcome measures, no real effect in any of them, each tested at 0.05. The chance that at least one comes out significant is .
Simmons, Nelson and Simonsohn demonstrated this in public in 2011. They played subjects either the Beatles' "When I'm Sixty-Four" or a control track, then asked their date of birth, and reported that the song made people a year and a half younger (20.1 against 21.5 years, , ). Music does not reverse ageing. That was the point: with a little undisclosed latitude over when to stop collecting data, which covariates to include and which measures to report, anything at all can be made significant — their simulations put the false-positive rate of such a study above 60%. (The authors have since noted that this particular effect is not robust: drop one covariate and climbs to .33. Which is, again, the point.)
Nobody faked anything. Every test is textbook-correct. What breaks is looking at the data first and choosing what to report afterwards — the practice now known as p-hacking: the 5% guarantee is gone, and nothing in the write-up will show that it left.
The bill arrived a decade later. In 2015, 270 researchers repeated 100 psychology experiments: 97% of the originals had been significant, and 36% of the replications were. This is what came to be called the replication crisis.
Said right, said wrong
The same p value has a few phrasings worth keeping and a few misreadings almost everyone has fallen into. Greenland and colleagues listed twenty-five of the latter in 2016. One side below collects the precise sentences, the other the misreadings you meet most often, each paired with the fix.
Said this way, it is right
The direction of the condition is right: it is P(data | null), a sentence about the data, not about the hypothesis. Almost every misreading below starts by flipping that condition, so holding on to the correct direction is the cheapest defence there is.
Null is about refutability, not about nothing happening. That is this article's spine, and it is exactly where the Chinese translation 虛無 most easily misleads.
It keeps 'no evidence against' apart from 'evidence for'. A small sample, high variance, or a coarse measure can all hide a real effect; the null is never proved.
The significance level is a parameter of a decision rule, living at the level of 'how often you err in the long run'; it says nothing about the probability that this result is true. Reading a decision rule as evidential strength is exactly where the welded-together textbook logic cracks.
Said this way, it is wrong
A p value is the probability of data like these given that the null holds. This reads the condition backwards. How far apart the two numbers are depends on how many hypotheses in the field were true to begin with, and the p value knows nothing about that.
To ask how likely a finding is to be real, you have to bring in prior odds and power (see the 2x2 table above).
Same reversal. The alternative hypothesis never enters the calculation of a p value; only the null does.
Weighing the credibility of the alternative needs a Bayes factor or a posterior, which is a different tool.
Failing to reject means the data gave you no good reason to reject. A small sample, a noisy measure, or high variance will hide a real effect. The null is never proved.
Report the effect size and interval. If you want to claim the difference is negligible, run an equivalence test and say in advance what negligible means.
The p value moves with both effect size and sample size. With a large enough sample, a trivial difference produces a tiny p.
Size lives in the effect size — a difference, a ratio, a standardized effect — not in p.
Fisher picked one in twenty as a convenient line, partly because it made the tables usable. There is no theory behind it.
Let the cost of the decision set the threshold: how expensive is a false positive, how expensive is a false negative.
Replication success depends on the true effect and the power of the study, not on this one p value, which carries very little information about it.
Talk about power and the precision of the effect estimate instead.
So what do we do
In 2016 the American Statistical Association did something it had never done before: issued a formal statement about a statistical method. The p value statement gives six principles, the bluntest of which is that scientific conclusions and policy decisions should not be based only on whether a p value passes a threshold.
It offers no replacement, because the tool was never the problem. The p value answers a well-defined question. Fisher's use of it holds together, and so does Neyman and Pearson's. What breaks is the weld between them, which lets people believe that finishing a test is the same as finishing an inference.
The practical advice is unglamorous. Report effect sizes and confidence intervals, so the reader can see how big, not just whether — in the aspirin trial the number worth arguing about is 0.03, not the string of zeros. Preregister the analysis, so the 5% actually means 5%. Publish the null results, or the literature will consist entirely of the 5% that cleared the bar. And when the null really does hold, say so with an equivalence test — "the difference is too small to matter" — instead of hiding behind "no significant difference".
Back to the word. Null was always about being nullifiable. It names the hypothesis the experiment sets up for the data to knock down: the starting point of an inference, not its conclusion. The moment "we failed to reject the null" gets announced as "we proved there is no effect", the word has been turned inside out.
Timeline
The working manual of significance testing. One in twenty is called a convenient line to draw; it becomes 0.05.
H0 paired with H1, errors of the first and second kind, a critical region. Testing is rewritten as a rule about long-run behaviour.
The lady tasting tea. The null hypothesis is defined as the one that may be disproved and is never proved.
They attack each other in journals and from lecterns. The quarrel is about what inference is for: weighing evidence, or governing behaviour.
Methods texts in psychology and the social sciences teach a hybrid: report the p value, and claim the long-run error guarantee. Neither man would have signed it.
A procedure performed mechanically, with no author, and no one willing to defend its logic.
Large replication projects fail to reproduce a great many published significant results. p-hacking, publication bias and low power are identified as structural causes.
Six principles. The bluntest: scientific conclusions should not rest on whether a p value clears a threshold. The same year, Greenland and colleagues list twenty-five misreadings.
Sources
- 1.Ronald A. Fisher (1935). The Design of Experiments Edinburgh: Oliver and Boyd (1st ed.). Ch. II §5 'The Statement of Experiment', p.13 (the tea-tasting setup); §8 'The Null Hypothesis', p.18 (1935 first edition; some reprints give p.19).Read it ↑1↑2
- 2.Gerd Gigerenzer, Stefan Krauss, and Oliver Vitouch (2004). The Null Ritual: What You Always Wanted to Know About Significance Testing but Were Afraid to Ask In D. Kaplan (Ed.), The Sage Handbook of Quantitative Methodology for the Social Sciences, pp.391–408. Thousand Oaks, CA: Sage. p.392 (Gigerenzer et al.'s interpretive account of Fisher's original intent behind the term "null"; not a dictionary-grade etymology entry).Read it
- 3.John T. E. Richardson (2021). A Closer Look at the Lady Tasting Tea Significance, vol. 18, no. 5, pp.34–37. Full article — re-examines Fisher's original account of the tea-tasting experiment in The Design of Experiments against two later retellings: Joan Fisher Box's 1978 biography of her father, and David Salsburg's 1968 account relaying Hugh Fairfield-Smith's decades-later recollection.Read it
- 4.Jerzy Neyman and Egon S. Pearson (1933). On the Problem of the Most Efficient Tests of Statistical Hypotheses Philosophical Transactions of the Royal Society of London, Series A, vol. 231, pp.289–337. p.291 (the 'rule of behaviour' passage, the original statement of the decision-theoretic logic); pp.296-297 (the notation H0, and the first systematic statement of what became Type I and Type II error). The word 'power' itself is not yet used as a fixed term in this paper.Read it
- 5.Gerd Gigerenzer (2004). Mindless Statistics The Journal of Socio-Economics, vol. 33, no. 5, pp.587–606. pp.591-592: works through a metal-plate manufacturer's quality-control scenario to show how Neyman-Pearson decision logic (setting Type I/II error rates, choosing sample size by cost-benefit trade-off, "accepting" a hypothesis without believing it) grew out of industrial acceptance-sampling logic.Read it
- 6.Gerd Gigerenzer, Stefan Krauss, and Oliver Vitouch (2004). The Null Ritual: What You Always Wanted to Know About Significance Testing but Were Afraid to Ask In D. Kaplan (Ed.), The Sage Handbook of Quantitative Methodology for the Social Sciences, pp.391–408. Thousand Oaks, CA: Sage. p.392 (the three-step definition of the 'null ritual'; six true/false items demonstrating six common misreadings of the p-value, all in fact false).Read it
- 7.Ronald A. Fisher (1925). Statistical Methods for Research Workers Edinburgh: Oliver and Boyd. 13th ed. (1958 printing series), p.44 (some editions give p.47, reflecting differences between printings).Read it
- 8.John P. A. Ioannidis (2005). Why Most Published Research Findings Are False PLoS Medicine, vol. 2, no. 8, e124. Full article — uses a positive predictive value (PPV) model to argue that in fields with low prior probability, small samples, and flexible analysis practices, most published statistically significant findings are in fact false positives.Read it
- 9.Steering Committee of the Physicians' Health Study Research Group (1988). Preliminary Report: Findings from the Aspirin Component of the Ongoing Physicians' Health Study New England Journal of Medicine, vol. 318, no. 4, pp.262–264. 22,071 male physicians randomized to aspirin (n=11,037) or placebo (n=11,034); 104 myocardial infarctions occurred in the aspirin group versus 189 in the placebo group, relative risk approximately 0.56 (a 44% reduction) — the effect was extreme enough that the trial's aspirin component was stopped early.Read it
- 10.Robert Rosenthal (1990). How Are We Doing in Soft Psychology? American Psychologist, vol. 45, no. 6, pp.775–777. Reinterprets the aspirin/myocardial-infarction trial's effect size using the Binomial Effect Size Display (BESD): a correlation of r approximately .03 is small in magnitude, but translated into "heart attacks prevented per N people," the practical effect is not negligible.Read it
- 11.Joseph P. Simmons, Leif D. Nelson, and Uri Simonsohn (2011). False-Positive Psychology: Undisclosed Flexibility in Data Collection and Analysis Allows Presenting Anything as Significant Psychological Science, vol. 22, no. 11, pp.1359–1366. Study 2: 20 University of Pennsylvania undergraduates listened to either the Beatles' "When I'm Sixty-Four" or a control song ("Kalimba"), then reported their birth date and father's age; adjusted mean age was 20.1 in the Beatles condition versus 21.5 in the control condition, F(1,17) = 4.92, p = .040. The paper's simulations separately show that allowing flexibility in data collection and analysis can push a single comparison's nominal 5% false-positive rate up to 60.7%.Read it
- 12.Open Science Collaboration (2015). Estimating the Reproducibility of Psychological Science Science, vol. 349, no. 6251, article aac4716. Replicated 100 empirical psychology studies (originally published in three journals): 97% of the original studies had reported statistically significant results (p < .05); only 36% of the replications reached statistical significance, and replication effect sizes were on average about half the magnitude of the original effect sizes.Read it
- 13.Sander Greenland, Stephen J. Senn, Kenneth J. Rothman, John B. Carlin, Charles Poole, Steven N. Goodman, and Douglas G. Altman (2016). Statistical Tests, P Values, Confidence Intervals, and Power: A Guide to Misinterpretations European Journal of Epidemiology, vol. 31, no. 4, pp.337–350. §2, the list of 25 common misinterpretations, corresponding to the section headed "Twenty-five misinterpretations."Read it
- 14.Ronald L. Wasserstein and Nicole A. Lazar (2016). The ASA Statement on p-Values: Context, Process, and Purpose The American Statistician, vol. 70, no. 2, pp.129–133. The six principles, listed starting p.131 (from "Principle 1").Read it