StatisticsLab

How to Show There Is No Difference

A test returns p = 0.56, and you cannot conclude the two are the same — a non-significant result may just mean the study had no power to see the difference. This article uses drug bioequivalence to explain statistical power (how large an effect a study can see) and equivalence testing (TOST: bounding the effect inside a margin fixed in advance), and why "not significant" has to be split into "we can rule out any meaningful difference" and "we simply could not tell".

On this page

Before a generic blood-pressure drug can go on sale, its maker has to show it works the same as the brand-name original. So they run a trial: two groups of patients, one on the brand, one on the generic, and they compare how far each group's systolic pressure drops. The two groups differ by 0.6 mmHg. A t test gives p0.56p \approx 0.56. Not significant.

The maker concludes: no significant difference, so the two drugs are the same.

That conclusion is wrong, and it is wrong in the way the first article was about. A is never established by your failure to overturn it. All p0.56p \approx 0.56 tells you is that this data could not reject "the two are identical" — and the same data cannot reject "the two differ by 2 mmHg" either. A non-significant result can mean there really is no difference, or it can mean the study had no power to see one. The two situations look identical, and the pp value cannot tell them apart.

So what should you do? If you genuinely want to claim "the two differ by nothing that matters in practice", you have to ask a different question. This article covers two things: how large an effect the study can even see (), and how to bound the effect inside a margin on purpose ().

What a non-significant result is compatible with

Start by opening up that p0.56p \approx 0.56. Underneath it sits an estimate and its uncertainty: the two groups differ by 0.6 mmHg, with a 95% running from -1.4 to +2.6 mmHg.

Suppose the two really differ by 2.0 mmHg — this batch is compatible with that.

Point estimate 0.6 mmHg, 95% confidence interval [-1.43, 2.63]. two-sided p = 0.56

0 (identical)95% CIyour value-6-4-20246difference between groups (mmHg)
Fixed data: the two group means differ by 0.6 mmHg, 120 per group, pooled SD 8 mmHg. The bar is the 95% confidence interval, the filled dot the point estimate, the hollow dot the true difference you assume with the slider. As long as it lands on the bar, the data are compatible with it.

The interval contains 0, which is why the test is not significant — another way of saying "we did not detect a difference". It also contains +2 mmHg, and -1 mmHg. Drag the slider and you will see that this data is just as compatible with "the generic is actually 2 mmHg better" as it is with "the two are identical". : it says only that 0 lies inside the interval, not that 0 is any more credible than the other values in there.

How wide the interval is depends on how many patients the study enrolled. More people, a narrower interval, a clearer view. That property has a name.

How large an effect the study can see

asks: if an effect of some given size is really there, what is the chance this study detects it as significant? Turned around, the smallest effect a study can reliably detect is its minimum detectable effect. For this trial — 120 patients per group, a standard deviation of 8 mmHg — only a true difference larger than about 2.9 mmHg has an 80% chance of coming out significant. A difference smaller than that, the design was never able to see.

So "not significant" splits into two situations that have nothing in common. In one, the study is powerful enough to see any difference that matters, and it genuinely saw none — here you have earned the right to say "no meaningful difference". In the other, the study had no power, and would have missed a real difference anyway — here all you can honestly say is "I don't know". To make the first statement, you need a tool built for it.

Equivalence testing: bounding the effect

: instead of asking "is there a difference", first say plainly "how large a difference would matter", then ask "can the effect be bounded inside that range".

Fix a margin Δ\Delta. In the blood-pressure example, clinicians already agree that a gap below some size makes no practical difference to a patient. That Δ\Delta is a medical judgement, and it has to be set before you look at the data.

Then run two one-sided tests (TOST for short): one testing whether the difference exceeds the upper bound +Δ+\Delta, one testing whether it falls below the lower bound Δ-\Delta. Reject both, and you have said that the effect is neither above +Δ+\Delta nor below Δ-\Delta — it is boxed inside ±Δ\pm\Delta. Now you can declare equivalence, as a positive claim.

There is a cleaner way to say the same thing. (90%, not 95%: two one-sided tests at 5% each correspond to a 90% two-sided interval).

margin = 3.0 mmHg: the whole 90% interval sits inside the band — equivalence declared.

Two one-sided tests p1 = 2.9e-4, p2 = 0.010, TOST p = max = 0.010 < 0.05 → equivalent

The smallest margin that can declare equivalence is 2.31 mmHg (the interval's half-width).

margin ±3.0090% CI-6-4-20246difference between groups (mmHg)
The pale band is the equivalence margin you set; the bar is the 90% confidence interval (not 95%: two one-sided tests of 5% each make 90%). Equivalence is declared when the whole bar is inside the band. Drag the margin below the interval's half-width and the bar pokes out. Same fixed data as the previous figure.

Drag the margin Δ\Delta: when the 90% interval shrinks entirely inside ±Δ\pm\Delta, both one-sided pp values drop below 0.05 and equivalence holds. Drag Δ\Delta narrower than the interval's half-width and you can no longer declare it — the "sameness" you are demanding is tighter than the study can guarantee.

Two non-significances, two conclusions

Back to that p0.56p \approx 0.56. Its 90% confidence interval runs from -1.1 to +2.3 mmHg.

If clinicians agree that "within 3 mmHg is no difference" (Δ=3\Delta = 3), the interval sits entirely inside ±3\pm 3, both one-sided tests pass, and you can state it: the two drugs are clinically equivalent. This non-significance is the first kind — the study had power, and it genuinely saw no difference that mattered.

If the clinical standard is stricter, requiring "within 2 mmHg to count as the same" (Δ=2\Delta = 2), the interval's upper edge at +2.3 pokes past it. The same data now cannot declare equivalence: the study lacks the precision to guarantee a difference below 2 mmHg. This non-significance is the second kind — could not detect is not the same as does not exist.

The discipline lives here: Δ\Delta has to be fixed before the data, never chosen after seeing where the interval landed and picking one that just barely contains it. That is drawing the target around the arrow, the same disease as the from the first article.

Why it matters, and where it breaks

Generic-drug review is where equivalence testing is most fully institutionalised. for the two to count as the same. That is precisely the logic above: nothing gets approved on the strength of "we found no difference"; the difference has to be boxed inside a range fixed in advance.

Three limits, all concrete. First, Δ\Delta is a value judgement, not a statistic — how large a difference counts as "unimportant" is decided by knowledge of the field, and statistics cannot stand in for it. Second, Δ\Delta has to be set in advance. Third, with too small a sample both directions fail: you can neither reject "there is a difference" nor declare "equivalent", and all you can report is that the study settled nothing. Saying "I don't know" is more honest than forcing a side.

The first article closed on one line: when the null really holds, use an equivalence test to say outright that "the gap is too small to matter", instead of vaguely reporting "no significant difference". This article is that line unfolded. "We did not detect a difference" is a statement about the study; "the difference is too small to matter" is a statement about the two drugs. Passing the first off as the second is the most common substitution in statistics.

Sources

  1. 1.Daniël Lakens (2017). Equivalence Tests: A Practical Primer for t Tests, Correlations, and Meta-Analyses Social Psychological and Personality Science, 8(4), pp.355-362. The paper as a whole: a practical primer bringing TOST out of pharmacology into general social science. It shows how to set a smallest effect size of interest as the equivalence bound in advance, how to run equivalence tests for t tests and correlations, and stresses that 'not significant' cannot be read as 'no effect'. DOI 10.1177/1948550617697177.Read it
  2. 2.Donald J. Schuirmann (1987). A Comparison of the Two One-Sided Tests Procedure and the Power Approach for Assessing the Equivalence of Average Bioavailability Journal of Pharmacokinetics and Biopharmaceutics, 15(6), pp.657-680. The paper as a whole: the original proposal of the two one-sided tests (TOST) procedure. Equivalence is split into two one-sided tests — one against the lower bound -Delta, one against the upper bound +Delta — and declared only if both are rejected at level alpha; the paper also shows this gives the same decision as checking whether the (1-2alpha)×100% confidence interval lies inside ±Delta. DOI 10.1007/BF01068419.Read it
  3. 3.Roger L. Berger and Jason C. Hsu (1996). Bioequivalence Trials, Intersection–Union Tests and Equivalence Confidence Sets Statistical Science, 11(4), pp.283-319 (with discussion). The paper as a whole: a rigorous treatment of the duality between TOST and the confidence-interval inclusion rule. The key result is that deciding equivalence at alpha = 0.05 by checking whether a 90% confidence interval falls inside ±Delta gives exactly the same decision as TOST (the Berger–Hsu duality); it also places equivalence testing inside the intersection–union framework. DOI 10.1214/ss/1032280304.Read it
  4. 4.U.S. Food and Drug Administration, Center for Drug Evaluation and Research (2001). Statistical Approaches to Establishing Bioequivalence (Guidance for Industry) U.S. Department of Health and Human Services, FDA/CDER. Average bioequivalence criterion: on the log-transformed AUC and Cmax, the 90% confidence interval for the test/reference ratio of geometric means must lie entirely within 80.00%–125.00%. The bounds are symmetric on the log scale (ln 0.80 = -ln 1.25). See 21 CFR Part 320 (§320.24 lists the accepted ways to demonstrate bioequivalence).Read it