StatisticsLab

Statistics Lab

Textbooks teach you how to compute. They rarely tell you what problem a method was invented to solve, or where it started being misused. This site takes one such idea at a time — one that looks self-evident and has a questionable history — and asks why it has the shape it has, who built it and for what, and what exactly breaks when it is misapplied. The simulations are yours to turn, because some things you only believe after re-running them ten thousand times yourself.

Inference

Hypothesis tests, estimates, intervals — the most used and the most misread part of the subject.

Where Do the Justices Stand: Estimating Ideal Points from Votes

A Bayesian graded-response model turns real votes into an axis of tendency to strike down, and a permutation test asks whether it lines up with the nominating president — a disciplined null, plus one collaboration-layer signal that is real

Where does the axis come from when we place a justice on a scale of tendency? Using the Taiwan Constitutional Court's official per-item vote records since 2022 as real votes, this piece estimates each justice's tendency to strike down (theta) with a Bayesian graded-response model, then tests whether that axis aligns with the nominating president. Nothing is detectable on four dimensions, and the null carries power, equivalence, and specification-curve discipline; per-justice credible intervals are wide and all adjacent ones overlap, so in this sparse-dissent term the justices are mostly indistinguishable. A second layer, the co-signing network of opinions since 1949 modelled with a hierarchical binomial, does find a moderate and robust same-nominee effect. The voting layer detects nothing; the collaboration layer detects something; the two speak to different things. No prior Taiwanese study has used official per-item real votes — this piece fills that gap, and reserves its novelty claim against one closely overlapping study whose full text remains unavailable.

2026-07-19 · 1 min read · ideal points, item response theory, graded response model, Bayesian estimation, permutation test, judicial behavior, non-significant result

How to Show There Is No Difference

Power and equivalence tests: turning "we found no difference" into "any difference is too small to matter"

A test returns p = 0.56, and you cannot conclude the two are the same — a non-significant result may just mean the study had no power to see the difference. This article uses drug bioequivalence to explain statistical power (how large an effect a study can see) and equivalence testing (TOST: bounding the effect inside a margin fixed in advance), and why "not significant" has to be split into "we can rule out any meaningful difference" and "we simply could not tell".

2026-07-18 · 6 min read · equivalence testing, TOST, statistical power, confidence interval, bioequivalence, non-significant result

What a Confidence Interval Actually Is

A poll's '±4 points', and the question of whom that 95% is really about

A poll puts a candidate at 45% with a margin of error of ±4 points, and the paper writes that there is a 95% chance the true support lies between 41% and 49%. Almost everyone says this, and it is wrong. This piece derives a confidence interval by hand from a single mean, and watches the 95% lose its footing at the exact step where the observed numbers are plugged in. A submarine, a ratio interval, and a six-statement quiz then show that the experts fail at the same spot. The frequentist tool is not broken — it answers a well-defined question correctly. That question is just not the one you meant to ask.

2026-07-18 · 12 min read · confidence interval, margin of error, frequentist inference, Neyman, credible interval, misinterpretation

Why It Is Called the Null Hypothesis

From a lady tasting tea to a signed statement: a short history of the p value quarrel

Null was never about nothingness. Fisher's null is the hypothesis that can be nullified — set up so the data have something to argue against. This piece follows one line: how a tea party produced the first p value, how Neyman and Pearson turned it into a decision rule, how textbooks welded the two together into something neither man designed, and where that weld cracks.

2026-07-13 · 7 min read · null hypothesis, p value, history of ideas, Fisher, Neyman-Pearson, replication crisis

How Large Are the Differences Between Justices?

Small-sample proportions and partial pooling in a Bayesian hierarchical model

Justice A files a separate opinion in 3 of 4 participating cases; justice B does so in 40 of 100. The two proportions describe the observed cases, but they do not justify a ranking. This article uses a binomial model, posterior distributions, and partial pooling to show how small-sample proportions can be estimated and reported.

2026-07-13 · 4 min read · Bayesian statistics, hierarchical models, partial pooling, small samples, Constitutional Court data, uncertainty

Randomness and convergence

How many times a procedure has to repeat before its output counts as random: how the distance is defined, and how fast it closes.

Every example has a source, and the source hangs beside the sentence. The simulations run on fixed seeds, so the same button gives the same numbers on any machine. The whole site is in both languages; the switch is at the top right.

About this site: method, sources, and the discipline behind the simulations