Anikaay.Independent reporting and practical explainers

Science

What a Confidence Interval Actually Tells You

When a study reports a 95% confidence interval, the most common misreading is that there is a 95% chance the true value lies inside it. That is not what it means, and the difference matters when you are deciding what to do.

The shorthand everyone uses

A study reports: “the effect was 4.1 percentage points, 95% CI [1.2, 7.0]”. The honest translation is much less quotable: if we repeated this entire study many times, and each time computed an interval in the same way, about 95% of those intervals would contain the true value.

Everything below follows from noticing that the 95% refers to the long-run success rate of the method, not to the confidence any particular result has.

Why the common misreading is wrong

Almost every news story about a study contains some version of “there is a 95% chance the result lies between 1.2 and 7.0 percentage points.”

This is a subtle error. The frequentist framework does not assign probabilities to fixed parameters — the true value is what it is, and we do not know it. The interval is a random quantity that varies from study to study, and the 95% describes how often intervals constructed this way capture the truth.

Where does this bite in practice? Consider a sequence of studies on the same question, each finding a significant effect and none replicating. Each study individually had an interval excluding zero. But with enough studies, at least one would be expected to miss purely by chance. This is the publication-bias problem that interval arithmetic makes concrete: a significant result is less surprising than it looks, because the studies that found nothing are disproportionately absent from the literature.

The Bayesian alternative is a genuinely different framework, in which the parameter does carry a probability distribution and “there is a 95% probability the true value lies in this interval” is meaningful. Converting between the two requires a prior, so the two framings do not produce identical numbers for identical data.

What the interval tells you about precision

Read the width first, before the point estimate.

A wide interval means the study could not pin down the answer precisely. This happens when the sample is small, the measurement is noisy, or the effect varies a lot across cases. All three are common, and a wide interval on a well-motivated question is a normal research outcome rather than a failure.

Width also tells you something about sample size directly. Halving the width requires roughly quadrupling the sample. This is why the cost of precision is usually the strongest argument for a larger study.

The standard error, and why it can mislead

The standard error of a mean is the standard deviation of individual observations divided by the square root of the sample size. Dividing by the square root makes the standard error shrink with sample size even when the underlying variation is enormous.

The practical consequence: a large sample can produce a very narrow interval around a statistically precise but substantively meaningless estimate. A trial with 50,000 participants can determine an effect size of two-tenths of a percentage point with high confidence. Whether that effect matters to anyone is a separate question that no amount of sample size will answer.

This is why a confidence interval should always be read next to the effect’s practical significance. An interval that excludes zero but sits entirely within a range of effects nobody would act on is a precise answer to a question you did not mean to ask.

Overlapping intervals

A widespread shortcut: if two groups’ intervals overlap, the difference is not significant.

This is not a valid test. It compares two interval estimates when what you want is an interval for the difference between them. The two are different quantities.

The failure is real in both directions. Two intervals can overlap substantially while the difference between groups is statistically significant — common with large samples, where both intervals are narrow and both are far from a meaningful effect. And two intervals can barely touch while the difference is not significant, particularly with small samples and wide intervals.

If you want to know whether two groups differ, the correct analysis estimates the difference directly and produces an interval for that difference. The shortcut is the reason “significantly different from” claims in published research are frequently wrong, and it appears often enough in press coverage that recognising the pattern is worth the effort.

Non-significant does not mean no effect

This is the most consequential mistake in both directions. A result whose interval includes zero means the study was not precise enough to establish the effect in either direction.

It does not mean the effect is absent. It means the experiment could not distinguish a meaningful effect from none. These are very different claims, and reporting them interchangeably leads directly to the claim that a therapy “does not work” when the accurate statement is that it was never tested adequately.

The symmetric error is just as bad: treating a non-significant result as evidence of no effect also ignores the possibility of harm.

Heterogeneity and prediction

Population-level intervals answer “what is the average?” For an individual case, the question is “what would happen to this person?”, and these are answered by different quantities.

An effect can be real on average and apply to only a small subgroup — highly effective for some patients, ineffective for others. A single average interval conceals this entirely, and reading it as a statement about a given individual is a category error.

Prediction intervals are the appropriate tool for individual cases. They are wider than confidence intervals, because they combine uncertainty about the underlying effect with the natural variation between individuals.

How to read one in the wild

When a study reports an interval:

  1. Check whether it is an interval for the effect or for the raw rate. Many reported intervals are for a proportion — a 40% rate with a 95% CI of [37%, 43%] — which is not the same object as an interval for a difference between groups.
  2. Look at the width before the point estimate. Width tells you whether the study could answer the question.
  3. Compare the interval to a meaningful effect size, not to zero. Ask whether the whole interval sits above the smallest effect worth acting on.
  4. Check what was controlled for. An interval for a raw comparison and an interval for an adjusted comparison of the same data can point in different directions.
  5. Ask whether the interval is percentile or bootstrap. These methods do not always behave well with small samples, and confidence limits can extend outside the possible range of the underlying measure — a proportion interval running past 100% is a signal of that.

What to take away

Confidence intervals quantify how precisely a study pinned down its answer. They are the right tool and they are widely reported.

They are not, on their own, a measure of how important a result is, how likely it is to replicate, or whether it applies to the case in front of you. Those questions need the study design, the effect size and the population — none of which fits inside the parentheses.

  • statistics
  • research-methods
  • probability

Frequently asked questions

Is a confidence interval the same thing as a prediction interval?

No, and the distinction matters in forecasting. A confidence interval describes uncertainty about an estimated population parameter — the average effect, a true rate. A prediction interval describes the range of a single future observation, such as an individual patient outcome. Prediction intervals are wider, because a single outcome carries both parameter uncertainty and individual variation.

Why do overlapping confidence intervals not tell me whether two groups differ?

Because the overlap test is about two separate interval estimates, each of which has its own uncertainty, and it is not the same comparison as testing the difference directly. You can get overlapping intervals with a significant difference and non-overlapping intervals without one, particularly with small samples. The correct approach is a direct test of the difference between groups, which produces its own interval.

If the interval includes a value with no plausible effect, is the study useless?

Not necessarily. An interval that includes zero when you care about effects away from zero means the study was not precise enough to rule out a practically important effect. That is a finding about precision, not about the world, and it is a reason to be cautious about drawing a conclusion rather than a reason to believe the opposite.

Sources and references

  1. The Interpretation of Confidence Intervals — Biometrika (via TU Delft repository), accessed 2026-09-15
  2. Statistics Done Wrong — on confidence intervals — Daniel Lakens, accessed 2026-09-15

Technology

What Actually Happens When You Get a 404

Not all missing pages behave the same way, and the difference between a real 404 and a "soft" one is invisible in the browser but very visible to a crawler. It is also the most common self-inflicted SEO problem on otherwise healthy sites.