A confidence interval is the range of values a study's result could plausibly take given its sample, and it tells a reader two things at once: the best estimate and the honest uncertainty around it. A treatment reported as cutting risk by a quarter might carry an interval running from a large benefit to nearly none — and that range, printed in every abstract from clinical trials to public-health surveys, is where the real finding lives. The U.S. National Cancer Institute's materials for consumers treat intervals as a standard part of communicating study results, not a technical afterthought.
Engevity News publishes information, not medical advice. This piece explains how to read a statistical range; it does not evaluate any treatment, screening test, or diagnosis.
What does the range actually mean?
A 95 percent confidence interval means that if a study were repeated many times with the same method and new samples, about 95 percent of such intervals would contain the true value. It is a statement about the procedure's long-run reliability, not a probability that the truth sits inside this particular range. For a reader, the operational meaning is simpler and fair: values inside the interval are compatible with the data, and values outside are largely not. If a relative risk of 1.3 comes with an interval of 1.05 to 1.8, the data are also consistent with a small association; if the interval were 0.9 to 1.9, the data would be consistent with no association at all — and with a large one.
Why is a wide interval a warning sign?
Width measures imprecision, and imprecision usually traces back to sample size: small studies produce wide intervals because few observations leave much room for chance. A wide interval is not fraud and not failure — it is an honest confession. The problem is what happens downstream: a small study's best estimate can look dramatic simply because noise has nowhere to hide it. This is why a finding of "risk doubled" from 40 people, with an interval stretching from slightly protective to massively harmful, deserves far less weight than a finding of "risk up 15 percent" from 400,000 people with a tight range. Small studies do not produce small claims; they produce unstable ones.
How does the interval relate to statistical significance?
The two are secretly the same information. For ratios — relative risks, odds ratios — a result is conventionally "statistically significant" when its 95 percent interval excludes the value 1.0, the marker of no difference; for differences in averages, the no-difference marker is zero. This gives readers a shortcut that skips the significance vocabulary entirely: find 1.0 inside the interval and the study has not ruled out "no effect", whatever the headline said. Finding 1.0 just barely outside the interval tells a subtler story — an effect that is detectable but fragile. The interval turns a binary verdict into a spectrum, which is what the evidence actually supports.
How do we know intervals communicate better?
The preference for intervals over bare significance testing is a long-standing recommendation of statisticians. The American Statistical Association's 2016 statement on p-values explicitly urged researchers to report effect sizes with confidence intervals rather than rely on thresholds, and medical journals have required intervals in abstracts for decades. The reasoning is empirical as much as philosophical: bare verdicts hide magnitude and precision, and readers given ranges make better-calibrated judgments about how firm a finding is. In fields from epidemiology to psychology's replication debates of the 2010s, wide intervals around early small studies proved to be early warnings that later, larger work did not hold the effects up.
What can a confidence interval not tell you?
The interval covers only the uncertainty that comes from sampling — from having measured a sample rather than everyone. It cannot see bias: a flawed questionnaire, a sample that does not resemble the population, dropouts, or a design that cannot isolate causes all produce confident-looking intervals around wrong numbers. It also cannot see the future: intervals describe this study's data, not the range of results later studies will find, though intervals that fail to overlap across studies are a strong hint that something — population, method, or chance — differs. Precision is not accuracy, and a tight interval around a biased estimate is a sharp picture of the wrong thing.
What do overlapping intervals between studies mean?
Readers comparing two studies often meet intervals that overlap and wonder whether the findings conflict. Overlap alone does not settle it: two estimates can have overlapping ranges and still differ meaningfully by formal test, because the comparison depends on the joint uncertainty of both, not on a visual inspection. The practical lesson runs the other way — intervals from different studies that fail to overlap are a strong signal that the true effects differ across populations, methods, or eras, and that at least one study's assumptions deserve a closer look. A cluster of studies whose intervals all cover a modest effect, with one early study's wide interval straying far, usually resolves toward the cluster as data accumulate. The interval, in other words, is not just a per-study caution label; it is the unit by which science stitches separate results into a cumulative answer.
How should a reader use intervals in practice?
A short reading routine turns intervals into judgment.
- Locate the interval in the abstract — it is usually in parentheses beside the main result.
- Find the no-effect marker: 1.0 for ratios, zero for differences.
- Ask whether both ends of the range would matter if true.
- Check the sample size when the interval is wide.
The third step is the one that decides practical meaning: an interval running from "barely any benefit" to "enormous benefit" is compatible with the data, but only the low end may justify changing behavior. The range, read carefully, is the study telling its reader exactly how much to believe — and how much to wait.
For more context, read How to check a health claim in five minutes.
For more context, read how to read a scientific paper.
For more context, read Why correlation is not causation in headlines.
