Sample size determines a study's statistical power, the chance of detecting a real effect of a given size, and underpowered studies routinely miss real effects while inflating the ones they catch. Conventional medical research designs aim for about 80 percent power, which means roughly one credible finding in five is expected to be missed by design. The landmark 1948 Medical Research Council streptomycin trial needed just over one hundred patients because tuberculosis at the time was often fatal, so its effect was large; most modern questions are far less obliging.
This is an explainer about research methods. It describes how studies are planned and interpreted, not what any reader should do about a medical condition.
What is statistical power?
Power is the probability that a study will detect an effect of a particular size if that effect truly exists. It is set mostly by three quantities: sample size, the threshold of statistical significance, and the size of the effect worth finding. Increase any one and power rises. Researchers conventionally target 80 percent power at the 0.05 significance level, a pair of habits that quietly budget for missing one real finding in five, and for a false alarm rate of one in twenty when nothing exists.
Power is not a verdict on quality; it is a budgeting decision made before data collection. A protocol states the effect size it is designed to detect, often called the minimum clinically important difference. A study powered to detect a large effect may be perfectly adequate for its question, and useless for a smaller one, which is why the target should be printed in every report.
Why do small studies overestimate effects?
Because of how significance filtering works. To reach significance with few participants, an observed effect must be large, so small studies that publish are systematically the lucky ones, those in which noise happened to swell the estimate. Methodologists borrow the term winner's curse from economics for this pattern. The true modest effect fails to clear the bar and stays unpublished, while its exaggerated twin goes to press.
The consequence is predictable: when a large trial follows a batch of promising small studies, the effect usually shrinks. This is one reason early-phase results deserve patience rather than headlines, and why a striking effect from a small sample is more often a warning sign than a discovery.
How do researchers calculate the sample size they need?
Before enrollment begins, statisticians run a calculation that works backwards from the question: given the expected event rate in the control group, the effect size worth detecting, the significance threshold and the target power, how many participants are required? Adjustments follow for dropouts, for unequal arm sizes, and for analyses at several time points.
- State the outcome and the event rate expected without the intervention.
- Choose the minimum effect size that would change practice.
- Set significance threshold and power, conventionally 0.05 and 80 percent.
- Compute the required participants per arm, then inflate for expected attrition.
- Register the plan before the first participant is enrolled.
When the calculated number is impractical, options include choosing a more common outcome, extending follow-up, standardizing measurements to reduce noise, or collaborating across centers. What does not work is enrolling whoever is available and calling it a study.
| Design factor | Effect on required sample size |
|---|---|
| Larger target effect | Fewer participants needed |
| Rarer outcome events | Many more participants needed |
| Noisier measurements | More participants needed |
| Higher power target | More participants needed |
| Multiple comparisons | More participants, or stricter thresholds |
Is a bigger sample always better?
Not automatically. A huge sample detects differences too small to matter, which is the mirror image of the small-study problem: significance without consequence. Size also cannot repair a flawed design; a randomized trial with biased outcome assessment stays biased at any enrollment, and a survey of an unrepresentative crowd is wrong at scale. The goal is adequacy for a stated question, not raw headcount.
Practical costs also grow faster than numbers, which is why multi-center collaboration and pooled analyses have become standard for uncommon diseases, where no single site sees enough patients. The United States Food and Drug Administration describes the staged structure of clinical research, small safety cohorts first, larger efficacy trials later, in its public materials for patients.
How do we know a study was adequately sized?
The published record allows a direct check, and it takes under a minute. Find the protocol's target effect size and the achieved sample, and compare them with the reported result. If a study of 40 patients reports a moderate effect with a wide interval that includes both trivial and impressive values, the study was too small to settle its own question. If a study reports a power calculation for one outcome but headlines another, the headline is under-powered by construction.
Confidence intervals carry this information compactly: a narrow interval around a clinically meaningful effect is the signature of an adequately powered study, and a wide interval is an honest admission that the data cannot yet discriminate. Readers who check the interval before the p-value will mislead themselves far less often.
What about rare diseases, where big samples are impossible?
Then design compensates for numbers. Crossover designs let each participant serve as their own control; registries pool patients across countries; adaptive designs stop early only when evidence is overwhelming and continue when uncertain. Small and rare is a real constraint, and honesty about it, wide intervals, provisional conclusions, replication before adoption, is the correct response, rather than confident claims built on thin data.
For more context, read How a finding earns the label reproducible.
For more context, read meta-analysis.
