Researchers check reproducibility by repeating a study's procedure, or its analysis, under conditions agreed in advance, and then comparing what appears. The largest direct test, published in Science in 2015 by the Open Science Collaboration, repeated 100 experimental psychology findings and saw about 36 percent of the original significant effects reach significance again, with effect sizes roughly halved on average. Replication is not an insult to the original; it is the only mechanism science has for separating durable findings from lucky ones.
What follows is a description of research practice, not advice about any medical condition or treatment decision.
What is the difference between reproducing and replicating?
The terms split cleanly. Reproducing a result means re-running the same data and the same analysis, checking that the numbers hold; this catches computational errors, miscoded variables and selective reporting. Replicating means collecting new data under a matching design and testing whether the finding appears again, which is the stronger test. Direct replications keep the method as close to the original as possible; conceptual replications test the same idea through a different route, which is informative but easier to explain away if it fails.
Both forms matter, and they answer different questions. Reproduction asks whether the reported analysis is faithful; replication asks whether the phenomenon is dependable. A finding can reproduce perfectly, clean code, clean data, and still fail to replicate because the original sample was lucky.
What did the big replication projects actually find?
They found wide variation, and that is the headline. The 2015 psychology project described above was followed by efforts in other fields. Surveys of researchers themselves point the same direction: a 2016 Nature survey of more than 1,500 researchers, published in the journal's comment section, found that some 70 percent had tried and failed to reproduce another scientist's experiment, and about half had failed to reproduce their own.
The pattern does not mean the fields are broken or that most published work is false. It means that a single study, however elegant, is a provisional data point, and that fields relying on small samples and flexible analysis should expect shrinkage under repetition. Some subfields, including several preclinical research areas, have since organized systematic replication programs, and results there have been mixed rather than uniformly grim.
Why do findings shrink when repeated?
Several forces push in the same direction. Small samples overestimate effects that happen to clear the significance threshold, the winner's curse described by methodologists. Flexible analysis, choosing among outcomes, measures and subgroups after seeing the data, manufactures publishable patterns from noise. Publication bias hides the unsuccessful attempts, so the record looks stronger than the evidence. And some effects are genuinely conditional, they depend on population, timing or context that the original authors did not fully specify.
- Small samples: published small studies are the lucky tail, and luck does not repeat.
- Flexible analysis: many plausible analyses guarantee some significant results by chance.
- Publication bias: failures vanish from the record, inflating apparent reliability.
- Under-specified methods: when the procedure is ambiguous, the repeat differs from the original in ways nobody can list.
How is a replication designed to be fair?
A credible replication is planned like a trial, before results exist. The team, often including the original authors, fixes the procedure, sample size and analysis in advance, often with public registration. Power is set high enough to detect the original effect size, because an underpowered replication that fails proves nothing. Criteria for success, significance, effect size comparison, interval estimates, are stated before data collection, so no one can relabel the outcome afterward.
This pre-commitment is the whole game. A replication that failed could have succeeded under a different analysis, and one that succeeded could have failed, so the discipline lies in freezing choices while still ignorant of the result.
How do we know a failed replication means the original was wrong?
Not always, and honest reporting says so. A replication can fail because the original was a fluke, because the repeat was flawed or underpowered, or because the effect truly depends on conditions that changed. The reader's check is symmetric: was the replication registered, was it powered for the original effect, were the original authors involved in specifying the procedure, and were the success criteria fixed in advance? When those boxes are ticked, a failure carries real weight; when they are not, both studies remain in suspension.
The mature conclusion is cumulative. Two well-run studies in tension define an open question; three or four converging results settle it. Fields that treat replication as routine, with registered reports and published protocols, move faster precisely because they waste less time on lucky single findings.
| Term | What is repeated | What it tests |
|---|---|---|
| Reproduction | Same data, same analysis | Computational and reporting fidelity |
| Direct replication | New data, same procedure | Dependability of the finding |
| Conceptual replication | New data, related method | Generality of the underlying idea |
| Registered replication | Multi-lab direct repeats, pre-planned | Precision of the effect estimate across settings |
What should a reader take from all this?
One study is a question, not an answer, and the word replicated should prompt a follow-up question: by whom, how closely, and was the repeat registered before its data existed? Findings that have survived independent, pre-registered repetition deserve real confidence. Findings that have never been repeated deserve patience. That gap, between the demonstrated and the merely published, is what reproducibility research exists to measure.
Different fields have inherited different baselines. Laboratory disciplines with standardized instruments tend to replicate more reliably than those measuring behavior or subjective outcomes, where context moves the measurement itself. The reader's rule travels well across all of them: weight a finding by the number of independent, well-powered attempts that have met it, and treat a lone striking result as a hypothesis priced for disappointment.
For more context, read What a meta-analysis actually combines.
For more context, read How does peer review actually work before publication?.
