Skip to content
About Contact
Engevity NewsScience & health

The replication crisis explained: when famous findings fail

Why landmark results sometimes collapse when other scientists repeat them, and what open science is doing about it.

The replication crisis explained: when famous findings fail
The replication crisis explained: when famous findings fail

The replication crisis is the ongoing discovery that some famous scientific findings fail when other scientists try to repeat them. A team reports an effect. A second team runs the same experiment and gets a much weaker result, or nothing at all. The problem showed up across psychology, medicine, and other fields, and it forced researchers to change how they publish and check their work.

The crisis does not mean science is broken. It means science is doing its job, slowly and awkwardly. A single study is one measurement, not a fact set in stone. As basic biology teaching notes put it, reliability is enhanced by increasing the number of measurements in an experiment or test — one copy of a result is never as trustworthy as many. Replication is simply that principle applied to whole studies.

What does replication actually mean?

Replication means running a study again to see whether the result holds. There are two broad kinds. A direct replication copies the original method as closely as possible: same kind of participants, same task, same analysis. A conceptual replication tests the same idea with a different method — a new population, a new measurement, a new design.

Both matter, but they answer different questions. A direct replication asks whether the original result was real. A conceptual replication asks whether the underlying idea survives a change of scenery. When a direct replication fails, the original claim is in trouble. When a conceptual replication fails, the idea may still be fine, but the evidence for it just got weaker.

Think of it like weighing yourself. One reading on one bathroom scale tells you something. A second scale that agrees makes the number believable. Ten scales that disagree tell you the first scale was the problem.

Why do published findings fail when repeated?

Several pressures push weak results into print, and none of them require anyone to cheat.

  • Small samples. A study with few participants is noisy. By chance, noise can look like an effect. Repeat the study with more people and the noise averages out.
  • Flexible analysis. Researchers make many small choices — which participants to include, which measure to highlight, when to stop collecting data. Some choices, made after seeing the data, can turn a null result into a publishable one. This is the territory our earlier piece on the p-hacking problem in statistical testing walks through.
  • Publication bias. Journals have historically preferred positive results. Studies that find nothing often sit in a drawer. That means the published record overrepresents findings that worked once.
  • Low prior odds. Some proposed effects are just unlikely to be real. When the starting odds are bad, even a statistically striking result is more likely to be a fluke than a discovery.

Put together, these pressures mean the file of published papers can look far stronger than the underlying evidence. Replication is the audit that reveals the gap.

How do we know famous findings actually failed?

Because researchers organized large, collaborative efforts to repeat whole batches of published experiments, and published the outcomes whatever they were. In psychology, multi-lab projects re-ran sets of well-known studies and reported that a substantial share of the effects came out much smaller than originally claimed, or did not appear at all. Similar checks in other fields found the same pattern: early striking results often shrink when tested again with bigger samples and locked-in analysis plans.

Two caveats keep this honest. First, a failed replication does not automatically prove the original was wrong. The repeat study could have its own flaws, or the effect could be real but smaller and context-dependent. Second, replication projects have their own limits — they cannot rerun every study, and some effects genuinely depend on the setting. What the projects established is a base rate: this happens often enough that no single unreplicated finding deserves full trust.

That is why careful readers ask for the same thing coaches ask of a new training method: not one impressive before-and-after, but whether it holds across people, gyms, and months. Consistency beats optimization.

What is open science doing about it?

The response has been structural, and it is gaining ground. The changes aim to make the process visible before results exist, so flexibility cannot quietly manufacture a finding.

  • Registered reports and preregistration. Researchers state their hypothesis, sample size, and analysis plan before collecting data. Journals can accept the design in advance, so a boring result no longer threatens publication.
  • Open data and open code. Sharing the raw data and analysis scripts lets other scientists check the work line by line, and rerun it.
  • Multi-lab collaborations. Dozens of labs run the same protocol at once. This produces large samples fast and tests whether an effect holds across settings.
  • Emphasis on effect size over significance. Reporting how big an effect is, with uncertainty, gives readers more usable information than a yes-or-no threshold.

Funders and journals have pushed in the same direction, and replication studies — once treated as dull — now get published in top journals. The culture shift is real, though incomplete. Incentives still reward novelty, and checking old work earns less credit than announcing new work. For related coverage, see How do scientists know how old a fossil is?.

What this means for how you read science news

Practical steps are simple. Treat one study as one data point, never a conclusion. Look for whether the finding has been repeated, ideally by a group independent of the original team. Notice the sample size and whether the analysis plan was fixed in advance. Be most skeptical when a striking result arrives from a small study with no replication — and least skeptical when large, independent teams keep landing on the same number.

Our analysis, reading the replication literature with a coach's eye: the crisis is less a scandal than a correction. The fields that audited themselves are producing sturdier results now, because the checks run before publication instead of after. The same discipline applies to reading any health claim you encounter, from a headline about a trial to a supplement ad. If you make health decisions based on research you , bring the questions in this piece to a qualified clinician rather than acting on a single .

Where replication fits in the bigger picture

Replication sits alongside the other self-correction tools this site covers: peer review, retraction, and good trial design. None of them works alone. Peer review filters obvious flaws but cannot verify results. Retraction removes papers that are wrong or fraudulent, usually long after publication. Randomization and blinding, as explained in our piece on how randomization strengthens clinical trial design, prevent bias inside a single study. Replication is the check that operates across studies — the one that catches everything the others miss. This connects to our earlier piece, Papers that get harsher peer review may end up more cited.

For readers who want to build this habit, our research literacy section collects these tools in one place, and the research desk follows new studies as they arrive. The evidence so far supports a modest, useful conclusion: science that expects to be checked produces claims worth trusting. What remains unknown is how quickly the incentives catch up with the ideals.

Sources

  1. DNA Replication - BioNinja

More from our brands

Part of the VUGA Network

Frequently Asked Questions

Does a failed replication prove the original study was wrong?
Not by itself. The replication could have its own flaws, or the effect might be real but smaller or limited to certain conditions. What a failed replication does establish is that the original result cannot be taken at face value. Confidence should scale with how consistently independent teams reproduce a finding.
Is the replication crisis only about psychology?
No. High-profile checks began in psychology, but weak replication has been documented in medicine, cancer biology, and other lab sciences. Any field that publishes small, flexible, single-lab studies is exposed. Fields with large preregistered trials, such as late-stage drug research, are less affected.
What is preregistration and why does it help?
Preregistration means documenting your hypothesis, sample size, and analysis plan in a public registry before collecting data. It prevents researchers from quietly trying many analyses and reporting only the one that worked. Registered reports go further by having journals accept a study based on its design, before the results exist.
How can a non-scientist tell if a finding is trustworthy?
Ask four questions: How big was the sample? Was the analysis plan fixed in advance? Has an independent team replicated it? Do the authors report effect sizes with uncertainty? A finding that survives those questions deserves more confidence than a striking result from a single small study.