Skip to content
Saturday, August 22, 2026
Engevity NewsScience & health
Research · Learning · Evidence
literacy

Peer review may catch less bad science than readers assume

A 2025 analysis in PNAS finds reviewers often disagree with each other — and shows which specific fixes actually help.

Peer review may catch less bad science than readers assume

Peer review is supposed to filter flawed research before it reaches print, but a 2025 analysis in the Proceedings of the National Academy of Sciences found that two reviewers looking at the same manuscript agree on its quality only weakly — evidence that the filter is looser than most readers assume, though the same body of research also shows a few specific fixes that measurably tighten it.

What is peer review actually checking for?

Before a study is published, editors send it to independent experts who are supposed to judge whether the work holds up. A 2025 methods guide published in the journal mBio, written to train early-career scientists to write their first reviews, lays out what that judgment is meant to cover: whether the experimental design and controls are adequate, whether the statistics support the stated conclusions, whether the study is honestly built on the existing literature, and whether the result is significant enough to justify publication.

The guide frames peer review as quality control rather than verification. Reviewers are not expected to repeat the experiments or audit the raw data line by line. They read the manuscript, the figures, and the methods, and flag what looks wrong or unsupported. That distinction matters for what comes next: a process built to catch sloppy reasoning is not the same thing as a process built to catch fabricated results.

The mBio guide also lays out how a review is typically structured once a reviewer has read the paper. It recommends working through the manuscript in a set order — abstract, figures, methods, results, then introduction and discussion — and organizing the eventual report into major points that would need to be fixed before the paper could run, moderate points that would strengthen it, and minor points that are technical corrections, along with confidential notes for the editor alone. The framework is meant for primary research articles specifically; it does not cover review articles, commentaries, or other formats that journals also send out for peer review.

How much do reviewers actually agree with each other?

Not much, according to the evidence. A 2025 paper in the Proceedings of the National Academy of Sciences, led by psychologist Balazs Aczel, pooled results from 45 earlier studies in which two or more reviewers independently rated the same manuscript. The average correlation between their ratings was 0.34 — a weak-to-moderate relationship on the scale researchers typically use, meaning reviewers evaluating identical work often reach different verdicts.

That low agreement does not by itself mean either reviewer is wrong. Reviewers bring different expertise, different tolerance for methodological gaps, and different views on what counts as "significant enough." But it does mean that whether a paper gets accepted, rejected, or sent back for revision can hinge heavily on which two or three people happened to be assigned to it.

What kinds of errors slip through anyway?

The same review cites evidence that reviewed, published papers still carry avoidable mistakes. Across studies of psychology journals it summarizes, roughly 18% of statistical results were found to be incorrectly reported, and inconsistencies in how p-values were reported affected the paper's stated conclusions in about one case in eight. Those are not exotic errors — they are the kind of arithmetic and reporting slips a careful check should catch before publication, and increasingly they surface only after publication, contributing to rising retraction rates.

The mBio guide points to one reason this happens: reviewing is a skill most scientists are never formally taught, and a checklist "cannot overcome deficiencies in graduate education" or make someone a statistics expert. A biologist asked to review a paper's methods may have no training to evaluate its statistical modeling, and journals do not always route papers to a reviewer who does.

How is grant review different from journal review?

The two are often confused, but they ask different questions. Journal peer review judges a finished manuscript. Grant peer review, as the National Institutes of Health describes its own process, judges a proposal for work that has not yet been done, scoring it on five criteria — significance, the investigators' qualifications, innovation, the technical approach, and the research environment — before a second panel weighs how well the proposal fits the funding agency's mission. NIH simplified this framework in January 2025 after reviewers reported that scoring complexity was pulling attention away from scientific merit.

Both systems rely on the same basic mechanism: unpaid experts making a judgment call about work they did not do themselves, under time pressure, without auditing the underlying data. That shared structure is part of why both are vulnerable to the same weaknesses — inconsistent standards between reviewers, and blind spots wherever a reviewer's expertise runs out. NIH's own description of its process frames the goal as reviews that are "fair, independent, expert, and timely" and "free from inappropriate influences" — a standard that names the same pressures peer review is generally trying to resist, whether the object under review is a manuscript or a funding proposal.

What reforms are being tested — and what actually helps?

The PNAS review also catalogs interventions that journals have tried, drawing on a systematic review that identified 24 randomized controlled trials testing changes to the review process. Adding a dedicated statistical reviewer to the process produced the largest measured improvement in manuscript quality among the interventions studied. Other reviewer-level changes — training, checklists, structured forms — produced smaller but still positive effects. Open review, in which reviewers' identities or comments are made public, was linked to markedly lower rejection rates, while double-blind review, which hides author identity from reviewers, showed only partial success at reducing gender bias in outcomes. None of these changes eliminated the underlying disagreement between reviewers; the trials measured smaller, individually modest gains, which is why the researchers describe the fixes as improvements to be layered together rather than a single solution.

InterventionWhat it changesEffect reported
Adding a statistical reviewerA dedicated statistics expert joins the review panelLargest quality improvement among interventions studied
Reviewer training or checklistsStructured guidance for what to checkPositive but smaller effect
Open reviewReviewer identity or comments made publicMarkedly lower rejection rates
Double-blind reviewAuthor identity hidden from reviewersPartial reduction in gender bias

How do we know this?

The central findings here come from a synthesis published in the Proceedings of the National Academy of Sciences in early 2025, which combined a meta-analysis of 45 reviewer-agreement studies with a separate systematic review of 24 randomized controlled trials testing peer-review interventions. Pooling many smaller studies this way gives more statistical weight than any single trial, but it inherits those studies' limits: most of the underlying agreement research comes from psychology and biomedical journals, so the 0.34 correlation figure may not describe every field equally, and the review's own authors note that far more randomized trials are still needed before firm conclusions can be drawn about which fixes work best. The paper does not claim peer review is broken beyond use — only that its filtering power is weaker and more unevenly applied than the phrase "peer-reviewed" tends to suggest to readers outside the process.

It also matters what these studies can and cannot show. A correlation between two reviewers' scores describes agreement, not accuracy — it cannot tell readers which of two disagreeing reviewers, if either, was right about a given manuscript. And a randomized trial that shows a statistical reviewer improves manuscript quality on average says nothing about how any single paper would have fared with a different reviewer assigned. Peer review, on this evidence, functions less like a fixed bar every study must clear and more like a probabilistic screen — one that catches some errors reliably, misses others depending on who happens to be reading, and is strongest exactly where journals have added the specific checks, like statistical review, that the trials tested directly.

Frequently asked questions

  • Does peer review catch fraud or fabricated data? Not reliably. As the mBio guide to reviewing explains, reviewers assess whether a study's design, statistics, and reasoning look sound on the page. They are not expected to audit raw data or repeat experiments, so deliberately fabricated results can pass through the process undetected unless something in the write-up itself looks inconsistent.
  • Why do two reviewers on the same paper sometimes disagree completely? A 2025 PNAS analysis pooling 45 earlier studies found the average agreement between reviewers rating identical manuscripts was weak — a correlation of 0.34. The gap reflects differences in reviewers' expertise, their tolerance for methodological gaps, and their judgment about what counts as significant enough to publish.
  • Is grant peer review the same as journal peer review? No. Grant review, as the National Institutes of Health describes its own process, scores funding proposals for future work on criteria like significance, innovation, and approach. Journal peer review evaluates a completed manuscript instead. Both rely on unpaid expert judgment, which is where their shared weaknesses come from.
  • What single change has shown the clearest benefit? Among interventions reviewed in the 2025 PNAS paper, which drew on 24 randomized controlled trials, adding a dedicated statistical reviewer to the review panel produced the largest measured improvement in manuscript quality — larger than training, checklists, or blinding changes tested on their own.

For a related science perspective, read Papers that get harsher peer review may end up more cited.

Sources

  1. Proctor, Abraham & Righi, "From novice to expert: preparing your peer review," mBio (2025)
  2. Aczel et al., "The present and future of peer review: Ideas, interventions, and evidence," PNAS (2025)
  3. NIH Office of Extramural Research, "Background — NIH Peer Review Process"