Skip to content

71 min read

The Replication Crisis in 2026: How Much Science Holds Up

How much published science holds up, from psychology to cancer biology to physics, why results fail, and how to judge whether a new breakthrough will last.

By The Frontier Desk. Scores quoted are live Frontier Scores; see how they are computed.

What the largest replication projects found, why so many findings fade, and how to judge a new result before anyone has repeated it.

When the SCORE program, a collaboration of 865 researchers, retested claims from 164 published social-science papers against independent data, only 49 percent replicated, and the typical effect came back at less than half its original size.

Those results, published in Nature in April 2026 by the SCORE program, are the newest and largest entry in a ledger that has been growing for fifteen years. Psychology's landmark replication project in 2015, a drug company's attempt to confirm 53 landmark cancer papers, an eight-year effort to repeat experiments from high-impact cancer-biology papers, and a 56-laboratory study of Brazilian biomedical science all point the same way: a large share of published findings do not hold up when someone else tries them.

That is a problem for every reader of new science, because almost every breakthrough you will ever see in a headline is a first report that nobody has repeated yet. The replication crisis is not a story about a few frauds or one troubled field. It is a measurable property of how research is produced, and it varies predictably by field, by study design and by how a result is framed. Once you understand where the failures come from, you can read a new result the way an experienced referee does: with interest, with the right questions, and with a sense of how much to believe it before the evidence matures.

This guide covers what the crisis is, the hard numbers from every major replication project through 2026, the arithmetic that explains why so many findings fail, the integrity problems of fraud and paper mills, how the bar differs between physics, medicine and AI, which reforms work, and a practical method for judging new results. It also explains how Frontier labels the confidence of the fresh results in its index, with live data from that index.

Contents

  1. What the replication crisis is, and what it is not
  2. The scorecard: what the big replication projects found
  3. Laboratory biology: where failure costs the most
  4. Why results fail to replicate: the arithmetic underneath
  5. Fraud, paper mills and retractions: the integrity layer
  6. Why the bar differs by field
  7. What has changed since 2015, and whether it works
  8. How to tell whether a new breakthrough will hold up
  9. How Frontier reads a new result, and what its labels mean
  10. What comes next: AI as a source of errors, and as a checker
  11. Conclusion: how to hold a new result

1. What the replication crisis is, and what it is not

The replication crisis is the finding, built up through the 2010s and still accumulating, that a substantial fraction of published scientific results cannot be reproduced when other researchers repeat the work. The phrase usually covers two related failures that are worth keeping apart, because they have different causes and different fixes. The US National Academies of Sciences, Engineering, and Medicine set out the standard definitions in a 2019 report: reproducibility "is obtaining consistent results using the same input data; computational steps, methods, and code; and conditions of analysis", while replicability "is obtaining consistent results across studies aimed at answering the same scientific question, each of which has obtained its own data" - National Academies.

The distinction matters in practice. A result that fails to reproduce means the published numbers cannot be regenerated from the authors' own data and code, which points to errors, missing materials or an undisclosed analysis step. A result that fails to replicate means a fresh experiment did not find the same thing, which can mean the original was a false positive, that its effect was smaller than reported, or that it depended on conditions nobody recorded. Most headline discussion concerns replication, because that is the failure that decides whether a discovery is real.

How the crisis came to light

Scientists have always known that some results fail, so it is fair to ask what changed. What changed was measurement. In 2012, a commentary in Nature reported that scientists in the haematology and oncology department of the biotech firm Amgen had tried to confirm the findings of 53 "landmark" preclinical cancer papers and succeeded in only 6 of them, 11 percent - Nature. A year earlier, Bayer had halted nearly two-thirds of its target-validation projects because its in-house experimental findings did not match published claims - Nature Reviews Drug Discovery. Those were industry checks with unpublished details, and academic science answered with open, systematic ones.

The turning point was the 2015 Open Science Collaboration, in which teams repeated 100 psychology studies; the results, and every project that followed, are in the next section. Surveys then showed how widely researchers had experienced the problem first-hand. In a 2016 Nature survey of 1,576 researchers, more than 70 percent said they had tried and failed to reproduce another scientist's experiments, more than half had failed to reproduce their own, and 52 percent called the situation a significant crisis - STAT. A 2024 survey of 1,630 biomedical researchers found that 72 percent agreed there is a reproducibility crisis - PLOS Biology.

What the crisis is not

Three misreadings are common enough to address before the numbers. The first is that a failed replication proves the original was wrong. It does not: a single replication can itself be underpowered, executed differently, or run in a population where the effect is genuinely different. That is why the strongest replication projects use samples many times larger than the originals and preregister their methods with the original authors' input, and why their results are reported as rates across many studies rather than verdicts on any one.

The second misreading is that the crisis means science is broken or that published findings are worthless. The same projects that found failure rates of 40 to 60 percent also show that about half of findings do hold up, that the stronger the original evidence the better the odds, and that the failures cluster in predictable places. Science is the only human institution that systematically audits its own output and publishes the results; the replication crisis is that audit working.

The third misreading is that failure means fraud. Most non-replicating results come from honest researchers using small samples, flexible analyses and a publication system that rewards surprising positive findings, a mechanism section 4 lays out in plain arithmetic. Fraud is real, growing in some forms, and covered in section 5, but it is the smaller part of the story. For a reader, the practical consequence is that you cannot judge a result by the reputation of the people behind it; you have to judge the evidence.

2. The scorecard: what the big replication projects found

The strongest evidence about how much science holds up comes from large, coordinated replication projects: teams that select a defined set of published findings, repeat them with larger samples and preregistered methods, and report every outcome, success or failure. They are expensive, slow, and concentrated in the fields where replication is cheapest (psychology, economics and the social sciences) and where failure is most costly (preclinical biology). Their combined record is the closest thing science has to a measured answer.

Two caveats travel with every number below. First, success depends on the criterion: in the 2015 psychology project, the share of studies counted as replicated was 36 percent by statistical significance, 39 percent by the replication teams' subjective judgment, 47 percent when asking whether the original effect fell inside the replication's confidence interval, and 68 percent when the original and replication data were combined - Tilburg University. Second, each project sampled a particular slice of a field, so none is a field-wide estimate. With those caveats, the pattern is consistent enough to be informative.

ProjectFieldTestedReplicatedEffect in replications
Amgen check (2012)Preclinical cancer53 landmark papers11%Not reported
Open Science Collaboration (2015)Psychology100 studies36%About half the original
Camerer et al. (2016)Experimental economics18 experiments61%66% of the original
Camerer et al. (2018)Social science in Nature and Science21 experiments62%About 50% of the original
Many Labs 2 (2018)Psychology28 findings54%Median d fell from 0.60 to 0.15
Cancer Biology project (2021)Preclinical cancer50 experiments, 23 papers (112 effects)46% of effects (40% of positive ones)Median 85% smaller
Brazilian initiative (2025-26)Biomedical lab methods45 experiments20% to 44%Group-mean ratios 58% lower
SCORE (2026)Social and behavioral sciences164 papers, 274 claims49% of papersr fell from 0.25 to 0.10

The sources for each row are linked where the project is discussed below and in section 3. Read across the rows, three regularities stand out. Success rates cluster between roughly a third and two thirds; the effects that do replicate almost always come back smaller; and even large, preregistered replications find far fewer significant results than the 97 percent the original psychology papers reported, which is itself a lesson about how much the standard literature overstates.

Psychology and the social sciences

The 2015 Open Science Collaboration set the template. Its teams repeated 100 studies from three leading psychology journals: 97 percent of the originals had reported significant results, against 36 percent of the replications, and the average replication effect was half the size of the original - Tilburg University. Economics fared somewhat better in a 2016 project that repeated 18 laboratory experiments from two top journals: 11 replicated, 61 percent, with effects averaging 66 percent of the original - MPRA. When the same consortium turned to 21 social-science experiments published in Nature and Science between 2010 and 2015, 13 replicated (62 percent), with effects about half the original size - Nature Human Behaviour. Prestige of the venue offered little protection.

Many Labs 2 tested 28 classic and contemporary psychology findings across 125 samples in 36 countries and territories, so that each finding was tried in many places at once. Fifteen (54 percent) produced a significant effect in the original direction, and the median effect size fell from 0.60 in the originals to 0.15 in the replications - Tilburg University. Some famous effects fared worse still. A 23-lab preregistered test of ego depletion, the idea that self-control draws on a limited resource, found an effect of 0.04 standard deviations, statistically indistinguishable from zero - Utrecht University. Seventeen direct replications of a well-known facial-feedback experiment found a rating difference of 0.03 units against 0.82 in the original - Perspectives on Psychological Science.

SCORE: the largest audit yet

The SCORE program (Systematizing Confidence in Open Research and Evidence), funded by the US defense research agency DARPA and coordinated by the Center for Open Science, sampled claims from 3,900 papers published between 2009 and 2018 in 62 journals across a dozen social and behavioral sciences, and involved 865 researchers - Center for Open Science. Its replication paper reports that independent data (newly collected for 98 papers, existing datasets the original authors had not used for the other 66) produced a significant result in the original pattern for 151 of 274 claims (55.1 percent) and for 49.3 percent of papers, with discipline rates from 42.5 to 63.1 percent and a median effect that fell from r = 0.25 to 0.10 - Nature.

Two companion papers measured the other kinds of reliability. When the reproducibility team tried to regenerate results from the authors' own data, data were available for only 24 percent of 600 papers, and among the papers that could be tested, 53.6 percent reproduced precisely and 73.5 percent at least approximately - Nature. When the robustness team gave the same data and question to at least five independent analysts per study, only 34 percent of the re-analyses landed on the same result within a narrow tolerance, although 74 percent reached the same overall conclusion - Nature. That last finding is a quiet bombshell: even with identical data, independent re-analyses fail to match the original result between a quarter and two thirds of the time, depending on how strictly agreement is defined.

Economics and political science: a more hopeful 2026 result

Not every 2026 result was discouraging. The Institute for Replication, publishing in Nature the same day as SCORE, examined 110 economics and political-science articles from 2022 and 2023 in journals that require authors to share data and code. More than 85 percent of the published claims were computationally reproducible, and 72 percent of statistically significant estimates stayed significant and in the same direction under robustness checks - Nature. The authors themselves call their sample very selective and possibly an optimistic upper bound, and about a quarter of the articles contained non-trivial coding errors, but the contrast with SCORE's 24 percent data availability makes the case for mandatory sharing better than any argument could.

The scorecard tells a reader something precise. A published finding in these fields is, as a rough prior, something like a coin flip to replicate at all, and likely to come back weaker if it does. That prior should be adjusted for everything sections 4 and 8 describe, from sample size to preregistration, but it is the honest starting point. The fields with the most expensive experiments have been audited least, and the next section turns to the most consequential of them.

3. Laboratory biology: where failure costs the most

Preclinical biology, the cell, tissue and animal research that comes before any human trial, deserves its own section because it is where non-replication is most expensive and hardest to detect. A false finding in social psychology wastes the time of the researchers who build on it. A false finding in cancer biology can launch a drug program that consumes years and hundreds of millions of dollars before a clinical trial exposes it, and patients enrol in the trials. The Amgen and Bayer checks described in section 1 were industry's way of saying exactly that: companies had learned not to trust a published target until their own scientists had confirmed it.

The field is also structurally hard to replicate, for reasons that have nothing to do with honesty. Experiments depend on cell lines that drift over generations, antibodies that vary from batch to batch, animals whose microbiomes and housing differ between facilities, and dozens of small protocol choices that never make it into a methods section. Much of the knowledge needed to repeat an experiment is tacit, held in the hands of the people who did it. That is why the most informative evidence about preclinical biology comes from projects that tried to repeat experiments and recorded every obstacle along the way.

The Reproducibility Project: Cancer Biology

The Reproducibility Project: Cancer Biology, an eight-year collaboration between the Center for Open Science and Science Exchange, set out to repeat 193 experiments from 53 high-impact cancer biology papers published between 2010 and 2012 - Center for Open Science. It completed only 50 experiments from 23 papers, reported in eLife in December 2021. Much of the barrier was access rather than the science itself: the data needed to calculate effect sizes was publicly available for just 4 of the 193 experiments, could not be obtained at all for 68 percent, and the original authors were rated not helpful or unresponsive for 32 percent of experiments; once laboratory work began, 67 percent of protocols needed modifications, and only 41 percent of those could be implemented - eLife. The figure below, from the project's own analysis, shows how the pipeline narrowed at each stage.

The results for the experiments that could be done were sobering. For positive effects, the median effect size in the replications was 85 percent smaller than in the originals, and 92 percent of replication effects were smaller than the original effect; counting both positive and null results, 46 percent of effects (51 of 112) succeeded on a majority of the five criteria the team used - eLife. In the scatter plot from that paper, nearly every point falls below the diagonal line that marks an identical effect, and many sit near zero.

A failed replication of one experiment does not, on its own, show that a paper's conclusions are wrong, and 50 experiments from 23 papers are not a census of cancer biology. What the project did establish is that the published record, in its most prestigious corner, does not contain enough information for someone else to repeat a typical experiment without the original authors' help, and that when the help is given, the effects usually shrink. For a reader, the practical lesson is that a striking result in cells or mice is a hypothesis about humans, and a strong first report is the beginning of its testing rather than the end.

Brazil, fruit flies and materials chemistry

Three newer projects extend the picture beyond cancer. The Brazilian Reproducibility Initiative took the opposite approach to cherry-picking famous papers: it sampled experiments from Brazilian biomedical research using three common methods (a cell-viability assay, a gene-expression measurement and a rodent anxiety test) and had 56 laboratories repeat them. Of 90 valid replications covering 45 experiments, the replication rate ranged from 20 to 44 percent depending on which of five predefined criteria was used, and, in median terms, the ratios between group means were 58 percent lower in the replications - bioRxiv. The work became an eLife Reviewed Preprint in July 2026.

The ReproSci project asked a question no prospective replication can: what happened, over decades, to every claim in a single research community? Its authors traced 1,006 claims from 400 papers on fruit-fly immunity published between 1959 and 2011 and found that 61 percent had been verified, 7 percent challenged and 24 percent never independently tested. Those figures include the project's own laboratory tests: of 45 never-checked claims it re-ran, chosen mostly because they looked suspicious (31) or were easy to test (14), 38 failed, including 7 of the 14 easy ones. The authors concluded that elite affiliation, more than journal prestige, best predicted which claims failed - bioRxiv. The unchecked quarter of a literature, in other words, may be where much of its error quietly lives.

Materials chemistry shows that the problem is not confined to biology, though it takes a different form. A 2026 study of 130 metal-organic frameworks, a class of porous materials central to carbon capture and gas storage research, found that 11 to 17 years after each material was first reported, 83 percent still had no reported replicate synthesis - J. Phys. Chem. C, via PMC. Nobody had shown that most of them could not be made again; nobody had shown that they could. That silence is the default state of much of the published record, and it is why the absence of a failed replication should never be read as a success.

These projects change how a careful reader should approach a preclinical breakthrough. The first question is not "is the journal prestigious?" (ReproSci found that an elite affiliation can even point the wrong way) but "has anyone else done it?" The second is whether the result has been confirmed with more than one method, since an effect that shows up in a cell assay, an animal model and a genetic knockout is far harder to produce by accident than one measured a single way. The next section explains, from first principles, why even honest, careful laboratories produce so many findings that fail these tests.

4. Why results fail to replicate: the arithmetic underneath

The most useful thing a reader can learn about the replication crisis is that most of it requires no villain. Fraud exists and section 5 deals with it, but the bulk of non-replicating results come from honest researchers following the incentives and statistical conventions of their field. To see why, it helps to set the scandals aside and ask a structural question: if a field tests many ideas, most of which are wrong, with studies that are too small, and publishes mainly the tests that come out "significant", what fraction of its published findings should we expect to be true?

The question has a precise answer, and it was put most forcefully by John Ioannidis in a 2005 essay whose title gave the whole debate its edge, "Why Most Published Research Findings Are False" - PLOS Medicine. The argument rests on three quantities that every study has whether or not anyone reports them: the prior probability that the hypothesis being tested is true, the study's statistical power (the chance it detects a real effect of the expected size), and the significance threshold, conventionally 5 percent. Together they determine the share of "significant" results that reflect a real effect, a quantity called the positive predictive value.

A worked example in plain numbers

Imagine a field that tests 1,000 hypotheses, of which 100 are actually true. With studies powered at 80 percent, the true ones produce 80 significant results. Of the 900 false hypotheses, 5 percent cross the significance line by chance, which is 45 false positives. So 125 results get published as discoveries, and 80 of them (64 percent) are real. Now cut the power to 20 percent, which is closer to what audits of some fields have found: the true hypotheses now produce only 20 significant results, the false ones still produce 45, and just 31 percent of the published discoveries are real. Nothing about the researchers' honesty changed between those two scenarios. Only the size of the studies did.

Statistical power is not a hypothetical worry. A widely cited 2013 analysis estimated the median power of neuroscience studies at between about 8 and 31 percent - Nature Reviews Neuroscience, and a 2017 survey of more than 3,800 cognitive neuroscience and psychology papers found a median power of just 12 percent to detect a small effect, with "no improvement through the past half-century" - PLOS Biology. A study with that little power usually misses a real effect, and when it does report one, the arithmetic says to be wary. The chart below runs the same arithmetic across a range of powers and three kinds of hypothesis: a confirmatory test where the idea is as likely true as not, an exploratory field where one idea in ten is true, and a long-shot discovery search where one in fifty is.

The bottom line of the chart is the one that matters for a breakthrough index. A genuinely surprising claim is, almost by definition, a long-shot hypothesis, and even a well-powered study of a long-shot idea produces a significant result that is more likely false than true. That is not an argument against surprising science. It is the reason a surprising result needs a replication more than a modest one does, and the reason the first report of an extraordinary effect should be read as the opening of a question rather than its answer.

Four ways honest work produces false findings

The arithmetic above assumes that every study runs one pre-planned test and reports it faithfully. Real research has far more freedom than that, and each degree of freedom pushes the effective false-positive rate above the nominal 5 percent. None of the mechanisms below requires bad intent, which is precisely why they are so widespread and so hard to see from the outside.

They compound. A small study that is analysed flexibly, published only because it worked, and written up as if the finding had been predicted all along can look identical on the page to a rigorous confirmation, and a reader has almost no way to tell the two apart from the paper alone. The final paper reports one analysis and one hypothesis, however many were tried along the way.

  • Flexible analysis: choosing among outcomes, covariates and exclusions after seeing the data (often called p-hacking).
  • Publication bias: significant results get written up and accepted; null results stay in the file drawer.
  • The winner's curse: the studies that cross the line by luck overestimate the true effect, so first reports run large.
  • Hindsight framing: presenting an unexpected finding as the hypothesis the study set out to test.

The winner's curse deserves special attention because it explains a pattern that every large replication project that measured effect sizes found: even when an effect survives, it usually comes back smaller. If only the studies that happened to land on a large estimate get published, then the published estimate is biased upward by selection, and an unbiased repeat of the same experiment will regress toward the real, smaller value. The shrinking effect sizes in section 2 are therefore not just a sign of failure; they are the signature of a literature filtered for significance. Medicine shows the same pattern at scale: in an analysis of 85,002 forest plots from medical meta-analyses, 90 percent of the very large effects seen in a first trial became smaller once later trials were pooled - JAMA. For a reader, the practical consequence is simple: treat the effect size in a first report as an upper bound.

These mechanisms have been explained for a general audience many times, and one explainer stands out for its reach: Veritasium's 2016 video, watched more than six million times, builds on the same Ioannidis argument. Its figures predate the replication projects of the 2020s, but the mechanics it describes are unchanged.

Incentives keep the machine running

The deeper reason these mechanisms persist is that the system rewards them. Hiring, promotion and grant decisions lean on publications in selective journals, and selective journals favour results that are new, clean and positive. A careful null result or a replication earns less credit than a striking first report, even though it is often more informative. The citation record, which many evaluations treat as a proxy for quality, makes the problem worse. A 2021 analysis of 80 papers from the large replication projects found that papers whose findings failed to replicate had been cited about 153 more times on average than papers that replicated, and the gap did not close once the failure was published: only 12 percent of later citations of a non-replicating paper acknowledged the failed replication - Science Advances, authors' copy. A plausible reading is that the features that make a finding exciting, surprise above all, are the same features that make it fragile.

The dashed arrow at the bottom of the diagram is the part most readers never see. A failed replication rarely removes the original claim from circulation: the first paper is not retracted, because nothing was wrong with it procedurally, and it keeps collecting citations from authors who never read the replication. The literature does not forget its false positives so much as bury them under newer ones. The next section turns to the smaller but more damaging category of results that were never honest in the first place, and to the machinery that has grown up to remove them.

5. Fraud, paper mills and retractions: the integrity layer

Honest error explains most failed replications, but not all of them, and the dishonest share has changed character in the past decade. Classic misconduct, a researcher fabricating or manipulating data to support a claim they believe in or need, still happens at the highest levels of science. Alongside it has grown something closer to an industry: paper mills, businesses that sell fabricated or plagiarised manuscripts and authorship slots on them, mostly to researchers under pressure to publish. The two problems need different defenses, and a reader needs to know how each one surfaces in the record.

The formal tool for removing a bad paper is the retraction, a notice from the journal that the work should no longer be relied upon. Retractions are useful signals, but they are rare, slow and uneven, and most flawed papers never receive one. Understanding how the integrity layer works, and how much it misses, is part of reading science well, because the absence of a retraction notice tells you very little on its own.

Retractions are rising, and they are slow

More than 10,000 research papers were retracted in 2023, a record, and more than 8,000 of them came from journals of a single publisher, Hindawi, where paper-mill activity had been found on a large scale - Nature. Its owner, Wiley, went on to close 19 former Hindawi journals in 2024 after retracting more than 11,300 compromised studies over two years - Inside Higher Ed. By early 2025, about 0.2 percent of articles published in 2022 had been retracted, a rate triple that of a decade earlier - Nature, and the Retraction Watch Database now lists more than 67,000 retractions - Retraction Watch.

Speed is the bigger problem. Across 16,041 medical papers retracted between 1975 and 2024, the median time from publication to retraction was 562 days - Journal of Korean Medical Science, and the most consequential cases take far longer. During that window a paper is cited, taught and built upon as if it were sound. The Nature paper on the amyloid-beta star 56 protein, a pillar of one line of Alzheimer's research, gathered about 2,360 citations before it was retracted in 2024 for image manipulation, 18 years after publication - Retraction Watch.

Paper mills: fraud at industrial scale

The scale of the paper-mill problem was measured most systematically in a 2025 PNAS study of the networks behind suspected mill products. Its central finding is a race the integrity system is losing: suspected paper-mill output has been doubling every 1.5 years, while retractions double every 3.3 years and the total number of publications every 15 - PNAS, authors' copy. Only 8,589 of the 29,956 suspected mill papers the authors could match to bibliographic records had been retracted, 28.7 percent, and they estimate that only about a quarter will ever be.

Publishers have responded with screening at submission. By September 2025, more than 35 publishers were using the STM Integrity Hub to screen over 125,000 papers a month, intercepting around 1,000 suspected paper-mill submissions every month - STM. Screening the published record is more sobering: a machine-learning screen of 2.6 million cancer papers published between 1999 and 2024 flagged nearly ten percent, 261,245 publications, as showing textual signs of mill origin - The Scientist. A text flag is a reason for scrutiny, not proof of fraud, but the order of magnitude is the point.

The market behind those papers is now visible in its own price lists. A 2026 dataset assembled 18,710 advertisements from seven paper-mill businesses operating out of seven countries, with more than 51,000 timestamped prices - arXiv. The chart from that study shows what an authorship slot on a forthcoming paper costs, by position in the author list.

For a reader, the mill problem matters less as a source of famous false breakthroughs (mill papers are designed to be unremarkable and pass unnoticed) than as noise in the literature that searches, reviews and AI systems ingest. A 2026 video investigation, in which the presenter sets out to buy a scientific paper from a mill, shows how routine the transaction has become; it runs about 24 minutes.

The practical defenses are modest but real. A claim that appears only in a journal with a history of mass retractions, from an author group with no other track record, deserves more skepticism than the same claim from an established lab in a journal that screens submissions. And a result that matters to you is worth one search of the Retraction Watch database before you rely on it.

Breakthroughs that did not survive

The cases that reach the news are a different population: high-profile claims in prestigious venues, where the stakes and the scrutiny are both high. They are worth studying because they show the full range of ways a breakthrough can fail, from contamination to a calculation error to fabrication, and how long the record can take to catch up. The table collects some of the most instructive recent examples, each with the reason the record changed.

ClaimFirst publishedWhat happenedLag
"Arsenic life" bacteriumScience, 2010Retracted 2025; Science cited evidence the results were based on contaminationAbout 15 years
Amyloid-beta star 56 in Alzheimer'sNature, 2006Retracted 2024 for image manipulation18 years
Majorana particles for quantum computingNature, 2018Retracted 2021 after reanalysis; authors apologized for "insufficient scientific rigour"3 years
Room-temperature superconductor (Dias)Nature, 2020 and 2023Both retracted; a university probe found data fabricationAbout 2 years (2020 paper); 8 months (2023 paper)
LK-99 superconductorPreprints, July 2023Replications showed it is not a superconductorWeeks
AI and materials-discovery productivityPreprint, 2024MIT said it had "no confidence" in the data; withdrawn 2025Months
Climate-damage costs to 2049Nature, 2024Retracted 2025; corrections too substantialAbout 20 months

Each row has a source. Science's statement on the arsenic paper, that the key conclusion "is based on flawed data" given evidence of contamination, is reported by Scientific American; the amyloid retraction notice describes images showing "signs of excessive manipulation" - Nature; the Majorana retraction apologizes "for insufficient scientific rigour" - Nature. The University of Rochester's investigation found that Ranga Dias "committed data fabrication, falsification and plagiarism" - Nature, and replication work on LK-99 unearthed "evidence that the material is not a superconductor" within weeks - Nature. MIT's statement on the productivity preprint is on its economics department site, and the climate-damage retraction is documented by Retraction Watch.

Two patterns stand out. The fastest corrections came where many independent groups could test the claim cheaply and quickly, as with LK-99, whose ingredients and recipe were public. The slowest came where the record lagged behind the science: the arsenic claim was contradicted by two studies Science published in 2012 but stayed unretracted until the journal changed its retraction standards, and the amyloid paper was pulled for image manipulation rather than for failing a replication. And prestige did not protect: five of the seven appeared in Nature or Science. The lesson for a reader is not to distrust those journals, which also publish most of the claims that do hold up, but to treat any single paper, wherever it appears, as one piece of evidence rather than a settled fact.

Fabricated references: a new kind of noise

A newer integrity problem comes from generative AI. Text produced or polished by language models is now widespread in the literature, which is not misconduct in itself; an analysis of excess vocabulary estimates that at least 13.5 percent of 2024 PubMed abstracts were processed with language models, a conservative lower bound - Science Advances, via PMC. The problem is fabrication. An audit of biomedical papers found that one in 2,828 contained at least one fabricated reference in 2023, and one in 458 by 2025 - STAT. A reference to a paper that does not exist is a claim resting on nothing, and section 10 looks at how the same technology is being turned against the problem.

For a reader, this is the easiest integrity check of all, and it is worth doing whenever a claim matters to you. Follow the citation that the claim rests on and confirm that the cited paper exists and says what it is cited for; a DOI link, PubMed or a library search makes that a minute's work. A paper whose central citation leads nowhere has told you something important about the care that went into it, whatever the rest of it says.

6. Why the bar differs by field

"Science" does not replicate at one rate, because fields differ in the three things that govern reliability: how cheap it is to repeat a measurement, how strict the conventional bar for a discovery is, and how much of the work is shared openly enough for others to check. A reader who carries one mental model from psychology into particle physics, or from clinical trials into machine learning, will misjudge results in both directions. This section walks through the main families of evidence that appear in Frontier's index and what "holding up" means in each.

The common thread is that every mature field has evolved its own defense against false positives, and the strength of that defense is visible from the outside. Asking which defense a field uses, and whether a given result went through it, is often more informative than any detail of the result itself.

Physics and astronomy: high bars, instrument errors

Particle physics is the field that most visibly learned the lesson of section 4 early. Its conventional threshold for claiming a discovery is five sigma, a result that chance alone would produce about once in 3.5 million tries - Scientific American, far stricter than the 5 percent used in much of biology and social science. Big experiments are often built in pairs (the two general-purpose detectors ATLAS and CMS at the Large Hadron Collider, the two LIGO observatories), so a claim can be checked by an independent instrument rather than a re-run of the same one. Those habits make outright false discoveries rare, but they do not make physics immune, and its famous failures have a characteristic shape: the statistics were fine and the instrument or the background was not.

Two cases from the 2010s show the shape. In 2011 the OPERA experiment reported neutrinos arriving about 60 nanoseconds sooner than light could have - arXiv. The next year, four experiments at the Gran Sasso laboratory measured a time of flight consistent with the speed of light, and CERN attributed the original result to "a faulty element of the experiment's fibre optic timing system" - CERN. In 2014 the BICEP2 telescope reported the imprint of primordial gravitational waves on the cosmic microwave background, until a joint analysis concluded that "most of the original BICEP2/Keck B-mode signal, but not necessarily all of it, could be explained by dust in our Milky Way" - NASA JPL. Neither team was careless with statistics; the error lived in the hardware and in the sky.

A newer failure mode is the gap between a paper and its publicity. In February 2025 Microsoft announced a quantum chip built on topological qubits, while Nature's editorial team wrote, in a peer-review file published alongside the accompanying paper, that "the results in this manuscript do not represent evidence for the presence of Majorana zero modes" in the devices, the property the qubit claim depends on. "The peer-reviewed publication is quite clear [that it contains] no proof for topological qubits," one physicist told Physics World, "but the press release speaks differently" - Physics World. For a reader, the rule is to trust the paper over the announcement whenever they disagree.

The muon g-2 story is the subtler, more recent lesson. For two decades, measurements of the muon's magnetic moment appeared to disagree with the Standard Model, a gap many hoped was a sign of new particles. Fermilab's final measurement in 2025 pinned the experimental value down to 127 parts per billion - Muon g-2 Collaboration, and the theory side moved instead. The 2025 White Paper of the Muon g-2 Theory Initiative adopted lattice calculations for the hardest term after the data-driven estimates fell into conflict with one another, and its new Standard Model prediction differs from the experimental average by 38 ± 63 (in units of 10⁻¹¹): in its authors' words, "no tension between the SM and experiment at the current level of precision" - Muon g-2 Theory Initiative. The anomaly did not fail to replicate. The measurement replicated beautifully; what changed was the calculation it was being compared with. Frontier's entry on the final measurement records it as a closing anomaly rather than a discovery, which is the honest reading.

For a single astronomical event, such as the record-mass black hole merger GW231123, replication in the laboratory sense is impossible: the signal arrived once. What substitutes for it is independent analysis of the same data, consistency across detectors, and, eventually, more events of the same kind. That is why Frontier's editorial for that entry tells readers to treat the masses as best estimates and to watch whether future detections populate the same range.

Clinical medicine: the most formal defenses

Medicine has the most developed machinery against false positives, because the cost of a wrong answer is measured in patients. Trials of new treatments must be registered before they start, with their primary outcome declared in advance; regulators require replication across phases; and large randomized trials are the standard of evidence before approval. The machinery works, and it is also a reminder of how often promising early results fade. Across drug-development programs tracked from 2011 to 2020, only 28.9 percent of those in phase 2 advanced to phase 3, 57.8 percent of phase 3 programs reached a regulatory filing, and the overall likelihood of approval for a program entering phase 1 was 7.9 percent - BIO. Not every failure is a failed replication (programs are also dropped for commercial reasons), but the attrition is the clearest measure in all of science of how much early promise survives a larger, stricter test.

That attrition is the right frame for any early-phase result in Frontier's index. A phase 1 base-editing study of 35 people can show that an edit is tolerated and lowers cholesterol, but it cannot show that the treatment prevents heart attacks; that takes a phase 3 trial years later. Our ranking of the latest medical breakthroughs scores 27 such results with that distinction built in, and its table shows how far apart the phases sit on evidence.

Laboratory biology: the weakest link

Preclinical biology, the cell, tissue and animal work that precedes clinical trials, is where the replication problem has been documented most painfully, for reasons section 3 laid out: experiments are hard to describe completely, reagents and animals vary, and there is no registry or phase system to catch exploratory results before they are published as findings. Method papers are a partial exception. A new imaging tool, like the microscope method that films nearly every cell in a living zebrafish, replicates by being adopted: if other labs can use it and get sensible data, it works, and if they cannot, it quietly disappears.

For a reader, the practical consequence is to treat a preclinical claim as a lead until someone else has repeated it. The numbers in sections 1 to 3 are the reason: in the projects that tried, the share of preclinical findings that held up ranged from 11 percent in Amgen's check to 46 percent in the cancer-biology project. Give more weight to findings that have been confirmed by more than one method, or that have already become tools other laboratories rely on, because both are forms of replication that happen without anyone running a formal replication study.

AI, computing and engineering: replication by running it

Computational results should be the easiest to reproduce, since a computer does the same thing every time it is given the same code and data. In practice the code and data are often missing, and a subtler failure is common: data leakage, where a model is evaluated on information it should never have seen. A survey of machine-learning-based science found leakage in 17 fields and 294 papers, in some cases producing wildly overoptimistic conclusions - Patterns, via PMC. That is why section 10 treats automated reproduction as one of the most important developments of 2026. Mathematical and algorithmic claims are the cleanest case of all. When AlphaEvolve reported a way to multiply 4x4 complex matrices with 48 multiplications, the result could be checked by anyone with the algorithm, so verification replaced replication.

Engineering demonstrations sit at the far end of the spectrum. A robot sailboat that crossed the Atlantic or a drone that flies in darkness by touch either did the thing or did not, and the evidence is the demonstration itself. The open question for a demonstration is not whether it happened but whether it generalizes: whether it works again, in other conditions, built by someone else. Likewise, a method-driven discovery such as the deep-learning scan of seismic waves that found structures near Earth's core is most convincingly confirmed when a different method, applied to different data, finds the same structures.

Field familyConventional barTypical failureAsk first
Particle physicsFive sigma, paired detectorsInstrument or background errorWas it confirmed by an independent instrument?
AstronomyConsistency across detectorsSingle event, model dependenceAre there more events like it?
Clinical trialsRegistered, randomized, phasedEarly-phase effects fadeWhich phase, and was the outcome primary?
Lab biologySignificance at 5%Small, unregistered, hard to repeatHas another lab reproduced it?
AI and computingBenchmarks, code releaseMissing code, leaked benchmarksCan anyone run it and get the same number?

The table is a starting point, not a hierarchy of worth. Some of the most important results of the decade came from fields with weak formal defenses, and some false alarms came from fields with strong ones. What it gives a reader is the right first question to ask, which is the question the field itself would ask before believing the result.

7. What has changed since 2015, and whether it works

The reforms that followed the replication crisis map neatly onto the mechanisms in section 4, which is the best sign that the field diagnosed itself correctly. Flexible analysis is countered by preregistration, declaring the hypotheses and analysis plan before seeing the data. Publication bias is countered by registered reports, where a journal agrees to publish a study on the strength of its design, before the results exist. Low power is countered by larger samples and multi-laboratory consortia, and irreproducible analyses by mandatory sharing of data and code. The incentive problem is the hardest, and it is being addressed, slowly, by funders paying for replications directly.

The question for a reader is whether these reforms change the reliability of what gets published, and the honest answer is that they work where they are used and are still used far too little. That distinction matters when you weigh a new result: a preregistered study with open data carries a different prior from the typical paper in the scorecard, and it is worth knowing how to tell the difference.

Preregistration and registered reports

The most striking evidence for preregistration comes from medicine, where it became a requirement. Among large trials of drugs and dietary supplements for cardiovascular disease funded by the US National Heart, Lung, and Blood Institute, 17 of 30 published before 2000 showed a significant benefit on their primary outcome, against only 2 of 25 published afterwards, when prospective registration on ClinicalTrials.gov had become the norm - PLOS ONE. The authors found that pre-registration "was strongly associated with the trend toward null findings". Psychology shows the same thing through registered reports: analysing the first hypothesis of each article, one study found 96 percent positive results in standard psychology papers against 44 percent in registered reports - Eindhoven University of Technology.

Neither drop means that registration makes real effects vanish. Neither study can prove cause (the trial authors write that their design "does not allow causal inferences"), but the most likely reading is that when researchers cannot choose their analysis after seeing the data, and journals cannot choose papers by their results, the published rate of positive findings falls toward the true rate, which is much lower than the unregistered literature implies. More than 300 journals now offer registered reports, either as a regular option or in a special issue - Center for Open Science. For a reader, the practical move is simple: look for a registration number (ClinicalTrials.gov for trials, the OSF or AsPredicted for other research) and check that the outcome in the headline is the one that was registered.

Bigger samples, but transparency still lags

Sample sizes have grown. Across 57,909 articles in 12 psychology journals, the median estimated sample size rose from 105 in articles published from 2010 to 2015 to 190 in those published from 2016 to 2021, although the median share of significant p-values per article stayed constant at 69 percent - PLOS ONE. Openness has improved much less. In psychology articles published in 2022, only 7 percent were preregistered, 14 percent shared raw data and 8.5 percent shared analysis scripts, and the authors of that audit concluded that "research transparency continues to be widely neglected in psychology" - Advances in Methods and Practices in Psychological Science.

That gap has consequences the 2026 results made concrete. Section 2 showed that when journals require authors to share data and code, as in the Institute for Replication's sample, more than 85 percent of claims could be reproduced computationally; across the broader SCORE sample, data were available for only a quarter of papers in the first place. Openness is not a guarantee of truth, but it is a precondition for anyone to check, and the contrast between the two samples, though not a controlled comparison, suggests that the requirement, not the encouragement, is what produces it.

Funding replication directly

The incentive problem from section 4 has a direct fix: pay people to replicate. The Dutch research council NWO ran three rounds of what was "the world's first funding programme specifically aimed at replication research" between 2016 and 2020, and in December 2025 Open Science NL, part of NWO, granted 5.2 million euros to 27 applicants to replicate earlier studies - Open Science NL. The UK set up a Metascience Unit in 2024 with an initial budget of £10 million to 2027, naming the replication crisis among its priorities - UKRI.

In the United States, a May 2025 executive order on "Gold Standard Science" listed reproducibility first among the standards federally funded research should meet, citing the view of "a majority of researchers" that science faces a reproducibility crisis - The White House. In September 2026 the National Institutes of Health announced plans for what Inside Higher Ed described as "a five-year, roughly $174 million effort" to encourage scientists to test each other's published findings - Inside Higher Ed. Money on that scale could change what a career in replication looks like, though it would still be small next to the budgets that fund new findings.

A cautionary tale from the reform movement itself

The reform movement has produced its own warning. In 2023 Nature Human Behaviour published a high-profile study reporting that new social-science findings, discovered and tested with rigour-enhancing practices including preregistration, replicated in 86 percent of attempts - Nature Human Behaviour. The paper was retracted in September 2024. The journal's concerns included "lack of preregistration for measures and analyses supporting the titular claim (against statements asserting preregistration in the published article)" and "selection of outcome measures and analyses with knowledge of the data" - Nature Human Behaviour. All authors agreed to the retraction because of the incorrect statements about preregistration, while disagreeing with other concerns.

The episode is not an argument against reform. It is an argument that reforms are not self-certifying: a study that says it was preregistered still has to be checked against its registration, and the claim "this was done rigorously" deserves the same scrutiny as any other claim. Readers can apply the same standard to every result in this guide.

So is the crisis over?

Not yet, and the evidence points both ways. On the improving side: samples are larger, registered reports are spreading, and journals that mandate sharing produce research that can be checked. The fruit-fly study in section 3 even found a higher replication rate than earlier assessments of other fields. On the other side, most of the literature a reader encounters was published before these reforms took hold, SCORE's sample of papers from 2009 to 2018 replicated about half the time, transparency remains rare, and the integrity problems in section 5 are growing faster than the defenses against them.

The first-principles conclusion is that the reforms fix the mechanisms they target, wherever they are adopted, and adoption is partial. So the scorecard's rough prior, about a coin flip for a typical published finding to replicate at all, still describes most of what you will read. A study that is preregistered, adequately powered and openly shared deserves a meaningfully better prior. The next section turns that into a working method.

8. How to tell whether a new breakthrough will hold up

The most encouraging finding in the whole replication literature is that failure is partly predictable. When researchers were asked to bet on which findings would replicate before the replications were run, their collective bets beat chance by a wide margin: in the first such study, prediction markets correctly called 71 percent of 41 psychology replications - PNAS, via PMC, and in the SCORE program's 2026 results, two independent human forecasting methods reached 76 and 78 percent on their best measures - EurekAlert. If informed people can see failure coming, the signals must be visible in the papers, and a careful non-specialist can learn to read most of them.

Two cautions keep that optimism honest. The same SCORE results found that none of three automated prediction methods was consistently effective, and that different measures of a paper's credibility were only weakly related to one another, which the authors of the program's cross-project analysis read as a sign that there are unlikely to be "broadly applicable shortcuts" to judging a finding. The signals below are therefore a way to set your confidence before the evidence matures, not a verdict: a result can fail every test and still turn out to be right, so the aim is to be neither the reader who believes every headline nor the one who dismisses everything new. They fall into three groups: what the study itself shows, how the claim around it has been framed, and what has happened since publication.

Questions about the study itself

The first group can be answered from the paper, and usually from its abstract. These are the features that the large replication projects found most strongly associated with success or failure, and they are the ones the arithmetic in section 4 says should matter. A study that scores badly here is not necessarily wrong, but it carries the profile of the results that most often were.

The point of asking them is calibration rather than suspicion. A large, preregistered, randomized study with a strong result is the closest thing to a reliable first report that science produces, and it deserves to move your beliefs a long way. A small, exploratory study with a borderline result deserves interest and patience in roughly equal measure.

  • How big is it? Dozens of participants, animals or samples is small for most effects; thousands is large.
  • How strong is the evidence? A p-value just under 0.05 or a confidence interval brushing zero is fragile.
  • Was the analysis declared in advance? Look for a trial registry number or a preregistration link.
  • Is the design controlled? Randomized, blinded, and compared with a real control group.
  • Is the effect plausible? A huge effect from a small intervention is a warning sign, not a bonus.

The strength of the original evidence is the best-documented of these signals. In the 2015 psychology project, 63 percent of original findings with a p-value below 0.001 replicated, against 18 percent of those with a p-value above 0.04, and "surprising effects were less reproducible" - Science, author copy. Taken together, these five questions describe the difference between a confirmatory study and an exploratory one, and that difference matters more than the journal's name. The NHS-Galleri trial answers every question well, which is why its negative primary result is so informative. The personalized kidney-cancer vaccine that kept all nine patients recurrence-free answers the first question poorly through no fault of its authors, since phase 1 trials are small by design, which is why our guide to it argues that "100 percent" in a nine-person study is the most dangerous number in the story.

Questions about the claim around the study

The second group concerns the gap between what a paper shows and what is said about it. Exaggeration usually enters not in the paper but in the press release and the headline, where caveats are trimmed, an animal result becomes a human one, and a secondary finding becomes the main event. Readers who only ever see the headline inherit every one of those distortions without knowing it.

The fix is to read one level closer to the source than the coverage you encountered. For a news story, that means the press release or the abstract; for a press release, the paper's own conclusions; for a paper, the registered protocol. Most of the distance between hype and evidence can be closed in five minutes by asking these questions.

  • Primary or secondary? Is the headline about the outcome the study was designed to test, or a subgroup?
  • Which organism? A result in cells, flies or mice is a lead for humans, not a finding.
  • Reviewed or not? A preprint has not faced peer review; say so when you share it.
  • Who paid? Industry funding is normal and not disqualifying, but it is context.
  • What do the authors say? The limitations paragraph is often more honest than the coverage.

Attention is not evidence, and sometimes it points the wrong way. A machine-learning census of 14,126 psychology papers found that media attention went with lower predicted replicability, while a paper's citations and the prestige of its authors' university were unrelated to it - PNAS. The organism question catches some of the most common over-readings. The Lsp2 study in Frontier's index is a clean, well-controlled piece of work showing how early-life protein restriction extends lifespan in fruit flies, and its own authors frame the open question as whether a comparable mechanism operates in mammals. A headline that says "eating less protein as a child extends life" would turn a good fly study into a bad human claim. The same discipline applies to secondary end points: the Galleri trial's stage IV result is a legitimate secondary finding, and reporting it as the trial's answer would invert what the trial found.

Questions about what happened next

The third group can only be answered with time, which is why a fresh result cannot pass it yet and why Frontier labels such results Early signal. These are the signals that, over months and years, turn a claim into knowledge or quietly retire it. Checking them takes a few minutes per paper with free tools, and the habit pays off precisely on the results that matter most to you.

The most valuable single check is for independent confirmation, meaning a different group, with different hands and equipment, finding the same thing. Everything else here is a proxy for that, useful because independent confirmation is slow and often never published at all.

In practice, ask five things. Has another group reproduced the result, ideally in a registered study? Did the journal publish a News & Views or perspective alongside it, and was that commentary skeptical? Has any correction, expression of concern or retraction notice appeared? Are the data and code open for anyone to re-analyse? And are independent groups building on the work, or only the original lab? A result that collects good answers to these questions over a year or two has earned a level of trust that no first report can claim.

For the post-publication record, start with the journal page and with PubPeer, where scientists comment on published papers, often anonymously. The Retraction Watch database, which Crossref acquired in September 2023 and made publicly available, records retractions gathered from publisher websites and is updated every working day - Crossref. Frontier surfaces the formal notices automatically in each entry's integrity line, but a critical comment thread can appear long before any formal notice does, and it is worth reading when it exists.

The signals at a glance

The table below collects the signals in one place, with what each tells you and where to look. It is a summary, not a scoring system: the right weight for each signal depends on the field, and section 6 explains why a physics result and a psychology result deserve different questions.

SignalWhat it tells youWhere to check
Sample sizeWhether the study could detect a real effectAbstract, methods
Strength of resultBorderline results replicate less oftenAbstract, results tables
PreregistrationWhether the analysis was fixed in advanceClinicalTrials.gov, OSF, the paper
Primary vs secondaryWhether the headline matches the designRegistered protocol
Organism and settingWhether a human claim is supportedMethods
Peer review statusWhether experts have checked it yetJournal page or preprint server
Independent replicationWhether it survives other handsLater literature, citing papers
Integrity recordCorrections, concerns, retractionsJournal page, PubPeer, Crossref

The questions are cumulative. A result that is large, preregistered, honestly framed, openly shared and independently replicated is about as reliable as science gets. A result that is small, flexible, overstated, closed and unreplicated may still be right, but you should hold it loosely and be unsurprised if it fades. Most new breakthroughs, including most of the ones in Frontier's index, sit between those poles, and the honest answer for them is "promising, unconfirmed".

9. How Frontier reads a new result, and what its labels mean

Frontier is an index of new science, which puts it in an awkward spot for a guide like this one: almost everything it lists is, by construction, at the stage where nobody has had time to replicate it. A paper published three weeks ago has not been repeated by an independent lab, has barely been cited, and may not yet have drawn its first critical letter. Pretending otherwise would commit the exact error this guide describes, so the index is built to show the state of the evidence next to the excitement and to keep the two from blending into one number.

The Frontier Score has three pillars, and only one of them asks whether a result is established. Evidence, a quarter of the score, grades whether the work cleared peer review and how selective its venue is, how many of five independent citation indices record it, whether anything beyond citations corroborates it (a registered clinical trial, released data or code, public funding on record), how openly it can be checked, and its integrity record. Impact and Novelty make up the rest, and Novelty carries half the weight because the index exists to find the edge, not to rank settled classics. The full recipe is on the methodology page, and every entry shows its own breakdown with each number linked to the record it came from.

Three rules that encode the replication problem

Three design rules matter most for a reader who has absorbed the history in this guide. They are not attempts to predict replication, which no public signal can do reliably for a three-week-old paper. They are guardrails that stop a fresh claim from borrowing a certainty it has not earned, written into the scoring code so they apply to every entry the same way.

The rules are deliberately mechanical. A judgment call made entry by entry would drift with the mood of the week, and it would be impossible for a reader to check. A rule in code can be read, criticised and audited, and its effect on the whole index can be counted, which is what the next subsection does.

  • Retraction is a hard gate. A retraction recorded by OpenAlex or Crossref sets the entire score to zero, whatever the citations.
  • Preprints lose evidence. An unreviewed preprint scores 45 out of 100 on peer review, against 78 for a mega-journal and 100 for a flagship.
  • Confidence caps the label. An entry rated Early signal cannot be labelled above Notable, however high its score.

The first rule handles the integrity layer from section 5: a correction or expression of concern lowers the integrity sub-score to 75 rather than collapsing it, because a correction is often a sign of a healthy record rather than a broken one. The second rule encodes the gap between a claim that has faced reviewers and one that has not, without pretending that peer review is replication. The third rule is the one that does the most work. Confidence is computed separately from the score, from citation maturity, about a year of time to settle, agreement across citation indices and the breadth of independent corroboration, and it decides how loudly an entry is allowed to present itself.

What the index looks like through that lens

The effect is easiest to see in the index as a whole. On 5 October 2026 the index held 206 published entries. Of those, 111 (54%) were rated Early signal, 36 (17%) Firming up and 59 (29%) Settled. Every one of the 103 entries published within the previous three months was an Early signal, including all 44 results in the 21-30 September Dispatch, and no entry in the index currently wears the top Landmark label, because the confidence cap holds back even the highest-scoring fresh work until its record matures. Nine entries are preprints, none is retracted, and six carry a correction notice from Crossref. The data behind those counts are open in the Frontier dataset.

The shape is the point. Confidence is mostly a function of time and uptake, which is exactly what a fresh result lacks, so the newest half of the index is uniformly provisional and the oldest third is mostly settled. A reader who wants to know how much weight to put on an entry should read the confidence label before the score, in the same way that a careful reader of a medical headline asks about the trial phase before the effect size.

Two September trials, read the way this guide recommends

Two large clinical trials, published in NEJM in late September and added to the index on 5 October, show why "Early signal" is a statement about the record, not about the quality of the study. The NHS-Galleri trial, published in NEJM, randomized 142,250 people aged 50 to 77 in England to have a multicancer blood test or to have their blood stored. After three annual screening rounds, the primary end point (the incidence of stage III or IV cancer across 12 prespecified cancer types) did not differ between the groups, with an incidence rate ratio of 1.03 (95% confidence interval 0.92 to 1.14). A secondary end point, stage IV cancer alone, came out at 0.86 with an interval reaching 1.00, the value that means no difference. This is a well-designed, preregistered, very large trial, and its central answer is negative. Any headline built on the stage IV figure is building on the secondary end point, which is precisely the move section 8 warns against.

The TRIUMPH-1 trial of retatrutide is the opposite case: 2,339 adults with obesity, randomized and double-blind for 80 weeks, with mean weight loss of 25.0% on the highest dose against 3.9% on placebo. The design is about as strong as clinical evidence gets, and it still sits at Early signal because the paper is days old and has almost no citation record. A reader has further reasons for patience: it is one trial funded by the drug's maker, and its abstract leaves open the questions that decide its real-world value (detailed safety data, what happens to weight after 80 weeks, and people with diabetes, who were excluded). Contrast both with the one-time base-editing infusion for LDL cholesterol, a phase 1 study of 35 participants that is rated Firming up because months of uptake have accrued: maturity of the record and strength of the design are different axes, and a reader needs both.

Where the score is blind

It would be easy to oversell what any index can do here, so it is worth being plain about the limits. The Frontier Score cannot see replication directly: no public database records, for a paper published this year, whether an independent lab has quietly failed to reproduce it. The integrity gate catches retractions, but section 5 showed that retraction is slow and rare relative to failure, and a result that simply does not hold up, without misconduct, usually keeps its place in the literature with no flag at all.

Citations are a second blind spot. The Impact pillar is field-normalized, which removes the crudest biases, but a citation does not say whether the citing paper agreed. The finding from section 4 that papers which failed to replicate kept attracting citations is a direct warning about any citation-based measure, including this one. That is why Impact is kept separate from Evidence, why confidence is a separate axis, and why the score of a fresh entry should be read as "how strong and new does this look today" rather than "how likely is this to be true". The same caution applies to our ranking of the biggest scientific breakthroughs of 2026: its order is a measure of evidence, impact and novelty at the time of writing, and its confidence split says how much of that evidence had matured. When an entry moves from Early signal to Settled over the following year, it is the record that has changed, and that is information a reader can use.

How this index and this guide are kept current

The index changes every day. At 06:00 UTC a run takes in up to four new papers from the venues where breakthroughs land, and every published score is re-checked against its live sources once a week, so a retraction notice recorded by OpenAlex or Crossref reaches the score within about a week. Since late September, each new entry's editorial (the summary, why it matters, and what to watch) has been drafted by AI agents from the paper's own abstract and checked claim by claim against that abstract by a separate pass before it goes live, and most of those checks have caught wording that overstated what a paper showed. The whole operation runs on Founden, a platform that Frontier's founder also runs, and the failure it has to be engineered against is the one this guide keeps returning to: in the words of that platform's write-up of AI data-corruption incidents and the fixes that prevent them, such systems "fail by acting confidently and reporting falsely".

Researching this guide turned up a small instance of the problem inside the index itself. The entry on Fermilab's final muon g-2 measurement, written in August, printed the measured value with the wrong power of ten, which made it ten times smaller than the figure in the collaboration's own abstract - arXiv. It was corrected on 5 October, the day it was found, and the entry now matches the paper. Errors of this kind are why every figure on the site links to its record: a reader can catch one in seconds, and so can the next check.

This guide was assembled the same way and is dated at the end. Every figure in it links to the page where it was checked, and the figures about Frontier come from the index's own live data on the day of writing, so they will drift as the index grows. The About page describes the pipeline step by step, and each entry's sources panel shows exactly which record each number came from.

10. What comes next: AI as a source of errors, and as a checker

The structural reason replication is rare is cost. Repeating a study means redoing the work: buying the reagents, recruiting the participants, booking the beam time, and spending months on a result that, if it confirms the original, most journals will not want to publish. Anything that changes the cost of producing claims, or the cost of checking them, changes the replication problem at its root. Artificial intelligence changes both at once, in opposite directions, and the balance between those two effects is the most important open question about the reliability of science in the next few years.

On the production side, the cost of a plausible-looking paper has collapsed. The paper-mill problem from section 5 predates language models, but generated text makes fabrication cheaper and harder to spot, and a newer failure has appeared alongside it: references to papers that do not exist. A Nature analysis published in April 2026 suggested that tens of thousands of publications from 2025 might include invalid references generated by AI - Nature. An invented citation is a small thing in any one paper, but it is the literature's equivalent of a broken chain of custody: a claim that appears to rest on evidence and rests on nothing.

The checkers are arriving

The same technology is being turned on the literature as an auditor, and the early results are concrete. In late 2024, a widely covered study warned that black plastic cooking utensils carried worrying levels of a flame retardant; a mathematical error meant the real exposure was ten times below the safe limit, and researchers showed afterwards that an AI model could have spotted the error in seconds - Nature. That episode launched an open-source effort, the Black Spatula Project, which had checked around 500 papers by March 2025 with a false-alarm rate of about ten percent, while a second tool, YesNoError, had scanned more than 37,000 articles in its first two months - heise. By August 2026, AI agents were finding faults in decades-old papers and reference databases, including some trusted boiling-point values in one reference database that chemists had relied on for decades and that had been wrong all along - Nature.

Error-spotting is the easy half. The harder test is whether an AI system can take a paper's data and code and reproduce its results, the computational form of replication. The benchmarks built to measure this tell a story of fast but uneven progress, and they are worth reading closely because each one measures a slightly different thing.

  • CORE-Bench (2024): 270 tasks from 90 papers in computer science, social science and medicine; the best agent solved 21% of the hardest tasks - arXiv.
  • PRBench (March 2026): 30 tasks from published physics papers; the best agent averaged 34%, and none fully reproduced a paper end to end - arXiv.
  • SocSci-Repro-Bench (June 2026): 221 social-science tasks; two leading coding agents reproduced "a large share" of findings when given the original data and code - arXiv.

Read together, the three benchmarks say that reproducing a computation from shared materials is becoming routine work for software, while rebuilding a method from a paper's description alone is still mostly beyond it. The physics benchmark adds a warning that should be printed on every AI research tool: among the systematic failure modes its authors found was the fabrication of output data, an agent producing numbers that look like the paper's instead of computing them. The social-science benchmark adds a subtler one. Its authors showed that agents "can be nudged toward confirmatory specification search through subtle prompt framing", which is a precise description of automated p-hacking.

What this means for a reader in 2026

The first-principles conclusion is that AI lowers the cost of checking only for the layer of science that is already computational. Re-running shared code on shared data is getting cheap; repeating a mouse experiment, a clinical trial or a telescope campaign is not. Fields and papers that publish their data and code are therefore about to be audited far more often than those that do not, and errors in them will surface faster. That makes openness a stronger signal than it already was: a paper whose materials can be fed to an automated checker is a paper that is going to be checked.

The second conclusion is a caution. The same tools that can catch a statistical error can generate a hundred analyses and keep the one that works, and a language model asked to support a hypothesis will often find a way. The defenses against that are the ones section 7 described for human researchers, and they transfer unchanged: preregistered analysis plans, registered reports, shared code, and independent replication. AI does not retire those reforms. It makes them the minimum.

11. Conclusion: how to hold a new result

Fifteen years of measurement have turned the replication crisis from an accusation into a body of evidence, and the evidence is more useful than alarming. Across the audited fields, something like half of published findings replicate, by the usual test of a significant effect in the same direction, almost all of those come back smaller, and the failures cluster where the arithmetic says they should: small studies, surprising claims, flexible analyses, closed data and headlines that outrun their papers. None of that makes new science untrustworthy. It makes it provisional in a predictable way, and a reader who knows the pattern can hold each result at the right strength.

A practical framework follows directly from the sections above, and it works for a headline, a press release or an entry in an index like this one. It does not require specialist training, only the habit of asking where a result sits before deciding how much to believe it.

Start from the base rate. A typical published finding is roughly a coin flip to replicate at all, and likely to come back weaker if it does; preclinical biology does worse, and preregistered, openly shared work deserves a better prior.

Adjust for the study. Size, strength of evidence, advance registration, controls and plausibility move the odds more than the journal's name does.

Check the claim against the paper. Ask whether the headline is the primary or a secondary outcome, which organism it concerns, whether it has been peer reviewed, and what the authors themselves say about its limits.

Wait for the record. Independent replication, expert commentary and integrity notices are what turn a claim into knowledge, and they take months to years.

Update in both directions. Confirmation by another group should raise your confidence most; years of silence should not raise it at all.

The fifth step is the one most readers skip. A result that nobody has tried to repeat is not confirmed by default, as the fruit-fly project's unchecked claims showed, and a result that has been independently confirmed deserves more confidence than any first report, however prestigious. That is the logic behind the confidence label on every entry in Frontier's index: new results are shown as Early signal because their record is young, and they earn Firming up and Settled only as the evidence accrues. The excitement of a breakthrough and the reliability of a finding are different things, and the most useful habit a reader of science can build is to keep them apart until the evidence brings them together.

This guide reflects the evidence available on 5 October 2026. Replication projects publish new results regularly, and the figures about Frontier in section 9 come from the live index on that date, so check the linked sources for the latest numbers.

The Frontier Brief

The week's frontier, scored.

Every Monday, the highest-scoring new breakthroughs in the index, each with its score and a one-line read on why it matters. What you get

One email a week. The top new breakthroughs, scored. Unsubscribe anytime.

The Replication Crisis in 2026: How Much Science Holds Up | Frontier