The reproducibility problem is usually argued as a matter of scientific standards. I want to argue it as a matter of capital, because the waste has been measured, and in one field in one country it runs to about 28 billion dollars a year.

Every figure below has a source.

What fails

The most cited survey is Baker’s 2016 poll of 1,576 researchers for Nature. More than 70 percent of respondents had tried and failed to reproduce another scientist’s experiment. More than half had failed to reproduce one of their own. 83 percent agreed there is a reproducibility crisis, with 52 percent calling it significant. Selective reporting was the factor most often named as contributing, at 66 percent.

Surveys measure opinion. The replication attempts measure outcomes.

The Open Science Collaboration attempted to replicate 100 studies from three leading psychology journals and published the result in Science in 2015. 36 percent of replications produced statistically significant results, against 97 percent of the originals. Replication effect sizes were about half the originals.

In preclinical biomedicine, Begley and Ellis reported in Nature in 2012 that researchers at Amgen could confirm the findings of only 6 out of 53 landmark oncology studies, which is 11 percent. Prinz and colleagues at Bayer had reported a similar picture the previous year, with roughly a quarter of published findings reproducible in-house.

These are different fields with different methods and different failure modes. The numbers cluster anyway.

What it costs

Freedman, Cockburn and Simcoe published an economic analysis in PLOS Biology in 2015. Their estimate: approximately 28.2 billion dollars per year is spent on preclinical research in the United States that is not reproducible, with around half of all US preclinical research falling into that category.

Their decomposition of causes matters as much as the total. The largest single contributor was biological reagents and reference materials, which includes contaminated and misidentified cell lines. Not statistics. Not fraud. Materials.

That number covers one country and one stage of one field. It excludes clinical research, excludes everything outside the United States, and excludes the largest category of cost, which is not measurable at all.

The cost that is not in the number

The 28 billion figure is direct expenditure. The consequential costs are bigger and harder.

Work built on a false foundation. A research programme that starts from a published result and spends three years extending it has lost three years if the result was noise. This does not appear in any budget line. It appears as a career.

Screening cost. Because you cannot trust a result by default, you replicate before building. Every group that does this pays the replication cost, and they pay it independently, and none of them publish the outcome. I wrote about where those results go and the answer is nowhere.

Drug development attrition. The clinical failure rate for candidates entering trials is high, and some fraction of that traces to preclinical findings that did not hold. The cost per failed programme runs into hundreds of millions.

Trust decay. Harder to price and probably the largest. If a working researcher’s prior on a novel published finding is around fifty percent, the literature functions as a set of leads rather than a body of knowledge, and every reader pays verification cost separately.

Retractions

The number of retractions issued in 2023 passed 10,000 for the first time, according to Nature’s analysis of the Retraction Watch database. Over 8,000 of those came from journals run by Hindawi, a Wiley subsidiary, largely in special issues with weaker editorial oversight. Among large research-producing nations, Saudi Arabia, Pakistan, Russia and China had the highest retraction rates over the preceding two decades.

Retraction growth is ambiguous evidence. It partly reflects better detection, which is progress. It also reflects paper mills operating at industrial scale, which is not. Integrity researchers quoted in that reporting describe the retracted set as the tip of the iceberg, and the reasonable reading is that retractions bound the problem from below.

A retraction also arrives late. Papers are typically cited for years before withdrawal, and citation of retracted work continues afterwards, because the citation graph does not update itself.

Why the incentives produce this

Nobody in the chain is behaving irrationally.

A researcher is evaluated on publications in high-impact venues. Replication does not produce those. Careful negative results do not produce those. The rational allocation of a limited career is toward novel positive findings, which is what the system asks for and gets.

A journal competes on impact factor, which rewards novelty and surprise. Surprising results are, on average, more likely to be wrong, because a result that contradicts a well-supported prior is more often noise than discovery. Selecting for surprise selects partly for error.

An institution competes on rankings built from publication counts and citations. None of the inputs measure whether anything reproduces.

A funder wants demonstrable output within a grant cycle. Replication and validation are neither novel nor fast.

Every actor is optimising sensibly against their metric, and the aggregate is a literature nobody fully trusts. Anyone deploying capital against that literature, a funder, a founder, an institution, is buying an asset whose failure rate is not disclosed anywhere on it. That is a structural result, not a moral one, and it will not be fixed by asking scientists to be more careful.

What actually moves

The interventions with evidence behind them share a feature: they change what is rewarded or when a decision is made.

Registered reports, where the publication decision happens after review of the design and before the result exists. Mandatory data and code deposition, which makes computational work checkable at near zero marginal cost. Reagent authentication requirements, which target the largest single cause Freedman identified and are cheap relative to the failure they prevent. Funded replication programmes, which have to be funded specifically because no individual has an incentive to run them.

And preprints, which at least make the timeline and the raw claim public early.

The version I care about

I work in theoretical physics and in quantitative finance, and neither has a wet lab. The failure mode transfers anyway.

In theory work, the equivalent of an unreproducible experiment is a derivation nobody has checked. Long calculations with many steps, published without the intermediate work, in a form where verification means redoing it. I recomputed the beta function in my own Stueckelberg extension twice with different gauge choices for exactly this reason, and I would not have believed the first result on its own.

In quantitative finance the failure mode is backtest overfitting, and it is arguably the worst case in any field, because the search is cheap, the data is fixed, and the incentive to report the best run is direct. A strategy selected after examining the data has an expected out-of-sample performance far below its backtest, and the entire literature of published anomalies is subject to this. That is why the robustness testing has become the substance of my own work rather than an appendix to it.

The general principle is the same everywhere. A result you have not tried to break is not a result. The cost of skipping that step has been measured, and it is about 28 billion dollars a year in one field in one country. Everything else on that bill is still unpriced.

Related: