The Negative Results Nobody Publishes
The most expensive thing in science is the experiment that has already failed somewhere else.
Somebody in a lab tries an approach. It does not work. They spend three months establishing that it does not work, convince themselves the failure is real rather than a mistake, and then move on to something else.
That knowledge now exists in one person’s head and one lab notebook. It is not in the literature. Fifteen other groups with the same idea will spend the same three months over the next five years. Nobody invoices for that and everybody pays for it.
This is the most straightforwardly wasteful thing in science and it is nobody’s fault in particular, which is why it persists.
How much is missing
The numbers are not subtle.
Daniele Fanelli analysed over 4,600 papers published between 1990 and 2007 and found the proportion reporting positive results exceeded 80 percent after 1999, peaking at 88.6 percent in 2005. In some fields, including psychology and ecology, it exceeded 90 percent. A world in which nine out of ten tested hypotheses are correct is not a world in which anyone is testing anything difficult.
The cleanest evidence comes from Franco, Malhotra and Simonovits, published in Science in 2014. They used the Time-sharing Experiments in the Social Sciences archive, which is unusual because it records every study that was approved and run, including the ones that vanished. That gives you the denominator, which is normally the thing you cannot see.
Of the null-result studies in the archive, only 10 out of 48 were published. And the larger part of the loss was not journals rejecting them. Most null results were never written up at all. Researchers looked at the result, judged it unpublishable, and did not submit.
The file drawer is not primarily a gatekeeping problem. It is a self-censorship problem, produced by an accurate read of the incentives.
Why it is worse than it sounds
If negative results were missing at random, the literature would be incomplete. They are not missing at random. They are missing conditional on the outcome, which means the published record is a biased sample of the work that was done.
Concretely: run twenty independent tests of a hypothesis that is false. At the conventional threshold, one comes back significant by chance. Nineteen go into drawers. The one gets published. The literature now contains one paper supporting a false claim and no papers contradicting it.
Meta-analysis, which is supposed to be the correction for single-study noise, inherits the bias. It can only pool what it can find. There are statistical tools for detecting publication bias in a body of work, funnel plots and their relatives, and they help, but they are inference about missing data rather than the data.
There is a modelling result on this that I found clarifying. A 2016 paper in eLife by Nissen and colleagues on the canonisation of false facts showed that under plausible publication bias, a false claim can become established in a field through the ordinary accumulation of evidence, with every individual study conducted honestly. No fraud required. The bias in what gets published is sufficient on its own.
The costs, itemised
Repeated work. The direct one. Every group that reruns a failed approach pays the full cost of the failure. Nationwide, across a field, over a decade, this is a large number that nobody bills. In any other sector, selling someone three months of work already known to be worthless would be a scandal. Here it is the default operating condition, funded out of public money.
Bad priors. A researcher deciding what to work on reads the literature and sees an approach that appears promising. It appears promising because the failures are invisible. The decision to spend two years on it is made on filtered information.
Reproducibility. Freedman, Cockburn and Simcoe put the annual cost of irreproducible preclinical research in the United States at roughly 28 billion dollars in a 2015 PLOS Biology analysis, with around half of preclinical research not reproducible. Publication bias is not the whole of that number, but a literature that only reports successes is a literature that is systematically hard to reproduce, because the successes include an unknown share of noise.
Young researchers. A doctoral student whose project produces a clean negative result has done real work and has nothing that counts. That is a career penalty for an honest outcome, and it teaches exactly the wrong lesson at exactly the wrong stage.
What actually changes it
The interventions that work all move the decision point earlier, before the result is known.
Registered reports. You submit the introduction, hypothesis and methodology. Reviewers evaluate the design. If it passes, the journal commits to publishing the outcome whatever it is. The publication decision is now made in ignorance of the result, which is the entire mechanism. Where these have been adopted the proportion of null results published rises steeply, which is what you would expect if the design was always fine and the outcome was the filter.
Preregistration. Weaker, since it does not carry a publication commitment, but it creates a public record that the study happened. A registry entry with no corresponding paper is itself evidence.
Preprints. The marginal cost of posting a negative result to arXiv or bioRxiv or Zenodo is close to zero, and it is citable and timestamped. No editor has to agree. If you have a clean negative result and no journal wants it, this is available today and most people do not use it.
Journals for null results. These exist. They have low prestige, which limits how much they can fix, because the problem was prestige.
The version of this in computational work
I want to name the form this takes in the fields I work in, because it is less discussed than the clinical version.
Computational and modelling work has an equivalent of the file drawer that is arguably worse, because iteration is cheap. You try forty model architectures, or forty feature sets, or forty parameter regions. Thirty-nine do not work. You report the fortieth.
Nothing about that is dishonest at any single step, and the resulting paper can be entirely accurate about what it did. But the reader cannot tell whether the reported result is a finding or the best of forty draws, and often neither can the author, because the search was interactive and nobody kept count.
This is the thing I worry about most in my own topological finance work, and it is why the robustness testing has become the whole project rather than a section at the end. The rule I have settled on is that if I would not be willing to publish the list of everything I tried, I do not trust the result I kept.
That list is also the negative result. It is the most useful thing I could give another person working on the same problem, and there is currently no venue for it, so I intend to put it in the paper. The absence of a venue is not a reason to withhold a result. It is a reason to publish it anyway and improve the record by one paper.
Related: