Peer Review Is Unpaid, Invisible, and Nobody Is Measuring It
Every claim in science is supposed to pass through peer review before it enters the record. It is the mechanism the whole system points to when asked how it maintains standards.
It is performed by volunteers, in time they are not compensated for, with no record that they did it, no assessment of whether they did it well, and no consequence for doing it badly. If you designed a quality assurance function this way in any other industry you would be asked to explain yourself.
The interesting failure here is not the missing money. It is the missing instrumentation. Nobody can run a function they do not measure, and nobody is measuring this one.
The size of the donation
Aczel and colleagues estimated this in Research Integrity and Peer Review in 2021, in a paper titled “A billion-dollar donation.”
Reviewers worldwide worked more than 100 million hours on peer review in 2020, which is over 15,000 years of labour. The monetary value of time donated by US-based reviewers alone exceeded 1.5 billion dollars. China-based reviewers contributed over 600 million and UK-based reviewers close to 400 million, putting those three countries together at roughly 2.5 billion dollars in a single year.
The authors note these are likely underestimates, since they cover only a portion of journals worldwide.
That labour is donated into a publishing industry with operating margins that are, by the standards of most sectors, remarkable. A supplier base that hands over its core input at zero cost is about as durable a position as a business can hold, and nobody ever negotiated it. I still do not think the margin is the most interesting part of the story, and it is the part that gets all the attention. The interesting part is what the absence of measurement does to quality.
What is not recorded
Consider what the system knows about a reviewer.
It does not know how many reviews they have done. There is no portable, verifiable count. Publons attempted this and was folded into Web of Science, coverage was always partial, and it depended on journals and reviewers opting in.
It does not know how long they spent. It does not know whether the review was two paragraphs or six pages. It does not know whether they caught the error that a later reader found, or whether they missed it. It does not know whether their recommendation correlated with anything about the paper’s subsequent fate.
An editor at a single journal builds a private impression of who is reliable. That impression does not leave the journal, does not transfer when the editor moves, and is not available to anyone else.
So the entire quality function operates without instrumentation. Nobody can answer the question “is peer review at this journal getting better or worse” with data, because the data does not exist.
The consequences
Reviewer fatigue concentrated on the conscientious. Editors go back to people who respond and do a good job. Those people receive more requests. The reliable reviewer’s reward for reliability is more unpaid work, so eventually they decline, and the load shifts to whoever says yes.
Variance nobody sees. Two papers of identical quality can receive completely different treatment depending on who was available that month. Studies of inter-reviewer agreement generally find it is low. In an instrumented system that variance would be a tracked metric with someone accountable for reducing it. Here it is folk knowledge.
Delay as the visible symptom. Months to a first decision is normal. The cause is a scarce volunteer resource allocated without any market or scheduling mechanism.
The gate is weaker than the system claims. The retraction and reproducibility numbers are what a weak gate looks like from the far side. Over 10,000 retractions in 2023, with more than 8,000 from one publisher’s journals, is not a story about reviewers being lazy. It is a story about a filter with no feedback loop on its own performance.
Paper mills exploit exactly this. An unmeasured, unaccountable, high-volume process staffed by volunteers is straightforwardly attackable, and it has been attacked at industrial scale.
Why it has not been fixed
The arguments against paying reviewers are real and worth stating properly.
Payment introduces incentive to accept work you are not qualified for. Payment sets a price that small journals and low-income institutions cannot match, which would stratify review quality by wealth. Payment may crowd out intrinsic motivation, which the behavioural literature suggests is a genuine risk for tasks currently done from professional obligation. And review is often argued to be part of what a salaried academic is already paid for, though that argument sits uneasily with the fact that the beneficiary is a commercial publisher.
I find these arguments partially convincing on payment specifically. I find none of them relevant to measurement.
There is no coherent argument for not recording who reviewed what, how long it took, and whether the review was any good. That is not a payment question. It is instrumentation, and the objection to it is inertia.
What would help
A portable reviewer record. Verified, cross-publisher, attached to an ORCID. Count, field, and ideally an editor’s quality rating. Reviewing then becomes visible in hiring and promotion, which is a form of compensation that does not distort the way cash might.
Open review where the field tolerates it. Publishing the reports, signed or unsigned, alongside the paper. Several journals do this and it works. A review that will be read is a review that gets written more carefully, and it lets readers judge the filter rather than trusting it.
Actual measurement. Turnaround, agreement rates, and correlation between review outcomes and later citation or retraction. Publish it per journal. This is the reform I would prioritise, because it is the one that makes every other reform evaluable.
Registered reports, which shift review to the design stage before results exist, and which I have written about in the context of negative results. Reviewing a design is a better-defined task than judging a finished claim, and it is one reviewers report finding more satisfying.
Credit that counts. Until review appears on a CV in a way a committee recognises, it will remain the thing you do after the work that gets counted.
The part I keep returning to
Science has spent a decade building infrastructure for reproducibility. Data repositories, code archives, preprint servers, persistent identifiers, registries. Almost all of it targets the artefacts.
The evaluation layer has received almost none of that attention, despite being the part everyone points at when asked why published work should be believed. We have better provenance for a dataset than for the judgment that let a paper through.
That imbalance is worth fixing, and the first step pays nobody and costs almost nothing. It is writing down what happened.
Related: