

THE MEASUREMENT TRAP
The same testimony carried a second figure that matters more. The organisation that maintains the retraction database stated its own confidence that the correct rate should be around two percent — ten times what is actually being recorded.
14 min red

THE MEASUREMENT TRAP
What happens to a system that can see everything and cannot tell when it is wrong
Article 5 of 8 · Series I of III · Published 9 September 2026 · Analysis → Forecast → Recommendations
How this series measures things. Every article applies the same three questions to its subject. Concentration: how many genuinely independent alternatives exist, once shared upstream origins are traced rather than counted. Criticality: what stops if this fails, and how quickly. Substitution time: how long until an alternative actually functions. Article 4 ended on a question. This article applies the instrument to the answer: what happens to the channel through which a system finds out it is wrong.
1. The Signal
In April 2026, a United States congressional hearing on scientific publishing received a figure worth sitting with.
The retraction rate across the scholarly literature stood at roughly 0.2 percent for 2025, up from about 0.02 percent in 2016. A tenfold increase in a decade.
The same testimony carried a second figure that matters more. The organisation that maintains the retraction database stated its own confidence that the correct rate should be around two percent — ten times what is actually being recorded.
Take that pair seriously. An institution's own measurement of its error rate, produced by the people best placed to know, is understood by those same people to be off by roughly a factor of ten. The measurement is not disputed by outsiders. It is disclaimed by its own custodians.
The composition is more revealing than the level. Of retractions recorded in 2025, compromised peer review and paper mills accounted for around 43 percent, with compromised peer review alone at about 37 percent. Citation manipulation accounted for a further 23 percent. Across 2023 to 2025, more than 6,400 retractions were attributed to fake peer review, and roughly 2,100 to AI-generated content, with notices citing tortured phrasing, non-standard phrasing, and text identified as generated by a language model.
The database records 10,409 retracted paper-mill articles from its inception to 2024, with a pronounced spike in 2023.
And the timing: average time from publication to retraction ran near one year in 2019, stretched to about two and a half years by 2021 and 2022, dropped sharply in 2023 when one publisher's mass retraction compressed it, and returned to around two years in 2024.
Read those together and the picture is specific. The dominant failure is not bad content getting through. It is the verification mechanism itself being the thing that was compromised — and the system's report on its own condition takes about two years to arrive and understates the problem by an order of magnitude.
This article is not about scientific publishing. Science is simply the one large system that publishes its own error corrections, which makes the mechanism visible there and invisible almost everywhere else.
2. The Mechanism
A measurement inside an institution is not a description. It is a control signal — something produced in order to be acted on, by parties who know it will be acted on.
That distinction is the whole subject. A thermometer reports temperature whether or not anyone is watching. A performance metric is generated by people and systems who understand what happens when the number moves, and it is therefore always a joint product of the underlying reality and the incentive to represent it in a particular way.
The two production costs
Every measured system has two costs sitting side by side, and their ratio governs everything that follows.
— The cost of producing the thing being measured. Doing the research, delivering the service, building the product, making the loan perform.
— The cost of producing the appearance of it. Generating the paper, formatting the report, hitting the reportable threshold, arranging the transaction so it books correctly.
While the second cost is close to the first, the measurement holds. Faking it is roughly as expensive as doing it, so most parties do it, and the signal tracks reality closely enough to steer by.
When the second cost falls sharply and the first does not, the signal decouples from what it was measuring. Not gradually and not everywhere at once, but wherever the ratio crosses, and the crossing is invisible from inside the signal.
This is the same structure as Article 2, moved one layer. There, the cost of generating a claim fell while the cost of resolving it did not. Here, the cost of generating the evidence of performance falls while the cost of performing does not. Both are asymmetries created by the same technology, and both produce a gap that the institution cannot see using the instruments it already has.
Why a failing measure does not look like one
The property that makes this dangerous is not the decoupling. It is that decoupling improves the reported numbers.
A metric under successful optimisation goes up. Output rises, targets are met, dashboards turn green. If anything, the trend looks better than it did when the measure was working, because the constraint that used to hold it down was the difficulty of the underlying task and that constraint has been removed.
So the signature of measurement failure is improvement. There is no natural point at which someone looking at the numbers becomes concerned, because the numbers are exactly what success looks like.
The scholarly case shows this precisely. Publication volume rose. Citation counts rose. Every metric an institution uses to assess research productivity improved. The failure was visible only through a separate mechanism — retraction — which is expensive, slow, and maintained substantially by volunteers and a small number of specialised outfits rather than by the institutions being measured.
Applying the instrument
The three questions transfer, and the answers are worse than in any previous article in this series.
Measure | Applied to an institution's knowledge of its own error | Typical answer |
Concentration | How many independent channels report that something is wrong? | Frequently one, and often that one is generated by the party being assessed |
Criticality | What happens if the channel fails? | Nothing observable. The system continues, confidently, in the wrong direction |
Substitution time | How long to build an independent channel? | Years, and it must be funded by the party it will embarrass |
That last cell is the structural problem, and it has no clean solution. An independent verification channel costs money, produces bad news, and is paid for by the organisation the bad news is about. It is the first thing cut under budget pressure and the last thing anyone defends, because its output is indistinguishable from a problem it created.
3. Subtheme One — The Half-Life of a Metric Is Collapsing
That a measure degrades once it becomes a target is not new. What has changed is the speed, and speed converts a manageable problem into an unmanageable one.
The old rate
Optimisation against a measure was historically limited by human effort. Someone had to notice what the metric rewarded, work out how to produce it more cheaply than the underlying thing, and then actually do that at scale. Each step took people and time.
A measure therefore had a useful working life, typically several years, sometimes a decade. Institutions could revise metrics on roughly that cycle: introduce, observe, notice the gaming, replace. The revision cycle and the decay cycle ran at comparable speeds, and the system stayed roughly in balance without anyone designing it to.
What changed
Each step of that optimisation loop is now substantially cheaper. Identifying what a measure rewards is a pattern-recognition task. Producing output that satisfies it is a generation task. Doing it at scale is a distribution task. All three have fallen in cost by more than an order of magnitude, and none of the corresponding institutional processes has.
So the decay time of a metric has collapsed while the revision cycle has not. An institution that could previously replace measures as fast as they degraded is now replacing them more slowly than they degrade, and the gap compounds with every cycle.
The scholarly numbers show the shape. Retraction time rose from about one year in 2019 to two and a half by 2022, and stood near two years in 2024. Detection is running two years behind production. In a domain where production volume is rising and generation cost is falling, a two-year detection lag means the record being corrected today reflects a generation process that has since become substantially cheaper.
Why replacing the metric does not fix it
The intuitive response is to design a better measure. It works once and then stops working, faster each time.
Any measure specific enough to be actionable is specific enough to be targeted. Any measure vague enough to resist gaming is too vague to steer by. And the effort required to find the new measure's weak point has fallen along with everything else, so each replacement has a shorter working life than its predecessor.
The consequence is that metric design is no longer the leverage point. An institution optimising its choice of measure is solving the problem it had in 1995. The live question is not which number to watch but whether any number can survive contact with cheap optimisation — and increasingly the answer is that numbers survive only where producing them remains genuinely difficult.
Which points at where the remaining leverage is: not in the measure but in the provenance of it. Not what the number says, but what had to be true for the number to exist.
4. Subtheme Two — The Verification Layer Is the Target
The second force is more specific than general gaming, and it is what makes this a systemic rather than a local problem.
The attack moved
Of retractions recorded in 2025, compromised peer review alone accounted for roughly 37 percent. Across 2023 to 2025, more than 6,400 retractions were attributed to fake peer review. The single largest category of recorded failure is not the content. It is the check on the content.
This is a rational reallocation of effort. Producing convincing content is one cost; compromising the mechanism that would catch unconvincing content is another. Where the second is cheaper, effort moves there, and it moves there without anyone coordinating.
And the consequence is qualitatively different from ordinary gaming. A compromised output is one bad item in a good system. A compromised verification layer is a good-looking system whose reports about itself are produced by the compromised component. The institution's confidence in its own quality is now generated by the part that failed.
Why this is invisible by construction
A working verification layer produces a stream of rejections, corrections and challenges. A compromised one produces approvals.
From the outside, and from the management layer above it, a compromised checking function is indistinguishable from an excellent underlying process. Both produce few problems. Both make the metrics look good. Both reduce the friction that everyone above them experiences as cost.
The two are distinguishable only by a mechanism outside the loop — which is what a retraction database is, and which is why its existence is doing more work in this article than any single figure it contains.
The layers compound
Each analytical layer between a decision-maker and the underlying reality performs a compression, and compression is lossy. What gets lost is not random.
Aggregation removes variance. Summarisation removes qualification. Dashboards remove the cases that do not fit the categories. At every step, the anomalous is the most expensive thing to carry forward and the first thing dropped — and the anomalous is precisely the signal that something is wrong.
So a system with more analytical layers is more confident and less corrigible than one with fewer, and the confidence is a product of the same process that removed the correction. Each layer is individually justified. Nobody added one in order to lose information. The loss is a property of the stack rather than of any decision in it.
This is where Article 4 connects. Acquired capacity removes the intermediaries who carried correction as a by-product, and it adds analytical layers that compress it out. The two effects run in the same direction, and neither is visible in any evaluation of the change, because evaluations are conducted using the compressed output.
What the scholarly case proves and what it does not
It proves that a large, sophisticated, self-aware system with strong professional norms and an explicit correction mechanism can run a measured error rate an order of magnitude below its own estimate of the true one, with a two-year detection lag, while the compromise sits in the verification layer.
It does not prove that the same is true elsewhere. It proves something more useful. Science is the domain with the best correction apparatus, not the worst. It publishes its retractions, funds databases that track them, and holds congressional hearings about them. The 0.2 percent against an estimated 2 percent is the performance of a system that is genuinely trying. Domains with no retraction mechanism are not doing better. They have no number at all, and the absence of a number is routinely read as the absence of a problem.
5. What Most Analysis Gets Wrong
That the answer is better metrics
Metric design was the right response when decay took years and revision took months. Both timescales have moved, and in opposite directions. A better measure now buys a shorter reprieve than the last one, and the effort spent designing it is effort not spent on the thing that still works, which is provenance.
That this is about dishonesty
The mechanism needs no bad actors. A system optimising against a proxy will drift from the underlying goal even when every participant is acting in good faith, because the proxy is what the environment rewards and the goal is not directly observable. Attributing the drift to misconduct produces enforcement responses to a design problem, and enforcement is the most expensive and least effective available answer.
That more data improves the picture
More data through the same compressing stack produces more confidence in the same distortion. The binding constraint is not volume, it is independence — whether any channel exists that the measured party does not produce. One independent channel is worth more than a hundredfold increase in instrumented data through the existing one, and it is the thing nobody funds, because it produces bad news about its own sponsor.
That the fix is to distrust the numbers
Blanket scepticism is as useless as blanket credence and considerably more comfortable, because it requires no work. The numbers are not uniformly compromised. They are compromised in proportion to how cheap they are to produce relative to the thing they measure, and that ratio is assessable case by case. The useful posture is neither trust nor distrust but a specific question about production cost, which is set out in section 12.
6. Base, Stress and Extreme
Four paths, with our probability assessment and the condition that would falsify each. Probabilities sum to one hundred.
Path | P | What it looks like | What would falsify it |
Metric churn | 50% | Institutions cycle measures faster, each replacement degrading sooner than the last. Reported performance stays strong throughout | A measure introduced after 2026 holding its relationship to outcomes for more than five years under active optimisation |
The provenance turn | 25% | Verification shifts from inspecting output to attesting process: preregistration, workflow records, machine-verifiable chains of custody | Provenance schemes remaining voluntary and unadopted by any major funder or regulator through 2030 |
Measurement retreat | 20% | Institutions stop publishing figures that embarrass them. Opacity replaces bad news, and comparison becomes impossible | Publication of self-critical statistics expanding rather than contracting across major institutions |
Restored verification | 5% | Independent checking is funded at a scale that closes the gap between measured and actual error | This is the falsifier for the article rather than a scenario needing one |
The third path deserves attention it will not get. A retreat from measurement produces no headline, no incident and no identifiable victim. It looks like a change in reporting practice, and it removes the only evidence that would have shown the first path operating.
7. Forecast — One Year, to mid-2027
The recorded error rate keeps rising, and remains far below the estimated true rate
Probability 0.75 · Confidence: Medium-High
We expect the measured retraction rate to hold at or above its 2025 level and probably to rise, while remaining several multiples below the roughly two percent that the people maintaining the database believe to be correct.
Both halves of that matter. A rising measured rate will be read as a worsening problem; it is at least as likely to be improving detection working through a backlog. A rate that rises while the gap to the estimated true value stays wide indicates detection improving more slowly than production, which is the actual finding.
Second-order effect. Rising recorded error becomes an argument against the institutions doing the recording. The bodies that publish their own failure rates look worse than those that publish nothing, and the incentive that creates is the mechanism behind the third scenario.
What would weaken it. The measured rate falling while submission volume continues rising, which would indicate either genuine improvement or a detection process that has stopped working, and the two would need separating.
8. Forecast — Three Years, to 2029
Provenance requirements arrive in at least one major system
Probability 0.60 · Confidence: Medium
We expect at least one major funder, publisher or regulator to require machine-verifiable evidence of process — preregistration, workflow attestation, chain of custody — as a condition of acceptance in a named programme, rather than relying on inspection of the output.
The logic that forces this is simple and applies well beyond research. Where output can be generated cheaply and convincingly, inspecting output stops discriminating. What remains discriminating is evidence about how the output came to exist, because that evidence has to be produced continuously while the work happens rather than assembled afterwards.
Second-order effect. Provenance requirements favour well-resourced institutions with existing infrastructure and disadvantage individuals and smaller groups, whose work is not less genuine but is less instrumented. This is a new threshold in the sense of Article 3, and it will subtract variety in exactly the way described there.
What would weaken it. Continued reliance on output inspection with better detection tools, and provenance schemes remaining voluntary and unadopted.
9. Forecast — Five Years, to 2031
The pattern is documented outside research
Probability 0.55 · Confidence: Medium
By the early 2030s we expect the same structure — cheaply generated submissions, a compromised or overwhelmed verification layer, and a measured failure rate far below the estimated one — to be documented in at least one large non-research domain with published figures. Insurance claims, credit assessment, benefits administration, procurement and content moderation are all candidates.
The forecast is really about disclosure rather than about occurrence. The mechanism does not require research-specific conditions and there is no reason to expect it to be confined there. What research has, and these domains do not, is a public correction record. The forecast resolves on whether any of them acquires one.
Second-order effect. The first domain to publish comparable figures will look worse than all the others, for the same reason science does now. That penalty on transparency is the single strongest force preventing the measurement from existing, and it will need to be offset by regulation rather than by good intentions.
What would weaken it. Non-research domains publishing verification-failure statistics that show low and stable rates under rising submission volumes, verified independently.
10. Forecast — Ten Years, to 2036
Verification splits from assessment and becomes a separate function
Probability 0.45 for a pronounced version · Confidence: Medium-Low
Over a decade we expect verification to separate institutionally from the bodies whose work it checks — funded differently, reporting differently, and structurally unable to be cut by the party it embarrasses. The auditor model, extended into domains that currently self-certify.
The mechanism forcing it is the one in section 2: an independent channel is the first thing cut and the last thing defended, so it survives only where it does not depend on the goodwill of its subject. Every durable verification institution in history has that property, and every one that lacked it was eventually defunded.
We hold this at lower confidence than the probability suggests. Institutional separation of this kind has historically followed a scandal rather than an argument, and scandals are not forecastable. The mechanism is sound and the timing is not.
What would weaken it. Verification remaining embedded in the assessed institutions while measured error rates converge on independent estimates, which would indicate that internal checking can work under cheap optimisation after all.
11. Signals to Watch
— The gap between measured and estimated error rates, wherever both exist. The level tells you little; the gap tells you whether detection is keeping pace
— Detection lag: time from occurrence to recognition. Two years means today's report describes a process that has since changed
— Which layer failures are attributed to. A shift from output failures to verification failures is the transition described in section 4
— Funding and staffing of independent checking functions, measured against the volume they check rather than in absolute terms
— Institutions ceasing to publish self-critical statistics. The moment of cessation is more informative than any figure published before it
— Adoption of provenance requirements: preregistration, workflow attestation, chains of custody replacing output inspection
— Whether any metric introduced after 2026 holds its relationship to outcomes for five years. We are keeping this list ourselves and it is currently empty
12. Recommendations — Individuals
The practical problem is that you are surrounded by numbers of wildly varying reliability, presented identically. Sorting them is a skill, and it reduces to one question.
Immediate — 30 days
For any figure you rely on to make a decision, ask what it cost to produce the number relative to what it cost to produce the thing the number describes. Where producing the number is much cheaper, treat it as a claim. Where the two are close, treat it as evidence.
This single question sorts most of what you encounter, and it does not require expertise in the subject. A rating that anyone can generate, a credential anyone can print, a review anyone can post and a metric a party reports about itself all fail it. An audited figure, a physically verifiable measurement and a number that took real work to fake all pass.
Build — 12 months
Maintain one independent channel for anything that matters to you. Not a second source that draws on the first — that is the nominal diversification this series keeps warning about — but a channel with a different production process. For a professional field, someone who does the work rather than someone who writes about it. For a financial position, a figure produced by a party with no stake in it.
Then check the provenance question on your own credentials and records. As verification shifts from output to process, evidence of how you did something becomes worth more than evidence that you did it, and process evidence cannot be reconstructed after the fact.
Position — 3 years
Assume the value of unverifiable claims falls toward zero and the value of verifiable ones rises sharply. Where you can arrange your work so that it leaves a trail — a record, an attestation, a public commitment made in advance — that trail is becoming the asset, and it costs almost nothing to create while the work is happening and cannot be created afterwards.
Avoid. Responding with blanket scepticism. It feels rigorous and it is a way of avoiding the work of sorting. Everything becomes equally doubtful, nothing is assessed, and decisions get made on instinct instead — which is a worse instrument than a mediocre number you have correctly discounted.
Why this works. You cannot repair anyone else's verification layer. You can grade the numbers you rely on by production cost, hold one channel that is genuinely independent, and leave a trail while it is still cheap to leave one.
13. Recommendations — Business
Most organisations have a measurement stack and no view of its reliability. The measures are reviewed for whether they are the right ones and almost never for whether they still work.
Immediate — 60 days
Take the measures your decisions actually run on and grade each by production cost. What does it cost the reporting party to produce this number, against what it costs to produce the underlying performance? Where the ratio is far below one, the number is a claim about intent rather than a measurement of outcome, and it should be labelled as such wherever it appears.
Then locate your verification layer and ask who funds it, who it reports to, and whether it can be reduced by the function it checks. If the answer to the last is yes, you do not have a verification layer. You have a quality-assurance function that survives at the pleasure of its subject.
Build — 12 months
Establish one channel that is structurally independent of the reporting line. Sampled direct observation, an external audit with genuine access, a customer channel that does not pass through the team being assessed. It does not need to be large. It needs to be outside the loop, and it needs a budget that the assessed function cannot influence.
Instrument detection lag as a metric in its own right. How long from a problem occurring to someone with authority knowing? This is the single most diagnostic number available about a measurement system, and almost nobody computes it, because it requires reconstructing the history of problems that were eventually found.
Position — 3 years
Shift from inspecting output to attesting process wherever the output can be cheaply generated. Supplier claims, contractor qualifications, compliance submissions and candidate credentials are all in this category already. The question is moving from is this document genuine to what had to happen for this document to exist, and building the systems for the second takes years.
Plan on the assumption that your own measures will be optimised against faster than they were. Set a review cycle for each measure and diary it, rather than reviewing when a number looks wrong — by then the measure has been decoupled for some time and the decisions taken on it are already made.
Avoid. Cutting the independent channel because it keeps producing problems. Its output is indistinguishable from a problem it created, which makes it the easiest line to cut and the most expensive to have cut. An organisation whose quality function reports no issues is either exceptional or blind, and the two look identical from the board.
Why this works. The concentration reading on institutional self-knowledge is usually one, and the single channel is frequently produced by the party being assessed. Adding one genuinely independent channel takes concentration from one to two, and Article 1's rule applies here as everywhere: the move from one to two changes what is survivable, and the move from two to ten adds very little.
14. Recommendations — Capital
Every position rests on numbers produced by parties with a stake in them. That is not a scandal, it is the normal condition, and it is manageable only if the grading is explicit.
Immediate — this quarter
Grade the inputs to your largest positions by production cost. Audited financials, physically verifiable output and settled transactions sit at one end. Self-reported operating metrics, forward guidance, sustainability disclosures and any figure generated by the party it flatters sit at the other. Both currently arrive in the same deck, in the same font.
Build — 12 months
Look for exposure to metric decay specifically. Where a valuation depends on a measure that is newly important, cheap to influence and central to how a business is assessed, that measure is under active optimisation and its relationship to underlying performance is degrading on a schedule.
Then check verification concentration across the portfolio. Positions in different sectors can rest on the same auditor, the same rating methodology, the same data vendor or the same certification body. That is a single upstream origin behind nominally independent assessments, and it is the same failure this series has described in supply chains, in institutions and in its own forecast portfolio.
Position — 3 years
Watch for the verification premium. As unverifiable claims lose value, assets whose performance is expensive to fake should widen their spread over assets whose performance is cheap to report. Physical output, settled cash flows and independently audited positions are on one side; self-reported engagement, pipeline and impact metrics are on the other.
And watch for the retreat. An issuer or sector that stops publishing a self-critical statistic has told you something, and it will not appear in any dataset because the datum is the absence.
Avoid. Treating a rising recorded error rate as deterioration. It is at least as likely to be improving detection, and the party that improved its detection is now penalised against peers who did not. Reading that penalty backwards is a common and expensive error.
Why this works. Capital's advantage is reallocating before a repricing. Metric decay is unusually forecastable — it follows the production-cost ratio, which is assessable from outside — and it reprices only when a failure becomes visible, which is years later.
15. What Would Change Our Mind
Each forecast carries its own weakening condition. Three developments would undermine this article's argument as a whole.
— A measure introduced after 2026 holds its relationship to outcomes for more than five years under active optimisation, in a domain where generating the appearance of performance is cheap. That would falsify the collapse in metric half-life directly.
— Verification failures decline as a share of recorded failures while volume rises. That would indicate the check is holding and the attack has not moved to it.
— Institutions demonstrably fund independent verification against their own performance at a scale that closes the measured-to-estimated gap. The structural argument in section 2 says this does not survive budget pressure; two durable counter-examples would show it can.
A correction to a limitation recorded in Article 4, and its result. The previous article noted that seven of eight forecasts across Articles 3 and 4 resolved against European institutions, and that the fix was to find non-European case material of comparable documentary quality rather than to add weaker examples for balance. This article's evidence and all four of its forecasts come from the international scholarly record and a United States congressional proceeding. The concentration is corrected for this article and not yet for the series, where the running total now stands at seven of twelve. We will keep reporting it until it is no longer worth reporting.
Founder's Lens
[ EDITORIAL GATE — WRITTEN BY HAND BEFORE PUBLICATION. Never generated. Replace this marker with the founder's text, or record a suspension. ]
16. Bottom Line
A measurement inside an institution is a control signal, produced by parties who know it will be acted on. It holds while producing the appearance of performance costs roughly what producing the performance costs. When the first falls sharply and the second does not, the signal decouples — and the decoupling makes the numbers look better, which is why nothing triggers concern.
Two forces are driving it. The working life of any given measure is collapsing, because every step of optimising against it has become cheap while the institutional revision cycle has not. And the attack has moved from the output to the check, which means the system's report on its own condition is now produced in part by the component that failed.
The one large system that publishes its own error corrections records a rate of about 0.2 percent against its own custodians' estimate of two percent, with roughly two years between occurrence and detection. That is the performance of the domain that is trying hardest. Domains without a correction mechanism are not doing better; they have no number, and no number reads as no problem.
The response is not better metrics, which buy a shorter reprieve each time, and not blanket scepticism, which is a way of avoiding the work. It is grading by production cost, holding one channel that is genuinely independent of the party being assessed, and shifting from inspecting output to establishing how the output came to exist.
Article 4 asked whether a system with acquired capacity can still find out that it is wrong. The answer is that it can, and only through a channel it does not control, cannot cut, and will have no reason to build until it has already been wrong at scale.
The most dangerous institution is not the one with poor information. It is the one whose information is excellent, abundant, current, and produced by the thing it is measuring.
Forecast record
Four forecasts, one per horizon, each with a threshold, a named verifier and a resolution date, recorded before the outcome is known.
Horizon | Forecast, resolving yes or no | P | Resolves |
1 year | The recorded retraction rate for 2026 across the scholarly literature is reported at or above 0.2 percent | 0.75 | 31 December 2027 · Retraction Watch database or successor |
3 years | At least one major research funder or publisher requires machine-verifiable process evidence — preregistration, workflow attestation or chain of custody — as a condition of acceptance in a named programme | 0.60 | 31 December 2029 · funder or publisher policy documents |
5 years | A large non-research domain publishes a comparable verification-failure statistic covering submissions or claims at volume | 0.55 | 31 December 2031 · regulator or industry body publication |
10 years | A verification function in at least one major domain has been separated institutionally from the bodies whose work it checks, with independent funding established in law or regulation | 0.45 | 31 December 2036 · legislation or regulatory framework |
Correlation, recorded rather than assumed away. The first and third share a parent cause in whether verification-failure statistics get published at all, which is itself one of the four scenarios. The second and fourth share a parent cause in the provenance turn. Two pairs rather than four independent observations. The set is not concentrated by jurisdiction, which corrects the limitation recorded in Article 4, and is concentrated by mechanism, which is a different weakness and one we are recording here rather than discovering later.
Directional statements elsewhere in this article carry no threshold and are deliberately excluded from the record.
Sources
Figure | Class | Source |
Retraction rate around 0.2% in 2025, up from about 0.02% in 2016 | Measured | Retraction Watch, reported to a US congressional hearing, April 2026 |
Retraction Watch's stated confidence that the correct rate should be about 2% | Estimate, by the database custodians | Same source |
2025 retraction reasons: compromised peer review and paper mills about 43%, compromised peer review alone about 37%, citation manipulation about 23% | Measured | Publisher analysis of 2025 retraction data, July 2026 |
Over 6,400 retractions attributed to fake peer review across 2023–2025; around 2,100 associated with AI-generated content | Measured | Bibliometric analysis of the Retraction Watch database |
10,409 retracted paper-mill articles from database inception to 2024, with a spike in 2023 | Measured | Scientometrics, bibliometric study of the Retraction Watch database |
Retraction time: near one year in 2019, about 2.5 years in 2021–22, sharply lower in 2023, around two years in 2024 | Measured | Same source |
One figure in this table is of a different kind and is marked accordingly. The two percent is an estimate by the people who maintain the database, not a measurement. We use it because an explicit statement by custodians that their own figure understates the phenomenon is unusual, informative and central to this article's argument. It is not a substitute for a measurement, and the article's case does not require the specific number — only that a substantial gap is acknowledged by the parties best placed to assess it.
In this series
— Previous: Article 4, When Capacity Concentrates — what changes when the historical limit on the centre stops binding.
— Next: Article 6, Adaptive vs Rigid — why the useful question about a system of government is not the one usually asked.
— The method behind the Chaos Index and this series: /methodology
THRIVE IN CHAOS
Decision Intelligence for an Uncertain World
Analysis → Forecast → Recommendations · Signal → Meaning → Action → Stability
Signal Over Noise · thriveinchaos.ai
AI intelligence system with human editorial oversight.
Forecasts are probability-based analytical assessments, not certainties. This material supports independent judgment and does not constitute financial, legal or investment advice.
Join the newsletter
Be the first to read our articles.


