The most disruptive early effect of AI on science may not be a flood of discoveries. It may be a subtraction.
Models can already search literature, translate notation, generate tests, formalize parts of proofs, reconstruct data pipelines, and compare claims across thousands of papers. As these tools improve, they may expose dependencies that were never checked as deeply as later users assumed. The trusted corpus could contract before it expands.
That possibility needs a name: epistemic debt. It is the gap between the confidence downstream work places in a claim and the verification the claim has actually received.
It is not an estimated error count. A claim can carry high epistemic debt and still be true; the debt is the unsupported confidence and downstream dependence that must be retired by better evidence.
The phrase does not imply misconduct. Debt can accumulate through ordinary specialization, expensive replication, missing artifacts, compressed peer review, or reasonable trust in prior work. The danger is that publication status is silently upgraded into certainty as citations accumulate.
The number we should not claim
There is a tempting statistic that roughly one third of mathematics papers contain wrong theorems. The available evidence does not support that sentence.
Leslie Lamport examined a record described by one unusually careful algebra reviewer. Among 84 papers, 28 reviews reported an incorrect statement in a proof or result, while 11 reported an incorrect result. Lamport explicitly warns that the sample was non-random, field-specific, and subjectively classified. The 28 category includes proof errors that may be repairable; it is not “false theorem.” D
The result is important precisely when bounded correctly. It demonstrates that nontrivial errors can pass publication and that a careful reader can find them. It does not estimate the error rate of all mathematical papers, all fields, or all theorems.
Christian Greiffenhagen’s study of mathematical peer review, based on 95 interviews and more than 100 referee reports, supplies the institutional explanation. Refereeing adds confidence but is not a proof certificate. Difficult results become trusted through a longer social process of use, scrutiny, correction, and time. D
Science routinely collapses at least five epistemic states into the word known:
- published after editorial and peer review;
- checked closely by an appropriate specialist;
- formally verified or computationally reproduced against available artifacts;
- independently replicated with a genuinely independent path; and
- robust across alternative assumptions, measurements, and models.
The states are not a universal ladder. A pure theorem does not require physical replication; an experimental claim cannot be settled by syntactic proof checking. The point is to name what kind of confidence exists.
A retraction is not a field verdict
The 2018 Nature paper “Quantized Majorana Conductance” offers a documented experimental case. Nature’s 2021 retraction note records the retraction, links underlying data, and identifies a Microsoft Station Q Delft affiliation for one author. Nature’s contemporaneous report described the work as led by researchers at a Microsoft laboratory in the Netherlands and reported the authors’ concern about insufficient rigor in the original analysis. D
This record supports a specific statement: one prominent paper was retracted. It does not support an accusation against every author, Microsoft Research, quantum computing, or Majorana research as a whole.
Sergey Frolov’s Nature commentary discusses failed confirmations, alternative explanations, selective-data concerns, and papers in Science. It is an expert argument, not a publisher adjudication of every cited claim. Three categories must remain separate:
- retracted result: a publisher has formally withdrawn the paper;
- failed or contrary replication: another effort did not reproduce the claim under its conditions; and
- disputed interpretation: experts disagree about what the evidence establishes.
Conflating them creates scandal, not epistemology.
AI’s accusation changes nothing by itself
A language model can produce a fluent critique of a correct paper and a fluent defense of an incorrect one. Its output is another claim. AI becomes epistemically useful when it lowers the cost of producing a checkable witness.
For a theorem, the witness may be a Lean or Coq proof whose formal statement has been matched carefully to the human theorem. For computation, it may be executable code bound to immutable data, dependencies, parameters, and expected outputs. For an empirical paper, it may be provenance-preserving extraction, full sensitivity analysis, or an independent protocol executed against new measurements.
The distinction is severe:
Agreement among models is additional opinion unless the models produce evidence that an independent process can check.
Shared training data and benchmarks create common-mode failure. Ten agents repeating the same hidden assumption are not ten replications.
What formalization proves—and what it does not
The Liquid Tensor Experiment shows that frontier mathematics can be formalized in Lean. It also shows the cost: sustained collaboration among domain mathematicians and formalization experts. The resulting checker establishes that a formal statement follows from formal premises inside the system. It does not automatically establish that the formal statement perfectly captures every intended informal claim.
FormalMATH documents both progress and present limits. The authors assembled 5,560 Lean 4 problems with a human-in-the-loop validation process. Under the reported budget, their strongest evaluated prover solved 16.46 percent. The benchmark is not “all mathematics,” and later systems may improve quickly. The result shows that formal AI work can be measured against a deterministic checker—and that current capability is far from automatic verification of the literature. D
Its role in this article is institutional: it makes audit capacity measurable by showing what portion of a bounded formal corpus a system can solve under stated conditions. Article 16 uses the same record for a different point—the separation between scientific-looking form and a proof object accepted by an independent checker.
Formalization may also reveal that an informal proof omitted a condition while the main theorem remains repairable. That is a success, not a scandal. A good audit system distinguishes repair from collapse.
The epistemic balance sheet
A paper-by-paper correctness score would be crude and harmful. A better institution maintains a dependency-aware balance sheet for consequential claims:
- the exact claim and version;
- downstream papers, systems, standards, or investments that depend on it;
- available data, code, proof objects, and provenance;
- specialist checks, computational reproductions, and independent replications;
- known disputes, boundary conditions, and corrections;
- the cost and priority of further verification;
- the actions required if the claim changes status.
Audit priority should follow consequence, not prestige. A modest result embedded in medical software, cryptographic infrastructure, or a billion-dollar experimental roadmap may deserve more checking than a famous but isolated conjecture.
This is where institutional incentives become decisive. Verification is a public good. A verifier may spend months producing no new headline result, and a successful audit may conclude that the original work was sound. Current career systems often reward the original claim more than the confidence infrastructure around it.
The National Academies’ integrity report treats research quality as a system property shaped by stewardship, publication pressure, and institutional practice. AI can reduce verification cost. It cannot create the career reward, artifact custody, or willingness to publish negative evidence. D
The strongest counterargument
Science is already self-correcting. Most errors are repairable, irrelevant to later work, or discovered through ordinary use. A massive audit apparatus could freeze exploration, encourage adversarial gotcha work, and spend scarce experts on old claims instead of new ones.
That objection rules out universal verification. It supports risk-based triage.
Exploratory work should be allowed to be exploratory and labeled accordingly. Settled or high-consequence claims should earn their status through stronger witnesses. Institutions should reward informative failed replications and corrections so researchers do not have to choose between honesty and survival.
The audit system itself must be audited. Models can optimize for apparent errors. Formalizers can translate the wrong statement. Replicators can lack tacit technique. Public scores can punish the fields that report uncertainty most honestly. Every finding needs versioning, appeal, and a distinction between error and intent.
Five different verdicts
- Scientific success: An audit improves knowledge even when it subtracts a claim or narrows a theorem.
- Technical success: Proof assistants, reproducible environments, and provenance tools create checkable objects, with domain interpretation still required.
- Transition success: Corrections succeed only when downstream papers, standards, software, and investments update.
- Institutional success: A laboratory succeeds when it preserves the original record, rewards correction, and makes dependencies visible rather than erasing embarrassment.
- Public-value success: Verification can prevent duplicated error and unsafe deployment, but indiscriminate auditing can consume more value than it creates.
What the successor must learn
An AI-native laboratory should establish an epistemic-audit group with prestige and independence comparable to discovery teams. Its job is not to declare papers wrong. It is to make consequential claims cheaper to check, track dependencies, preserve correction history, and produce witnesses others can inspect.
The first visible result may look like scientific regression because the trusted corpus shrinks. That is the wrong accounting. Removing unsupported certainty is capability formation.
The open question is the one our current indexes cannot answer: how much of what we call knowledge is verified, how much is merely unchallenged, and who is institutionally rewarded to discover the difference?