In 2022 I began two short working notes. One asked what Feynman and Shannon had missed about a scientific system connected by the Internet and pushed by hypercompetitive incentives. The other tried to separate mathematical illusions, collective hallucinations, delusions, and fraud. They were intellectual seeds, not evidence. They were fragmentary and sometimes used language too intuitive or accusatory to carry a serious argument.
This article keeps the question and replaces the shortcuts.
Large language models make the surface of expertise cheap: fluent prose, disciplinary vocabulary, citations, equations, code, diagrams, peer-review language, and the ritual structure of an experiment. That can expand human capability. It can also reduce the information carried by scientific form. A polished paper once weakly signaled that someone had paid a substantial production cost. The signal was never reliable. Now its cost approaches zero.
The institutional response cannot be “ban the tool.” It must be to make the evidence state explicit.
Article 15 asks which inherited claims deserve audit and how institutions should retire epistemic debt. This article asks a different production question: when machines can emit scientific form at negligible marginal cost, what evidence-bearing object distinguishes a claim from a performance of a claim?
Shannon: the bandwagon becomes a generator
In “The Bandwagon”, Claude Shannon warned that information theory’s success had attracted weakly justified applications and relabeling. He called for rigorous research and for the field to retain contact with the problems its formalism could actually solve. D
The warning concerns label inflation. A powerful field supplies language, prestige, venues, and funding categories. Researchers have incentives to describe adjacent work through that label even where the technical connection is thin.
An LLM changes the cost structure. It can borrow the vocabulary, citation pattern, and argumentative shape of every fashionable field simultaneously. A human bandwagon still required researchers to learn enough of the style to participate. A model can generate the style on demand.
This does not mean the resulting work is false. It means vocabulary carries less evidence of conceptual contact. Institutions must test the mapping rather than reward the label.
Feynman: the correction loop, not the costume
In “Cargo Cult Science”, Richard Feynman described inquiry that reproduces the outward apparatus of science while omitting the discipline of reporting facts that could make the claim wrong. His standard was not merely methodological ritual. It was an unusually complete honesty about alternative explanations, prior failures, and conditions under which a result should not be trusted. D
An LLM has no personal integrity to exercise and no private motive to conceal. It produces under an objective supplied by training and evaluation. The responsibility moves outward—to the workflow, reward rule, provenance system, and accountable humans.
If a system rewards plausible answers and penalizes abstention, fluent guessing is adaptive. Kalai and colleagues argue that next-word pretraining lacks truth labels for many low-frequency facts and that accuracy-only evaluations can reward guessing rather than calibrated uncertainty. Human publication systems can create an analogous pressure: confident novelty earns credit, while negative results, replication, and “we do not know” struggle for space. D
The machine does not introduce the incentive. It scales the response to it.
Wigner: valid mathematics still needs a referent
Eugene Wigner’s “The Unreasonable Effectiveness of Mathematics in the Natural Sciences” celebrates a genuine mystery: mathematical structures developed in one setting can later describe physical phenomena with astonishing accuracy. Recruiting the essay into a general warning against mathematics would reverse its meaning. D
Wigner’s mystery nevertheless exposes three separate tests:
- Is the mathematical derivation valid?
- Do its objects and assumptions identify the relevant physical system?
- Do observations support the model within stated conditions?
An LLM can generate mathematical-looking prose that fails the first test. It may generate a valid derivation that fails the second. A beautiful model can pass the first two as a hypothesis and fail the third. Mathematics with no current physical application is not defective; pure mathematics is judged by its own questions. The error is claiming empirical authority from formal elegance alone.
The data turn was real
Halevy, Norvig, and Pereira’s “The Unreasonable Effectiveness of Data” argued that very large datasets with comparatively simple methods could outperform smaller, more elaborately modeled approaches in language tasks. Modern AI vindicated much of that scaling intuition. D
The success does not imply that predictive effectiveness, causal explanation, mathematical validity, and scientific truth are the same achievement. A system may predict useful text without representing why the claim is true. It may reproduce the consensus accurately where the consensus is wrong. It may combine sources that share one hidden assumption and present their agreement as independence.
The inversion is now complete:
- Shannon worried that people would borrow one successful field’s language; a model can borrow every field’s language.
- Feynman worried that people would reproduce scientific form without its correction discipline; a model reproduces form without possessing intentions.
- Wigner marveled that mathematics could map nature; a model can generate a mathematical map before anyone establishes the territory.
A taxonomy that prevents accusation by metaphor
Not every epistemic failure is a hallucination, and error is not evidence of fraud.
- A mistake is a corrigible false statement or invalid step.
- A mathematical illusion predictably exploits intuition or an omitted condition while appearing valid.
- A model hallucination is plausible output unsupported by the model’s available evidence.
- Cargo-cult science is a process that preserves scientific form while lacking a reliable error-correction mechanism.
- A bandwagon effect expands a rewarded label beyond its evidentiary warrant.
- Mathematical overreach uses valid formal reasoning with unsupported assumptions or an unestablished mapping to the world.
- Fraud is intentional deception and requires evidence of intent.
A retraction can result from mistake, misconduct, or other causes. A failed replication can expose falsehood, boundary conditions, tacit technique, or incompatibility between protocols. AI use establishes none of these by itself.
This taxonomy is a control against personal bashing. Criticism should identify the destructive behavior or missing correction mechanism and let the documentary record carry the story.
From form to witness
The scarce output in AI-assisted science will not be hypotheses or polished manuscripts. It will be witness-bearing claims: claims attached to artifacts that let an independent process check an important part.
Witnesses are domain-specific:
- theorem → proof object plus a checked match between formal and informal statements;
- software claim → executable environment, tests, and failure conditions;
- data claim → immutable data, provenance, transformations, and sensitivity analysis;
- security claim → explicit model, reduction, exploit, or adversarial evaluation;
- empirical claim → protocol, calibrated instruments, complete selection record, and independent measurement.
FormalMATH demonstrates the difference. The benchmark contains 5,560 formally verified Lean 4 problems, and a deterministic checker can reject an invalid proof regardless of rhetorical quality. The best reported prover solved 16.46 percent under the authors’ budget. Formal success is real evidence; failure to solve is not evidence that the theorem is false.
Here the benchmark is not a census of mathematical reliability or an audit of published literature. It isolates the form/checker boundary: fluent mathematical text and a checker-accepted proof object are different evidence states.
My own work under review on composed generative models belongs in the weakest evidence class here. The author-supplied record reports calibration-closure results for several composition operators, suggesting that combining models in natural ways does not automatically escape a hallucination floor. Until the manuscript and proof are independently audited, it is provisional evidence, not an established theorem. Its useful boundary is constructive: a checkable witness changes the evidence state; another confident model opinion may not. D
A multiplicative diagnostic
One conceptual model is deliberately unforgiving:
Epistemic value = formal validity × empirical adequacy × provenance × adversarial exposure.
This is not a numerical formula or a universal definition of knowledge. It expresses a bottleneck. Excellence in one dimension cannot compensate for zero in another when the claim requires all four. Beautiful mathematics does not rescue an ungrounded physical identification. Perfect provenance does not rescue invalid reasoning. Repeated agreement does not substitute for adversarial exposure when every evaluator inherits the same corpus.
Exploratory work may legitimately have unknown empirical adequacy or incomplete adversarial exposure. Its label should say so. The standard strengthens when a claim becomes settled, safety-critical, or deeply depended upon.
The strongest counterargument
Scientific conventions are not empty costumes. Shared form lets experts inspect complex work. AI can make notation consistent, find missing citations, generate counterexamples, write tests, and bring formal tools to researchers who could not otherwise use them. Demanding witnesses for every statement would crush speculative thinking.
Agreed. The design needs two channels.
The exploratory channel welcomes conjecture, analogy, simulation, and generated possibilities, with visible uncertainty and provenance. The evidentiary record contains claims whose status is tied to domain-appropriate witnesses and correction history. Work can move between channels as evidence changes.
Peter Denning’s “Cargo Cult AI” already applies Feynman’s critique to the distinction between generating convincing forms and conducting falsifiable inquiry. The useful move is not another insult. It is institutional architecture that makes the correction loop observable. D
Five different verdicts
- Scientific success: AI can expand conjecture and expose error; success depends on separating proposal from verified result.
- Technical success: Formal checkers, executable artifacts, provenance systems, and automated experiments can make more claims inspectable.
- Transition success: Witnesses matter only if journals, funders, standards bodies, and deployed systems update when evidence changes.
- Institutional success: The laboratory must reward abstention, replication, correction, and preservation of failed tests—not output volume alone.
- Public-value success: Faster generation can democratize expertise or flood the commons; the outcome follows governance and incentive design.
What the successor must learn
An AI-native laboratory should not measure intelligence by how much scientific form it produces. It should measure how efficiently it turns uncertainty into witness-bearing, adversarially exposed claims—and how gracefully it retracts or repairs them.
That requires named human accountability, audit trails, model and dataset dependency maps, independent checkers, and rewards for discovering that an attractive answer is unsupported.
When the form of expertise becomes nearly free, the open question becomes the institution’s defining choice: what evidence will remain costly enough to trust, and who will be rewarded for producing it?