The successor to Bell Labs is not Bell Labs with GPUs.
It cannot be recreated by naming a building, aggregating several grants, or buying a frontier model. The twentieth-century laboratory depended on an economic and regulatory settlement that no longer exists in the same form. Rebuilding its surface without rebuilding its incentives would create an expensive stage set.
The objective is different: construct an institution that can carry important uncertainty across projects and transitions while remaining publicly accountable, technically corrigible, and connected to use.
AI changes what is possible. It does not decide why the institution should exist.
It is an amplifier, not a directional force. Under incentives that reward correction, provenance, and hard transition, AI can scale productive search and verification. Under incentives that reward volume, confidence, and short-cycle visibility, it can scale plausible noise and common-mode error. Institutional design selects which effects become durable.
Start from the failure modes
The preceding cases identify a set of constraints.
- A corporate parent funds what it can capture and is exposed to its business cycle.
- A government program can coordinate extraordinary temporary capability but does not automatically preserve it.
- A nonprofit institute can connect projects while its people remain governed by chargeability and funding cliffs.
- A university produces open knowledge and talent but does not own the complete systems-and-transition stack.
- An FFRDC can preserve sponsor memory and engineering capability while risking task-order narrowing or incumbent capture.
- A startup can productize a focused invention but cannot rationally maintain a broad research commons.
- Output metrics can remain healthy after the capability graph has fractured.
- AI can scale either witness-bearing correction or scientific-looking noise.
The proposed institution must not declare one of these forms the winner. It must combine their complementary strengths while keeping their failure modes visible.
The proposed form
Create a federally chartered, independently governed nonprofit consortium with a long but reviewable mission charter. Federal anchor funding supplies continuity and public purpose. Industry members supply domain problems, engineers, data under controlled conditions, manufacturing, and transition paths. Universities supply open inquiry, students, and disciplinary depth. Philanthropic capital protects a bounded speculative portfolio. Allied participation expands talent and interoperability where openness and security permit.
The charter should last long enough to support careers and facilities—longer than a program or fund—but contain five-year external reviews and a ten-year reauthorization decision. Durability without review becomes entitlement. Review without horizon recreates the proposal treadmill.
No payer should control the whole portfolio. A single dominant sponsor would gradually turn the laboratory into its engineering arm. Funding diversity is not a virtue by itself; it is protection against one clock becoming the institution’s constitution.
Existing components prove feasibility, not completion
The United States already operates relevant pieces.
NCSES reports that 42 FFRDCs performed $31.7 billion of R&D in fiscal 2024, up from $17.7 billion in 2014. Federal sources supplied 98.5 percent of the 2024 total. The 2024 increase includes a reporting-method change at one laboratory, so it should not be read as pure capability growth. The figures establish scale and durability, not whether every center thinks long. D
The February 2026 master list contains 41 current centers, while the expenditure survey reports 42 FY2024 performers. These are different dates and populations, not conflicting counts. The current centers also span 26 R&D laboratories, 10 study-and-analysis centers, and five systems-engineering-and-integration centers; a successor should learn from each category without pretending they have one charter. D
The Federal Acquisition Regulation defines the special sponsor relationship in terms of long-term need, continuity, independence, access, and work that cannot be met as effectively by ordinary resources. GAO’s historical account emphasizes institutional memory and trusted objectivity. These are essential design elements.
The concern about FFRDC decline must remain specific. GAO’s DHS review documents a bounded case in which task-order overlap and weak tracking of sponsor use created management problems. It does not prove that all FFRDCs are “engineering houses at best.” Systems engineering, test, integration, and acquisition support can be core public capabilities. The failure is allowing deliverable service to eliminate original technical judgment, independent challenge, apprenticeship, or reusable research state.
DARPA’s program-manager model demonstrates empowered, fixed-term technical leadership. The National Academies’ ARPA-E assessment demonstrates a related active-management model. The successor should borrow authority and urgency without becoming another temporary network.
NAIRR demonstrates shared public–private access to compute, data, models, tools, and training. At its two-year mark NSF reported more than 600 research teams and 6,000 students across every state, the District of Columbia, and Puerto Rico. NSTC is designed as a public–private semiconductor consortium spanning research, prototyping, investment, and workforce. DOE’s Technology Commercialization Fund demonstrates cost-shared lab-to-market partnership. Each solves part of the graph. None owns the whole path from uncertain question to retained public capability. D
Six layers, one accountable institution
1. Institutional memory
The laboratory records hypotheses, decisions, alternatives, experiments, failures, code, proofs, data provenance, and transition outcomes as versioned state. Final papers remain important but are not the archive.
Every major claim has a dependency record. Every completed project has a continuity disposition. Every critical tool has a named custodian or an explicit end-of-life decision.
2. Scientific agents
AI systems assist literature synthesis, conjecture generation, software, proof, simulation, experimental planning, anomaly detection, and adversarial review. They do not receive authority by fluency. Consequential outputs must cross domain-appropriate checkers or produce witnesses.
Model and dataset dependencies are recorded so agreement among agents sharing one corpus is not mistaken for independent replication. Abstention and calibrated uncertainty receive credit.
3. Secure shared resources
The institution maintains compute, data access, reproducible environments, privacy-preserving analysis, identity and authorization, export and security controls, and auditable model use. Resources should be portable enough that no one vendor becomes the institution’s hidden sovereign.
Controls must be tiered. Public data should not inherit classified handling because one program is sensitive. Controlled work should not be forced into false openness. The default is the least restrictive tier consistent with evidence and law.
4. Physical interface
AI confined to text can optimize representations without contact with reality. The laboratory needs automated experiments, testbeds, semiconductor access, fabrication, robotics, sensors, and measurement. Physical and operational evidence are correction channels.
Not every site needs every instrument. A federated system can work if access, scheduling, calibration, data provenance, and technical staff are durable rather than reconstructed by project.
5. Human apprenticeship
Junior and senior researchers share responsibility for consequential outcomes over multiple years. Seniority does not confer veto power. The point is transfer of problem selection, experimental taste, systems judgment, and ethics through work that can fail.
Technical careers must not force every excellent researcher into people management. Program leaders who choose major bets need practiced discovery and stewardship records, while fiduciary, legal, safety, and public-accountability authority remains independent.
6. Transition
Product engineers, mission users, standards experts, procurement officers, manufacturing partners, licensing staff, and venture pathways enter before the final demonstration. Every program identifies its complementary assets and who has authority to adopt the result.
Transition is funded separately. It is neither a promise in the proposal nor a technology-transfer office encountered at the end.
Governance: couple authority without concentrating it
The governing board should contain public-mission representatives, active technical practitioners, industry and transition expertise, workforce and public-interest voices, and independent fiduciary and safety members. No constituency should hold a majority.
Technical portfolios should be led by people with demonstrated research judgment and multi-year program stewardship. Counts of papers or patents can signal experience but cannot serve as a universal gate. Fixed terms, conflict disclosure, plural fields, external replication, protected dissent, and retrospective decision audits limit expert capture.
The board sets mission, risk appetite, and resource boundaries. It should not select individual technical approaches. Program leaders select portfolios within the charter and must document the evidence that changed their decisions.
An independent verification office reports both to technical leadership and an audit committee, with authority to publish bounded correction records. Security and legal review can block unsafe release but must record the category of restriction so secrecy does not become an unreviewable veto.
The decision-rights constitution
“Good governance” is not a control surface. The charter must say which body can make which decision, what evidence it owes, and which decisions it cannot make.
| Body | May decide | May not decide alone | Required trace |
|---|---|---|---|
| Governing board | Mission boundaries, capital envelope, risk appetite, executive appointment, reauthorization, and dissolution | Individual scientific conclusions or preferred technical approach | Public mission rationale, conflicts record, minority views, and five-year capability verdict |
| Technical council | Portfolio hypotheses, program creation, major technical resource allocation, and program termination recommendations | Its own verification grade, security law, or member-specific commercial terms | Alternatives considered, evidence that changed the decision, and retrospective calibration |
| Program lead | Team formation, staged bets, milestones, competing paths, and bounded stops within an approved program | Suppressing a correction, extending the charter, or granting exclusive institutional rights | Versioned decision log, kill criteria, uncertainty register, and transition assumptions |
| Verification office | Reproduction status, replication protocol, artifact sufficiency, dependency warnings, and correction notices | Mission value, program continuation, or allegations of intent without evidence | Typed witness, provenance, common-mode analysis, appeal record, and downstream notification |
| Security and legal office | Access tiers, statutory restrictions, safety holds, and release conditions | Indefinite secret veto without a review category and date | Authority, threat model, narrowest restriction, review date, and appeal route |
| Transition owner | Adoption plan, operational acceptance, manufacturing or procurement path, and maintenance obligation | Scientific truth status or control of unrelated research directions | Complement map, adoption evidence, full lifecycle cost, and reasons for rejection |
| Public-interest trustee | Enforcement of public-use, access, march-in, and asset-transfer terms | Technical selection or day-to-day management | Distributional impact, access terms, exercise or waiver rationale, and public-value account |
The separations matter. Verification can lower a claim’s evidence status without automatically killing an exploratory program. A transition partner can reject a system as unusable without declaring its science false. A security office can delay release without erasing the existence of a result. The board can dissolve an institution without rewriting its technical record.
Every consequential decision receives a time-bounded appeal to a body that does not report solely to the original decision maker. Appeals should expose reasons, not create endless process. Emergency safety authority remains possible, but expires into an ordinary evidentiary review.
The capital stack
Different money buys different obligations.
- Federal anchor funding buys core teams, facilities, archives, and public mission continuity.
- Competitive agency programs buy ambitious bounded objectives and external technical pressure.
- Industry membership buys access to precompetitive capability, problem definition, rotations, and transparent licensing options.
- Procurement and advance commitments buy a transition path for validated capability.
- Philanthropic funding buys a protected fraction of unusually speculative work within the charter.
- Licensing and venture returns recycle some captured value without becoming the laboratory’s sole survival condition.
These streams must be reported separately. Otherwise project revenue can masquerade as capability investment and license revenue can quietly redefine the research agenda.
Capability accounts beside financial accounts
The laboratory needs two ledgers. The financial ledger records legal stewardship of money. The capability ledger records whether the institution is renewing the state on which future work depends.
Capability margin = verified renewal of reusable capability − observed capability depreciation.
The expression is an accounting identity to organize evidence, not a promise that unlike capabilities can be collapsed into dollars or one score. Each major capability keeps a record across seven accounts:
- Teams: continuity of critical role combinations, voluntary exits, internal mobility, and independently capable successors.
- Adjacency: time to route cross-layer failures, recurring collaboration, independent challenge paths, and external weak ties.
- Assets: practical access, calibration, utilization, portability, maintainers, and replacement lead time for compute, data, instruments, and facilities.
- Memory: reuse of code, provenance, negative results, design rationales, and prior decisions—not archive volume.
- Epistemic correction: computational reproducibility, independent replication where applicable, formal alignment, red-team findings, correction latency, and downstream repair.
- Transition options: named adopters, complementary assets, procurement or manufacturing path, standards position, maintenance owner, and credible alternative users.
- Reconstitution cost: estimated time, scarce roles, dependencies, and expenditure required to rebuild after a break, with post-event estimates compared against actual recovery.
Reported papers, patents, prototypes, deployments, and revenue remain essential flow measures. The capability accounts ask whether those flows left the institution better able to attempt the next uncertain problem. A paper can renew memory; a prototype can renew a testbed; a failed experiment can renew judgment. Their capability value depends on reuse and transfer, not their label.
The accounts should remain plural. A single “capability score” would invite the proxy optimization this series criticizes. Each five-year review publishes the tradeoffs and uncertainty: which capability grew, which was consumed, which was deliberately allowed to end, and why.
Constitutional floors as testable pilot hypotheses
A charter without numbers invites waiver; universal numbers invite gaming. The pilot should therefore predeclare provisional ranges, compare them across sites, and revise them only through a recorded hypothesis test. A credible starting design would include:
- 15–25 percent of technical program resources protected for reusable tools, shared facilities, archives, apprenticeship, and researcher-initiated sponsor-relevant uncertainty rather than current deliverables;
- 5–10 percent of program resources reserved for independent reproduction, replication where meaningful, formalization, red teams, artifact repair, and negative-result preservation, separate from routine quality assurance;
- at least two materially independent technical paths through the first decisive test for high-consequence, high-uncertainty programs, unless the decision log explains why diversity would be performative;
- a named transition owner and complement map before scale-up, while allowing early exploration to proceed without a fictional customer;
- no private member majority and no single essential compute, data, or tool vendor without a tested exit path; and
- published technical-time and administrative-load distributions, with mandatory redesign if burden rises for two consecutive years without a corresponding change in risk, legal obligation, or outcome quality. A
These are not empirically optimized rates. They are deliberately falsifiable starting hypotheses. Too much protected base can shelter weak work; too much verification can starve exploration; forced duplication can waste scarce talent. The pilot compares alternative bands and reports the marginal capability gained or lost, rather than laundering the initial percentages into permanent doctrine.
The same logic applies to AI common-mode risk. A multi-agent review does not count as independent when agents share a base model, retrieval corpus, benchmark, or grader. High-consequence claims require a dependency map across models, evidence, methods, and incentives. The National Academies distinction between computational reproduction and new-data replication supplies a minimum vocabulary; the laboratory must extend it to model lineage and synthetic-data dependence.
The IP and public-value compact
Foundational tools, standards, benchmark infrastructure, and nonsensitive results should enter a research commons, sometimes after a limited lead period. Precompetitive work is shared among members under uniform terms. Privately funded company-specific development may remain proprietary.
Exclusive licenses are field-limited, time-limited, and milestone-based when public funds created the option. Rights revert or march in when a licensee warehouses the result. Dual-use work receives a controlled tier with periodic declassification or release review rather than permanent default secrecy.
The objective is not maximum openness or maximum appropriation. It is sufficient private incentive for transition plus sufficient public return to replenish the commons.
A ten-year falsifiable pilot
Begin with two or three mission areas where the coordination failure is visible, shared infrastructure matters, and outcomes can be evaluated—such as trustworthy AI-assisted science, privacy-preserving data systems, or cross-layer semiconductor and quantum assurance.
Predeclare the tests.
Annual operational review: integrity incidents, artifact reproducibility, staff continuity, resource utilization, security, and administrative burden.
Five-year capability review: new technical abilities, apprentice independence, cross-program reuse, competing approaches, transitions, corrections, and capabilities deliberately ended.
Ten-year mission review: scientific and public value, deployed systems, standards, spillovers, concentration risk, and counterfactual evidence against comparable existing programs.
Comparison cannot wait until year ten. Each pilot site should be matched before launch to existing programs or organizations with similar mission, maturity, field, and resource intensity. Where the portfolio permits, sites should randomize or phase in selected governance mechanisms—protected base time, independent verification budgets, apprenticeship overlap, or transition funding—rather than changing everything at once. Some outcomes will remain nonrandom and path-dependent; the evaluation should state that limitation instead of presenting a synthetic control as an experiment.
The comparison set must include successful current institutions, not only failures: frontier corporate laboratories, strong university centers, FFRDC R&D laboratories, national laboratories, DARPA- or ARPA-style programs, open-source scientific communities, and mission-driven startups. Current business R&D concentration makes industry indispensable to that control group. The proposal earns expansion only if it creates capability those arrangements do not, or creates comparable capability with materially greater access, resilience, correction, or public value.
The pilot fails if it becomes a captive contractor, an academic grant distributor, a venture studio, a permanent procurement incumbent, or an AI-content factory. It fails if administrative labor consumes the protected technical horizon. It also fails if it protects research without producing hard external evidence or credible paths to use.
Those are exit criteria, not rhetorical cautions. Before launch, the charter should set measurable ceilings or comparison rules for administrative time, sponsor concentration, vendor concentration, unreproduced major claims, capability reuse, apprentice independence, and the share of programs with accountable transition owners. Crossing a threshold triggers redesign, leadership review, funding reduction, or dissolution rather than an automatic plea for permanence.
The strongest counterargument
This proposal may simply aggregate every institutional virtue and assume away every conflict. Patient capital reduces pressure but can reduce urgency. Industry transition improves relevance but can narrow openness. public governance protects legitimacy but invites politics. Security enables mission work but restricts diffusion. AI improves scale but concentrates technical power.
There is no architecture that eliminates these tensions. The proposal is valuable only if it makes them governable and testable. Multiple pilot sites or teams should use different designs. Funding should be staged. External groups should replicate important results. The charter should contain a dissolution and asset-transfer plan so institutional survival is not the hidden objective.
The strongest control is competition among institutional hypotheses, not competition that forces every project to rebuild from zero.
Five prospective verdicts
- Scientific success: Does the laboratory produce important knowledge and correct itself faster than comparable arrangements?
- Technical success: Does it create working, independently testable capability across theory, software, hardware, and experiment?
- Transition success: Do mission users, standards bodies, firms, and public systems adopt and maintain selected results?
- Institutional success: Do teams, tools, memory, apprenticeship, and independent problem selection strengthen across project boundaries?
- Public-value success: Do benefits, diffusion, security, and cost justify the public and private capital without concentrating unaccountable power?
The final proposition
The United States cannot recreate the twentieth-century corporate laboratory through admiration, appropriation, or executive decree. It can still transfer living knowledge into a new incentive system.
The nearly trillion-dollar R&D economy supplies enormous activity. The missing object is a durable institution accountable for making the pieces cohere across a generation. AI makes the design more urgent because it increases both productive search and plausible noise. It can preserve memory or fossilize error; widen participation or centralize authority; accelerate proof or accelerate performance theater.
An AI-native laboratory is therefore an institution, not a model with a building. Its intelligence lies in what it remembers, what it can test, whom it trains, how it corrects itself, and whether it can carry a true but inconvenient result all the way to use.
The open question is no longer whether the old Bell Labs can return. It is whether we can build a successor before we discover that money can purchase equipment, models, and projects—but not the lost ability to make them cohere.