Karim Eldefrawy

Cryptography, Cybersecurity, Privacy Computer Science

Co-founder & CTO of Confidencial.io
2017-2021: SRI (Internet Origin)
2011-2016: HRL (Part of IBM Research)
2006-2010: PhD@UC Irvine

Scientific curiosity

Article 17 of 17 · Draft

An AI-Native Laboratory for the Public Good

The successor to the old corporate laboratory should combine durable public mission, industry transition, practiced technical judgment, and AI-enabled verification without recreating monopoly or bureaucracy.

A public-interest laboratory connects institutional memory, scientific agents, secure resources, physical testbeds, human apprenticeship, and transition pathways around a governed mission core.
Conceptual institutional architecture; it is not an existing organization or a funding commitment.

Existing institutional scale

FFRDC R&D spending grew substantially over a decade

The data show that the United States already funds durable government-linked R&D institutions at scale. Spending growth does not establish that every center preserves exploratory research, independent judgment, or transition capability.

Current dollars; the 2024 increase partly reflects a reporting-methodology change at Idaho National Laboratory. Expenditure is not a capability score. Inspect the source record.

Source-linked argument map

The successor is a governance and incentive system, not a model with a building

Current programs demonstrate components; none alone supplies the complete institutional stack.

  1. Documented observation

    The building blocks already exist

    FFRDCs provide durable sponsor relationships, DARPA provides empowered program leadership, NAIRR provides shared AI resources, NSTC provides a public–private prototyping model, and DOE funds transition partnerships.

    U.S. General Servic... DARPA U.S. National Scien... National Institute ... U.S. Department of ...

  2. Measurement gap

    No component owns the whole path

    Shared compute, programs, universities, centers, and commercialization mechanisms remain governed by different clocks and success measures.

    National Center for... U.S. Government Acc...

  3. Design mechanism

    Align six durable layers

    Institutional memory, scientific agents, secure resources, physical testbeds, apprenticeship, and transition must share a mission and capability balance sheet.

    U.S. National Scien... National Institute ... U.S. Government Acc...

  4. Bounded conclusion

    Run a falsifiable ten-year pilot

    Start with a few mission areas, independent governance, staged capital, competing technical paths, public-interest rights, and predeclared failure criteria.

    National Academies ... U.S. Department of ...

The proposal should be rejected or redesigned if a pilot cannot outperform existing arrangements on capability formation, correction, transition, and public value.

The successor to Bell Labs is not Bell Labs with GPUs.

It cannot be recreated by naming a building, aggregating several grants, or buying a frontier model. The twentieth-century laboratory depended on an economic and regulatory settlement that no longer exists in the same form. Rebuilding its surface without rebuilding its incentives would create an expensive stage set.

The objective is different: construct an institution that can carry important uncertainty across projects and transitions while remaining publicly accountable, technically corrigible, and connected to use.

AI changes what is possible. It does not decide why the institution should exist.

It is an amplifier, not a directional force. Under incentives that reward correction, provenance, and hard transition, AI can scale productive search and verification. Under incentives that reward volume, confidence, and short-cycle visibility, it can scale plausible noise and common-mode error. Institutional design selects which effects become durable.

Start from the failure modes

The preceding cases identify a set of constraints.

  • A corporate parent funds what it can capture and is exposed to its business cycle.
  • A government program can coordinate extraordinary temporary capability but does not automatically preserve it.
  • A nonprofit institute can connect projects while its people remain governed by chargeability and funding cliffs.
  • A university produces open knowledge and talent but does not own the complete systems-and-transition stack.
  • An FFRDC can preserve sponsor memory and engineering capability while risking task-order narrowing or incumbent capture.
  • A startup can productize a focused invention but cannot rationally maintain a broad research commons.
  • Output metrics can remain healthy after the capability graph has fractured.
  • AI can scale either witness-bearing correction or scientific-looking noise.

The proposed institution must not declare one of these forms the winner. It must combine their complementary strengths while keeping their failure modes visible.

The proposed form

Create a federally chartered, independently governed nonprofit consortium with a long but reviewable mission charter. Federal anchor funding supplies continuity and public purpose. Industry members supply domain problems, engineers, data under controlled conditions, manufacturing, and transition paths. Universities supply open inquiry, students, and disciplinary depth. Philanthropic capital protects a bounded speculative portfolio. Allied participation expands talent and interoperability where openness and security permit.

The charter should last long enough to support careers and facilities—longer than a program or fund—but contain five-year external reviews and a ten-year reauthorization decision. Durability without review becomes entitlement. Review without horizon recreates the proposal treadmill.

No payer should control the whole portfolio. A single dominant sponsor would gradually turn the laboratory into its engineering arm. Funding diversity is not a virtue by itself; it is protection against one clock becoming the institution’s constitution.

Existing components prove feasibility, not completion

The United States already operates relevant pieces.

NCSES reports that 42 FFRDCs performed $31.7 billion of R&D in fiscal 2024, up from $17.7 billion in 2014. Federal sources supplied 98.5 percent of the 2024 total. The 2024 increase includes a reporting-method change at one laboratory, so it should not be read as pure capability growth. The figures establish scale and durability, not whether every center thinks long. D

The Federal Acquisition Regulation defines the special sponsor relationship in terms of long-term need, continuity, independence, access, and work that cannot be met as effectively by ordinary resources. GAO’s historical account emphasizes institutional memory and trusted objectivity. These are essential design elements.

The concern about FFRDC decline must remain specific. GAO’s DHS review documents a bounded case in which task-order overlap and weak tracking of sponsor use created management problems. It does not prove that all FFRDCs are “engineering houses at best.” Systems engineering, test, integration, and acquisition support can be core public capabilities. The failure is allowing deliverable service to eliminate original technical judgment, independent challenge, apprenticeship, or reusable research state.

DARPA’s program-manager model demonstrates empowered, fixed-term technical leadership. The National Academies’ ARPA-E assessment demonstrates a related active-management model. The successor should borrow authority and urgency without becoming another temporary network.

NAIRR demonstrates shared public–private access to compute, data, models, tools, and training. At its two-year mark NSF reported more than 600 research teams and 6,000 students across every state, the District of Columbia, and Puerto Rico. NSTC is designed as a public–private semiconductor consortium spanning research, prototyping, investment, and workforce. DOE’s Technology Commercialization Fund demonstrates cost-shared lab-to-market partnership. Each solves part of the graph. None owns the whole path from uncertain question to retained public capability. D

Six layers, one accountable institution

1. Institutional memory

The laboratory records hypotheses, decisions, alternatives, experiments, failures, code, proofs, data provenance, and transition outcomes as versioned state. Final papers remain important but are not the archive.

Every major claim has a dependency record. Every completed project has a continuity disposition. Every critical tool has a named custodian or an explicit end-of-life decision.

2. Scientific agents

AI systems assist literature synthesis, conjecture generation, software, proof, simulation, experimental planning, anomaly detection, and adversarial review. They do not receive authority by fluency. Consequential outputs must cross domain-appropriate checkers or produce witnesses.

Model and dataset dependencies are recorded so agreement among agents sharing one corpus is not mistaken for independent replication. Abstention and calibrated uncertainty receive credit.

3. Secure shared resources

The institution maintains compute, data access, reproducible environments, privacy-preserving analysis, identity and authorization, export and security controls, and auditable model use. Resources should be portable enough that no one vendor becomes the institution’s hidden sovereign.

Controls must be tiered. Public data should not inherit classified handling because one program is sensitive. Controlled work should not be forced into false openness. The default is the least restrictive tier consistent with evidence and law.

4. Physical interface

AI confined to text can optimize representations without contact with reality. The laboratory needs automated experiments, testbeds, semiconductor access, fabrication, robotics, sensors, and measurement. Physical and operational evidence are correction channels.

Not every site needs every instrument. A federated system can work if access, scheduling, calibration, data provenance, and technical staff are durable rather than reconstructed by project.

5. Human apprenticeship

Junior and senior researchers share responsibility for consequential outcomes over multiple years. Seniority does not confer veto power. The point is transfer of problem selection, experimental taste, systems judgment, and ethics through work that can fail.

Technical careers must not force every excellent researcher into people management. Program leaders who choose major bets need practiced discovery and stewardship records, while fiduciary, legal, safety, and public-accountability authority remains independent.

6. Transition

Product engineers, mission users, standards experts, procurement officers, manufacturing partners, licensing staff, and venture pathways enter before the final demonstration. Every program identifies its complementary assets and who has authority to adopt the result.

Transition is funded separately. It is neither a promise in the proposal nor a technology-transfer office encountered at the end.

Governance: couple authority without concentrating it

The governing board should contain public-mission representatives, active technical practitioners, industry and transition expertise, workforce and public-interest voices, and independent fiduciary and safety members. No constituency should hold a majority.

Technical portfolios should be led by people with demonstrated research judgment and multi-year program stewardship. Counts of papers or patents can signal experience but cannot serve as a universal gate. Fixed terms, conflict disclosure, plural fields, external replication, protected dissent, and retrospective decision audits limit expert capture.

The board sets mission, risk appetite, and resource boundaries. It should not select individual technical approaches. Program leaders select portfolios within the charter and must document the evidence that changed their decisions.

An independent verification office reports both to technical leadership and an audit committee, with authority to publish bounded correction records. Security and legal review can block unsafe release but must record the category of restriction so secrecy does not become an unreviewable veto.

The capital stack

Different money buys different obligations.

  • Federal anchor funding buys core teams, facilities, archives, and public mission continuity.
  • Competitive agency programs buy ambitious bounded objectives and external technical pressure.
  • Industry membership buys access to precompetitive capability, problem definition, rotations, and transparent licensing options.
  • Procurement and advance commitments buy a transition path for validated capability.
  • Philanthropic funding buys a protected fraction of unusually speculative work within the charter.
  • Licensing and venture returns recycle some captured value without becoming the laboratory’s sole survival condition.

These streams must be reported separately. Otherwise project revenue can masquerade as capability investment and license revenue can quietly redefine the research agenda.

The IP and public-value compact

Foundational tools, standards, benchmark infrastructure, and nonsensitive results should enter a research commons, sometimes after a limited lead period. Precompetitive work is shared among members under uniform terms. Privately funded company-specific development may remain proprietary.

Exclusive licenses are field-limited, time-limited, and milestone-based when public funds created the option. Rights revert or march in when a licensee warehouses the result. Dual-use work receives a controlled tier with periodic declassification or release review rather than permanent default secrecy.

The objective is not maximum openness or maximum appropriation. It is sufficient private incentive for transition plus sufficient public return to replenish the commons.

A ten-year falsifiable pilot

Begin with two or three mission areas where the coordination failure is visible, shared infrastructure matters, and outcomes can be evaluated—such as trustworthy AI-assisted science, privacy-preserving data systems, or cross-layer semiconductor and quantum assurance.

Predeclare the tests.

Annual operational review: integrity incidents, artifact reproducibility, staff continuity, resource utilization, security, and administrative burden.

Five-year capability review: new technical abilities, apprentice independence, cross-program reuse, competing approaches, transitions, corrections, and capabilities deliberately ended.

Ten-year mission review: scientific and public value, deployed systems, standards, spillovers, concentration risk, and counterfactual evidence against comparable existing programs.

The pilot fails if it becomes a captive contractor, an academic grant distributor, a venture studio, a permanent procurement incumbent, or an AI-content factory. It fails if administrative labor consumes the protected technical horizon. It also fails if it protects research without producing hard external evidence or credible paths to use.

Those are exit criteria, not rhetorical cautions. Before launch, the charter should set measurable ceilings or comparison rules for administrative time, sponsor concentration, vendor concentration, unreproduced major claims, capability reuse, apprentice independence, and the share of programs with accountable transition owners. Crossing a threshold triggers redesign, leadership review, funding reduction, or dissolution rather than an automatic plea for permanence.

The strongest counterargument

This proposal may simply aggregate every institutional virtue and assume away every conflict. Patient capital reduces pressure but can reduce urgency. Industry transition improves relevance but can narrow openness. public governance protects legitimacy but invites politics. Security enables mission work but restricts diffusion. AI improves scale but concentrates technical power.

There is no architecture that eliminates these tensions. The proposal is valuable only if it makes them governable and testable. Multiple pilot sites or teams should use different designs. Funding should be staged. External groups should replicate important results. The charter should contain a dissolution and asset-transfer plan so institutional survival is not the hidden objective.

The strongest control is competition among institutional hypotheses, not competition that forces every project to rebuild from zero.

Five prospective verdicts

  • Scientific success: Does the laboratory produce important knowledge and correct itself faster than comparable arrangements?
  • Technical success: Does it create working, independently testable capability across theory, software, hardware, and experiment?
  • Transition success: Do mission users, standards bodies, firms, and public systems adopt and maintain selected results?
  • Institutional success: Do teams, tools, memory, apprenticeship, and independent problem selection strengthen across project boundaries?
  • Public-value success: Do benefits, diffusion, security, and cost justify the public and private capital without concentrating unaccountable power?

The final proposition

The United States cannot recreate the twentieth-century corporate laboratory through admiration, appropriation, or executive decree. It can still transfer living knowledge into a new incentive system.

The nearly trillion-dollar R&D economy supplies enormous activity. The missing object is a durable institution accountable for making the pieces cohere across a generation. AI makes the design more urgent because it increases both productive search and plausible noise. It can preserve memory or fossilize error; widen participation or centralize authority; accelerate proof or accelerate performance theater.

An AI-native laboratory is therefore an institution, not a model with a building. Its intelligence lies in what it remembers, what it can test, whom it trains, how it corrects itself, and whether it can carry a true but inconvenient result all the way to use.

The open question is no longer whether the old Bell Labs can return. It is whether we can build a successor before we discover that money can purchase equipment, models, and projects—but not the lost ability to make them cohere.

Argument under pressure

Two rounds of objection—not a ceremonial counterargument

These are simulated steelman exchanges. The second objection responds to the first answer; the conclusion is narrowed where the objection survives.

Claim under test

A federally anchored public–private laboratory can preserve research capability better than the fragmented system.

  1. First-level objection

    Such an institution will be captured by incumbent firms, political priorities, security bureaucracy, or permanent internal constituencies.

  2. Response

    Those are central design risks. Funding diversity, fixed governance terms, conflicts rules, transparent portfolios, protected dissent, external replication, and periodic mission reauthorization are required.

  3. Second-level objection

    The safeguards themselves create administrative load and may select safe, legible work instead of frontier research.

  4. Bounded conclusion

    Separate fiduciary controls from technical micromanagement, apply reporting proportionally to risk, and judge the pilot by whether technical teams retain time and authority for uncertain work.

U.S. Government Accou... National Academies of... U.S. Government Accou...

Claim under test

AI-native infrastructure can extend institutional memory and accelerate science.

  1. First-level objection

    Centralized models, data, and compute can amplify common-mode errors, surveillance, vendor dependence, and concentration of scientific authority.

  2. Response

    Yes. The architecture needs provenance, model diversity, reproducible environments, privacy-preserving access, independent red teams, and portable artifacts.

  3. Second-level objection

    Those controls do not solve who sets public priorities or bears responsibility when an AI-assisted result causes harm.

  4. Bounded conclusion

    AI receives no governing authority or legal accountability. Named humans and chartered bodies own priorities, risk acceptance, correction, and deployment decisions.

U.S. National Science... National Institute of...

Revision ledger

Corrections and feedback incorporated

Inspectable claims

Source Ledger

  1. Federal statistical table

    Indicators 2026: U.S. R&D expenditures by type, selected years 2000–2024

    National Center for Science and Engineering Statistics. Supports: R&D totals, constant-dollar trends, and the basic/applied/development composition.

  2. Agency primary source

    Become a program manager

    DARPA. Supports: Fixed program-manager tenure, portfolio authority, milestones, and performer coordination.

  3. Federal oversight report

    Federal Research—Management and Oversight of FFRDCs

    U.S. Government Accountability Office. Supports: Long-term government relationships, institutional memory, objectivity, and special access in the FFRDC model.

  4. Agency progress report

    NAIRR at two years

    U.S. National Science Foundation. Supports: Shared AI resources, federal/private participation, research access, and the move toward durable infrastructure.

  5. Agency primary source

    National Semiconductor Technology Center

    National Institute of Standards and Technology. Supports: The public–private consortium, prototyping, investment, collaboration, and workforce model.

  6. Agency primary source

    Technology Commercialization Fund

    U.S. Department of Energy. Supports: Cost-shared lab-to-market partnerships and commercialization infrastructure.

  7. Federal statistical report

    R&D Spending at Federally Funded R&D Centers Surpassed $31 Billion in FY 2024

    National Center for Science and Engineering Statistics. Supports: The number of FFRDCs, expenditure growth, funding concentration, and FY2024 R&D composition.

  8. Federal regulation

    Federal Acquisition Regulation 35.017—Federally Funded Research and Development Centers

    U.S. General Services Administration. Supports: The special long-term need, continuity, independence, access, sponsorship, and periodic-review requirements that define the FFRDC relationship.

  9. Federal oversight report

    Federal Research Centers: DHS Actions Could Reduce the Potential for Unnecessary Overlap among Its R&D Projects

    U.S. Government Accountability Office. Supports: A bounded example of FFRDC work organized through task orders and weaknesses in tracking how sponsors used resulting deliverables.

  10. Consensus study report

    An Assessment of ARPA-E

    National Academies of Sciences, Engineering, and Medicine. Supports: ARPA-E's use of technically exceptional program directors with substantial authority over program design, active management, milestones, and portfolio decisions.

Public record

Version history

Published versions are immutable snapshots. Corrections create a new version rather than silently changing the historical record.

No public snapshot has been archived yet. The first snapshot is created when the article is published.

Read the correction, withdrawal, and versioning policy

← Return to the series map

This essay distinguishes firsthand memory, documentary evidence, interviews, and analysis. Published versions remain available even after revision or withdrawal.