My website lists 79 research works, 31 granted U.S. patents, and 12 funded R&D awards. It also records software, talks, awards, and a company built from earlier research. These are not invented metrics imposed from outside; I chose to curate them because they record real work. D
This disclosed record is an audit object, not a representative benchmark. It is useful because the claims and omissions can be inspected against a known ledger. Nothing in the counts establishes what another researcher, field, institution, or career stage should produce.
They also demonstrate the measurement problem.
The lists remain after a project ends. They remain after colleagues move. They remain if code no longer builds, an instrument is dismantled, a result is superseded, or no organization can reproduce the integrated system. That durability makes outputs useful historical evidence and dangerous substitutes for capability.
What each count actually certifies
The publication record establishes that the site curates 79 works. It does not establish that every result is correct, equally important, independently replicated, or still supported by an active research line.
The patent record links 31 granted inventions to public patent documents. A granted patent has passed a legal examination for patentability under the relevant process. The USPTO’s guidance makes clear that this is patent examination, not scholarly peer review. A patent does not certify that the invention was built, adopted, profitable, secure under every model, or institutionally preserved.
The project record lists 12 funded awards. An award establishes that a sponsor selected and funded work under stated conditions. It does not establish that the program met every goal, transitioned, or left reusable capability.
These are not criticisms of the outputs. They are type checks. A system becomes confused when evidence of one kind is accepted as a verdict of another.
An output type system
R&D evaluation needs the equivalent of a type checker. Each artifact licenses some inferences and rejects others.
| Artifact | Strongest ordinary inference | Independent check | What it does not certify |
|---|---|---|---|
| Peer-reviewed paper | A venue accepted a disclosed claim as worth publishing after its process | Specialist scrutiny, formal checking, computational reproduction, or new-data replication as applicable | Universal correctness, importance, or use |
| Granted patent | A patent office allowed claims under legal patentability requirements | Claim construction, prior-art challenge, implementation, and technical evaluation | Peer-reviewed science, product performance, adoption, or freedom to operate |
| Prototype or demonstration | A configuration worked under stated conditions | Reproduction, adversarial test, boundary analysis, and operation by another team | Generality, manufacturability, reliability, or maintainability |
| Funded program | A sponsor selected activity under a budget and mission | Milestone evidence, independent test, transition, and post-program reuse | Technical success or public value |
| Citation | Another work referenced the artifact | Context and purpose of the citation; delayed influence | Agreement, correctness, or causal impact |
| Award | An awarding body made a retrospective judgment under stated criteria | Transparent criteria, field context, and later evidence | Complete comparison or institutional continuity |
| Deployment or revenue | A user or buyer accepted a system under some conditions | Retention, reliability, safety, public impact, and lifecycle evidence | Scientific novelty or broad social benefit |
The National Academies distinction is an example of type discipline: rerunning the original code and data is computational reproduction, while a new study with new data addresses replicability. Calling both “verified” deletes the independence boundary. D
The type checker should not make output less valuable. It makes success statements more credible. “This patent was granted,” “this artifact reproduced,” “this system deployed,” and “this team still possesses the capability” are four useful claims precisely because they are not treated as synonyms.
Outputs are projections
Imagine the actual research system as a high-dimensional state:
- people and complementary roles;
- working relationships and mentorship;
- code, data, instruments, fabrication, and test environments;
- tacit knowledge and negative results;
- authority to choose problems;
- access to users and transition partners;
- independent mechanisms for finding error.
A publication count projects that state onto one axis. A patent count projects it onto another. Funding totals, prototypes, citations, and awards are additional projections. None is false. Each discards information.
The danger appears when the projection becomes the objective. Teams learn which work produces countable units. Institutions allocate support to what renews grants, improves rankings, fills an IP ledger, or produces a demonstration before review. Maintenance and replication remain undercounted because they often preserve value rather than create a new unit.
The Leiden Manifesto does not reject quantitative indicators. It argues that indicators should support qualitative expert assessment, respect field differences, and remain open to scrutiny. Research on delayed recognition and novelty adds another warning: novel work can have more variable impact and take longer to be recognized, so short windows systematically distort selection. D
The capability-lag problem
Output and capability can move in opposite directions for a time because output lags accumulated state.
A mature team may publish and patent vigorously while losing junior hiring, maintainers, or discretionary time. Previously initiated projects continue yielding results. The visible ledger looks healthy. The missing investment becomes apparent only when the next generation of problems fails to start.
This is analogous to a factory meeting shipments while deferring maintenance. The analogy has limits—research is not a production line—but the accounting insight holds. Current output draws on past capability. It does not measure the replacement rate.
A minimal institutional report should therefore pair flows with stocks:
- Flow: papers, patents, prototypes, dollars, milestones, hires, releases.
- Stock: intact teams, maintained artifacts, functioning facilities, mentorship depth, active transition relationships, and time-to-reconstitute lost capability.
- Depreciation: departures, obsolete environments, inaccessible data, broken interfaces, lost problem authority, and unrecorded failures.
Without the last two, rising output may conceal capability debt.
Do not repair one proxy with a larger proxy
The obvious response is a composite score: combine papers, citations, patents, funding, deployments, team retention, and public value with weights. That creates a more sophisticated target and a less interpretable failure.
Unlike quantities can compensate in the arithmetic even when they cannot compensate in reality. A large publication count can cancel an irreproducible system. A profitable deployment can cancel a collapsed apprenticeship chain. An intact team can cancel a decade without a hard result. Once weights determine resources, actors optimize the exchange rate among categories.
The capability balance sheet should remain a dashboard with noncompensable constraints. A mission may declare that some conditions are necessary—safety evidence, a reproducible build, a transition owner, or an independent successor—and report the rest as visible tradeoffs. Different fields will choose different constraints. What matters is that failure in one dimension cannot be hidden inside aggregate success without an accountable override.
A useful review has three layers:
- Typed outputs: what artifact exists and what inference does it license?
- State transitions: what became known, usable, corrected, maintained, teachable, or accessible because of it?
- Capability renewal: which people, tools, relationships, and options now make the next important problem less costly or more possible?
The third layer can be negative even after a valuable result. A team may rationally consume itself to solve an urgent mission. The account should record that trade rather than recode it as institutional success.
AI destroys the production-cost signal
Before generative AI, a polished manuscript, codebase, or technical survey imposed substantial effort. That cost never guaranteed quality, but evaluators sometimes treated it as a weak signal that serious work lay behind the form. AI makes form, variants, and plausible supporting artifacts cheaper.
This is good when it makes a small team more capable. It is dangerous when an institution keeps rewarding unit counts. The number of proposals, papers, experiments, code commits, or agent evaluations can grow while scarce human attention, independent data, physical tests, and accountable maintenance remain fixed.
The metric response should move downstream toward harder state changes:
- Can another team reproduce or operate the artifact?
- Did a new measurement discriminate among hypotheses?
- Did a correction propagate through dependencies?
- Did users accept responsibility for maintenance?
- Did an apprentice become independently capable?
- Did the work reduce the reconstitution cost of a public capability?
AI can help collect this evidence. It can also generate the evidence narrative. Consequential state transitions therefore need bound artifacts, named custodians, and independent checks rather than self-reported prose alone.
A retrospective signal and its boundary
The NDSS Test of Time Award is a rarer kind of evidence. It recognizes influence after a long interval rather than optimizing for immediate attention. In 2024 it recognized work I coauthored on an Internet-censorship system. That is meaningful evidence of enduring scholarly and technical influence. D
It still does not answer every question. Did the system deploy? Which ideas were reused? Did the original team remain connected? Did the institution preserve the capacity to pursue the next problem? What public benefit followed? A retrospective award strengthens one verdict without collapsing the other four.
That discipline protects the work from both inflation and erasure. We need not pretend an award proves everything in order to say that it proves something important.
The strongest counterargument
Research capability is latent. Unlike a paper count, it cannot be audited directly. Managers who dislike an objective metric can invoke “tacit knowledge” or “future potential” to protect favored teams. A capability ledger could become a vocabulary for institutional rent-seeking.
That is a real danger. The answer is not to retreat to one-dimensional counts. It is to operationalize the stock.
For each claimed capability, ask for observable tests:
- Can an independent team build or reproduce the artifact?
- Which later projects reused it?
- How many critical roles have credible successors?
- What is the time and cost to restore the environment?
- Can the group initiate a new problem rather than only service inherited tasks?
- Has a user, sponsor, or product team accepted responsibility for transition?
- Which claim was corrected because the institution’s error-detection process worked?
These questions do not eliminate judgment. They make judgment contestable.
What would falsify the capability-lag account?
The thesis should weaken if output vectors predict future research capacity as well as direct capability measures. If publication, patent, program, and deployment histories reliably forecast team regeneration, artifact reuse, new-problem initiation, transition, and low reconstitution cost, the additional balance sheet may not justify its burden.
It should also weaken if capability measures add mostly narrative noise: reviewers cannot agree even on bounded proxies, organizations manipulate them more easily than counts, or teams classified as healthy repeatedly fail external technical tests.
The opposite evidence would be a leading-indicator result. Track teams before a funding cliff, closure, merger, or leadership change. If loss of critical roles, maintenance, adjacency, or question authority predicts later output and transition decline after controlling for field and resources, capability accounting has earned explanatory value.
The personal ledger offers only a pilot. For each item, the site can eventually attach an evidence state: claim status, artifact status, independent use, transition, current custodian, and whether a living team can extend the work. The exercise may reveal strong continuity or more archival remnants than expected. Either result is useful; the framework should not guarantee its own conclusion.
Five different verdicts
- Scientific success: Publications and retrospective recognition provide evidence of new and influential knowledge, with correctness and importance remaining claim-specific.
- Technical success: Patents, code, and prototypes provide evidence of disclosed or working mechanisms, not universal system performance.
- Transition success: A spinout or deployment requires separate evidence of adoption, operation, and maintenance.
- Institutional success: Output records say little by themselves about whether teams, tools, and apprenticeship survived.
- Public-value success: Citations, standards, safer systems, commercial use, and trained people may each carry public value that the originating organization captures only partly.
What the successor must learn
The successor laboratory should make every major review two-sided. The output ledger asks what was produced. The capability balance sheet asks what the institution can now do that it could not do before—and what it can no longer do despite the outputs it retains.
No composite score should hide the answer. A high patent count cannot compensate for an irreproducible system. A maintained team cannot compensate for years without a hard external result. The dimensions should remain visible so governance must confront the tradeoff.
The uncomfortable open question is personal and institutional: which item on my own output lists represents a living capability today, and which is now only a durable record that the capability once existed?