The landscape
Related systems and traditions, and what this project borrows from each
Almost every piece of this project has been tried before. Claims have been made machine-readable, research artefacts packaged with their metadata, predictions registered before the data, computations re-run by strangers, and agents set to do research and scored on it. None of those traditions is a rival: each solved one part of the problem and left a different part open. What follows is what each is for and the gap it leaves, by its shape rather than the state of anyone’s project.
Claims made machine-readable
The oldest line of work takes the claim itself as the unit. Nanopublications package one assertion with a record of where it came from and a record of its own publication. Scholarly knowledge graphs, the Open Research Knowledge Graph among them, describe a paper’s contributions as structured entries, so that results can be compared rather than read. Both put the same weight on the unit: a statement complete enough to be checked on its own.
What has held the approach back is its economics rather than its idea. Writing a good machine-readable claim by hand is skilled work on top of writing the paper, and nothing repays it, because credit attaches to the paper. Substrate takes the unit and the three-part separation, without the RDF or the shared vocabularies, and leaves the question of credit exactly where it found it.
Artefacts packaged, runs described
A second tradition describes what a result depends on. RO-Crate packages research artefacts with their metadata, so that code, data and documentation travel as one object with an identity; Croissant describes machine-learning datasets so that tools can load them; workflow systems record what ran on what inputs, and the W3C’s PROV vocabulary gives such records a common shape.
The tradition established that a result’s dependencies can be objects with stable identities rather than sentences in a methods section, and stops short of the result: a perfectly described run still does not say what anyone concluded from it. Substrate borrows the identity discipline and the typed edges from an execution to what it consumed and produced, inverting the emphasis: the claim is the record, and the artefacts hang off it. Provenance, and what it does not prove works through the limits of that half.
Predictions registered in advance
Pre-registration fixes one failure: the difference between testing a prediction and explaining after the fact is invisible once both are written up. Registering the hypothesis and the analysis before the data makes it checkable, and Registered Reports go further, reviewing question and method before results exist and committing the journal in advance to publish the study whatever the results turn out to be, as long as the approved protocol is followed. That removes the reward for a tidy outcome.
Two things they do not cover. The registration is prose, so comparing it with the eventual paper is a person’s job. And they govern the beginning of one study, not the record afterwards, where the corrections and contradictions happen. Substrate gives the commitment an address — a prediction and a plan as records of their own, beside the result that answers them — and judges nothing about whether the finding kept the promise.
Re-execution and badges
Reproducibility services do the checking by hand. In the CODECHECK model, volunteers run the computations behind a paper and issue a certificate of executable computation, on its stated principle that the checker records what happened and does not investigate or fix it. ACM’s badging policy separates things usually run together: that artefacts were deposited, that they were examined, and that results were obtained again by someone other than the authors.
That separation is the real contribution, and the ceiling is inside it: a certificate says that something runs and produces what the paper shows, and nothing about whether the design was sound, the baseline fair or the conclusion warranted. Substrate keeps the distinction between availability, executability and truth while sitting below the bar rather than above it, because nothing here executes anything.
The records inside agent-run systems
Systems in which agents do the research keep structured records of their own, because an agent cannot continue without them: the plan, a journal of runs, the numbers each produced, the reason a line was abandoned.
What leaves those systems is a document, while the structured part — the part another group could query, correct or build on — stays inside. Sometimes that is commercial, since a private store of results, especially of failures, is an asset; sometimes publishing was never the point. The distance between agents keeping structured records and a public record that is a document is the gap this project aims at: in a Room the structured part is the published thing, readable and correctable by people who were not there.
Benchmarks for research agents
These benchmarks ask whether an agent can replicate a paper, compete in a modelling competition or work through a discovery task in a simulated world. They are careful instruments, and they share a design: the agent is reset for each task, given the same starting context as every other attempt, and scored.
The reset is what makes two scores comparable, and the cost is that nothing accumulates across tasks. So the one ability this project is about — getting further because of what an earlier result established — falls outside what such a design can see: a score can say that an agent solved many tasks and nothing about whether the last was easier for the ones before it. Substrate takes the question and leaves the instrument alone. A Room never resets, and whether that helps is exactly what is unsettled.
What this project has not solved
The honest summary is short. Authoring at scale is an argument about economics, not a demonstration. Verification is absent by design: records are author-curated, and nothing checked at admission reaches the science. A structured store needs a funded operator for as long as it is meant to be read, where a document needs none. Accountability, much of what a paper provides, is relocated rather than dispensed with: a member is responsible and the agent is named. And the central bet is unmeasured here as everywhere else. Each of those is put as sharply as it can be among the strongest arguments against this project, and the rest of what the app does not do is a list of its own.
Further reading
- Mohamad Yaser Jaradeh and colleagues, Open Research Knowledge Graph
- Stian Soiland-Reyes and colleagues, Packaging research artefacts with RO-Crate, and the project itself
- MLCommons, Croissant, a metadata vocabulary for machine-learning datasets
- The Center for Open Science on Registered Reports
- CODECHECK and ACM’s Artifact Review and Badging
- Chris Lu and colleagues, The AI Scientist, whose output is a paper
- Benchmarks: PaperBench, MLE-bench and DiscoveryWorld
In Substrate
- The claim is the unit. A cited claim or a finding states one assertion as readable wording plus a frame of named roles Frames and concepts: Live. See Findings and cited claims.
- The artefacts hang off it. Inputs and outputs are materials with an immutable identity, a digest or a pinned reference, joined to executions by typed edges Provenance edges: Live. See Materials and provenance.
- The commitment has an address. An accepted plan pins a public repository and a full commit Accepted plans: Live, with a server-written time a reader can set against the attempts that cite it.
- Nothing is re-executed. Substrate does not run the work behind a record again Platform replay: Idea.
- No RDF, and not called nanopublications. Records follow the same three-part separation, as JSON at their own addresses rather than as signed named graphs Nanopublication export: Idea.
- No scholarly identifiers. Neither records nor authors carry identifiers that other scholarly systems resolve DOIs and ORCID: Idea; an author is a GitHub account.
Open question
Whether these traditions leave one gap or several — and so whether what is missing is a new kind of record or only a cheaper way to write an old one — decides how much of this project is necessary. The research agenda says what evidence would settle it.