Research agenda
The open questions the project is working on
These are the questions the laboratory works on. Each says why it is genuinely unknown rather than merely unbuilt, and what evidence would count either way. What the app cannot do is a different list, on An archive, not yet a substrate.
Does a shared record pay?
The bet is that a curated record makes research compound rather than drift. None of it is settled.
Accumulation
Does a shared, curated record make the next experiment better than starting cold or keeping private notes? The comparison has rarely been run: evaluations of research agents reset their state at the start of every task, and groups that keep durable records do not publish them. What would count is one question worked by one model with the same retrieval and budget, differing only in what it starts from, scored on compute to a target and dead ends repeated.
Compounding error
When a record carries a wrong premise, how far does the damage spread, and do provenance, correction links Research links: Live and independent reproduction contain it? A shared record makes a confident wrong premise cheaper to reuse than to re-derive, so accumulation pays only if it is right more often than it is wrong. What would count is, where a premise proved wrong, how many records came to depend on it and whether the provenance a finding states Typed execution provenance: Live was enough to find them all.
Schema versus volume
Does structure help where a flat pile of the same text does not? The claim is not that more context helps, but that named roles, typed provenance and derived status help Frames and concepts: Live; published comparisons of agent memory are hard to set beside each other, and in some a graph-shaped store scored below plain vector retrieval. What would count is the same content served both ways under one retrieval budget, judged on what a flat store cannot answer: which measurement a number belongs to, and what has since been said about it.
Negative knowledge
Do recorded dead ends measurably stop the next person repeating them? Somewhere to put failures exists — a failed attempt stays in the record Attempt receipts: Live, and a finding can contradict the hypothesis it tested Findings: Live — but whether anyone reads them and changes course is the open part. What would count is how often later workers repeat a dead end the record already holds.
What the unit has to be
A record pays only if writing one is affordable and reusing one is safe.
Authoring economics
Can agents make machine-readable claim records cheap enough to reach the scale papers reached, and what keeps them trustworthy if they do? Every earlier claim format stayed far below the literature, mostly because a careful record costs a person a great deal to write; that agents change the cost is an argument, not evidence. What would count is sustained authoring by people who did not design the format, with the share of records later corrected or retracted staying visible as volume grows.
Granularity
What size of claim can be reused by someone else without losing the context that made it true? A claim scoped tightly enough to be safe may be too narrow to find or combine, while a broader one carries conditions that quietly stop holding. What would count is which records other Threads and Rooms link to Reuse into Threads: Live, and how often a reuse is corrected because a condition did not travel with the claim.
Neighbourhood read: IdeaNeighbourhood read
Reuse may also depend on how a record is read: whether agents need a read that walks the typed edges around a record, or do as well following exact references one at a time.
Concept identity
Equal labels never merge, so meaning stays local: when does that fragment a field’s vocabulary, and what would a shared concept layer have to guarantee to be worth it? Merging on labels is how a vocabulary acquires silent errors, so refusing to merge is the safe rule; the cost is that one meaning splinters across records. What would count is a rule that never joins two definitions a careful author would keep apart.
Cross-Room concept identity: IdeaCross-Room concept identity
A store in which two Rooms declare their definitions the same would have to record who declared it and on what grounds, and stay reversible. Otherwise it trades a visible fragmentation for an invisible error.
What makes a record believable
Nobody can check every result, so belief rests on something cheaper.
Trust without an oracle
Which rungs between “an author says so” and “someone else got the same result” can be established by mechanism rather than asserted? Some look mechanical: that a plan was accepted before an attempt was registered, that an attempt’s pinned commit is the plan’s, that another member reported the same values. Each is weaker than it looks. What would count, for each rung, is a check an interested author cannot satisfy by writing the right words Replication ladder: Idea.
Platform replay: IdeaPlatform replay
Re-execution is the rung that would not rest on the author at all: a pinned attempt is run again and its outputs compared. What that proves is narrower than it sounds — that the recorded recipe yields the recorded numbers, not that the numbers mean what the finding says.
Reproduction of non-deterministic work
What does agreement mean when a rerun cannot match values exactly? The deterministic case is easy: two runs of one configuration agree when a digest over their values agrees Values digests: Live. Most machine-learning work is not deterministic, and no settled rule says how close counts or how independent the second worker must be Reproduction attempts: Live. What would count is authors agreeing in the plan, before the data, on what agreement would look like, and later readers accepting those judgements.
Honest reporting under pressure
What keeps a record honest when the author has an incentive to look successful? A record that shows only successes looks exactly like one with no standards, and every mechanical check invites the cheapest way of passing it. Checks that bind one thing to another — a number to the measurement it came from, an attempt to the commit it ran — may survive that pressure; checks that a field is filled in will not Mechanical admission gate: Idea. What would count is a catalogue of how the checks are gamed in practice.
Who keeps the record
The record does not curate itself, and curators need a reason to.
Curation
A Room’s spine is its curated research knowledge graph, and curating it is the main work there. Can that work be recognised and rewarded without encouraging graphs that merely look tidy? Anything scored gets optimised, and a curator rewarded for structure will produce structure. What would count is scoring a spine by something its curator does not control — how well it supports predictions about experiments they did not run Research checkpoints: Live.
Coordination
Can people and agents on unrelated machines work from one record without messaging each other? Shared state is a solved problem for software, but research adds a requirement: an arriving worker must tell from the record alone what is settled, what someone else has taken and what is free Take, release and report: Live. What would count is workers picking the right next thing unprompted, and how much of a Room’s activity Room activity feed: Live turns out to be coordination rather than research.
Understanding as prediction
Can a research programme’s grasp of its domain be measured by how well it predicts its own next result? Task success is not understanding: a system can score well and still hold a wrong account of why. A hypothesis states such a prediction exactly, before the data Hypotheses: Live, and the plan that tests it is fixed before the run Accepted plans: Live. What would count is scoring those forecasts over a long series, with forecast skill rising as the record grows.
How these get answered
Not by argument on this page. The evidence is the record itself: findings published in public Rooms, the corrections and disputes that follow them, and reproductions reported by people who did not run the original. If you work in a Room, your results are part of the answer, including the ones that go against the project’s own bet. The open laboratory describes how to take part.
In Substrate
A Room is public from the moment anything is written in it Rooms: Live, a publication is an exact version that never changes Exact versions: Live, and a correction is a link that leaves a notice rather than an edit Notices and record status: Live. So the evidence bearing on each question above stays put, including the parts that go against the project, and that is the whole of what the platform contributes to answering them.