Documentation

Sign in with GitHub
DocumentationThe question

What an experiment leaves behind

The code, data, attempts and failures a result depends on, and what is usually lost

An experiment is an event. Its trace is a document. The event takes days and touches hundreds of files; the trace is a few paragraphs, a table and, with luck, a repository. That ratio is normal, and calling it a scandal helps nobody, because the alternative — writing everything down — is worse. The narrower question is which parts of what the event produced the next piece of work actually needs, and whether those parts survive.

What the work produces

Set aside the result for a moment and list what exists at the end of one experiment that did not exist at the start.

  • Code at a particular state, including the changes made while the run was in flight and never committed.
  • Data at a particular state: the subset actually used, after the filtering script nobody wrote down.
  • Configurations: seeds, hyperparameters, library versions, the machine, the flag that was flipped on Wednesday and left.
  • Runs that failed, in several different ways, only one of which was interesting.
  • Intermediate results: the plot that was opened once, understood, and closed.
  • Decisions taken, with the options dropped and the reason each was dropped.
  • An understanding, held by whoever did the work, of which parts of it are fragile.

The last item is unlike the others. It is not a file. It is the sense of which number would move if the split changed, which failure was a defect in the code and which was the phenomenon, and which part of the setup is load-bearing in a way the write-up will make sound incidental. It is usually the most valuable thing produced and the least transferable, which is why the standard advice for taking up someone else’s work is to go and talk to whoever did it.

The loss leaves no hole

What makes this hard to correct is that nothing looks wrong. A finished paper is a coherent artefact: it raises questions and answers them, and a reader who has not done the work cannot see the questions that were raised and dropped. A tidy repository has the same property. Ten runs of which two are reported leave a record indistinguishable from a record of two runs. Absence has no notation.

The person best placed to notice is the one who cannot. For the author nothing is missing: it is in memory, a month old and still vivid, so the account reads as complete because its first reader is silently supplying the rest. A year later the author loses that advantage too.

The cost lands on the next person

The loss is invisible partly because it is deferred. It is real, but it is paid somewhere else, by someone who does not know what they are paying for.

A constructed example. A group publishes a selection rule that reaches a target accuracy on far fewer labels. While developing it they had found that the saving disappears when the pool of unlabelled examples is small, and had worked out why: the rule then keeps choosing from one narrow region of it. The diagnosis never appeared anywhere. A year later another group takes the method up on a problem with a small pool — the case where labels are expensive and the method looks most attractive — and six weeks later has a null result and, if it is diligent, the same diagnosis. They were not misled. They repeated a piece of work that had already been done and thrown away.

The pattern has three parts, each a separate cost: the successor re-derives reasoning that was already sound, re-runs configurations that were already abandoned, and rediscovers the fragile parts of the setup by breaking them one at a time. Each rediscovery is individually cheap, which is exactly why it persists — nobody pays enough for it to be worth fixing — and the total arrives as a slightly slower field rather than as a bill anyone receives.

When there is nobody to ask

All of this assumed a colleague. Much of what a research group knows is held by people who are still there and can be asked, and that fallback has quietly carried the weight the written record does not. Machine-run work removes it.

When an agent does the work, the understanding described above sits in a working context that ends when the session does. Nothing of it persists to be questioned later. Whatever was not written down at the time did not merely go unrecorded: afterwards it does not exist anywhere. The trace is not a summary of the work. It is the whole of what the work leaves.

Two things follow. The share of research whose only representation is its trace rises, because agents produce more traces and none has anyone standing behind it who remembers. And an agent reading the record inherits only what the record says: it will take the obvious next step exactly as often as the record fails to mention that the obvious next step does not work.

Recording is not free either

The conclusion is not that everything should be recorded. Recording the whole list above would cost more than the experiment, and a system demanding it would fill with whatever is cheapest to supply: fields completed to get past them. Deciding what is worth keeping is itself a research judgement, and not one a platform can make for the author.

So the useful question is which few things repay writing down. Four candidates recur: what was predicted before the data arrived; what failed and how; the exact state of the inputs and outputs, so a later reader is arguing about the same objects; and the decision that changed direction, with the option it beat. None is expensive on its own, and each is obvious at the time and unrecoverable a year later. Each is also a place where a record can make the careful account the cheap one to write rather than leaving it to whoever is writing up.

In Substrate

Substrate widens what can be recorded around a result and compels none of it: each item below is something an author or their agent chooses to write, and nothing here checks what it is told.

  • Executions, including the ones that failed. An attempt is registered before it runs and reports its events as it goes, and a failed attempt stays in the record Attempt receipts: Live. See Attempts.
  • Reported from where the work happened. A local command runs a pinned commit and delivers that attempt’s events Capture adapter: Live.
  • Inputs and outputs as objects. A material’s identity is a digest of its bytes, or a pinned reference when nothing was hashed, rather than a path that can move Materials: Live.
  • Where a copy can be found. Members report locations, including when a material can only be asked for from whoever reported it Location reports and obtainability: Live.
  • What produced what. Typed edges record that an attempt consumed one material and produced another, and that a material is evidence for a finding Provenance edges: Live. See Materials and provenance.
  • Reasoning, when someone writes it. A checkpoint is an attributed synthesis of a Thread up to a point, with its open issues and a next action Research checkpoints: Live — the nearest the record comes to the judgement that otherwise lives in one person’s head.
  • Results that went the other way. A finding relates to its experiment with a verdict that may contradict the hypothesis it tested Findings linked to experiments: Live, so a negative outcome is a publication rather than a silence. See Negative results and gaps.
  • Nothing outside the author attests to any of it. A receipt is the registrant’s own report, unsigned by the environment that ran the work Signed receipts: Idea.

Substrate also re-runs nothing Platform replay: Idea: the work happens in your own tools, and what reaches the record is what you report about it.

Open question

Whether recorded dead ends actually stop the next worker repeating them, or simply accumulate unread, is not something argument settles. The research agenda states what evidence would count, and the known limits of the app list what it does not record at all.