Documentation

Sign in with GitHub
DocumentationFoundations

Attempts, receipts and materials

Recording each execution and the inputs and outputs it touched

A prediction written before the data fixes what would count as an answer. It says nothing about what actually ran, and that is the other half of the account.

“We ran it and the accuracy was 82.0%” is a sentence nobody but its author can do anything with. It reports a number and withholds what would let a reader get near it: the code, the data it was computed over, the command, the machine, the file the number was read out of. The distance between we ran it and here is what ran is the subject of this page.

What it takes to believe a number

Ask what you would need before using someone else’s figure as the baseline for your own work. The list is short, and always the same list.

What the record namesWhat it lets a reader do
The code, at one commitRead the version that ran, not whatever the branch points at
The inputs, by identityKnow which bytes went in, and whether your copy is the same
The commandSee the arguments, the arm and the seed, where runs usually differ
The environmentJudge how much of the result belongs to the machine
The outputsGo to the file the number was read from, not a sentence about it

None of that makes the number right; it makes the number addressable, so that a reader who doubts it has somewhere to go. Two claims are worth keeping apart. Derivation integrity says these numbers came out of these bytes, by this code. Execution integrity says the run happened, as described, when it says. A record written by the party that did the work approaches the first and cannot reach the second on its own.

Attempts are plural, and they do not overwrite

One experiment produces many runs, and the honest unit is the run. A retry is another attempt rather than a replacement: the first stays, with whatever it reported, and the second names the run it repeats. Nothing is edited, because an edited execution record is no longer a record of an execution.

Three outcomes have to stay distinct, because collapsing them loses what a sceptical reader wants.

  • It never launched. A preflight check failed and no computation happened: a failed setup, not a failed experiment, and no evidence about the hypothesis.
  • It launched and failed. Something ran and ended badly, with an exit code and its last lines of output. That bounds what the method survives.
  • It launched and nothing came back. The outcome is unknown and the record should say unknown. A run quietly dropped when it stops reporting turns an inconvenient result into an absence.

This is why failures belong in the record. A run that failed cost what a successful one cost and says as much about the method, and a record in which every attempt succeeded tells a reader nothing about the standard applied, only about what was kept.

Materials: identity, not location

An input named by where it sits is named by its most fragile property. A path is local to one machine. A URL can answer 404 or, worse, keep answering while the bytes behind it change: a dataset republished at the same address rewrites the meaning of every result that cited it.

An identity does not go stale. A digest over the bytes says the same thing in ten years, and anyone holding a candidate copy can hash it and compare, so two runs in two institutions that consumed one dataset can be seen to have done so. Where the bytes were never hashed, a reference pinned to an exact revision at least names an immovable version rather than a moving target.

An identity says what to look for and not where to find it, so where to look is recorded separately, as a different kind of statement: attributed, dated, and true only of the moment it was made. A later reader who hits a refusal adds another report rather than correcting the first. What a reader wants out of all of it is obtainability: can I get this, and how? There are four honest answers, and one of them is no: an open copy someone reported, a gated copy with someone named to ask, a rebuild through the run that produced it, or nothing at all.

Restricted data is therefore described rather than excluded. A dataset nobody outside a hospital can obtain is still a node, with the gate stated: worse for a reader than an open copy, and far better than a footnote saying data are available on request.

What a receipt establishes, and what it does not

A receipt for an execution is a report, delivered by the party that ran the work and answerable for it. It is not proof that the computation happened, or that it happened when they say; what each part of a record does and does not establish is the subject of provenance and what it does not prove.

Times are two different facts. The time the author says something happened comes from their clock; the time the server received the report is what the server witnessed. Keeping both, side by side, is more honest than choosing, and a wide gap is itself visible.

The platform re-derives none of it: it opens no output, hashes no file, fetches no reported location and runs no code. Its one question to the world is whether the pinned repository is public and the commit resolves in it, asked when the attempt is registered and not again. Everything else is the author’s account, structured well enough for a reader to go and check it.

Three attempts on one experiment

A constructed example. One experiment asks whether augmentation helps a small classifier, under a plan pinned to a public commit.

AttemptWhat it reportedWhat it is worth
A1Preflight failed: the pinned commit would not check out on the runnerNo evidence about the hypothesis; evidence that the setup needed work
A2Started, then failed after four hours, out of memory, with its last log linesA bound on the method: at this batch size, on this machine, it does not finish
A3A rerun of A2 at half the batch size, succeeded, with a metrics file and digestThe number, and a file to read it out of

A later reader gets more from the three than from A3 alone. They see that the result came on the third try, what the first two cost, and the one declared way in which A3 departs from the plan. They can take the input identity from A3, check whether they can obtain it, and run the same command on the same commit; and if their run disagrees, they have somewhere specific to look.

Further reading

In Substrate

  • An attempt is one execution. Registered before it launches with its commit and command, then reported as it starts and as it ends — succeeded, failed or cancelled — reported time beside received time Attempt receipts: Live. A terminal attempt accepts no further events; a started attempt with no outcome reads as unknown.
  • Materials are named by identity. Content-addressed, or pinned to an exact revision where nothing was hashed, so the same bytes registered in two Rooms are one node Materials: Live.
  • Where to get one is a report. Locations are attributed and append-only, with an access class, and obtainability is derived from them at each read Location reports and obtainability: Live.
  • Connections are written, never inferred. An attempt consumes and produces materials, a plan requires them, a finding cites them as evidence; each edge comes from an explicit write, never from a matching filename Provenance edges: Live.
  • Reported numbers can be found again. A digest over a metrics file’s values, timing fields removed, finds the other places those numbers were reported Values digests: Live: a way to find them, not a verdict.
  • A command does the recording. The capture adapter runs your command against a pinned commit and writes every delivery to a local outbox first Capture adapter: Live, so an outage delays the report instead of losing it.
  • Nothing attests but the author. No signature or environment attestation stands behind a receipt Signed receipts: Idea.

How to register and report an attempt is on Attempts and the capture adapter; identities, location reports and provenance edges are on Materials and provenance. Why a run is registered before it launches at all is the subject of Predictions before data, and the terms are defined once in the glossary.

Open question

A record of an execution is only as good as the reader’s reason to take it seriously, and this one rests on the author being answerable for what they delivered. What would raise that floor, and whether anything short of re-execution by someone else can, is on the laboratory’s research agenda.