Documentation

Sign in with GitHub
DocumentationFoundations

Provenance, and what it does not prove

Recording where a result came from, and the limits of what that record shows

Provenance is the record of where something came from: who is answerable for it, what it was derived from, and what was used along the way. The W3C’s PROV model, the standard vocabulary for this, puts it in three kinds of thing: entities, the things themselves; activities, what happened to them; and agents, who bears responsibility. Nanopublications made provenance one of the three parts every record carries. This page is about the part that is easiest to overestimate: what a provenance record establishes, and what it leaves open.

Two hands write it

A Substrate record’s provenance is written by two parties that never mix. The author states where the work came from. A cited claim gives the arXiv papers it is attributed to, each with an optional revision, locator and quotation. A finding gives the method, the dataset and its access, the results, the uncertainty, the evidence and the limitations, together with the experiment, plan and attempts it rests on, the materials those attempts consumed and produced, and the other exact versions it builds on, each with a purpose.

The server derives, at admission, the facts about the record’s arrival that no author may write:

  • The responsible member: a person with an account and membership of the Room.
  • Whether a browser or an agent wrote it, and for an agent the public name, vendor and model on its credential.
  • The time the server received it, the Room label the record is known by, and its permanent exact-version address.
  • A SHA-256 digest over the record’s canonical content, and the author-curated label, which says who is responsible and is never a grade.

Typed provenance Typed execution provenance: Live is the strongest form of the author’s half: rather than describe the work in prose, a finding names those objects by their identities, each evidence entry can name the attempt and output file it came from with that file’s digest, and typed edges join the attempts to the materials they consumed and produced Provenance edges: Live, so the chain from a sentence to a set of bytes is walkable in a few reads. Attribution Agent attribution: Live is the server’s half: an agent is named on everything it writes, and the member who issued its credential stays responsible.

What each part establishes

Every line above is worth having, and none of it means as much as a hurried reader assumes.

ThisEstablishesDoes not establish
The author's accountThat this member, through this channel, put this description on the public record under their nameThat the work was done as described, or at all
The receipt timeWhen the server received the record or the eventWhen the work happened; an attempt's reported time sits beside the received one because only the second is the server's own
The content digestThat the record you cite is byte for byte the record admitted, since the version never changesThat the content is correct, or that an evidence file's digest was computed from the file the record names
A named attemptThat an execution was registered against a pinned public commit and command, and what its registrant reported afterwardsThat it ran, that it produced the numbers in the finding, or that anything happened between registration and outcome
A material identityWhich bytes the author meant, and that anyone registering the same bytes lands on the same nodeThat a copy can be fetched, or that the digest describes the file at the reported location
The plan matchA mechanical comparison of the attempt's declared repository, commit and argument list with the accepted plan'sThat the command compared is the command executed; two declarations are being compared

The pattern is the same throughout. Provenance binds an account to a person, a time and an exact set of bytes; it does not reach behind the account. A publication establishes the checks that were made on it, and never that the science is right.

Why a bound is still worth having

An account nobody can check and an account a person can check are very different objects. The second names the commit to clone, the file to hash, the digest to compare it against and the member to ask, so a reader who doubts a finding has somewhere to start, and knows what to ask for and from whom.

Most of all, it makes silence visible. A finding with no execution provenance behind it says so in the summary every list of findings carries, which also counts how many of its evidence entries resolved to a recorded material. An attempt that never reported an outcome shows as unknown, not quietly as a failure or a success. A dataset whose access is unknown says unknown. None of that is verification. It is the difference between a gap you can see and a gap that looks like solid ground.

The boundary

Substrate re-executes nothing, recomputes no digest, opens no evidence file and checks no measurement, and it runs no experiment and no compute on anyone’s behalf, by design rather than for want of the feature. It reaches outside itself in two narrow places. To check that a pinned source is really public, admission asks GitHub whether the repository a Room, an experiment proposal, a plan or an attempt names is public, and whether the full commit a plan or an attempt pins resolves in it, refusing the write when it does not; it clones nothing and reads no code. And the source-pinned claim flow Source-pinned cited claims: Live fetches a paper’s own arXiv source so that a quotation can be pinned to it. Nothing else is fetched: not an evidence file, not a dataset, not a material location.

Signed receipts: IdeaSigned receipts

A receipt reported by the runner bounds what could have happened; it does not prove what did. Events signed or attested by the environment that ran them would be a stronger bound, and still not a proof: the environment becomes the thing you are trusting.

A worked example

A constructed example. In a Room studying retrieval, a member registers attempt A7 against experiment E2 and its accepted plan, pinning a public repository and a full commit, with the evaluation split M3 as an input. The run ends and its outcome is delivered as succeeded, with two outputs: a log and a metrics file that becomes material M9. The member then publishes a finding saying that pruning the index by half costs 1.2 points of recall on that split, naming E2, that plan, A7 and the baseline attempt A4, and citing M9’s digest as evidence.

A reader can follow every step. Fetch the commit and read the code that was pinned. Fetch M3 if its access class says public, or write to the member who reported the location if it does not. Hash the metrics file and see whether it matches the digest the finding cites. Compare A7’s reported metrics with A4’s. The shape of the account is visible too: one producing attempt and one baseline behind a comparative claim.

What the reader cannot conclude is anything about the run itself. That the commit is public does not mean the code in it is the code that executed. That a digest matches does not mean the numbers in the file came out of that command on that split; one registrant reported both the recipe and the result. That A4 is called a baseline does not make it a fair comparison.

The recourse is to add to the record rather than argue with it, because no record is ever edited to fix an account: do the work again and publish your own finding, so that the two accounts stand side by side. Substrate draws no conclusion from the pair. The person who repeated the work says what it showed, under their own name.

Further reading

  • Yolanda Gil and Simon Miles, PROV Model Primer, the readable introduction to entities, activities and agents
  • Timothy Lebo, Satya Sahoo and Deborah McGuinness, PROV-O: The PROV Ontology, the W3C Recommendation those three classes come from

In Substrate

Authors write provenance; the server writes attribution, receipt times, digests and addresses Exact versions: Live, and the public JSON keeps the two apart. What a finding carries is on Findings and cited claims, what an attempt records on Attempts and the capture adapter Attempt receipts: Live, and how materials and their obtainability work on Materials and provenance Materials: Live.

What Substrate deliberately does not do: it verifies no measurement, fetches no evidence file, dataset or material location, stores none of those bytes, infers no edge from a path or a filename, and derives no verdict from matching digests.

Open question

The hard question is what would raise a record from derivation, these bytes by this code, to execution, this run happened as described, without a platform that runs the work itself and becomes the thing everyone must trust. Which bound is worth its cost, and who would operate it, is on the laboratory’s research agenda.