Documentation

Sign in with GitHub
DocumentationFoundations

Reproduction, replication, independence

The different ways a result can be checked again, and why independence matters

However fully an execution and the materials it touched are recorded, a result nobody has checked again is one person’s report. Checking it again covers several different activities, whose value lies in what each one kept the same and what it changed. Two respectable vocabularies use the main pair of words in opposite directions, so this page fixes them.

Four words, four different activities

ActivityWhat happensWhat it tests
RerunThe same person runs the same code on the same inputs again.Whether the pipeline is deterministic. It tests the instrument, not the claim.
ReproductionSomeone works from the recorded recipe — pinned commit, command, same materials — and sees whether the numbers come back.Whether the recipe is complete and yields what was reported.
ReplicationSomeone tests the same claim with a different setup: other data, another implementation, a different team.Whether the finding is a property of the world rather than of one pipeline.
ReplayA neutral party re-executes the recorded recipe mechanically.What a reproduction tests, with the author out of the loop.

Independence is not a fifth activity. It is a property any of the four can have more or less of, and it is about what the second run does not share with the first: the person, the code, the data, the comparator, the hardware, the analysis. Two results that share materials are a paired contrast, not two independent results. Change the person and one failure mode goes away, because a mistake in the analysis is one its own author would repeat; change the data and another goes. Change nothing and you have learned that your machine is consistent.

Respected sources disagree. The ACM’s badges award Results Reproduced when a later team obtains the main results using the authors’ own artefacts, and Results Replicated when they obtain them without those artefacts. The US National Academies fix reproducibility as the computational case — same input data, code and analysis — and replicability as consistent results from a study that collected its own data. Both agree with the table above; a good deal of older work uses the pair the other way round. So do not trust the label: whichever word you write, say what was held fixed and what varied.

Exact equality is the wrong test

Demand identical numbers and almost every honest reproduction fails. A seed drawn in a different order moves the result. A parallel reduction over a different number of devices adds the same values in a different order, and floating-point addition is not associative, so the totals differ in the last digits. A library upgrade changes a kernel; other hardware picks another algorithm. None of that is a defect, and a test that treats it as one rejects good work.

So agreement has to be defined: on which metric, over how many runs, summarised how, and within what margin. And the margin has to exist before the second number does. A tolerance chosen after seeing both numbers is a description rather than a test, and it will be generous in exactly the cases where the author wants it to be. The criterion belongs in the accepted plan Accepted plans: Live, beside the metric, its direction and the seeds: the same discipline as writing the prediction down before the data, applied to the comparison. Declared in advance, 0.8241 against 0.8238 is a pass anyone can check; without one it is an argument.

What a matching digest tells you

Sometimes the numbers do match exactly. A digest over the values in a metrics file Values digests: Live is equal when the same values were reported, timing fields aside. That is evidence about determinism and nothing else. It does not establish that the finding is valid, that the measurement was the right one, or that the evaluation data stayed out of training.

It is also silent about independence. If the second run consumed the same bytes by identity, both runs share whatever is wrong with those bytes — a contaminated split, a mislabelled column, a truncated file — both report the fault identically, and the agreement looks like confirmation. A matching digest across runs on shared inputs is a strong statement about the pipeline and a weak one about the world.

Who decides that something was reproduced

Substrate runs nothing and re-derives nothing: it executes no plan, opens no output file and fetches no reported location. What it does is keep the second attempt beside the first Attempt receipts: Live, with the intent its registrant declared; a member who did not take the experiment may register one with the intent reproduction Reproduction attempts: Live, and the assignment does not move. The server does compare that attempt’s repository, commit and argument list with the accepted plan’s Plan match: Live, which sets declarations against declarations and refuses nothing either way.

Whether any of that amounts to a reproduction is a judgement, and the place to state it is a finding Findings: Live, under the name of the person who made it: what was held fixed, what varied, what agreement was being tested and whether it held. The finding names the attempts it rests on Typed execution provenance: Live and can carry a verdict toward the hypothesis Findings linked to experiments: Live. A reader is left with an attributed claim, the attempts behind it and no platform verdict. Nothing promotes a claim as reproductions accumulate Replication ladder: Idea; you count the independent reports and judge how independent they were.

What makes a result reproducible in practice

Most of the recipe a second worker needs is what an attempt already records: the code at one commit, so that a reference names an exact version rather than the latest; the inputs by identity Materials: Live; the command argument by argument, including which arm of the plan it is; the environment, recorded only if your run records it; and the outputs the numbers were read from. That is the subject of attempts, receipts and materials.

Reproduction adds one requirement of its own, the expensive one: the inputs have to be obtainable by somebody else Location reports and obtainability: Live. A result whose inputs nobody outside one institution can get is not reproducible by a stranger, however complete the rest of the recipe. That is normal in medicine and in industry, and the answer is to say so: record the location as restricted, name who to ask, and put in the finding’s limitations what a reader cannot check.

A constructed example

A constructed example, in a Room studying a data-augmentation step. The experiment’s assignee registered attempt A3 against an accepted plan: a pinned commit, the command with the augmentation arm selected, one input corpus named by digest, three seeds, and an interpretation rule under which two runs agree when their mean accuracy differs by no more than 0.005. A3 reported 0.8241.

A second member fetches the corpus from its latest location report, which is public, checks out the commit and runs the same command on their own hardware. They register A7 with intent reproduction, naming the same corpus in its inputs; it reports 0.8238. The values digests differ, as they must when the numbers do, and the plan’s rule is satisfied.

They publish a finding saying the result reproduces on second hardware within the plan’s margin. Its results give both numbers; its uncertainty says both runs consumed the same corpus, so a fault in it would appear in both; its limitations say the original environment was not matched. The finding names A7 and A3 in its execution provenance and cites A7’s metrics file as evidence.

Nothing in the record reached that conclusion. The reproduction is the second member’s statement, with their name on it, and a third member who thinks 0.005 too generous can say so in a record of their own.

Further reading

In Substrate

Second attempts and second findings sit beside the first, and Substrate draws no conclusion from either. It re-executes nothing itself Platform replay: Idea, and nothing outside the author attests that an execution happened at all Signed receipts: Idea, so every rung above “the author says so” has to come from people. That deliberate boundary is stated in full on An archive, not yet a substrate; registering and reporting an execution is on Attempts and the capture adapter, and what a finding may say about one on Findings and cited claims.

A reproduction that came out against the original is a finding like any other; what makes such a record usable is on Negative results and gaps.

Open question

What agreement should mean when a rerun cannot match exactly, and how independent a second worker has to be before their agreement counts, are both unsettled; so is whether authors will fix a criterion in advance and readers accept the judgements afterwards. All of it is on the laboratory’s research agenda.