Predictions before data
Stating hypotheses and plans before results, and what that protects against
How one record stands to another is one thing a reader needs. The order in which they arrived is another, and it is the harder one to recover.
Two sentences can carry the same content and not the same weight. The augmented model will beat the baseline by at least three points on the held-out set is a prediction. The augmented model beat the baseline by 3.4 points, as we expected is an explanation. If the record holds only the second, a reader cannot tell whether the expectation preceded the number or was fitted to it.
That is not a formality of good manners. A prediction made before the data could have failed and did not; an account written afterwards came from someone who already knew which account would fit. Both may be true. Only one of them was ever at risk.
Why the order changes the evidence
Any analysis contains many small choices: which runs to include, which seeds, which metric, where an outlier stops being data. Fixed in advance they are a design; made while a number is on the screen they are a search. Andrew Gelman and Eric Loken called the second case the garden of forking paths: following a single path honestly still inflates the chance of finding something, because the path was chosen with the data in view.
Two related habits have their own names.
- Outcome switching: the result reported is not the one the work set out to measure. Five metrics were computed and the one that moved is the one written up.
- HARKing, hypothesising after the results are known: an explanation drawn from the data is presented as the prediction that the data then tested.
Machines enlarge the problem rather than change it. Automated search multiplies the paths, and the more paths there are to arrive by, the less a winning number means on its own.
What a useful prediction contains
A prediction is useful when someone reading it alone can say what would settle it. Four things have to be in it.
| Part | What it fixes |
|---|---|
| What is being compared | The treatment, its baseline, and how many times each runs |
| On what | The dataset, split or population, named precisely enough to obtain again |
| Measured how | The metric, its unit, which direction is better, how runs are aggregated |
| What would count as failure | The outcome that would make the author say the prediction did not hold |
The fourth is left out most often and does most of the work. A design that cannot fail is not evidence. So write the possible outcomes down, at least two of them, one an honest failure: the null result, the run that will not reproduce, the comparison that could not be made. If every outcome you can imagine would be written up as support, you have not described an experiment.
Two further parts belong beside them. A stopping rule gives the condition under which you halt and report rather than continue, usually a control that has to behave before anything else is readable; when it fires, the design has worked. An interpretation rule says how the number will be read. Both are easy to write beforehand and nearly impossible to write honestly after.
Exploratory and confirmatory work
Exploratory work looks for patterns. Confirmatory work tests a prediction fixed before the data existed. Both are legitimate and research needs both. They are not the same evidence, and a reader who cannot tell them apart reads the weaker one as the stronger.
So the line worth defending is exploratory against confirmatory, not recorded against unrecorded. Exploration belongs in the record; what must not happen is that the label is chosen after the result. The difference is measurable: in psychology, where Registered Reports have reviewers accept a design before any data are collected, Anne Scheel and colleagues found positive results in under half of such studies, against the large majority of the standard literature.
Nor is the standard a plan that never changes. Departures are common, and often undisclosed. What a record can ask instead is that the change is stated and given its own moment of record, so the original and the amendment stand side by side.
The same result, before and after
A constructed example. Written before any run:
On the held-out test set of a ten-class image benchmark, a small convolutional network trained with augmentation will reach at least 3 percentage points higher accuracy than the same network trained without it, averaged over five seeds. A gap under 3 points, or one that does not hold in four of the five seeds, counts against the prediction. If the unaugmented baseline misses its published accuracy by more than 1 point, we stop and report that.
The same work, described once the numbers were in:
Data augmentation improves accuracy by 12 percentage points on our benchmark.
The second sentence may be true. It is still the weaker evidence, for reasons visible in it. It names no comparison it could have lost: 12 points is whatever came out. It does not say how many seeds ran, so one lucky seed and an average of five read identically. And nothing was said in advance about what would count against it, so nothing in it can have failed. The first can be wrong in four stated ways, and that is what gives a prediction its force.
What a record shows about order, and what it cannot
Substrate keeps a hypothesis and a plan as records of their own, each admitted at a moment the server stamps; an attempt is registered before it launches and reports its outcome afterwards, in a separate delivery with its own received time Attempt receipts: Live. So the record holds an order, and it is not the author’s to rearrange.
Two witnesses do that work and neither is enough alone. A pinned commit shows the text did not change afterwards, because it is fixed by its own content. The server’s receipt order shows the sequence, because a position is assigned when a delivery arrives and cannot be moved behind an earlier one. A date inside a commit is neither: it is written by the machine that made it.
The plain limit is this. Receipt time is when the server was told, not when the work happened, and nothing here witnesses the work: Substrate runs no code and watches no machine, and its one question to the world is whether a pinned repository is public and the commit resolves in it. So an author can hold a plan back, run everything, and register the plan before reporting any attempt, and the order would look exactly like the honest one. What the record gives is the sequence of statements, the person answerable for each, and the fact that an earlier statement cannot be changed to fit a later number: a constraint on an author, not a proof, and no badge says a prediction came first.
Further reading
- Andrew Gelman and Eric Loken, The garden of forking paths
- Norbert Kerr, HARKing: Hypothesizing After the Results are Known
- Brian Nosek and colleagues, The preregistration revolution
- Anne Scheel, Mitchell Schijen and Daniël Lakens, An Excess of Positive Results
In Substrate
- A hypothesis is a record. An exact prediction with its scope and the exact versions it rests on as premises Hypotheses: Live, so a changed prediction is a new hypothesis naming the old one, not an edit to it.
- A plan is accepted before the work. An immutable statement of what will be run, pinned to a public repository and a full commit, with the protocol, the metric, the seeds and the rule by which the result will be interpreted Accepted plans: Live.
- Someone is answerable at each point. One member at a time holds an experiment and reports its status as their own report, never as observed process state Take, release and report: Live.
- The server compares a plan with a run. An attempt citing an accepted plan has its repository, commit and argument list compared with the plan’s entrypoint and each arm the plan declares, and the verdict recorded on it Plan match: Live. A mismatch refuses nothing; it is shown beside any declared deviation.
- Intent is a declaration. A plan says whether it is prospective, exploratory or retrospective, and that is the author’s word for it. Substrate checks structure, attribution and permission to publish, not the science.
The procedure is on Hypotheses and experiments, what a run records is on Attempts and the capture adapter, and the words this page uses for records are each defined once in the glossary.
Open question
How much of this can become a property of the medium rather than a promise the author keeps is unsettled: a record can order statements and pin text, and it cannot see the laboratory. Which further witnesses would be worth their cost is on the research agenda.