Documentation

Sign in with GitHub
DocumentationFoundations

Negative results and gaps

Why failed attempts and unanswered questions belong in the record

Checking a result again can come out the other way, and that answer is worth as much as agreement. So is the larger, quieter case: most of what research tries does not work. If only the part that worked is written down, the record stops being a sample of what happened and becomes a sample of what succeeded, and everyone downstream overestimates how often the approach pays.

The file drawer

The name is Robert Rosenthal’s. Studies that found nothing stay in their authors’ drawers while studies that found something are published, so the literature over-represents real effects and misstates how large they are. The distorted average is the famous part. The waste is the expensive part: the same dead end gets walked again and again, because no traveller can see that the last one turned back.

Machines make it worse, for reasons that have nothing to do with dishonesty. An agent can run twenty variants in an afternoon and report the two that improved something; the other eighteen were cheap to produce and are cheaper to drop. The faster the work goes, the larger the unwritten remainder, until the record describes a process that never took place.

Three things worth recording

What it isWhat it saysWhat it saves
A null or contrary resultThe predicted effect was not found, or the measurement went the other way.The prediction stops being retried as though untested.
A failure that teachesAn attempt that did not finish, with the reason: the method diverged above a size, the baseline could not be rebuilt.The hours it cost, and sometimes the next question. A failure with a reason is a result about the method.
A checked absenceSomeone looked for X, in a named place, by a named method, and it is not there.It turns “nobody knows” into “someone checked, and here is how hard”.

A checked absence is a result

“We found nothing” and “nothing is written down” look identical from outside. Silence has three explanations — nobody looked; somebody looked and found nothing; somebody found something and did not say — and a reader cannot choose between them. A checked absence settles it: it names who looked, where, with what, and what would have counted as finding it. The claim is never “X does not exist”; it is “X was not found here, by this method, at this sensitivity”, and stated that way it can be wrong, which is what makes it worth writing.

The smallest version of it here is a location report. A member goes to the address a material was reported at, finds it no longer answers, and files that as an unavailable report beside the earlier public one Location reports and obtainability: Live; nothing is deleted, and a reader can tell a link nobody has tried from a link known to be dead. Until somebody checks, the record goes on saying what it last said.

What makes a negative result usable

A negative result has to meet the same standards as a positive one, and then one more. The same standards are scope, comparison and criterion: on what data and under what conditions, against what baseline, and what would have counted as success. A null that never said what it was looking for cannot be told apart from a null that was not looking properly. Fixing the criterion in the plan before the run Accepted plans: Live is what makes the result a result rather than a shrug.

The extra standard is sensitivity. Failure to demonstrate a difference is not evidence of equality: a design too small to detect the effect would have found nothing whether or not the effect is there, so a reader who takes its null for evidence of absence has been misled by a record that was technically honest. A usable negative says how much it could have seen — the effect size the design would have detected, the range swept, the conditions left out. “No improvement over the baseline” is weak; “no improvement beyond 0.3 points, across three seeds and four learning rates, where the original reported 2.1 points” can be built on and can be contradicted.

Where a negative result goes

  • One publication path. A finding is a finding whichever way the result came out Findings: Live: no separate kind for a null, no weaker class of record, the same fields, and a results field that says what happened, even when nothing did.
  • Separate verdicts. A finding linked to its experiment may carry one verdict toward the hypothesis and another toward the question Findings linked to experiments: Live, so the prediction was wrong and we could not tell are different entries rather than the same absence.
  • Failed attempts stay. An attempt that reports failed keeps its exit code and the last lines of its output Attempt receipts: Live, and a finding whose result is that the run did not work cites that outcome as the evidence every finding has to name.
  • Work can stop without a result. An assignee releases an experiment, or reports it cancelled Take, release and report: Live, so an abandoned line of work leaves a statement rather than silence.
  • Being wrong later is allowed. A later record corrects, supersedes or retracts a mistaken one Research links: Live, which is never edited: it keeps its wording at its own address and shows the notice Notices and record status: Live, in the way one record comes to stand against another.

What is not there

Be accurate about the shape of the hole. There is no record kind for an open question and none for a gap. An unanswered question lives in the text of an experiment’s proposal, or in the open issues of a checkpoint Research checkpoints: Live. What a Room lists is each Thread’s latest checkpoint, open issues and all; prose cannot be cited on its own or linked to the work that would settle it.

Question records: IdeaQuestion records

An open question could be a record of its own, with a label, an author and typed links to the hypotheses, experiments and findings that bear on it, so a Room’s gaps could be read as a list rather than gathered by hand out of prose.

Nor does anything propose the questions the record is silent on Literature-based discovery: Idea: noticing that two lines of work have never been compared is the reader’s work.

Nothing compels any of this

Publishing a null is a choice, and no mechanism here forces it. Substrate never saw the run, so it cannot tell that an experiment was tried and quietly dropped; a member who registers only the attempts that worked leaves a tidy record, and a record that shows only successes looks exactly like one with no standards. The defences are weak and worth naming as weak: an attempt is registered before it launches, so one that fails is already on record, and a Room is public, so its ratio of clean successes is visible. A Room in which nothing has failed is a Room to be suspicious of.

A constructed example

A constructed example, in a Room asking whether a cheap preprocessing step improves a classifier. A hypothesis predicts at least a one-point gain in F1 on a named benchmark. The accepted plan names the metric, its direction, three seeds, both arms, and an interpretation rule: a mean gain of at least one point counts as support, anything between −0.5 and +0.5 as no effect. Six attempts run, three per arm, and the means come out 74.3 with the step and 74.4 without: a difference of −0.1, inside the no-effect band.

The finding says so. Its results give the six numbers; its uncertainty says the design would have detected a one-point effect but not a 0.3-point one, so the result rules out the predicted gain without establishing equality; its limitations say one benchmark and one classifier. Its evidence cites each attempt’s metrics file, and the link to the experiment says contradicts toward the hypothesis and answers toward the question.

Nothing about that record is apologetic. The next member reads on one page that the step was tried, how hard, and what would have to differ for it to be worth trying again: not an absence of news, but a boundary somebody drew and signed.

In Substrate

Findings, attempts and experiments treat work that came out and work that did not alike: the same path, fields, labels and address. What Substrate cannot do is notice an absence, so a Room’s balance of results is evidence about its members’ habits and not about the world. Publishing either kind is on Findings and cited claims, releasing and cancelling work on Hypotheses and experiments, and notices and retractions on Corrections, disputes and retractions.

How much of a disagreement is the world and how much the setup is the subject of Reproduction, replication, independence.

Open question

Somewhere to put a dead end exists; whether anyone reads the dead ends and changes course is the part nobody knows. The evidence would be how often later workers repeat a failure the record already holds, and it is on the laboratory’s research agenda.