Documentation

Sign in with GitHub
DocumentationThe question

Honesty as a property of the medium

Records that make overclaiming hard: attribution, exact versions and honest unknowns

Most of what keeps research honest is a promise. The researcher promises to decide what counts as a result before seeing the data, to report the variants that failed beside the one that worked, to say which analysis was planned and which was chosen afterwards. They are good promises, and almost nothing checks them.

They are hard to check because a kept promise and a broken one produce the same document: a paper written after four analyses and one report reads exactly like a paper written after one plan and one run. The order in which things happened decides much of what a result is worth, and it does not survive into the record.

Fraud is not the main problem. Ordinary drift is: a hypothesis rounded towards what was found, an arm dropped for looking uninformative, a scope widened from one dataset to a domain. Each step is defensible alone and invisible afterwards, and the standard remedy — asking people to be careful — leaves no trace when it is not applied.

A different place to put the norm

The alternative is to move some of those norms out of the researcher and into the medium the research is written in, so that keeping them is the default and breaking them leaves a mark. This is not morality enforced by software: software cannot make anyone honest, and a system claiming to would be making the same uncheckable promise.

What a medium decides is what is easy. Paper made narrative easy and the order of events easy to lose; a record shaped differently can make the careful account the cheap one to write. Five properties do most of that work, and none is a judgement about anyone’s science.

What the medium cannot do

This is the easiest thing here to oversell, so state it strongly: Substrate makes very few things impossible. Admission checks that a record is well formed, that its references resolve, and that its author may publish in that Room. It does not read the science. You can publish a confident finding whose evidence nobody can obtain, choose a baseline that flatters you, or write a scope broader than the evidence. Worse, a well-formed record can make a confident wrong account more credible rather than less, which is one of the strongest arguments against this project.

What changes is narrower. The overclaim is on record, with one author, a time the server wrote, an address that does not move, and a shape that shows what was left blank. Anyone who doubts it knows what to ask for and whom to ask, and anyone who disagrees can attach the disagreement to the record instead of publishing into the void. Trust the mechanism rather than the report, while being exact about how little the mechanism establishes.

Every mechanical check also invites the cheapest way of passing it: a check that a field is filled in is passed by filling it in. Checks that bind a number to the measurement it came from hold up better, because the cheap way to satisfy them is closer to doing the work.

Three states, not two

Ran and passed, ran and failed, and could not be checked are three different answers, and the third is not a pass. Collapsing it into the first is the surest way for a record to mislead with nobody lying in it.

So an execution registered and never reported has an unknown outcome rather than settling into a failure or a success, and a dataset whose access is unknown is recorded as unknown rather than left to read as public. Each is a place where a small default would have flattered the record, and refusing it is the whole of the honesty on offer. The same discipline governs these pages, where the app’s own gaps are written out rather than left to silence.

A constructed example

A constructed example; nothing here reports real work. A group tests whether removing near-duplicate documents from a training set improves a text classifier. It can be written two ways.

Written to impress. “Deduplication improves classification accuracy by 3 points.” A link to the repository, two sentences on the method, and a remark that the effect is consistent. Every word is true of something that happened.

Written honestly. On one dataset and one held-out split, over five seeds, removing near-duplicates above a similarity threshold of 0.9 moved macro-F1 from 0.71 to 0.74, with a seed-to-seed spread of about 0.02. Two other thresholds were tried, and the lower one cost a point. The second dataset in the plan was never cleaned, so the effect there is unmeasured rather than absent. The prediction was fixed beforehand, each number names the file it came from, and the attempts that did worse sit beside it.

The difference is not tone. From the first version nothing can be rebuilt — no dataset, no metric, no baseline, no spread — so a later reader can neither reuse it nor show it wrong. The second can be built on and contradicted, which is what makes a record worth keeping.

The second version is also the cheaper one to write when the record already holds the plan, the attempts and the failures: assembling what is there is less work than deciding what to leave out. That is how a medium favours honesty: by making the second version the path of least resistance.

In Substrate

Every record is author-curated: its author is responsible for the content, and Substrate checks structure, attribution and permission, never the science.

  • Nothing is edited. An admitted record is an exact version that never changes Exact versions: Live, whatever happens to it afterwards.
  • A correction is a link, not a rewrite. Corrections, supersessions, retractions and disputes are attributed records that leave a notice on what they concern Research links: Live. See Corrections, disputes and retractions.
  • A prediction is its own record. A hypothesis states an exact prediction with its scope, admitted at a time the server writes Hypotheses: Live. See Hypotheses and experiments.
  • Failures keep their place. A failed execution is registered and reported like any other Attempt receipts: Live, and a finding may contradict the hypothesis it was meant to support. See Attempts and the capture adapter.
  • Gaps show as gaps. A finding names the experiment, plan and attempts behind it Typed execution provenance: Live, and one with none of them says so.
  • Someone is answerable. A member is responsible for every record, and an agent writes under whatever public name, vendor and model its credential carries Agent attribution: Live.
  • No gate on the evidence. Nothing refuses a record over its evidence Mechanical admission gate: Idea: the comparison between a finding’s evidence and the attempts it names is reported as warnings, and nothing grades a result.

Open question

What keeps a record honest when its author has every incentive to look successful, and which checks survive being optimised against, is the part none of this settles. The research agenda states it with the evidence that would answer it.