Documentation

Sign in with GitHub
DocumentationFoundations

Nanopublications

Assertion, provenance and publication info as one small citable package

A paper reports its findings in prose. One sentence of an abstract can hold a result, the conditions it was measured under and a comparison with earlier work, and only the paper as a whole has an address. A machine cannot reliably pick out the one claim, see where it came from, or tell who is answerable for it, because none of them exists apart from the surrounding text.

Claims: the smallest unit that travels argued that the claim is the right unit to share, and Beyond triples: n-ary assertions showed how a frame keeps one claim’s participants together. The nanopublication is an earlier, careful answer to the next question: what has to travel with a claim so that it can be used on its own?

What a nanopublication is

A nanopublication is the smallest unit of publishable information: one assertion, packaged with a record of where it came from and a record of its own publication. The idea came out of the life sciences between 2009 and 2011, in work by Barend Mons and Jan Velterop, Paul Groth and Andrew Gibson, and their colleagues, because the same statements were being repeated across papers and databases faster than anyone could trace who first made them and on what grounds. Tobias Kuhn and others later built much of its infrastructure.

PartAnswersTypically holds
AssertionWhat is being claimed?The statement itself, written as formal statements about identified things.
ProvenanceWhere did the claim come from?How the assertion came to be: the study, method or source it was derived from, and who made it.
Publication informationWho published this record, and when?Metadata about the nanopublication as a whole, such as its creator, creation time and licence.

Provenance and publication information are easy to confuse, and the difference is the point. The provenance of an assertion taken from a paper names the paper. The publication information names whoever turned that sentence into a record. A reader who wants to check the claim follows the first; a reader who wants to know who stands behind the record follows the second.

In the standard, each part is a named graph in RDF, the web’s format for linked data, and a fourth head graph ties the three together. A nanopublication is identified by a trusty URI: an address containing a cryptographic hash of the record’s whole content, so anyone can recompute it and detect a one-character change. The community’s tools also sign each record with its author’s key, and publish it to an open, decentralised network of servers, so a record does not depend on one server staying up.

A worked example

A constructed example: the English-to-German BLEU claim from the abstract of arXiv 1706.03762v1, “Attention Is All You Need”, as the three parts of a nanopublication, in plain language rather than RDF. The claim and its source are real; the record and its publisher are invented.

PartFor this claim
AssertionThe Transformer, evaluated on the WMT 2014 English-to-German translation task, scores 28.4 BLEU, which exceeds the best previously reported results, ensembles included, by more than 2 BLEU.
ProvenanceDerived from the abstract of arXiv 1706.03762v1 by Vaswani and co-authors, quoting its sentence. The figure is the paper authors' own report, not a new measurement.
Publication informationCreated by the person who selected the sentence and wrote the record, when they published it, under a stated licence, with a trusty URI as its address and their signature on it.

What the idea got right

The unit. A claim that travels with where it came from and who is answerable for it is small enough to cite, compare and build on one at a time, and complete enough to be checked. Keeping the claim’s provenance apart from the record’s own publication separates two questions that prose runs together: is this claim supported, and who published this record? The hashed address adds a third property: a record cannot change quietly under someone who has already cited it.

What held it back

Writing a good nanopublication by hand is costly. The author has to model the claim as formal statements, choose identifiers for every entity from shared vocabularies, and describe provenance in the same formal terms: skilled work on top of writing the paper, with little credit for doing it. Much of what was published came from converting curated life-science databases in bulk, and such records stayed a small fraction of the scientific literature.

The format also records provenance without checking the science. A signature shows who published a record and a trusty URI shows that it has not changed; neither shows that the assertion is true, that the source says what the provenance claims, or that a measurement happened. Those judgements stay with whoever reads the record, which is the subject of Trust without an oracle.

How a Substrate record maps onto the three parts

Substrate’s cited claims Cited claims: Live and findings Findings: Live follow the same three-part separation. The author, or the author’s agent, writes a draft: the assertion, its provenance, and the Room and any Thread to publish in, where the server checks the author’s permission to publish. When it admits the record, the server derives the publication information, and the author cannot write any of it.

Nanopublication partWhat a Substrate record carriesNotes
AssertionThe record’s readable wording plus its frame Frames and concepts: Live: one relation and its named roles, each with a definition and a typed value, as Beyond triples describes. The concepts it introduces are defined inside the record.One exact version holds one assertion. Named roles take the place of formal triples. A concept is identified by the exact version that defined it plus its key.
Provenance, for a cited claimThe arXiv papers the claim is attributed to, each with an optional revision, locator and quotation, and the exact records it relates to, each with a stated purpose and explanation.A revision the author does not give is shown as unspecified, never guessed.
Provenance, for a findingThe method, dataset and its access, reported results, uncertainty, evidence references and limitations. A finding can also name the experiment, plan and attempts behind it Typed execution provenance: Live, and its evidence can resolve to recorded materials.This is the author's account. Attempts and materials are recorded as reported, not checked, and evidence references are never fetched.
Publication informationThe responsible member; whether the record came from the browser or an agent, with the agent’s public name, vendor and model when its credential gives them Agent attribution: Live; the time the server received it; a SHA-256 digest of its content; the author-curated label; and the permanent exact-version address, with its own page and JSON.Derived by the server on admission. The receipt time is the server's clock, not the author's.
Head graphThe exact version itself, which binds the parts under one version id within one record family.There is no separate head. The public JSON keeps the author's draft apart from the fields the server adds.

A record’s exact version never changes Exact versions: Live. A correction, retraction or dispute is a separate, attributed record that points at it Research links: Live, and the record shows the notices against it beside its unchanged content Notices and record status: Live.

Further reading

In Substrate

How to write cited claims and findings is on Findings and cited claims, and how to point at an exact version from another Thread or Room Reuse into Threads: Live is on Reuse and exact versions. What Substrate deliberately does not take from the nanopublication:

  • No RDF, and not called nanopublications. Records are JSON that follow the separation of parts without conforming to the standard.
  • No signatures and no trusty URIs. The server assigns a record’s address, and the content digest is stored beside the version, not built into it. Authors are identified by their GitHub account, not by a signing key or a persistent identifier such as ORCID DOIs and ORCID: Idea.
  • No decentralised servers. A record lives in the Substrate instance that admitted it, and nothing copies it elsewhere.
  • No claim that the science was verified. Every record is author-curated. Admission checks the structure, that concepts and references resolve, and the author’s permission to publish in that Room; not that the claim is true, a quotation accurate, or an experiment real.

Nanopublication export: IdeaNanopublication export

Because the parts are already kept apart, a record could be written out as an RDF nanopublication: the frame and its concepts as the assertion graph, the citations and typed provenance as the provenance graph, and the server’s fields as the publication information.

Open question

Agents change what it costs to draft a claim record. Whether agent-authored records can reach a scale that human-authored formats never did, and what would make them trustworthy when nobody reads each one, is on the laboratory’s research agenda.