Documentation

Sign in with GitHub
DocumentationFoundations

Beyond triples: n-ary assertions

Statements with many participants, and how named roles express them

Claims: the smallest unit that travels asked what a statement has to carry to be read on its own. This page asks what becomes of those parts once the statement is data.

Most structured data on the web is written as triples: a subject, a predicate and an object. “The Transformer uses attention” is one. A triple is a good shape for a fact with two participants, and a scientific result rarely has only two.

Take the sentence from the abstract of arXiv 1706.03762v1, “Attention Is All You Need”: the model “achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU.” It names a model, a score, a metric, a task, a comparison and a margin: six participants, held together by one sentence. This page is about keeping them together once the sentence becomes data.

Why triples lose the claim

A constructed example. Write that sentence, and the next one in the abstract, which reports 41.0 BLEU on the WMT 2014 English-to-French task, as triples. Each sentence gives three: the Transformer achieves a score, is evaluated on a task, and is measured in BLEU. That makes six triples, and every one of them is true.

Now put them in one graph, as a triple store does. A graph is a set, so the two copies of “Transformer measured in BLEU” become one, and five edges remain: the Transformer achieves 28.4, achieves 41.0, is evaluated on English-to-German and on English-to-French, and is measured in BLEU. Which score belongs to which task? Nothing in the graph says. Six true edges went in, and neither claim can be read back out.

The failure is quiet. Every edge looks right on its own, and nothing records that a grouping was ever there, so a later reader or agent can pair 41.0 with the German task without contradicting anything stored. What the triples dropped is the binding between the participants, and the binding is the claim.

Step 1 of 3

Triples lose the claim

Two claims from one abstract, split into subject–predicate–object triples. Merged, each edge is still true, but nothing says which score belongs to which task.

The diagram scrolls sideways.

TWO CLAIMS, ONE ABSTRACTClaim 1 · arXiv 1706.03762v1“Our model achieves 28.4 BLEU on the WMT2014 English-to-German translation task …”Transformerachieves28.4Transformerevaluated onWMT 2014 En–DeTransformermeasured inBLEUClaim 2 · arXiv 1706.03762v1“On the WMT 2014 English-to-Frenchtranslation task, our model establishes anew single-model state-of-the-art BLEU scoreof 41.0 …”Transformerachieves41.0Transformerevaluated onWMT 2014 En–FrTransformermeasured inBLEUTHE MERGED TRIPLE SETachieves28.4evaluated onWMT 2014 En–Demeasured inBLEUachieves41.0evaluated onWMT 2014 En–FrTransformerWhich score was reported on which task?Six true triples, five distinct edges, no answer.The sentences are quoted from the abstract of arXiv 1706.03762v1; the triples are constructed for illustration.
Text description
  • Claim 1, from the abstract of arXiv 1706.03762v1: “Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task …”
  • Claim 2, from the abstract of arXiv 1706.03762v1: “On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.0 …”
  • As triples: Transformer achieves 28.4; Transformer evaluated on WMT 2014 En–De; Transformer measured in BLEU; Transformer achieves 41.0; Transformer evaluated on WMT 2014 En–Fr; Transformer measured in BLEU.
  • Merged into one set, the six true triples become five distinct edges, and nothing says which score was reported on which task.

What an n-ary relation is

The arity of a relation is the number of participants it takes. “Uses” is binary: a system uses a method. The German claim is one relation with six participants, which makes it n-ary, and what is true or false is the whole statement: drop any participant and you have a different claim, often a false one.

The participants are not interchangeable either. Each plays a role: 28.4 is the result, not the score of the comparator. A sentence carries roles in its grammar; a data structure has to name them.

Three standard remedies

Semantic web languages met this problem early, because in RDF and OWL “a property is a binary relation”, in the words of the W3C note on n-ary relations. Three remedies are in common use, and each one turns the relation into something that can be pointed at.

RemedyHow it worksWhat it keeps
ReificationMakes a statement itself a resource. RDF's reification vocabulary describes a triple by its subject, predicate and object, and RDF 1.2 adds triple terms, so a triple can be the object of another.Statements about statements: who made this one, and who disputes it. Each still describes one binary triple, so on its own it does not bind six participants.
A relation nodeCreates a new node for each instance of the relation and links every participant to it by a named property. This is the first pattern in the W3C note.All the participants of one statement, each with its role. The node needs no meaningful name: it exists to hold the participants together.
QualifiersKeeps a main property and value and attaches further pairs to it. Wikidata's help page gives Austria's religion as Catholic, qualified by a percentage.Context for one main pair. It fits when one participant is primary and the rest refine it.

The German claim has no primary pair: the score means nothing without the task and the metric. That makes the relation node the natural fit, and it is the pattern Substrate follows.

A frame with typed roles

A frame is a relation plus its named roles, each role with a definition and a typed value: a relation node whose roles are declared. The relation names the kind of situation, such as “compares reported results”, and each role’s definition says what its participant is, such as “comparator: the baseline method”. The word comes from linguistics, where FrameNet, founded by Charles Fillmore, describes situations such as cooking as frames with named participants: the cook, the food, the container and the heating instrument.

A constructed example: the German claim as a Substrate frame, built from the Compares reported results starter template, with two roles its author adds.

RoleTypeValue
relationconceptCompares reported results
subjectconcept, reused by referenceTransformer
comparatorconceptExisting best results
metricconceptBLEU
scopetextWMT 2014 English-to-German translation task
result, addeddecimal with its unit28.4 BLEU
margin, addedtextover 2 BLEU

The margin stays text: the abstract says “over 2 BLEU”, and the number 2 would claim a precision the source does not give. A value the source does not state is left out, never written as zero.

Why a number stays on its frame

Concepts are things other records can point at: the Transformer, BLEU. A value is different. A number is not a thing to be shared; it is a property of this measurement. If 28.4 were a node of its own, every record reporting 28.4 of anything would meet on it, a BLEU score on another task or a percentage in another field, and those meetings would mean nothing.

So the value stays on the frame, with its unit, as a literal that belongs to this one assertion. Substrate writes it as a decimal string, “28.4”, beside the unit “BLEU”, so the digits the author wrote are the digits every reader gets, with no floating-point rounding in between.

How Substrate writes a frame

The assertion of every cited claim Cited claims: Live and finding Findings: Live is its readable wording plus its frame Frames and concepts: Live: wording a person reads, and a frame an agent can read.

  • Relation. One predicate, which is itself a concept.
  • Roles. Two to twelve, each with a name unique in the frame, a definition, and one typed value. Their order carries no meaning.
  • Value types. A concept defined in this record, a concept reused from another, an exact record that is itself a participant, text, a decimal with its unit, or true or false.
  • Concept definitions. A label and a definition for each concept the record introduces, whose identity is the defining version’s address plus its key.

The manual form on the Publish page Publish page: Live offers starter templates such as Compares reported results, and an agent drafting a record for you follows the same conventions. The author renames, adds and removes roles as the claim needs.

What a frame does not settle

Roles are named by authors, not drawn from a closed list. Two authors can call the same participant a comparator and a baseline, and nothing lines the two up. A closed vocabulary would make records easier to compare, and it would force claims into slots that do not fit them. Open names express more and compare less. The templates encourage shared names without enforcing them.

Frames are author-curated. Substrate checks the structure: role names are unique, values have valid types, and every concept the frame names is defined in the record or resolves to an existing one. It does not check that the frame says what the wording says, or that the wording says what the paper says. A frame can misstate a claim as easily as a sentence can, and it looks more authoritative doing it, so read a frame against its wording and its source.

One frame holds one assertion. A sentence that makes two claims becomes two records. The abstract’s German and French results are two frames, and that is exactly what keeps their numbers apart.

Further reading

In Substrate

Cited claims and findings carry their assertion as readable wording plus a frame with one relation and named, defined roles with typed values. Writing a frame by hand is on Concepts and frames, and having your agent draft one and preview it in your chat Local publication preview: Live is on Publish with your agent. The next page, Hyperedges and their bipartite shadow, shows what happens when many frames share their concepts.

Open question

Open role names let an author say exactly what each participant is, at the cost of records that do not line up role by role. Whether agents drafting frames converge on shared role names without a closed vocabulary, or whether records from different Rooms will need an explicit mapping, is on the laboratory’s research agenda.