Beyond triples: n-ary assertions
Statements with many participants, and how named roles express them
Claims: the smallest unit that travels asked what a statement has to carry to be read on its own. This page asks what becomes of those parts once the statement is data.
Most structured data on the web is written as triples: a subject, a predicate and an object. “The Transformer uses attention” is one. A triple is a good shape for a fact with two participants, and a scientific result rarely has only two.
Take the sentence from the abstract of arXiv 1706.03762v1, “Attention Is All You Need”: the model “achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU.” It names a model, a score, a metric, a task, a comparison and a margin: six participants, held together by one sentence. This page is about keeping them together once the sentence becomes data.
Why triples lose the claim
A constructed example. Write that sentence, and the next one in the abstract, which reports 41.0 BLEU on the WMT 2014 English-to-French task, as triples. Each sentence gives three: the Transformer achieves a score, is evaluated on a task, and is measured in BLEU. That makes six triples, and every one of them is true.
Now put them in one graph, as a triple store does. A graph is a set, so the two copies of “Transformer measured in BLEU” become one, and five edges remain: the Transformer achieves 28.4, achieves 41.0, is evaluated on English-to-German and on English-to-French, and is measured in BLEU. Which score belongs to which task? Nothing in the graph says. Six true edges went in, and neither claim can be read back out.
The failure is quiet. Every edge looks right on its own, and nothing records that a grouping was ever there, so a later reader or agent can pair 41.0 with the German task without contradicting anything stored. What the triples dropped is the binding between the participants, and the binding is the claim.
Step 1 of 3
Triples lose the claim
Two claims from one abstract, split into subject–predicate–object triples. Merged, each edge is still true, but nothing says which score belongs to which task.
One hyperedge
A hyperedge joins any number of participants at once. The whole claim is one edge, and each participant carries the role it plays in it.
The frame node
The claim becomes one node, the assertion, with a typed row for each role. A concept is named by reference; a literal such as 28.4 BLEU stays on the frame and belongs to this assertion alone.
The diagram scrolls sideways.
Text description
- Claim 1, from the abstract of arXiv 1706.03762v1: “Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task …”
- Claim 2, from the abstract of arXiv 1706.03762v1: “On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.0 …”
- As triples: Transformer achieves 28.4; Transformer evaluated on WMT 2014 En–De; Transformer measured in BLEU; Transformer achieves 41.0; Transformer evaluated on WMT 2014 En–Fr; Transformer measured in BLEU.
- Merged into one set, the six true triples become five distinct edges, and nothing says which score was reported on which task.
- One hyperedge, labelled Compares reported results, encloses six participants:
- subject: Transformer (concept reused by reference)
- comparator: Existing best results (concept)
- metric: BLEU (concept)
- scope: WMT 2014 English-to-German translation task (text)
- result: 28.4 BLEU (decimal with unit)
- margin: over 2 BLEU (text)
- Wording: “Vaswani et al. report that the Transformer achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles, by over 2 BLEU.”
- Assertion node with relation Compares reported results (concept): A reported comparison of two methods under a stated evaluation protocol.
- subject (concept reused by reference): Transformer. The method being evaluated.
- comparator (concept): Existing best results. The baseline method.
- metric (concept): BLEU. The measurement and its interpretation.
- scope (text): WMT 2014 English-to-German translation task. The dataset and evaluation conditions.
- result (decimal with unit, added by the author): 28.4 BLEU. The score the subject reports on the metric in this scope.
- margin (text, added by the author): over 2 BLEU. The stated improvement over the comparator, in the source's words.
- Literal values stay on this assertion and are never shared with another.
Hover, focus or tap a participant to pick out its role.
Hover, focus or tap a row to read its definition.
What an n-ary relation is
The arity of a relation is the number of participants it takes. “Uses” is binary: a system uses a method. The German claim is one relation with six participants, which makes it n-ary, and what is true or false is the whole statement: drop any participant and you have a different claim, often a false one.
The participants are not interchangeable either. Each plays a role: 28.4 is the result, not the score of the comparator. A sentence carries roles in its grammar; a data structure has to name them.
Three standard remedies
Semantic web languages met this problem early, because in RDF and OWL “a property is a binary relation”, in the words of the W3C note on n-ary relations. Three remedies are in common use, and each one turns the relation into something that can be pointed at.
| Remedy | How it works | What it keeps |
|---|---|---|
| Reification | Makes a statement itself a resource. RDF's reification vocabulary describes a triple by its subject, predicate and object, and RDF 1.2 adds triple terms, so a triple can be the object of another. | Statements about statements: who made this one, and who disputes it. Each still describes one binary triple, so on its own it does not bind six participants. |
| A relation node | Creates a new node for each instance of the relation and links every participant to it by a named property. This is the first pattern in the W3C note. | All the participants of one statement, each with its role. The node needs no meaningful name: it exists to hold the participants together. |
| Qualifiers | Keeps a main property and value and attaches further pairs to it. Wikidata's help page gives Austria's religion as Catholic, qualified by a percentage. | Context for one main pair. It fits when one participant is primary and the rest refine it. |
The German claim has no primary pair: the score means nothing without the task and the metric. That makes the relation node the natural fit, and it is the pattern Substrate follows.
A frame with typed roles
A frame is a relation plus its named roles, each role with a definition and a typed value: a relation node whose roles are declared. The relation names the kind of situation, such as “compares reported results”, and each role’s definition says what its participant is, such as “comparator: the baseline method”. The word comes from linguistics, where FrameNet, founded by Charles Fillmore, describes situations such as cooking as frames with named participants: the cook, the food, the container and the heating instrument.
A constructed example: the German claim as a Substrate frame, built from the Compares reported results starter template, with two roles its author adds.
| Role | Type | Value |
|---|---|---|
| relation | concept | Compares reported results |
| subject | concept, reused by reference | Transformer |
| comparator | concept | Existing best results |
| metric | concept | BLEU |
| scope | text | WMT 2014 English-to-German translation task |
| result, added | decimal with its unit | 28.4 BLEU |
| margin, added | text | over 2 BLEU |
The margin stays text: the abstract says “over 2 BLEU”, and the number 2 would claim a precision the source does not give. A value the source does not state is left out, never written as zero.
Why a number stays on its frame
Concepts are things other records can point at: the Transformer, BLEU. A value is different. A number is not a thing to be shared; it is a property of this measurement. If 28.4 were a node of its own, every record reporting 28.4 of anything would meet on it, a BLEU score on another task or a percentage in another field, and those meetings would mean nothing.
So the value stays on the frame, with its unit, as a literal that belongs to this one assertion. Substrate writes it as a decimal string, “28.4”, beside the unit “BLEU”, so the digits the author wrote are the digits every reader gets, with no floating-point rounding in between.
How Substrate writes a frame
The assertion of every cited claim Cited claims: Live and finding Findings: Live is its readable wording plus its frame Frames and concepts: Live: wording a person reads, and a frame an agent can read.
- Relation. One predicate, which is itself a concept.
- Roles. Two to twelve, each with a name unique in the frame, a definition, and one typed value. Their order carries no meaning.
- Value types. A concept defined in this record, a concept reused from another, an exact record that is itself a participant, text, a decimal with its unit, or true or false.
- Concept definitions. A label and a definition for each concept the record introduces, whose identity is the defining version’s address plus its key.
The manual form on the Publish page Publish page: Live offers starter templates such as Compares reported results, and an agent drafting a record for you follows the same conventions. The author renames, adds and removes roles as the claim needs.
What a frame does not settle
Roles are named by authors, not drawn from a closed list. Two authors can call the same participant a comparator and a baseline, and nothing lines the two up. A closed vocabulary would make records easier to compare, and it would force claims into slots that do not fit them. Open names express more and compare less. The templates encourage shared names without enforcing them.
Frames are author-curated. Substrate checks the structure: role names are unique, values have valid types, and every concept the frame names is defined in the record or resolves to an existing one. It does not check that the frame says what the wording says, or that the wording says what the paper says. A frame can misstate a claim as easily as a sentence can, and it looks more authoritative doing it, so read a frame against its wording and its source.
One frame holds one assertion. A sentence that makes two claims becomes two records. The abstract’s German and French results are two frames, and that is exactly what keeps their numbers apart.
Further reading
- Natasha Noy and Alan Rector, Defining N-ary Relations on the Semantic Web, a W3C Working Group Note
- RDF 1.2 Concepts and Abstract Data Model, for triple terms and reifiers
- RDF Schema 1.1, for the original reification vocabulary
- Help:Qualifiers, Wikidata’s guide to qualifying a statement
- FrameNet, frames and their participants in English, after Charles Fillmore’s frame semantics
In Substrate
Cited claims and findings carry their assertion as readable wording plus a frame with one relation and named, defined roles with typed values. Writing a frame by hand is on Concepts and frames, and having your agent draft one and preview it in your chat Local publication preview: Live is on Publish with your agent. The next page, Hyperedges and their bipartite shadow, shows what happens when many frames share their concepts.
Open question
Open role names let an author say exactly what each participant is, at the cost of records that do not line up role by role. Whether agents drafting frames converge on shared role names without a closed vocabulary, or whether records from different Rooms will need an explicit mapping, is on the laboratory’s research agenda.