Meaning without a master ontology
Defining concepts where they are used, and reusing them by exact reference
Two records use the word transformer. One means the architecture introduced in arXiv 1706.03762v1. The other might mean that architecture, or any network built from stacked self-attention layers, or a device in an electrical substation. A person reading both records can usually tell which. A program cannot, and neither can an agent assembling premises out of a hundred records it did not write.
Hyperedges and their bipartite shadow showed that records meet on concepts: a concept is a node that many assertions can point at, and two assertions pointing at one node are reachable from each other. This page is about the question that leaves open: when should two records point at the same node, and who decides?
The ontology answer
The established answer is to fix the vocabulary in advance. An ontology or controlled vocabulary gives each thing one identifier, one maintained definition and a place in a hierarchy, so software can tell that two records are about one thing. Wikidata, MeSH and the Gene Ontology work this way, and they work: the Gene Ontology is kept by a consortium with a team of editors and is, in its own words, constantly revised as biological knowledge accumulates.
That maintenance is the cost, and it is not incidental. Someone has to decide what a term means for everyone who uses it, and that decision is hardest exactly where the science is unsettled, which is where the interesting terms are. Using a new one means proposing it, waiting, and accepting an editor’s judgement about what it covers.
A vocabulary is therefore always behind the literature. A field coins a name in a preprint, a dozen groups adopt it in slightly different senses, and some of those names are abandoned before any registry has an identifier for them. A system that admits only terms an authority already holds cannot record the work that most needs recording.
Defined where it is used
Substrate takes the other route: the definition is local and the identity is exact. A concept is a defined term inside a record, a label and a definition the author writes where the term is first needed Frames and concepts: Live. On admission it takes a permanent identity made of the record’s exact version and the concept’s key, and that version’s definition never changes Exact versions: Live. Publishing a revised definition creates a new identity and leaves every earlier assertion as it was.
Nothing approves the term, so a word is usable the moment someone defines it and no author is asked to pick the nearest wrong identifier from a list. The definition is what a later reader inspects, so it is worth writing for a stranger. The price is that a label carries no authority at all: two records that both write transformer hold two concepts, and no operation joins them.
Reuse is the only way to share
Two records come to share a concept in exactly one way. The later record names the earlier concept by exact reference, giving the defining version and the key where it would otherwise define a term of its own, and it may name a definition from any Room. Admission resolves that reference and refuses the whole record if no such concept exists, so a stored reference always points at a real definition. It says nothing about whether the meaning fits. That judgement is the author’s, and the definition can be read first, on the defining record’s page or through a read of its own.
Reuse is a small, deliberate act, and it says one thing only: this role means what that definition says. It commits neither author to the other’s claim and endorses nothing.
The cost is fragmentation
Two authors who mean the same thing and never reuse each other’s definition leave two nodes that no query joins, and a Room that defines rather than reuses accumulates near-duplicates. Nothing declares two definitions the same, and there is no shared store in which concepts from different Rooms can be declared identical Cross-Room concept identity: Idea. Whether two definitions describe one meaning is an argument, and an argument belongs in a Thread, under the name of whoever makes it Threads: Live.
A wrong merge costs more than a missing one
A constructed example. One Room studies retrieval and defines a concept labelled coverage: the share of a collection’s relevant documents that a query set reaches. Another Room studies test generation and defines its own coverage: the share of a program’s branches that a test suite executes. Each Room publishes a finding that fills a role with its own coverage concept, and each reports that coverage rose from 0.71 to 0.86.
A system that merged on the label would make one node holding two incompatible meanings, and the two findings would meet on it. An agent asking what else is said about coverage would be handed the other finding with nothing to warn it, because a merged edge looks exactly like the edge an author makes by deliberate reuse. A reader comparing 0.71 with 0.86 would be comparing nothing at all. And the mistake travels: every later record that builds on the join inherits it, and the join has no author to ask about it.
Fragmentation fails in the other direction, and fails visibly: a search finds half of what exists and an agent misses work it should have read. Anyone who notices can repair that — read both definitions, reuse one, and the records meet from then on, joined by a named person who can be asked why. A wrong merge is repairable only by someone who spots it, and nothing points at it.
Finding a concept to reuse
Concepts are found by literal text in their labels Literal discovery: Live: a case-insensitive substring, no wildcards, optionally narrowed to the Room a definition originated in or to one defining version. Each result carries the full definition, the defining version and key, the author and the origin Room, so whoever is choosing can read the meaning before pointing at it.
What that cannot do is find a meaning. A search for one word never reaches a definition whose label chose another word for the same idea, there are no synonyms and no hierarchy, nothing searches the text of the definitions, and a search for coverage returns every unrelated coverage beside the one that fits. Discovery is literal because the alternative is a system that guesses at meaning and hands its guesses on as structure.
Compared with a controlled vocabulary
| Aspect | Controlled vocabulary | Defined where it is used |
|---|---|---|
| Where meaning is fixed | Centrally, in the vocabulary, by maintainers and editors | In the record that introduces it, by its author, permanently |
| What a match means | Two records on one identifier are about one thing, as the maintainers define it | Two records on one concept are reachable from each other, and nothing more |
| What it risks | A wrong merge nobody sees, and terms that arrive late or never | Fragmentation: several definitions of one meaning, and a search that finds some of them |
The two are not opposites in principle: vocabulary work usually reconciles them by letting a local term declare a match to an external identifier with a stated strength, such as exact, broader or related.
Further reading
- Antoine Isaac and Ed Summers, SKOS Simple Knowledge Organization System Primer, on how controlled vocabularies are modelled, graded match relations included
- Harry Halpin, Patrick Hayes, James McCusker, Deborah McGuinness and Henry Thompson, When owl:sameAs Isn’t the Same: An Analysis of Identity in Linked Data, on how strong an identity claim really is
- The Gene Ontology, an example of a maintained vocabulary and the editorial work that keeps one alive
In Substrate
A record defines its own concepts, each with a label and a definition and an identity made from the record’s exact version and the key; a later record shares one only by naming that identity. How to write definitions worth reusing, read one before you commit to it, and search labels is on Concepts and frames; what a concept does inside an assertion is on Beyond triples: n-ary assertions.
What Substrate deliberately does not do: it keeps no central vocabulary, grounds no concept to an external identifier, merges no two definitions however alike their labels, and draws no edge an author did not write. A Room that wants its terms used has one way to get them used, which is to define them well enough that a stranger reuses them.
Open question
Whether independently written definitions converge on their own, given good enough search and agents patient enough to read before they define, is something a record layer can measure. So is the harder half: what evidence should be required before anyone declares two definitions the same. Both are on the laboratory’s research agenda.