Questions where the record is silent
Finding open questions by reading what the record connects and what it leaves out
Every other page in this group is about writing something down carefully; this one asks what can be read from what nobody wrote down at all. A record can hold two well-attested claims and no statement joining them, where the join was not rejected but never considered — knowledge that is published and still unknown, because its pieces sit in literatures that never meet. It is the one place a structured record might propose a question rather than only answer one.
Undiscovered public knowledge
In 1986 the information scientist Don Swanson published two papers on one idea. The first argued that a literature larger than anyone can read becomes a place where facts are public and still inaccessible. The second worked an example. The clinical literature on Raynaud’s syndrome, in which blood flow to the extremities is episodically restricted, reported high blood viscosity, high platelet aggregation and altered vascular reactivity in those patients. The separate literature on dietary fish oil reported that its eicosapentaenoic acid lowered all three. The two bodies of work shared no articles, no authors and no citations in either direction, and nobody had written that fish oil might relieve Raynaud’s. Swanson wrote it down as a hypothesis, explicitly untested. Three years later a double-blind trial found that fish oil delayed the onset of vasospasm in primary Raynaud’s, though not where the condition was secondary to another disease.
The ABC pattern
The shape generalises. One literature relates A to B, another relates B to C, and nothing relates A to C. Where the two relations compose, A to C is a candidate hypothesis, a candidate precisely because the record is silent about it. Fish oil was A, the blood properties B, Raynaud’s C. The method that grew from it is literature-based discovery, and generating pairs was never its hard part; finding the few worth an afternoon is.
Why structure makes a silence findable
Prose hides the pattern: to see that two bodies of work never meet you must read both and hold two vocabularies at once, which is why the famous cases were found by people who went looking. A record made of claims does better, for reasons earlier pages set out.
- Assertions meet on concepts by exact reference rather than by wording, so “these two never meet” is a fact about the record and not an accident of vocabulary. The shape that makes the question a short walk is on Hyperedges and their bipartite shadow.
- Typed links Research links: Live say what one record does to another, from a closed vocabulary, so the absence of a link is as readable as its presence — though a link names records of its own Room and reaches no further.
- Every claim carries its scope, so a proposed join can be tested for whether its two halves apply to the same setting.
A pair of concepts a Room studies separately and never links — no assertion naming both, no link joining them, yet each reachable through a shared third — is a place the record is silent, and silences of that kind can in principle be enumerated.
A silence is not a discovery
A gap is a place to look, not a finding: nothing about its shape suggests anything fills it. A silence has at least three innocent readings: nobody looked; somebody did, in a literature this Room does not hold; or the two relations do not compose, and the pair is an artefact of naming — two labels for one meaning, or one label covering two.
So whatever a search proposes must be screened, and screening is the expensive half. A screen ends in an open question worth an experiment, a pointer to work that already answers it, or nothing — and a pair screened and found empty is worth recording, being a different thing from a pair nobody has looked at. Refuting a proposal is often the better result, because whoever destroys the novelty has usually found the paper the proposer missed.
Literature-based discovery: IdeaLiterature-based discovery
A search over a record’s silences would enumerate never-linked pairs, order them for screening, and hand each to a person or an agent. That order is screening order and nothing more, and would have to say so; otherwise it reads as a claim that the top pair is the most promising science, which no enumeration can know.
What such a search would need
Three things, against what stands in their place.
| What a search needs | What a reader has instead |
|---|---|
| Traversal Neighbourhood read: Idea | An agent follows exact references one read at a time. Nothing walks outward from a record into the neighbourhood around it. |
| Concept identity across Rooms Cross-Room concept identity: Idea | Equal labels never merge, on purpose. One meaning defined in two Rooms is two nodes, so a pair can look unstudied when it has been studied twice. |
| Questions asked of every Room Cross-Room queries: Idea | Concepts and publications can be searched without naming a Room, but every list read was written by hand for an expected question. An unexpected one is assembled by the reader. |
Density is the precondition
The method has a requirement easy to state and hard to meet: the graph must be dense enough that a never-linked pair is surprising. In a small record almost every pair is unlinked, and almost all of them because nobody would ever join them. A search over a record that thin returns mostly noise, and with few shared concepts its ordering is close to arbitrary. Density is the condition under which the idea works at all, not a tuning problem for afterwards, which is why it comes last here.
A constructed example
A constructed example, not a real Room. A Room on training stability for small language models holds two lines of work. One Thread reports that a tighter gradient-clipping threshold reduces loss spikes on a data mixture; another reports that a longer warmup reduces the same spikes on the same mixture. Each names a concept for the spike behaviour, and no assertion in the Room relates clipping threshold to warmup length.
Screening asks three questions. Has anyone outside the Room written the join? An optimiser paper treating clipping and warmup as two controls on one quantity would settle it, and that paper is the useful output. Do the claims apply together — do their scopes share a mixture, a model size and a schedule? Are the two spike concepts one node reused, or two definitions sharing a label? Only if the pair survives all three is there a hypothesis with a direction to predict, that the two controls trade off at fixed stability, and only then something worth an experiment.
Looking for a silence by hand
Nothing enumerates pairs for you, but in one Room the work is within reach.
- Read the definitions rather than the labels: concept search is literal text over labels Literal discovery: Live and hands back each definition with the version holding it. Two concepts the Room treats separately, with no assertion naming both, are a pair to think about.
- Screen it yourself, in the order above, and expect most pairs to fail at the first or the third question.
- Write down what you found. A pair that leads somewhere becomes a hypothesis and an experiment Experiments: Live. One that leads nowhere is worth recording for the same reason that failed attempts and unanswered questions belong in the record: the next reader should not spend your afternoon again.
An open question has no record of its own Question records: Idea, so the place to leave one is a checkpoint, which covers a Thread up to a cursor with its open issues and a next action Research checkpoints: Live.
Further reading
- Don R. Swanson, Undiscovered Public Knowledge, The Library Quarterly, 1986
- Don R. Swanson, Fish Oil, Raynaud’s Syndrome, and Undiscovered Public Knowledge, Perspectives in Biology and Medicine, 1986
- R. A. DiGiacomo, J. M. Kremer and D. M. Shah, Fish-oil dietary supplementation in patients with Raynaud’s phenomenon, The American Journal of Medicine, 1989
- Neil R. Smalheiser, Rediscovering Don Swanson: the Past, Present and Future of Literature-Based Discovery
In Substrate
Substrate proposes no questions, ranks no pairs and screens nothing. What it offers is a record legible enough to read for absences: concepts carrying their definitions, found by literal label text, on Concepts and frames; and a Room’s hypotheses, experiments, findings and open issues in one read A Room in one read: Live, the nearest thing to a view of what it leaves alone. Turning a screened pair into an experiment, and recording one that led nowhere, are on Hypotheses and experiments. What the record cannot be asked is on An archive, not yet a substrate.
Open question
Whether the gaps in a shared record yield genuinely new and productive questions, or only pairs that were never worth joining, is unsettled, and it cannot be settled until a record is dense enough to try it on. That question, and what would count as evidence either way, is on the research agenda.