What the Archive Does Not Know

The sequel to Archive-to-Royalties. Three ideas — one built, one being implemented, two proposed — about what it means when a system that answers from recorded expertise cannot answer.
What the Archive Does Not Know
Vincent Challier, Spine surgeon, M.D.
Nassim Dehouche, Computer scientist, Ph.D.
The sequel to Archive-to-Royalties. Three ideas across four categories — one built, one being implemented, two proposed — about what it means when a system that answers from recorded expertise cannot answer.
TL;DR. A refusal is not a failure, it is a measurement. Every time a queryable archive cannot answer a clinician's question, it has located a gap between what a specialty asks and what it teaches. Nobody is collecting those measurements, and they are the most valuable thing the system produces. Three ideas follow: refusal behaviour exists and works; refusal logging is being implemented; commissioned knowledge and countable expert opinion are proposals. This piece distinguishes each, because confusing them would be the worst error we could make.
Between morning clinics, a surgeon pulls up a patient's file. A sixty-two-year-old with L4-L5 degenerative spondylolisthesis, grade I slip, stable on dynamic films, MRI showing moderate lateral recess stenosis at L4-L5 with preserved disc height. The question is not whether to decompress — that is settled. The question is whether to fuse. The guidelines offer a recommendation but not a confident one, and the surgeon has seen colleagues go both ways. What she wants to know is what the recorded experience actually says: in patients with this specific profile, what happens to those who were decompressed alone versus those who were fused, at three and five years, for slip progression and reoperation?
She asks the archive. The system thinks for a moment and says it does not know.
The instinctive reaction is disappointment. The interesting reaction is curiosity. What the system just did was locate the exact boundary between what a body of recorded expertise covers and what a practitioner needs. That boundary is not an error. It is a measurement, and nobody is collecting those measurements. They are, we will argue, the most valuable thing the archive produces.
This is the second piece in a series. The first argued that a queryable archive should be metered and its contributors paid. This one asks the question that follows and that nobody has asked: if a system answers questions from a body of recorded expertise, what does it mean when it cannot?
Why refusal is hard
A system that answers everything is useless in medicine. A retrieval system that never refuses is either ignoring its own evidence or fabricating beyond it. The ability to say I do not know — to recognise that a question falls outside the corpus and to stop rather than extrapolate — is the engineering achievement, not the failure mode. This behaviour exists and it works. It is the first thing in this piece that is built rather than proposed.
Refusal is hard because it requires the system to distinguish between three states: the corpus contains an answer and the system found it; the corpus contains an answer and the system failed to find it; and the corpus does not contain an answer at all. Only the third is a true refusal, and only the third is valuable as a measurement. A production system will sometimes confuse the second for the third — a retrieval miss that looks like a knowledge gap. The noise is real and worth stating. But the signal underneath it — questions the corpus genuinely cannot answer — is real too, and it accumulates. Over enough queries, the noise averages and the signal persists, because the same gap surfaces from different users asking different phrasings of the same underlying question. A single refusal may be noise. A cluster of refusals around the same topic, from independent users, across a year, is structure.
Measured ignorance
Aggregate refusals across a year and you have a map of what a specialty asks that its teaching cannot answer. Each refusal record is a question a practitioner asked that the corpus could not resolve. Taken together, they form a demand-side map of a specialty's blind spots — not what a committee guessed people wanted to know, but what working clinicians actually needed and failed to find.
The counterfactual makes the point. How does a congress programme committee currently decide next year's sessions? By argument, seniority, and the previous year's programme. Nobody measures what the membership actually failed to find out. The eleven questions members asked most often last year that the archive could not answer, ranked — that is a programme, written by demand. It is near-term and concrete, and refusal logging to enable it is being implemented now.
The surgeon's question about the spondylolisthesis patient — decompress alone or fuse — is one such data point. She was not the only person to ask something in that territory; she was the one whose phrasing the system could not resolve. A year of those refusals, clustered, is a session at next year's congress — written not by the programme chair's judgement but by the membership's demonstrated need.
The refusal map measures what people thought to ask, not what they should have asked. It is a map of felt ignorance, not of actual ignorance. A practitioner who has been trained not to expect answers in a given area will not ask, and the map will be silent where it should be loudest.
Commissioned knowledge
Once a gap can be measured, it can be priced. A society could post a premium against an unanswered question: the talk that closes this gap earns at a higher rate. Compensation stops being purely retrospective — rewarding what was taught — and starts shaping what gets taught next.
Performing-rights royalties do not commission songs. Publishers do not commission papers by measured demand. A royalty system with a forward-looking channel — where the size of an identified gap sets the reward for filling it — does not, as far as we can establish, exist anywhere. We are proposing it. We have not built it.
Paying for content against a metric invites gaming the metric. Anyone can inflate demand for a topic they intend to teach. The mitigation is structural rather than technical: gap measurement and reward-setting should not be controlled by the same party. That is an argument for governance, not for cleverness, and it is the reason this idea depends on institutional design rather than a better algorithm.
Counting expert opinion
Expert opinion sits at the bottom of every evidence hierarchy, and one reason is mundane: it has never been countable. You cannot say how many experts independently hold a view, how often it has been relied upon, or whether anything in the record contradicts it. It has been a category, not a measurement. Systematic reviews refer to it as a level of evidence, but the level is a label — not a quantity.
With per-passage attribution over a consented corpus, some of that would become countable. Not "experts believe X" but a structured statement: this claim has been drawn upon this many times, asserted independently by this many contributors, unchallenged by any other passage in the corpus, across this period. The unit is not a poll or a vote. It is a traceable record of use, attribution, and the absence of contestation within the corpus.
This matters most where randomised evidence does not and will not exist: rare complications, intraoperative judgement, salvage after failure. These are the situations where accumulated practitioner experience is genuinely the best available knowledge. They are also, not coincidentally, what the talks people travel to hear are about — the sessions that fill rooms at congresses because the literature is silent and the stakes are high.
We do not position this as competing with trials or as a new evidence level. It is making the bottom of the hierarchy legible instead of leaving it as an unexamined residual category. This is a proposal. We have not built an evidence-grading system.
The objection
Retrieval frequency measures salience, not correctness. A confidently wrong claim, repeated by people who trained together, would score highly. The problem is correlated sources. Eleven surgeons who all trained in the same unit under the same chief are not eleven observations. They are one observation with eleven voices. Their convergence proves nothing beyond shared training, and any system that counted them as independent would be measuring echo, not evidence.
Independence is structural and it is measurable. Contributors who never trained together, never co-authored, never shared a faculty, and belong to different societies are structurally independent sources. That is computable from affiliation data — not perfectly, but usefully. The principle comes from outside medicine. Ronald Burt's work on structural holes (Burt, 1992; Burt, 2004) establishes that opinion and practice are more homogeneous within groups than between them. The direct consequence for us: convergence across a structural hole — between groups that do not share training, faculty, or professional networks — is evidence. Convergence within a cluster is an echo. Weight testimony accordingly.
We are borrowing a tool, not inventing one, and saying so is what makes the borrowing credible. Burt's argument was about innovation and brokerage in markets; we are applying it to the problem of counting expert opinion without double-counting a school of thought. The transfer is not original. The application is.
What this does not fix: independent sources can be independently wrong, especially where a whole field has inherited an error. Structural independence raises the evidential value of convergence; it does not make convergence proof. A field-wide blind spot remains invisible to this method, as it does to every method that works within the field's own assumptions.
What it would take
None of this is next quarter. Each idea has dependencies, and naming them concretely is more useful than gesturing at difficulty.
Measured ignorance requires a consented corpus large enough that refusal clusters reflect real gaps rather than the quirks of a small dataset. It requires a year of accumulated refusal data, collected in a way that distinguishes true refusals from retrieval misses. It requires the three-state separation described above to be reliable in practice, not only in principle. And it requires the consent and data-governance infrastructure that makes the corpus legitimate to query in the first place — material recorded under one set of expectations cannot be mined under another without the contributor's knowledge.
Commissioned knowledge depends on governance structures no society has built. Gap measurement and reward-setting must sit with different parties, and the reward formula must be published, contestable, and changeable only through a process that contributors trust. That is institutional design before it is technical design, and it cannot be shortcut by a better algorithm.
Counting expert opinion requires affiliation metadata collected at the point of consent rather than retrofitted later. It requires per-passage attribution — knowing which contributor's material produced which claim, at chunk granularity — over a corpus that does not yet exist at scale. And it requires a methodological framework that would have to survive peer review by the very methodologists most motivated to attack it: the people who built GRADE, who maintain the evidence hierarchies this proposal sits underneath, and who will rightly ask whether structural independence is sufficient, whether affiliation data is accurate enough, and whether the whole enterprise is not just sophisticated consensus polling dressed up as measurement.
The honest timeline is years, not months. What exists: a retrieval system that refuses when it should, and records that it did. What is being implemented: refusal logging at sufficient resolution to aggregate. What is proposed: everything else in this piece.
Close
What changes for a specialty that can see its own blind spots? Not the gaps themselves — those exist whether or not they are measured. What changes is the ability to decide, deliberately, what to do about them. A programme committee that knows where its members consistently fail to find answers is working from different information than one that guesses. A society that can price a gap before it commissions a talk is allocating its educational budget by demand rather than by precedent. A field that can count expert opinion, with its correlated sources and inherited errors made visible, is looking at the bottom of its evidence hierarchy with a clarity it has never had.
The surgeon asked her question and the system said it did not know. That moment — unremarkable, routine, repeated thousands of times across a year — is a measurement nobody is collecting. It tells you where the teaching stops. It tells you what the next congress should be about. It tells you, if you are willing to follow it far enough, something about how expert opinion might become countable instead of merely categorical.
If you work in graduate medical education, research methodology, or society governance, and you think this is wrong — or right — we would like to hear from you.
Sources
Burt, R.S. (1992). Structural Holes: The Social Structure of Competition. Harvard University Press. The foundational text establishing that opinion and practice are more homogeneous within social clusters than between them, and that brokerage across structural holes is a source of informational advantage.
Burt, R.S. (2004). "Structural Holes and Good Ideas." American Journal of Sociology, 110(2), 349–399. Extends the structural holes argument to the relationship between network position and the likelihood of generating ideas that are both novel and valued.
Evidence hierarchies and expert opinion: The placement of expert opinion at the lowest level of evidence hierarchies is conventionally attributed to the Oxford Centre for Evidence-Based Medicine levels and the GRADE framework. The observation that it functions as a category rather than a measurement is ours, offered as a proposal.



















%20Medium.png)


.png)