Prediction outcomes
question · corpus proposedAn earlier declaration was accepted by Zach Mainen; it has moved since — this is 826012b59a12, unaccepted. · no paper read under it
Did the paper's predictions come out, and does the tree say so?
A decision about meaning, and then the propagation of it through everything downstream.
zmainen/elife-claim-trees#28 — the decision is taken there; the propagation is taken here.
How it works
The rule is mechanical, and it is now rule v3. For each role: prediction claim it looks at the edges pointing in and out, and sorts the prediction by the shape of that wiring: a tests edge with an outcome (confirms or refutes) beside it, a tests edge alone, nothing, or a conjunction of several. It reads the graph and nothing else — not the paper, and not whether the prediction was in fact met.
That limit is the point of running it. The issue #28 ruling settled the wiring convention — a tests edge carries its outcome beside it — so the buckets now measure which trees have had the outcome filled in, not which convention a paper happened to follow.
What it found · added 2026-09-11
tests is neutral by design, so a tree can record that a prediction was tested and never whether it was met. Of 47 predictions, 23 now carry a recorded outcome and 23 carry a tests edge and no outcome yet. The issue #28 ruling gave the vocabulary its missing half — refutes, the negative outcome beside confirms — so a failed prediction can now be written down; what remains is filling the rest in.
Rests on
- claim-format · feature
- relation-vocab · question
How it is defined
What this layer reads besides its dependencies. Each is a declared input: its content is hashed into every run, so editing one makes those runs stale.
The declaration names this path and the repository does not have it. An input that does not exist hashes to nothing, so it cannot make a run stale — the layer is declared to depend on something it is not in fact tracking.
This layer is proposed, and what is proposed is the rule
Nothing on this page has been accepted, and no claim in the corpus has been changed by it. What is offered for review is a rule and its behaviour across all 47 predictions in the corpus — not 47 decisions to be taken one at a time.
So the useful way to read it is: does the rule sort these correctly, does it refuse the cases it should refuse, and would applying it leave the corpus in a better state than it is now? The 1 abstentions are listed in full below, because a rule that declines what it cannot settle is what makes reviewing a sample sufficient — attention goes to the margin instead of the confident middle.
The question
A prediction is a commitment made in advance: if this hypothesis holds, then this particular measurement should come out this particular way. It is the part of a paper's argument that can fail, which is what makes it worth recording separately from the result that tests it.
In this corpus a result is joined to the prediction it bears on by a relation called
tests. tests is neutral by design — it
records that a test happened, not which way it came out. That neutrality is correct: the
two facts are genuinely different, and collapsing them would make every test look like a
success.
The problem is that nothing else records the outcome either. There is a positive
relation, confirms, used 19 times
on predictions here. There is no negative counterpart. The two
oppositional relations do not fill the gap: contradicts and
rules-out both assert the target is false, and a refuted prediction
is not a false statement — it was a correct derivation from its hypothesis, and the
failure damages the hypothesis one edge upstream, not the prediction.
So the corpus can record that a prediction succeeded and cannot record that one failed. Every tree published from it will look like a paper whose predictions all worked out.
The rule
It classifies each prediction by the shape of the graph around it — which edges point at it, and whether the prediction's own words enumerate more commitments than it has edges. It does not judge whether a result actually matches the prediction it is wired to. That is a reading, it needs a person or an agent, and a rule that guessed at it would launder a judgement as a measurement.
| Bucket | n | What it means |
|---|---|---|
recorded | 23 | The graph says how the test came out. An outcome edge -- `confirms` or `refutes` -- points at the prediction from the result that settled it. |
unrecorded | 23 | A result is wired to the prediction with `tests`, which is neutral, and nothing records the outcome. The paper asserts the result, so a reader would infer the prediction was met -- but that inference is the reader's, not the graph's. |
untested | 0 | Nothing points at this prediction at all. It was derived from a hypothesis and then left unconnected to any result. |
abstain:conjunction | 1 | The prediction states more commitments than it has edges -- "should show (i)... (ii)... (iii)..." with fewer results wired than parts. One verdict cannot be right about all of them, so the rule declines rather than picking the majority. |
abstain:convention | 0 |
What it does to the whole corpus
23 of 47 predictions — 49% — are tested by a result the paper asserts and carry no outcome at all. A reader of the graph can see that each was tested and cannot see that any of them was met.
Three conventions, splitting by paper
This is not a gap in the format. The format can express an outcome, one paper does it for
every one of its predictions, another uses confirms instead of tests, and seven never record one. Nobody reconciled them, and no check
notices. It is per-paper variation in how a corpus-level convention was applied — so
approving one paper's tree is evidence about a convention another paper's tree silently
violates.
| Paper | tests only | tests+confirms | confirms only | neither |
|---|---|---|---|---|
| bouyeure-2026-fear-rsa | 5 | · | · | · |
| gadeke-2026-guilt-insula | · | · | · | · |
| headley-2026-inhibitory-rhythms | 6 | · | · | · |
| kammer-2026-foveal-feedback | 4 | · | · | · |
| kolb-2026-igabasnfr2 | 2 | · | · | · |
| meijer-2025-serotonin-additive-r1 | · | · | · | · |
| meijer-2025-serotonin-orthogonal | · | · | · | · |
| rozak-2026-neurovascular-dl | 2 | · | · | · |
| scheller-2026-self-prioritization | 4 | · | · | · |
| wengert-2026-kcnc1 | · | · | · | · |
Where the rule declines (1)
Listed in full, because they are the whole argument for trusting the rest. A rule that produced a verdict for all 47 would be guessing on these.
prediction-seizures-and-sudep abstain:conjunction wengert-2026-kcnc1 If the inhibitory failure produced by Kv3.1 LOF is sufficient to drive the encephalopathy phenotype, then Kcnc1-A421V/+ mice should exhibit (i) spontaneous convulsive seizures captured on continuous video-EEG, (ii) seizure-induced…
tests: spontaneous-seizures-and-sudep-kcnc1 · confirms: a421v-mice-die-before-122d, spontaneous-seizures-and-sudep-kcnc1 ·states 3 commitments (i ii iii)
The confident middle, sampled
One unrecorded case per paper, so the sample spans the corpus rather than
showing several variations of one paper's house style. If the rule is wrong about these it
is wrong about 23 claims.
prediction-context-specificity-increases-during-reversal bouyeure-2026-fear-rsa If PFC context-specific coding is the substrate for context-dependent fear renewal, then context representations should become more distinct (greater within-context vs between-context pattern similarity) in PFC during reversal com…
tested by context-specificity-increases-reversal · outcome not recorded
prediction-beta-optimal-distal headley-2026-inhibitory-rhythms If the optimal frequency of rhythmic inhibition at a compartment is set by matching the rhythm period to the local spike timescale, then distal inhibition — where apical Ca²⁺ and NMDA spikes lead the soma by ~20 ms and ~25 ms resp…
tested by beta-optimal-distal-dendritic-entrainment · outcome not recorded
prediction-cross-decoding-generalizes kammer-2026-foveal-feedback If foveal feedback shares a representational format with bottom-up sensory drive, a classifier trained on foveal V1 responses in the experimental (feedback) condition should generalize above chance to foveal V1 responses in the co…
tested by cross-decoding-experimental-to-control · outcome not recorded
prediction-improved-sensor-enables-new-measurements kolb-2026-igabasnfr2 If iGABASnFR2's improvements (4-fold ΔF/F, 7-fold on-cell affinity, faster rise kinetics, 2P-compatible spectra) cross qualitative capability thresholds, then side-by-side comparison with iGABASnFR1 in three demanding preparations…
tested by igabasnfr2-invivo-barrel-cortex, igabasnfr2-retina-direction-selectivity, igabasnfr2-single-bouton-hippocampus · outcome not recorded
prediction-pipeline-outperforms-baselines rozak-2026-neurovascular-dl If the DL segmentation + registration + radius-estimation stack genuinely improves on conventional baselines, then (a) the UNet/UNETR ensemble should significantly outperform ilastik on volumetric segmentation metrics (HD95, Dice,…
tested by novas3d-outperforms-ilastik, radius-estimation-r2-0p68, registration-doubles-vessel-count, unetr-outperforms-ilastik-hd95 · outcome not recorded
prediction-additive-effects-other-associated scheller-2026-self-prioritization If social and perceptual salience operate via independent mechanisms, then for stimuli that are simultaneously socially salient (other-associated) and perceptually salient (high local contrast) the observed processing rate change…
tested by self-social-additive-perceptual · outcome not recorded
The rule has a version, and that is the point
The manifest records rule_version: 3, because an approval is
granted to a rule at a version and not to a rule in general. The first version of this
one was wrong in a way worth showing.
v1 detected an enumerated conjunction — a prediction that commits to several things at
once — by looking for roman numerals, and found four. v2 also matches
(a)(b)(c) and finds six. The two it had been missing include the largest
conjunction in the corpus: a four-part prediction wired to nine testing results. Under
v1 that prediction was silently classified as though it made a single commitment.
Nothing about v1 looked broken. It ran, it produced a plausible number, and the only way to catch it was to look at what it did across the corpus and notice that a prediction with nine testers had not been flagged. That is the argument for reviewing behaviour in aggregate rather than approving items one by one: the items all looked fine.
What accepting this would not settle
This layer can say which predictions have no recorded outcome. It cannot supply the
outcomes, and it cannot record a failure, because the vocabulary has nowhere to put one.
That decision belongs to the relation vocabulary and is genuinely open: adding a
refutes relation under mira:opposes would assert that a
refuted prediction is false, which is not what happened to it, and the relation
checker already rejects edges that oppose a claim their own paper asserts.
There is a separate defect this layer makes visible without being able to fix. In the
MIRA export, tests is declared a subclass of mira:supports. So
every tests edge reaches eLife as a supporting relation:
the tree declines to say how a test came out, and the export answers anyway.
Inputs and outputs
- Reads, besides its dependencies
-
- scripts/prediction_outcome.py · declared, and not in the repository — it hashes to nothing, so it cannot make a run stale
- Produces
-
- review/prediction-outcome.json
- Views
-
- table — rendered above, over the 47 items in the artifact
Running it
The command comes from the declaration, so this text and what actually runs cannot
diverge. pipeline.py run also runs the unmet dependencies first.
python3 scripts/pipeline.py run <paper> prediction-outcome
Underneath, that runs python3 scripts/prediction_outcome.py --write.