Prediction outcomes

question · corpus proposed

An earlier declaration was accepted by Zach Mainen; it has moved since — this is 826012b59a12, unaccepted. · no paper read under it

Did the paper's predictions come out, and does the tree say so?

A decision about meaning, and then the propagation of it through everything downstream.

zmainen/elife-claim-trees#28 — the decision is taken there; the propagation is taken here.

How it works

The rule is mechanical, and it is now rule v3. For each role: prediction claim it looks at the edges pointing in and out, and sorts the prediction by the shape of that wiring: a tests edge with an outcome (confirms or refutes) beside it, a tests edge alone, nothing, or a conjunction of several. It reads the graph and nothing else — not the paper, and not whether the prediction was in fact met.

That limit is the point of running it. The issue #28 ruling settled the wiring convention — a tests edge carries its outcome beside it — so the buckets now measure which trees have had the outcome filled in, not which convention a paper happened to follow.

What it found · added 2026-09-11

tests is neutral by design, so a tree can record that a prediction was tested and never whether it was met. Of 47 predictions, 23 now carry a recorded outcome and 23 carry a tests edge and no outcome yet. The issue #28 ruling gave the vocabulary its missing half — refutes, the negative outcome beside confirms — so a failed prediction can now be written down; what remains is filling the rest in.

Rests on

Feeds — a change here disturbs these

How it is defined

What this layer reads besides its dependencies. Each is a declared input: its content is hashed into every run, so editing one makes those runs stale.

scripts/prediction_outcome.pythe script that runs itnot in the repository

The declaration names this path and the repository does not have it. An input that does not exist hashes to nothing, so it cannot make a run stale — the layer is declared to depend on something it is not in fact tracking.

This layer is proposed, and what is proposed is the rule

Nothing on this page has been accepted, and no claim in the corpus has been changed by it. What is offered for review is a rule and its behaviour across all 47 predictions in the corpus — not 47 decisions to be taken one at a time.

So the useful way to read it is: does the rule sort these correctly, does it refuse the cases it should refuse, and would applying it leave the corpus in a better state than it is now? The 1 abstentions are listed in full below, because a rule that declines what it cannot settle is what makes reviewing a sample sufficient — attention goes to the margin instead of the confident middle.

The question

A prediction is a commitment made in advance: if this hypothesis holds, then this particular measurement should come out this particular way. It is the part of a paper's argument that can fail, which is what makes it worth recording separately from the result that tests it.

In this corpus a result is joined to the prediction it bears on by a relation called tests. tests is neutral by design — it records that a test happened, not which way it came out. That neutrality is correct: the two facts are genuinely different, and collapsing them would make every test look like a success.

The problem is that nothing else records the outcome either. There is a positive relation, confirms, used 19 times on predictions here. There is no negative counterpart. The two oppositional relations do not fill the gap: contradicts and rules-out both assert the target is false, and a refuted prediction is not a false statement — it was a correct derivation from its hypothesis, and the failure damages the hypothesis one edge upstream, not the prediction.

So the corpus can record that a prediction succeeded and cannot record that one failed. Every tree published from it will look like a paper whose predictions all worked out.

The rule

It classifies each prediction by the shape of the graph around it — which edges point at it, and whether the prediction's own words enumerate more commitments than it has edges. It does not judge whether a result actually matches the prediction it is wired to. That is a reading, it needs a person or an agent, and a rule that guessed at it would launder a judgement as a measurement.

Bucket n What it means
recorded 23 The graph says how the test came out. An outcome edge -- `confirms` or `refutes` -- points at the prediction from the result that settled it.
unrecorded 23 A result is wired to the prediction with `tests`, which is neutral, and nothing records the outcome. The paper asserts the result, so a reader would infer the prediction was met -- but that inference is the reader's, not the graph's.
untested 0 Nothing points at this prediction at all. It was derived from a hypothesis and then left unconnected to any result.
abstain:conjunction 1 The prediction states more commitments than it has edges -- "should show (i)... (ii)... (iii)..." with fewer results wired than parts. One verdict cannot be right about all of them, so the rule declines rather than picking the majority.
abstain:convention 0

What it does to the whole corpus

23
23
1

23 of 47 predictions — 49% — are tested by a result the paper asserts and carry no outcome at all. A reader of the graph can see that each was tested and cannot see that any of them was met.

Three conventions, splitting by paper

This is not a gap in the format. The format can express an outcome, one paper does it for every one of its predictions, another uses confirms instead of tests, and seven never record one. Nobody reconciled them, and no check notices. It is per-paper variation in how a corpus-level convention was applied — so approving one paper's tree is evidence about a convention another paper's tree silently violates.

Paper tests onlytests+confirmsconfirms onlyneither
bouyeure-2026-fear-rsa 5 · · ·
gadeke-2026-guilt-insula · · · ·
headley-2026-inhibitory-rhythms 6 · · ·
kammer-2026-foveal-feedback 4 · · ·
kolb-2026-igabasnfr2 2 · · ·
meijer-2025-serotonin-additive-r1 · · · ·
meijer-2025-serotonin-orthogonal · · · ·
rozak-2026-neurovascular-dl 2 · · ·
scheller-2026-self-prioritization 4 · · ·
wengert-2026-kcnc1 · · · ·

Where the rule declines (1)

Listed in full, because they are the whole argument for trusting the rest. A rule that produced a verdict for all 47 would be guessing on these.

prediction-seizures-and-sudep abstain:conjunction wengert-2026-kcnc1

If the inhibitory failure produced by Kv3.1 LOF is sufficient to drive the encephalopathy phenotype, then Kcnc1-A421V/+ mice should exhibit (i) spontaneous convulsive seizures captured on continuous video-EEG, (ii) seizure-induced…

tests: spontaneous-seizures-and-sudep-kcnc1 · confirms: a421v-mice-die-before-122d, spontaneous-seizures-and-sudep-kcnc1 ·states 3 commitments (i ii iii)

The confident middle, sampled

One unrecorded case per paper, so the sample spans the corpus rather than showing several variations of one paper's house style. If the rule is wrong about these it is wrong about 23 claims.

prediction-context-specificity-increases-during-reversal bouyeure-2026-fear-rsa

If PFC context-specific coding is the substrate for context-dependent fear renewal, then context representations should become more distinct (greater within-context vs between-context pattern similarity) in PFC during reversal com…

tested by context-specificity-increases-reversal · outcome not recorded

prediction-beta-optimal-distal headley-2026-inhibitory-rhythms

If the optimal frequency of rhythmic inhibition at a compartment is set by matching the rhythm period to the local spike timescale, then distal inhibition — where apical Ca²⁺ and NMDA spikes lead the soma by ~20 ms and ~25 ms resp…

tested by beta-optimal-distal-dendritic-entrainment · outcome not recorded

prediction-cross-decoding-generalizes kammer-2026-foveal-feedback

If foveal feedback shares a representational format with bottom-up sensory drive, a classifier trained on foveal V1 responses in the experimental (feedback) condition should generalize above chance to foveal V1 responses in the co…

tested by cross-decoding-experimental-to-control · outcome not recorded

prediction-improved-sensor-enables-new-measurements kolb-2026-igabasnfr2

If iGABASnFR2's improvements (4-fold ΔF/F, 7-fold on-cell affinity, faster rise kinetics, 2P-compatible spectra) cross qualitative capability thresholds, then side-by-side comparison with iGABASnFR1 in three demanding preparations…

tested by igabasnfr2-invivo-barrel-cortex, igabasnfr2-retina-direction-selectivity, igabasnfr2-single-bouton-hippocampus · outcome not recorded

prediction-pipeline-outperforms-baselines rozak-2026-neurovascular-dl

If the DL segmentation + registration + radius-estimation stack genuinely improves on conventional baselines, then (a) the UNet/UNETR ensemble should significantly outperform ilastik on volumetric segmentation metrics (HD95, Dice,…

tested by novas3d-outperforms-ilastik, radius-estimation-r2-0p68, registration-doubles-vessel-count, unetr-outperforms-ilastik-hd95 · outcome not recorded

prediction-additive-effects-other-associated scheller-2026-self-prioritization

If social and perceptual salience operate via independent mechanisms, then for stimuli that are simultaneously socially salient (other-associated) and perceptually salient (high local contrast) the observed processing rate change…

tested by self-social-additive-perceptual · outcome not recorded

The rule has a version, and that is the point

The manifest records rule_version: 3, because an approval is granted to a rule at a version and not to a rule in general. The first version of this one was wrong in a way worth showing.

v1 detected an enumerated conjunction — a prediction that commits to several things at once — by looking for roman numerals, and found four. v2 also matches (a)(b)(c) and finds six. The two it had been missing include the largest conjunction in the corpus: a four-part prediction wired to nine testing results. Under v1 that prediction was silently classified as though it made a single commitment.

Nothing about v1 looked broken. It ran, it produced a plausible number, and the only way to catch it was to look at what it did across the corpus and notice that a prediction with nine testers had not been flagged. That is the argument for reviewing behaviour in aggregate rather than approving items one by one: the items all looked fine.

What accepting this would not settle

This layer can say which predictions have no recorded outcome. It cannot supply the outcomes, and it cannot record a failure, because the vocabulary has nowhere to put one. That decision belongs to the relation vocabulary and is genuinely open: adding a refutes relation under mira:opposes would assert that a refuted prediction is false, which is not what happened to it, and the relation checker already rejects edges that oppose a claim their own paper asserts.

There is a separate defect this layer makes visible without being able to fix. In the MIRA export, tests is declared a subclass of mira:supports. So every tests edge reaches eLife as a supporting relation: the tree declines to say how a test came out, and the export answers anyway.

Inputs and outputs

Reads, besides its dependencies
Produces
  • review/prediction-outcome.json
Views
  • table — rendered above, over the 47 items in the artifact

Running it

The command comes from the declaration, so this text and what actually runs cannot diverge. pipeline.py run also runs the unmet dependencies first.

python3 scripts/pipeline.py run <paper> prediction-outcome

Underneath, that runs python3 scripts/prediction_outcome.py --write.