Agents

This is a recorded run of the extraction pipeline on one paper: which role read which part of it, under which prompt, using which model, and what each proposed. It shows what the agents do. It is not the record of how this corpus's claim files were produced — see below.

No model output here becomes a finding by being produced. The role below called external-reviewer is a second model pass, not a person, so nothing in the extraction chain is checked by review in the ordinary sense. Where a person has ruled — on the scheme, on a layer version, on a proposed claim — it is recorded as an approval against the thing ruled on and shown where that thing appears. Everything else should be read as a model's output that has been recorded rather than as a finding that has been checked.

The agreement matrix

One paper — gadeke-2026-guilt-insula — reconciled: which reader surfaced each claim, the sentence each quoted, and the confidence that follows. This is reconcile's overlap view, and it is drawn on that paper's cell by the same component.

claim confidence results-readercaption-readerstructure-reader evidence verified
Responsibility for a social choice that yields a low (negative) outcome for a partner produces interpersonal guilt, experienced by the decision-maker as a larger decrease in momentary happiness than when the partner made the same choice. hypothesis single-source “This study investigated the neural mechanisms involved in feelings of interpersonal guilt and responsibility evoked by social decisions in h…” — — —
The anterior insula is the neural substrate of the guilt effect, increasing its BOLD response when participants are responsible for low outcomes affecting their partner. hypothesis single-source “Next, we sought to uncover the neural mechanisms associated with our guilt effect and those involved in tracking consequences of participant…” — — —
Functional connectivity between guilt- and responsibility-related outcome-phase regions and prefrontal cortex changes depending on whether participants decide for themselves alone or also for their partner, and on the type of choice (Safe or Risky). hypothesis single-source “We hypothesized that connectivity with regions that showed guilt- and responsibility-related responses during the outcome phase (see previou…” — — —
If responsibility for outcomes generates guilt, then participant happiness should decrease more after low lottery outcomes for the partner when the participant rather than the partner chose the lottery. prediction single-source “In our definition, guilt occurs due to responsibility for low lottery outcomes for the partner.” — — —
If the anterior insula tracks guilt, then insula BOLD should be higher in the Social than the Partner condition and show a significant Social-by-low-outcome interaction. prediction single-source “To identify regions likely to be involved in the guilt effect, we selected those satisfying two conditions: higher activity in the Social co…” — — —
Participants’ probability of choosing the risky option (lottery) increased with the difference between the expected value of the lottery and the value of the safe option (Study 1: t(4796) = 9.26, p < 3.1e–20, β = 0.074; Study 2: t(3829) = 10.62, p < 5.3e–26, β = 0.093). control · fig2a,fig2d single-source “As expected, participants’ probability of choosing the risky option (lottery) increased with the difference between the expected value of th…” — — —
Participants chose the lottery more often in the Solo condition than in the Social condition in Study 1 (t(4796) = 2.54, p = 0.011, β = 0.164) but not in Study 2 (t(3829) = 0.23, p = 0.82, β = 0.015). empirical · fig2a,fig2d high “Participants chose the lottery more often in the Solo condition than in the Social condition in Study 1 (t(4796) = 2.54, p = 0.011, β = 0.16…” “Participants chose the risky option slightly more often in the Solo condition than in the Social condition in Study 1 ( A ) but not in Study…” — —
There was no significant interaction between the difference in expected values and experimental conditions in either study (p > 0.52). control single-source “There was no significant interaction between the difference in expected values and experimental conditions in either study (p > 0.52).” — — —
Risk premiums did not differ between Social and Solo conditions in either study (Study 1: t(39) = 1.53, p = 0.134, d = 0.24, BF10 = 0.49; Study 2: t(43) = –0.21, p = 0.84, d = –0.03, BF10 = 0.17). control · fig2b,fig2e high “Risk premiums did not differ between Social and Solo conditions (Study 1: Figure 2B, t(39) = 1.53, p = 0.134, Cohen’s d = 0.24, BF10 = 0.49;…” “( B, E ) Risk premiums did not differ between Solo and Social conditions.” — —
The risk-aversion parameter ρ did not differ between gain and loss trials (Study 1: t(17) = 0.21, p = 0.84, d = 0.05; Study 2: t(15) = –0.61, p = 0.55, d = 0.15). control single-source “As ρ did not vary between gain and loss trials (Study 1: t(17) = 0.21, p = 0.84, d = 0.05; Study 2: t(15) = –0.61, p = 0.55, d = 0.15; paire…” — — —
Participants were slightly more risk averse (higher ρ) in the Social than the Solo condition in Study 1 (t(39) = 2.27, p = 0.03, d = 0.36, BF10 = 1.69) but not in Study 2 (t(43) = 1.40, p = 0.17, d = 0.21, BF10 = 0.41). empirical · fig2c,fig2f high “We found that participants were slightly more risk averse in the Social than in the Solo condition in Study 1 (Figure 2C, t(39) = 2.27, p = …” “Values of the risk aversion parameter ρ in the Solo and Social conditions were broadly consistent with Risk premium values, but showed that …” — —
Participants showed very similar risk preferences whether deciding only for themselves (Solo) or for themselves and their partner (Social), with only a tendency toward higher risk aversion in the Social condition in Study 1. synthesis single-source “In sum, participants showed very similar risk preferences when making decisions affecting only themselves (Solo condition) or themselves and…” — — —
In both studies, participant momentary happiness correlated with the rewards obtained in the current trial by both the participant and the partner. empirical · fig3a,fig3b,fig3e,fig3f high “Across all trials, in both studies, participant momentary happiness correlated with rewards obtained in the current trial by the participant…” “Happiness varied with rewards received by the participant ( A, E ) and by the partner ( B, F ).” — —
Rutledge and colleagues established that changes in momentary happiness during a probabilistic reward task are explained by recent reward expectations and the prediction errors arising from them. literature-context single-source “Following Rutledge and colleagues’ methodology, which considers that changes in momentary happiness in response to outcomes of a probabilist…” — — —
A likelihood ratio test showed the Responsibility model fitted the happiness data better than all other models, including the Responsibility Redux model (Study 1: all LR ≥ 47.36, p < 0.0001; Study 2: all LR ≥ 77.83, p < 0.0001). empirical · table1 single-source “a likelihood ratio test (Equation 9) revealed that the Responsibility model fitted better than all the other models, including the Responsib…” — — —
The Responsibility model yielded higher R2 values than all other models (Study 1: all t > 3.6, p < 0.007; Study 2: all t > 2.9, p < 0.034), except the Guilt-envy model in Study 1 (t = 2.19, p = 0.17). empirical · table1 single-source “The Responsibility model yielded higher R2 values than all the other models (Study 1: all t > 3.6, p < 0.007; Study 2: all t > 2.9, p < 0.03…” — — —
Participants’ own reward prediction errors (sRPE) influenced happiness more than the partner’s reward prediction errors (social_pRPE and partner_pRPE) (Study 1: all Z > 6.0, p < 0.001; Study 2: all Z > 3.7, p < 0.003). empirical single-source “weights for sRPE were higher than for social_pRPE or partner_pRPE (Study 1: all Z > 6.0, p < 0.001; Study 2: all Z > 3.7, p < 0.003).” — — —
The partner’s reward prediction errors resulting from the participants’ own choices (social_pRPE) had weights greater than 0 (Responsibility model: Study 1: Z = 2.85, p = 0.004; Study 2: Z = 3.26, p = 0.001), contributing to explaining participants’ momentary happiness. empirical single-source “weights for social_pRPE were greater than 0: Responsibility model: Study 1: Z = 2.85, p = 0.004, Study 2: Z = 3.26, p = 0.001; Responsibilit…” — — —
The stability of the estimated computational-model parameters was verified with a parameter recovery procedure. methodological · fig3s1 high “The stability of these estimated parameters was verified using a parameter recovery procedure (see Methods and Figure 3—figure supplement 1)…” “Stability of the estimated parameters of the temporal difference models was evaluated by attempting to recover parameters from synthetic dat…” “We then compared these new estimated parameters to the actual parameters from which the synthetic data were generated, as follows: For each …” —
The ‘Responsibility’ computational model was used to generate expected BOLD responses per participant for the model-based fMRI analysis. methodological high “We used the model to create expected BOLD responses for each participant (see Methods) and as a manipulation check searched for responses in…” — “In addition, we created another GLM (GLM2) with regressors designed to identify brain regions whose activation reflected the variables of th…” —
The Responsibility Redux model, taking into account expected, previous and current rewards, reward prediction errors for both participant and partner, and decision-maker, predicted the variations in participants' momentary happiness well. empirical · fig3c, fig3g single-source — “A computational model taking into account expected, previous and current rewards, reward prediction errors for both participant and partner,…” — —
Participant happiness was lower when the participant was the decision-maker (Social + Solo vs. Partner), independent of outcome (Study 1: t(3600) = –3.92, p < 0.0001, β = –0.14; Study 2: t(2870) = –6.07, p < 0.0001, β = –0.24). empirical single-source “we assessed whether happiness varied depending on the participant’s agency (Social + Solo vs. Partner), and found happiness to be lower when…” — — —
The lower happiness when the participant is the decision-maker may reflect responsibility aversion, a cost of the ‘weight of the responsibility’. interpretation single-source “This is interesting in itself and may reflect the drive behind responsibility aversion reported by Edelson et al.’s 2018 study: being assign…” — — —
The interaction between partner outcome and decision-maker was significant (Study 1: t(1180) = 3.52, p = 0.0004, β = 0.37; Study 2: t(937) = 2.85, p = 0.0045, β = 0.33): when the partner received the low outcome, participant happiness was lower when the participant rather than the partner had chosen the lottery. empirical · fig3d,fig3h high “Crucially, the interaction between partner outcome and decision-maker was significant (Study 1: t(1180) = 3.52, p = 0.0004, β = 0.37, 95% CI…” “Crucially, responsibility for low lottery outcomes for the partner decreased participant happiness more than the same outcomes following par…” — —
The linear mixed model containing all three two-way interaction terms (Model 5) best explained the happiness data in both studies, and its crucial partnerHigh:participantDecided (guilt) interaction was significant. empirical · app1table2 high — “In both studies, Model 5 ( Equation 9 in the Results section of the main text), which contained all three two-way interaction terms, explai…” “In both studies, the model reported in Equation 10 fitted the data significantly better (p < 2e−5) than the simpler models without interac…” —
The behavioural guilt effect (larger happiness decrease after low partner outcomes following participant rather than partner choices) is compatible with ‘simple guilt’. interpretation single-source “This behavioural effect (difference in happiness obtained when the partner received low lottery outcomes after participant rather than partn…” — — —
The guilt effect occurred whether the participant received the high lottery outcome (Study 1: t(39) = –3.58, p < 0.001, d = 0.56; Study 2: t(43) = –2.68, p = 0.01, d = 0.4) or the low outcome (Study 1: t(39) = –3.39, p = 0.002, d = 0.54; Study 2: t(43) = –3.58, p < 0.001, d = 0.54). control single-source “The ‘guilt effect’ occurred whether the participant received the high lottery outcome (Study 1: t(39) = –3.58, p < 0.001, d = 0.56, BF10 = 3…” — — —
Responsibility for choices did not influence happiness following positive (high) lottery outcomes for the partner (both studies, all |t| < 1.3, p > 0.2, BF10 < 0.2). control single-source “Responsibility for choices did not influence happiness following positive lottery outcomes for the partner (both studies, all |t| < 1.3, p >…” — — —
All BOLD/fMRI results derive from Study 2 (the fMRI study), whereas the behavioural results come from both Study 1 and Study 2. scope high “We analysed the BOLD responses of brain regions engaged during decision-making and at the time of receiving the outcomes of the choice using…” — “Forty healthy participants (14 male, mean age 26.1, range 22–31) participated in Study 1 (behaviour only study), and 44 healthy participants…” —
The bilateral ventral striatum was more active when participants chose the risky rather than the safe option (Cohen’s d = 0.72 left, 0.85 right), replicating previous findings. control · fig4a high “We searched for brain regions engaged more when participants chose the risky instead of the safe option and found such responses in the bila…” “( A ) Regions showing a greater response when participants chose the risky (lottery) rather than the safe option, irrespective of Social or …” — —
Decisions in the Social compared with the Solo condition engaged three clusters — the precuneus (d = 0.79), left temporo-parietal junction (d = 0.59), and medial prefrontal cortex (d = 0.54). empirical · fig4b high “Three significant clusters of voxels were identified (Figure 4B and Appendix 1—table 3), in the precuneus (d = 0.79), the left temporo-parie…” “( B ) Regions showing a greater response when participants chose for both themselves and their partner rather than just for themselves (Soci…” — —
Only the precuneus and TPJ showed positive Risky–Safe differences in both the Social>Solo and Social>Partner comparisons, being most active when participants chose the lottery in the Social condition. empirical · fig4c high “Only the precuneus and TPJ showed positive differences in both comparisons (Figure 4C), indicating that these regions were most active when …” “Coefficients of linear mixed models (LMMs) indicate that two of these regions, precuneus and TPJ, were most active when participants chose t…” — —
During receipt of lottery versus safe outcomes, clusters were more active in the bilateral anterior insula, dmPFC, right STS, bilateral ventral striatum, right dorsolateral prefrontal cortex, and bilateral inferior parietal lobe. empirical · fig4d high “A cluster of voxels more active during receipt of lottery outcomes than outcomes of safe choices was identified in the bilateral anterior in…” “( D ) Brain regions more active during receipt of the outcomes of lotteries than safe choices (all conditions).” — —
The insula ROIs responded more to low lottery outcomes for the partner in the Social than the Partner condition (even after subtracting responses to high outcomes), mirroring the behavioural guilt effect. empirical · fig4e high “Thus, activation in our insula ROIs increased in situations during which participants experienced guilt for low outcomes impacting their par…” “voxels here responded more to low lottery outcomes (L) for the partner when these resulted from participant’s rather than the partner’s choi…” — —
A mass-univariate voxel-wise analysis found a small left anterior insula cluster (peak T = 3.95, d = 0.59, 22 voxels) responding more to low partner outcomes following participant than partner choices, which survived small-volume family-wise-error correction (p = 0.024). empirical · fig4f high “We found a weak response in a small cluster within the left anterior insula (peak T = 3.95, d = 0.59, 22 voxels, peak intensity at [–28 24 –…” “A cluster of voxels within the left insula ROI showed higher responses to low lottery outcomes for the partner if these resulted from partic…” — —
Prior literature documents an association between the anterior insula and guilt. literature-context single-source “Given the documented association between anterior insula and guilt (see Introduction), we proceeded to test whether this result survived cor…” — — —
As a manipulation check, bilateral ventral striatum activation increased with expected certain rewards and the expected values of chosen lotteries (left: pFWE = 0.002, T = 5.63, d = 0.75; right: pFWE = 0.005, T = 5.46, d = 0.70). control · fig4g high “We found that activation in bilateral ventral striatum indeed increased with the amount of expected certain rewards and the expected values …” “( G ) Activation in bilateral ventral striatum explained by a computational model-based regressor coding participant rewards.” — —
One cluster in the left STS responded more to partner reward prediction errors resulting from participant rather than partner choices (pFWE = 0.022, T = 4.70, d = 0.53, 100 voxels, peak MNI [−52 –32 0]). empirical · fig4h high “We found this effect in one cluster within the left STS (pFWE = 0.022, T = 4.70, d = 0.53, Z = 4.57, 100 voxels, peak at MNI [−52 –32 0]; Fi…” “one cluster in the left superior temporal sulcus region showed a higher response to partner reward prediction errors resulting from particip…” — —
The left superior temporal sulcus cluster responded to model-based regressors coding participant reward prediction resulting from participant and partner choices across both sessions of the experiment. empirical · fig4i single-source — “( I ) Response in this cluster to the computational-model-based regressors coding participant reward prediction resulting from participant a…” — —
The authors suggest this left STS region tracks a partner’s unexpected outcomes less when they do not follow from the participant’s decisions. interpretation single-source “This finding suggests that this region of the left STS tracks a partner’s unexpected outcomes less when they do not follow from the particip…” — — —
Prior functional connectivity work has shown network differences between social and self-only choices, midbrain–anterior cingulate interactions during guilt compensation, and links between insula connectivity and responsibility aversion. literature-context single-source “Functional connectivity analyses have revealed differences in networks engaged by social and self-only choices (Jung et al., 2013; Ogawa et …” — — —
Left anterior insula connectivity with a right IFG cluster was highest when participants made Risky choices for themselves and Safe choices for both players (pFWE = 0.020, T = 4.34, d = 0.80, 115 voxels, peak MNI [46 16 22]). empirical · fig5 high “The first analysis revealed a cluster in the right IFG whose connectivity to the insula (the seed region) was highest when participants made…” “Changes in functional connectivity between the left anterior insula (seed) and a cluster in the right inferior frontal gyrus at the time of …” — —
A left IFG cluster showed the opposite pattern of connectivity with the left STS seed — highest for Safe-self / Risky-both-players choices — but did not survive correction for multiple comparisons (p uncorrected = 0.001, T = 4.44, 35 voxels). empirical · fig5s1 high “The second analysis revealed a smaller cluster in the left IFG that did not survive corrections for multiple tests, where connectivity with …” “connectivity was highest when participants made Safe choices for themselves and Risky choices for both players (p uncorrected = 0.001, T …” — —
Dot products between individual neural guilt responses and the Yu et al. (2020) guilt-related brain signature (GRBS) were overall positive (mean = 5.22, median = 6.97, sign test p = 0.017, Cliff’s Delta = 0.4). empirical single-source “The dot products between individual responses and the GRBS varied between –40.1 and 36.7, but overall these values were positive (mean = 5.2…” — — —
Individual GRBS dot-product values did not correlate with the behavioural guilt responses (Spearman’s Rho = –0.058, p = 0.725). control single-source “We assessed whether inter-individual differences in these dot product values correlated with the behavioural guilt responses, but did not fi…” — — —
Among the computational models fitted to momentary happiness data, the Responsibility Redux model achieved the best (lowest) AIC in both studies (Study 1 AIC –1499; Study 2 AIC –1195). methodological · table1 single-source — “Responsibility Redux 4 0.361 0.331 –999 –1499” — —
In mixed-effects regressions on choices, the Social condition significantly increased choice of the risky option in Study 1 but not in Study 2. empirical · app1table1 single-source — “Condition Social 0.14* 0.03^ 0.01 0.01” — —
During the outcome phase, responses to low lottery outcomes were higher in the Social than the Partner condition in both left and right insula and lower in the right middle temporal cortex. empirical · app1table9 single-source — “Social 0.41*** 0.18*** –0.12**” — —
The difference in response between low and high lottery outcomes was greater in the Social than the Partner condition in left insula, right insula, and right middle temporal cortex. empirical · app1table10 single-source — “Social 0.44*** 0.19*** 0.67***” — —
Participants rated their partners highly on sympathy, cooperation, honesty, openness, and sociability in both studies. control · app1table11 high — “How honest did they seem? 9.05 (1.11) 9.34 (1.10)” “Results of a brief questionnaire indicate that this approach was successful: participants’ average ratings of their partners in terms of sym…” —
On each trial participants chose between a safe monetary option and a risky lottery, under three conditions: deciding for themselves (Solo), for themselves and the partner (Social), and for the partner deciding for both (Partner). scope single-source — — “There were three kinds of trials: decisions by the participant only for themselves ( Solo condition), decisions by the participant for them…” —
Study 2 reproduced the Study 1 design inside the fMRI scanner with otherwise identical parameters, except for longer inter-stimulus intervals (3–11 s) and a fixed experimenter partner. scope single-source — — “In Study 2, participants performed two sessions of the experiment described above inside the fMRI scanner. All parameters were identical exc…” —
To hold the partner's behaviour constant across participants, the partner's decisions were simulated by an algorithm that always chose the option with the highest expected value. methodological single-source — — “In order to ascertain constant decisions by the partner, the partner’s decisions were simulated using a simple algorithm that always selecte…” —
Momentary happiness was modelled with five computational models (Basic, Inequality, Guilt-envy, Responsibility, and Responsibility Redux) sharing separate, exponentially decaying terms for certain rewards, expected value, and reward prediction errors. methodological single-source — — “All models contained separate terms for certain rewards, expected value for lotteries and reward prediction errors, with influences that dec…” —
Model selection among the happiness models used likelihood ratio tests comparing the Responsibility model pairwise against each of the other models. methodological single-source — — “To formally assess which of our models fitted the data best, we supplemented the AIC, BIC, R 2 and adjusted R 2 values reported in Tabl…” —

The 13 roles

Three readers work in parallel on deliberately different slices of the paper, so that agreement between them carries information. The rest run in sequence over their output. Every prompt is a committed file and is reproduced in full below.

Read the paper

results-reader 46 claims proposed task · 80 lines · + contract

Reads the Results section. Returns candidate claims, each with a panel, a role and the sentence it rests on.

extract/prompts/results-reader.md

# Results reader

You are reading one scientific paper to find what it claims. You are given its abstract, its
Introduction, its Results and its Discussion, preceded by a panel inventory — and nothing else:
no figure captions, no methods. Two other readers are given those, none of you sees the others'
work, and a reconciler compares the three afterwards. Agreement between independent readings is
the signal, so read your slice on its own terms and do not try to guess what the others will
find.

Your slice is numbered. A bracketed span id sits in front of every sentence — `[results-026]
Participants chose …` — and the panel inventory at the top lists every figure and table that
exists, its panel ids, and the first line of its caption. The Introduction states the
organising hypothesis, the questions the paper asks, and the prior findings it builds on; the
Results state the empirical argument; the Discussion states the interpretation and the
literature-context premises. The Discussion yields `interpretation` and `literature-context`
claims, and never an `empirical` claim that the Results do not also state — do not turn a
sentence that recalls a result in order to interpret it into a second empirical claim. The
inventory lists the panels that exist, so a claim the prose anchors to "Figure 3B" is `fig3b`;
never invent a panel that is not in the inventory. Put in `span` the id shown before the
sentence your `evidence` was quoted from; the quote itself is verbatim prose and must not
include the bracketed id.

## What this slice carries

The prose is the only place the paper states its argument *as an argument*. That is
what you are for.

Find the paper's **questions** first — what it set out to answer — from the abstract and the
opening of the Introduction. Then return the hypotheses as the answers the paper commits to:
each one carries `addresses`, the question it answers, as the paper states it or in one
sentence. A question the paper states but commits to no answer for is not a hypothesis, so do
not return one; and never invent a hollow hypothesis to stand in for a question.

Find:

- **The hypotheses.** The proposition the paper bets on, usually introduced in the abstract or
  the opening of Results. One to three per paper. Write it as a claim about the world, not as
  "the authors investigated".
- **The predictions.** What the paper says should be observed if a hypothesis holds. The prose
  often states these as "if … then", "the model predicts", "should be". Where the paper derives
  a prediction and then tests it, surface the prediction as its own claim.
- **Each empirical result as the paper states it**, with the panel where the prose names one
  ("Figure 3B" → `fig3b`), the numbers where the prose gives them, and the paper's own verb.
- **Controls**: results whose job is to rule something out or to show a manipulation worked.
  "No difference in risk premiums between conditions" is a claim, and usually a control.
- **Synthesis and interpretation** where the prose makes them: "taken together", "these results
  show", "suggests a mechanism whereby". Mark them by role; do not fold them into the results
  they rest on.

## The hypothesis, the prediction and the test are three claims

When one paragraph carries a hypothesis, the prediction deduced from it and the result that
tests it, return three claims. The empirical result alone loses the deductive structure, and
that structure is the thing the claim graph exists to record. A hypothesis is never demoted to
a prediction because predictions were also found; adding predictions never reduces the number
of hypotheses.

## Rules

- `evidence` is a verbatim quote from the text you were given, at most two sentences, with the
  bracketed span id stripped off. It is checked against the source. If you cannot quote, mark
  the claim `tentative` and say why.
- `span` is the id bracketed before the sentence the evidence came from, `results-026`. `null`
  when the evidence spans no single numbered sentence.
- `panel` only when the prose names it, and only an id the inventory lists. Otherwise `null`.
  Never invent a panel letter.
- Keep the paper's strength. "Consistent with" stays "consistent with"; "suggests" stays
  "suggests". Do not write "demonstrates" for a paper that did not.
- No number you did not read. "A large fraction" stays "a large fraction".
- Negative and null results are claims. Do not skip them.
- One proposition per claim. A sentence joining two findings with "and" is two claims. A finding
  the paper states once across conditions is one claim, not one per condition.
- A Results section typically yields 15–30 claims. Fewer than 8 usually means the hypotheses,
  predictions and synthesis were missed; more than 40 usually means one finding was split by
  condition.

## What to return

The JSON array described below, and nothing else: no prose before or after it, no code fence.
The vocabulary that follows defines every value you may use.
caption-reader 25 claims proposed task · 57 lines · + contract

Reads figure and table captions. Returns candidate claims anchored to the panel they describe.

extract/prompts/caption-reader.md

# Caption reader

You are reading one scientific paper to find what it claims. You are given its figure and
table captions, preceded by a panel inventory — and nothing else: no results prose, no methods,
and not the figures themselves. Two other readers are given those, none of you sees the others'
work, and a reconciler compares the three afterwards. Agreement between independent readings is
the signal, so read your slice on its own terms.

The panel inventory at the top lists every figure and table that exists, its panel ids, and the
first line of its caption. Anchor each claim to a panel id the inventory lists — a caption
naming "Figure 3B" is `fig3b` — and never invent a panel that is not in it. Your captions carry
no bracketed span ids, so leave `span` null, and keep every `evidence` quote verbatim with no
bracketed id in it.

## What this slice carries

Captions are where the panel-level results are written down precisely: which panel, which
condition, which measurement, which number. That is what you are for. For each figure, read
its caption panel by panel and return, for every panel that shows a result, one claim stating
what that panel shows. Some panels show two or three results; some show none.

Find:

- **One claim per result-bearing panel**, anchored to that panel's id and carrying the
  caption's numbers verbatim. The figure list you are given names the panel letters the caption
  uses.
- **Claims that span panels.** When a caption states one result across several panels
  ("A–C, comparing baseline, perturbation and recovery"), return one claim with the panels
  listed, `fig3a, fig3b, fig3c`, rather than three claims that each say a third of it.
- **Controls.** A panel showing no difference, no effect, or a manipulation check is a claim,
  and its role is usually `control`.
- **Methodological panels.** A schematic, a model architecture, a parameter sweep that sets up
  an analysis asserts nothing about the world. Return it with role `methodological` if a later
  result depends on it, and return nothing for a pure cartoon.
- **Table rows that carry a result nowhere else.** A table is a float; treat each row that
  reports a finding as a panel of that table (`table1`, `app1table4`).

## Rules

- `evidence` is a verbatim quote from the caption, at most two sentences. It is checked
  against the source.
- `panel` is required and lowercase, using the figure's own id as given: `fig2c`, `fig3s1a`
  for a figure supplement, `table1`. Assign the claim to the panel that shows the data, not the
  one that illustrates the setup.
- Numbers exactly as the caption gives them. "Approximately 5 Hz" stays "approximately 5 Hz".
- Where a caption describes both a measurement and a simulation of it, return two claims: a
  measurement and a model prediction have different epistemic standing.
- Read the caption's stated direction of an effect; inhibition and saturation reverse naive
  intuitions.
- Most papers yield 15–30 claims from their captions, roughly the number of
  result-bearing panels. Include figure supplements: they carry the controls the main figures
  abbreviate.

## What to return

The JSON array described below, and nothing else: no prose before or after it, no code fence.
The vocabulary that follows defines every value you may use.
structure-reader 10 claims proposed task · 66 lines · + contract

Reads the abstract and section structure. Returns candidate claims the paper frames as its own conclusions.

extract/prompts/structure-reader.md

# Structure reader

You are reading one scientific paper to find what it claims. You are given its Methods,
its appendices and its supplementary material, and nothing else: no abstract, no results
prose, no captions. Two other readers are given those, none of you sees the others' work, and
a reconciler compares the three afterwards. Agreement between independent readings is the
signal, so read your slice on its own terms.

Your slice carries no panel inventory and no bracketed span ids, so leave `span` null, name a
`panel` only where the methods name one, and keep every `evidence` quote verbatim.

## What this slice carries

The methods say what was actually done and what it rests on. Nothing else in the paper
states the boundary conditions on its results, so this slice carries what the others cannot:

- **Scope claims.** Where the results apply: the model class, the preparation, the species,
  the population and its size, the design. "All results come from a single-cell compartmental
  model." "Study 2 excluded four participants for head motion." These usually bound every
  empirical claim in the paper.
- **Methodological claims.** A capability or analytical commitment that later results depend
  on for their meaning: the null distribution, the model comparison that licenses a model-based
  analysis, the validation of a sensor, a pre-registration and what it fixed. Write what the
  commitment is and what it licenses.
- **Controls the methods document.** Where a control condition or manipulation check is
  described, return it as a claim about what that control establishes, with role `control`.
- **Load-bearing assumptions.** A parameter taken from the literature, an initialisation not
  derived, a threshold not sensitivity-tested. These are scope or methodological claims about
  what the results are conditional on.

## What is not a claim

Procedure for its own sake. Which software ran the task, how participants were recruited,
which scanner was used: these are facts, not claims, unless a result turns on them. The test
is whether a downstream result would mean something different if this were different. If not,
leave it out.

The vocabulary's **What is not a claim** section, under the `methodological` role, works this
test on real sentences: a normalisation, a measure's definition, a seed choice, a
significance threshold, a software choice — each returned by a reader and each not a claim,
set beside the procedure that a result does turn on and so is methodological. Read it before
you return a methodological claim, and check your candidate against the negative list.

One line is easy to cross: an exclusion count, the sample size and the description of the
design are components of the paper's scope claim, not claims of their own — fold them into the
scope claim that says where the results apply rather than returning each as a separate
methodological or scope claim.

## Rules

- **Never infer a result from a method.** "We computed the correlation between X and Y" says
  nothing about whether X correlates with Y. Without the results prose you cannot know what
  was found, only what was measured; return the methodological claim, not the result you
  imagine it produced.
- `evidence` is a verbatim quote from the text you were given, at most two sentences. It is
  checked against the source.
- `panel` only where the methods name one; usually `null`.
- Where the methods describe both an experiment and a simulation, say which a claim concerns.
- This slice typically yields 5–15 claims: a few scope claims, a few methodological warrants,
  the controls the methods document. A computational paper with several models may yield more;
  more than 30 means procedure is being returned as claims.

## What to return

The JSON array described below, and nothing else: no prose before or after it, no code fence.
The vocabulary that follows defines every value you may use.

Reconcile and review

reconciler task · 85 lines · + contract

Reads all three readers' output. Returns one draft table, each claim tagged by how many readers proposed it.

extract/prompts/reconciler.md

# Reconciler

Three readers have each read one slice of the same paper — the results prose, the captions,
the methods — and each has returned a list of candidate claims. None saw the others' work. You
are given all three lists. Fold them into one draft claim table in which every claim records
which readers surfaced it and the evidence each quoted.

## What "the same claim" means

Two candidates are the same claim when they assert the same proposition about the paper's
content: the same entities, the same outcome, the same direction, and the same panel or no
panel. One being a more precise statement of the other is the same claim; keep the more
precise wording.

Two candidates are different claims when any of these differs:

- the panel;
- the direction of the effect;
- the scope (one global, one conditional);
- the role class. A hypothesis, the prediction deduced from it, and the result that tests it
  are three claims about one piece of content, and only the results reader can surface the
  first two. Do not fold a prediction into the result that tests it, or a hypothesis into the
  results that support it: that erases the target of every `entails` and `tests` edge the
  graph would carry.

Decide between three outcomes by one test — what would verify each candidate:

- **Same**: the two sentences would be verified by the same computation on the same data,
  whatever their wording and whichever study each cites. Merge: keep the more precise wording,
  and carry both readers' evidence.
- **Part**: one states one comparison, one condition, one measure or one study of what the
  other states as a whole. Keep both, and write `part_of` (below).
- **Different**: a different computation, a different direction, region or condition. Keep both,
  with no relation between them.

Apply the test, rather than defaulting to "keep apart": a split you make out of caution hides
that two readers found one thing, and the merge and the `part_of` are what the confidence tag
and the part edge are for.

There is a third outcome between merging and keeping two independent claims: one is a **part
of** the other. A candidate that states one comparison, one condition, one measure or one
study of a proposition another candidate states as a whole is a component of it — not the same
claim, because it says less, and not an independent claim, because dropping it weakens the
whole rather than leaving it standing. Keep both, and put the whole's exact `claim` sentence in
the part's `part_of` field. Most restatements of one result across two studies are parts of the
claim that covers both, so this is where a pair you would otherwise have left apart out of
caution belongs: the insula-ROI result stated beside the voxel-wise result is a part of the
claim that the insula tracks the guilt effect, not a second copy of it. Only the whole may be a
target; a part points at one whole in the same table.

## Confidence

Confidence is a fact about agreement, and the readers' partition means most claims are
visible to at most two of them. Use the definitions in the vocabulary below: `high` when more
than one reader surfaced it and they agree; `contested` when more than one did and they
disagree; `single-source` when one did. Single-source is the expected case for panel-level
numerics, for scope and methodological claims, and for synthesis — it is where each reader's
slice is unique, not a mark against the claim.

## Which reader to believe

- **Wording**: the caption reader for panel-level results, the results reader for hypotheses,
  predictions, synthesis and interpretation.
- **Panel**: the caption reader.
- **Numbers**: the caption reader.
- **Role**: when the results reader says `synthesis` and the caption reader anchors the same
  proposition to a panel, it is `empirical` and the synthesis context goes in `notes`. When the
  structure reader says `methodological` and the caption reader shows it as a control panel, it
  is `control`.

## Rules

- Every claim keeps every reader's evidence quote, verbatim, under that reader's name.
- Drop nothing. A candidate you doubt stays in as `single-source` with the doubt in `notes`.
- Rewrite for clarity where readers phrased one claim differently, but only from their
  evidence: no new numbers, no stronger verb.
- Three readers of a typical research paper produce 40–70 candidates and reconcile to 25–45
  claims. If you merged more than half, you are over-merging; if you merged nothing, the
  overlap between the results reader and the caption reader on panel-level findings has been
  missed.

## What to return

The JSON object described below, and nothing else: no prose before or after it, no code
fence. The vocabulary that follows defines every value you may use.
external-reviewer task · 61 lines · + contract

Reads the draft table and the paper. Returns a revised table — this is what stands in for a human review step.

extract/prompts/external-reviewer.md

# External reviewer

Three readers have extracted candidate claims from one paper and a reconciler has folded them
into a draft claim table. You are given the paper's abstract, its Introduction, its Results, its
Discussion and its figure captions, and that draft. Revise the draft so that it carries the
paper's argument, not only its results.

You no longer have to infer the organising hypothesis or the cited premises from the empirical
sequence: the Introduction states the hypothesis and the questions the paper asks, and the
Discussion states the interpretation and the literature-context it builds on. Read them. The
Discussion yields `interpretation` and `literature-context` claims, and never an `empirical`
claim that the Results do not also state. Anchor every panel to an id the captions show, and
never invent one. Any `evidence` you add is a verbatim quote with no bracketed span id in it.

You stand in for a person here. Nothing you return has been read by one, and the pipeline
records that.

## What the readers systematically miss

The readers work from slices and return what each slice states. What they under-recover is
structure that the paper establishes across slices or leaves implicit:

- **Hypotheses.** A paper rarely says "we hypothesize"; it says "we asked whether" or "we
  sought to understand how". When the draft has empirical claims that together test one
  underlying proposition and no claim states that proposition, add it with role `hypothesis`
  and `panel: null`. Most papers have one to three. An existing hypothesis is never
  reclassified as a prediction because you are also adding predictions.
- **Predictions.** Where the paper's argument runs hypothesis → prediction → test and the draft
  has only the test, add the prediction: "if [hypothesis], then [observable]", role
  `prediction`, `panel: null`. One prediction per test is the usual shape.
- **Controls.** A result whose work is to rule out an alternative ("not due to", "no effect
  of", "regardless of", a null in a confound analysis) is `control`, not `empirical`, even
  though its content is a measurement.
- **Literature-context.** A claim that restates a cited prior finding the paper builds on is
  `literature-context`, whether or not the citation is explicit. A downstream layer resolves
  these against CrossRef, so the role matters.
- **Synthesis versus interpretation.** Synthesis stays inside the paper's own evidence;
  interpretation maps the results onto a framework outside it. The readers label both
  `synthesis`.
- **Claims that span panels.** Where the prose anchors one result to several panels and the
  draft carries a single panel, extend `panel` to the list.

## What you may and may not do

You may change a claim's `role`, its `claim_type` when the role change requires it, its
`panel`, and its `notes`; and add claims. You may not delete a claim, merge two claims, or add a
number or a panel that appears neither in the prose you were given nor in the draft. A claim
you think should go is kept; add an edit for it with `notes` set to `"[reviewer] doubtful: …"`.

Mark every change. A claim you edit needs only the fields you are changing — list them in the
`edits` array keyed by the exact `claim` sentence. A revised claim's `notes` begins
`[reviewer] role: empirical → control.` with the reason. An added claim in `additions` must
carry `confidence: "single-source"`, `sources: ["reviewer"]`, an `evidence_by_agent.reviewer`
entry saying what in the prose or the draft it was inferred from, and `notes` beginning
`[reviewer] added:`. Claims you do not change are omitted from `edits` — do not list them.

## What to return

A patch object with two keys — `edits` and `additions` — and nothing else: no prose before or
after it, no code fence. The runner applies the patch to the draft, so you need only state the
changes. The schema that follows defines every value you may use.

Infer relations

edge-inference task · 169 lines · + contract

Reads the reconciled claims. Returns typed relations between them — the deductive spine.

extract/prompts/edge-inference.md

# Edge inference

You are mapping the logical structure of one paper's claim table: the typed relations that hold
between its claims. This is the deductive spine — a hypothesis entailing its predictions, an
empirical result testing one back — together with the scope constraints, the dependency chains,
and the compositions that hold the argument together. The claims are given; the edges between
them are what you return.

Relations are named in the corpus's own vocabulary, defined in full in the vocabulary that
follows this task — not CiTO, not claimrel, just the bare relation names. Each relation is
directed, and the direction rule stated beside its definition is binding: an edge that points
the wrong way is wrong even when the two claims are genuinely related. Read the definitions and
the confusable-pair notes before you begin — `requires` is not `supports`, `entails` is not
`tests`, `part-of` is not `supports`, `rules-out` is not `refutes`, `confirms` is not
`supports`.

## What you are given

The paper's claims, numbered. Refer to a claim only by its number — never by a slug and never by
a paraphrase. Each claim is shown with its role, the panel it rests on, the paper's stance toward
it, its full sentence, the span id its evidence was drawn from where the draft records one, and
the verbatim evidence quotes the readers gave. Reason only from what is there. A relation you
cannot defend from this digest is one you should not write.

## What to return

One edge per line, each a single JSON object and nothing else — no prose before or after, no code
fence:

```
{"source": 12, "target": 3, "relation": "tests", "why": "one sentence, citing a span id where the digest shows one"}
```

- `source` and `target` are claim numbers. The edge runs from source to target in the direction
  the vocabulary states for that relation.
- `relation` is one of the corpus relation names. Never emit `derived-from`: it is written
  mechanically as the reciprocal of `entails`, so emitting it by hand double-counts the same
  deduction.
- `why` is one sentence saying why the edge holds, defensible from the digest, and citing the
  span id (for example `results-026`) wherever the claim it rests on shows one.

## Every tested prediction carries its outcome

A `tests` edge is neutral: it records that a result bears on a prediction, not which way the
test came out. That verdict is the point of the test, so it must be written down. For **every**
`tests` edge you emit, from a result to a prediction, emit an outcome edge beside it, from the
same result to the same prediction:

- `confirms` — the result came out the way the prediction said it would.
- `refutes` — the result came out against the prediction.

Read the evidence quotes to decide which; a `tests` edge with no `confirms` or `refutes` beside
it leaves the prediction a waypoint with no verdict on it, and the linter flags it.

Outcomes aim at **predictions only**. A result that bears on a hypothesis does not `confirms`
or `refutes` it — it `supports` the hypothesis (or `rules-out` an alternative). If a result
came out against a hypothesis, write the prediction the hypothesis entails and refute that
prediction; do not aim `refutes` at the hypothesis.

## Look for tensions between results the paper asserts

After the outcome edges, go back through the results the paper asserts and look deliberately for
**tensions**: two claims, both of which stand, that pull a shared implication in opposite
directions, or that cannot be jointly explained without a further claim. This is not a contrast,
where two results simply differ across a condition, region, population or measure and the
difference is the finding (that is `dissociates-with`). A tension is marked by “although”,
“however”, “despite”, “no correlation with”, “at the cost of” — a replication that holds at the
group level while the individual-level null bounds what it can mean; a gain bought at the cost of
a resolution; two metrics that disagree about which method is better. Where you find one, write:

```
{"source": 12, "target": 7, "relation": "in-tension-with", "why": "one sentence"}
```

`in-tension-with` is symmetric — write it once per pair. It holds only between two claims the
paper **asserts**; it is not `contradicts` (both claims hold), not `qualifies` (neither narrows
the other), and not `rules-out` (nothing is eliminated). Do not fail to surface a real tension:
a paper usually resolves one in its discussion, and an unresolved one is a gap worth seeing.

## A synthesis points outward, and a bounded result says what it bounds

Two edges are easy to leave off, and both hide a claim from the argument the paper is making.

**A synthesis or interpretation must carry an argument edge outward.** A claim that integrates
several results — a `synthesis` or an `interpretation` — earns its place by what it does to the
argument, not by what feeds it. So it must point at something: `supports` the hypothesis it
argues for, `rules-out` an alternative it eliminates, or `interprets` the member it reframes.
The results beneath it are its details and belong under it, but they are not its reason for
being on the page. A synthesis with only incoming edges is loose — it aggregates evidence and
then does nothing with it — so give it the outward edge that says what it settles.

**A result you will not confirm still gets the edge that says what it bears on.** When a result
was run to test a prediction but you judge it does not settle that prediction — an uncorrected
cluster, an effect that does not survive, a measure that only partly meets the commitment — do
not write `confirms`, and do not leave it wired to nothing. It still bounds a finding: write
`qualifies` from the result to the claim whose applicability it narrows. `qualifies` is the edge
for a result that bounds a finding without supporting it; declining to confirm is a verdict, and
a verdict that leaves no edge leaves the result invisible to the argument.

## Look for unsupported parts of the argument

Then look for parts of the argument that rest on nothing: a hypothesis with no prediction tested,
a prediction with no result testing it, an empirical claim with no evidence behind it. These are
not edges — they are the *absence* of one — so mark each on its own line, one object per claim,
alongside the edge lines:

```
{"unsupported": 4, "reason": "hypothesis with no tested prediction"}
```

`unsupported` is the claim's number (or slug); keep each reason to one clause. This is the other
half of the reading the ruling asks for: we do not want to fail to surface these.

## Complete the tree: every claim rests on something, or stands alone

Now go back over the whole tree one more time. This is a completion pass, not a fresh reading:
the spine is written, the outcomes are on it, the tensions are marked — what is left is to
account for every claim that the reading so far has not connected to any other.

For **every** claim that has no `supports`, `extends`, `validates`, `confirms`, `refutes`,
`tests`, `rules-out`, `part-of` or `interprets` edge in *either* direction — nothing pointing at
it and nothing it points at, across all of those relations — do one of two things:

- **Write the edge.** Say what the claim rests on, or what rests on it: the result that a
  methodological choice makes interpretable, the finding a claim supports, the whole a
  measurement is a part of, the prediction a result tests. Name the relation from the vocabulary,
  point it the way the vocabulary's direction rule states, and cite the span that shows it — a
  completion edge is held to the same standard as any other. Most unconnected claims have a real
  place in the argument that the spine reading simply did not reach: a control validating a
  result, a scope bounding one, an empirical finding supporting the hypothesis it was run under.
- **List it as unsupported.** If, having looked, the claim genuinely rests on nothing the paper
  says and nothing rests on it, mark it on its own line, exactly as in the section above, with a
  reason of one line saying it stands alone in the paper's argument:

```
{"unsupported": 9, "reason": "stands alone — no claim in the paper bears on it or rests on it"}
```

This is a pass to *find the edges the tree already has and did not write down*, not a licence to
invent connections. An edge you cannot cite a span for is an edge you should not write; when the
honest answer is that a claim stands alone, that absence is itself a finding, and listing it
under `unsupported` is the right move. It is better to mark a claim unsupported than to wire it
to a neighbour it has no real relation to.

## The rules the direction checks enforce

These are checked mechanically after you answer; an edge that breaks one is dropped rather than
written, so do not spend an edge on it.

- **`tests`** runs from an empirical result or a control to the **prediction** it checks. Never
  from the prediction, and never aimed at a hypothesis.
- **`confirms`, `refutes`** run from the empirical result or control to the **prediction** whose
  test they settle — `confirms` when it came out as predicted, `refutes` when it came out
  against. Both are aimed at a prediction only; aimed at a hypothesis they are dropped.
- **`entails`** runs from a **hypothesis** to a prediction it deductively implies. An empirical
  result never entails anything.
- **`scopes`** runs from a **scope** claim to the claims it bounds, or to `*` for every empirical
  claim in the paper.
- **`rules-out`, `contradicts`, `opposes`** may target only a claim the paper does **not** assert
  — one whose stance is `entertains` or `rejects`. A paper does not rule out what it claims; if
  every stance in the digest reads `asserts`, emit none of these, and let the alternative it
  eliminates get its own node later.
- **`part-of`** is composition, not evidence: the source states one comparison, condition, measure
  or study of a proposition the target states as a whole. An independent finding that supports a
  claim is not a part of it. A part points at exactly one whole, and the chain of wholes may not
  close a cycle.

It is better to miss an edge than to invent one. Return only the edges you can defend from the
digest.
questions task · 50 lines · + contract

Reads the abstract and the paper's hypothesis and alt- claims. Returns the research questions the paper states, and which claim answers each.

extract/prompts/questions.md

# Questions

You are reading one scientific paper to recover the **questions** it set out to answer, and to
say which of its claims answers each. A question is not a claim — a claim is a declarative
sentence — so it is recorded on the paper, not as a node in the graph. The vocabulary that
follows this task defines what a question is and how a hypothesis relates to it.

## What you are given

- **The abstract**, and the **Introduction** when it is available. (The prepared paper does not
  always carry an Introduction; when it is absent, read the questions off the abstract, which
  states them for a research paper.)
- **The paper's hypotheses and its rejected alternatives**, each as a slug and a sentence. The
  hypotheses are the answers the paper commits to; the `alt-` claims are the rivals it argues
  against. Both answer questions, and those are the only claims you may attach to a question.

## What to find

Find the questions the paper states — what it asked, in the abstract or the opening of the
Introduction. A paper usually asks one to three. State each as a question, in the paper's own
terms where it gives them, in one sentence. Then say, for each hypothesis and each rejected
alternative, which question it answers.

- A hypothesis is the answer the paper **commits to**; the alternatives it rules out are the
  **other answers** to the same question. So a hypothesis and the alternatives it competes with
  usually address the *same* question.
- Do not manufacture a question by hollowing out a hypothesis — "X involves neural mechanisms"
  is not a question the paper asked. Return the question it actually posed, or leave a claim
  unaddressed if none fits.
- Every `addresses` value must be one of the `id`s you returned, and every slug must be one you
  were given. Do not invent slugs or questions.

## What to return

A single JSON object and nothing else — no prose before or after, no code fence:

```json
{
  "questions": [
    {"id": "q1", "text": "Does the anterior insula encode responsibility-contingent guilt?"}
  ],
  "addresses": {
    "hypothesis-insula-tracks-interpersonal-guilt": "q1",
    "alt-agency-aversion-not-guilt": "q1"
  }
}
```

Number the questions `q1`, `q2`, … in the order you would present them. A claim that answers no
stated question is simply left out of `addresses`.
parts task · 49 lines · + contract

Reads the tree's claims, with their roles, panels and edges. Returns the `part-of` edges — which claims are components of other claims.

extract/prompts/parts.md

# Parts

You are reading one paper's claim tree to find its **parts**: claims that are components of
other claims. A claim tree has two grains. The coarse grain is the argument a reader of the
paper follows; the fine grain is every comparison, condition, measure and study the prose
states on the way to it. A **part** is a claim at the fine grain that belongs to one at the
coarse grain, and the format now carries both, so neither has to be thrown away. The vocabulary
that follows this task defines `part-of`.

In short: a claim is **part of** another when it states one comparison, one condition, one
measure or one study of a proposition the other states as a whole, and dropping it would weaken
that whole without falsifying it. The whole already says what the part establishes; the part
carries the panel and the number that establish it.

## What you are given

The paper's claims — each as a slug, its role, the panel it rests on, its sentence, and the
typed edges it already carries. The graph half-says composition already: a part usually carries
a `supports` or a `tests` edge into the claim it is a component of. Treat that as a cue, not the
test — an independent finding can support a claim without being a part of it, and that is the
distinction `part-of` exists to draw. Apply the definition, not the edge.

## What to find

For every claim that is a component of another under the definition, the whole it is a part of.

- The whole is another claim **in this same tree**, named by its slug. Do not invent slugs.
- A part points at exactly **one** whole. If a claim seems to compose two, it composes the
  narrower one; if neither is narrower, it is probably not a part.
- A part is not the same claim as its whole, and the edges may not form a cycle: follow the
  chain of wholes upward and it must end.
- A part may itself be a whole for a finer part. Do not force a tree to one level, and do not
  manufacture a second level where the prose states none.
- Most claims are wholes. A claim that states an independent finding, a hypothesis, a
  prediction, a scope condition or a control in its own right is not a part — leave it out.

## What to return

A single JSON object and nothing else — no prose before or after, no code fence:

```json
{
  "parts": [
    {"part": "<slug>", "whole": "<slug>", "why": "one sentence: the comparison, condition, measure or study this states of the whole"}
  ]
}
```

Return only the parts you find. A tree with no parts returns `{"parts": []}`.
stance task · 74 lines · + contract

Reads the tree's controls and empirical claims, the questions, and the Results prose. Returns the alternatives the paper rejects, and the `rules-out` edges from its controls.

extract/prompts/stance.md

# Stance

You are reading one paper to recover the **alternatives it argues against** — the rival
explanations, confounds and competing accounts it raises in order to reject, together with the
controls or results that eliminate each. An alternative explanation is a proposition, so it
gets a claim of its own; what marks it as a rival is its **stance**, not its role. The
vocabulary that follows this task defines stance, and the `rules-out` relation that points from
a control to the rival it kills.

Why this matters. A claim tree records what a paper's results *support*. Its eliminative
arguments — "this rules out explanation X of finding Y" — carry their warrant in the
elimination, and an elimination whose target is not a node in the graph has nowhere to point:
it gets wired to the nearest claim that does exist, which is usually one the paper asserts, and
then the graph says the paper contradicts itself. The rule the format enforces is that a
`rules-out` edge is **never** aimed at a claim the paper asserts. This layer gives each ruled-out
rival a node of its own so the edge has something true to point at.

## What you are given

- **The paper's controls and empirical claims** — each as a slug, its role, the panel it rests
  on, and its sentence. These are the claims that do the eliminating: a control exists to rule
  an alternative out, and an empirical result often does too. A control's sentence usually names
  its target in a closing clause — "evidence against …", "confirming …", "validating … as a
  manipulation check", "regardless of whether …".
- **The paper's research questions**, each with an id. An alternative is one of the *other
  answers* to a question the paper's own hypothesis answers.
- **Alternatives already recorded**, when the paper has any — each as its slug and its sentence.
  If you raise the same proposition again, reuse its slug so the existing claim is refined
  rather than duplicated.
- **The Results prose**, one sentence per line, each prefixed with the span id you cite in
  `span`.

## What to find

Every alternative the paper raises in order to reject, or raises and leaves open. State each as
the paper would have stated the proposition it argues against — the rival's own claim, the one
the paper is saying is false — not the paper's refutation of it.

- `stance` is `rejects` when the paper argues the alternative is false, and `entertains` when it
  raises the rival and does not settle it.
- `ruled_out_by` names the control or empirical claims — by slug, from the list you were given —
  whose result eliminates the alternative. A rejected alternative normally has at least one.
  Name only claims that do the eliminating, and only slugs you were given; do not invent them.
- `addresses` names the question the alternative answers, by id. Leave it out when none of the
  stated questions fits, or when the paper states no questions.
- `span` cites the Results sentence that states or names the alternative.
- `slug` is optional. Give one to reuse an alternative already recorded; otherwise it is derived
  from the proposition. It is written with an `alt-` prefix either way.
- Do **not** raise an alternative the paper actually asserts. A proposition the paper concludes
  is true is one of its claims with the default stance, not an alternative — and aiming an
  elimination at it is the one error the format forbids.

## What to return

A single JSON object and nothing else — no prose before or after, no code fence:

```json
{
  "alternatives": [
    {
      "slug": "alt-agency-aversion-not-guilt",
      "claim": "The happiness cost in the Social condition is general agency aversion — the unpleasantness of being the chooser as such — and is not contingent on responsibility for the partner's outcome.",
      "role": "hypothesis",
      "stance": "rejects",
      "addresses": "q1",
      "ruled_out_by": ["participant-happiness-lower-when-participant"],
      "span": "results-041",
      "why": "one sentence: how the named control or result eliminates this rival"
    }
  ]
}
```

Return only the alternatives you find. A paper that raises none returns `{"alternatives": []}`.

Adjudicate coverage

coverage-adjudicator task · 46 lines

Reads unmatched spans of the paper and the claim tree. Returns a covered / gap / no-assertion verdict for each span.

extract/prompts/coverage-adjudicator.md

You are sorting spans of a scientific paper that an automated coverage check could not match
to any claim in that paper's claim tree.

The automated check matched a span to a claim only when the claim literally restates a
statistic found in the span, or names the same figure panel. That test is deliberately
mechanical, so it produces two kinds of false alarm alongside the real gaps. Your job is to
separate them.

For each span, decide which ONE of these it is:

- `covered` — a claim in the tree does account for this span's content; the check could not
  see it, usually because the claim states the effect without repeating the numbers. Name the
  claim slug.
- `gap` — no claim in the tree accounts for this. A real hole in the corpus.
- `not-an-assertion` — the span states no result of its own. This covers pure cross-references
  ("see Appendix 1—table 4"), analysis narration, bare panel labels, graphical-encoding notes
  ("error bars represent SEM", "each coloured dot is one recovered parameter"), and raw table
  rows or column headers, which are data rather than a stated finding.

## Rules

Judge on meaning, not on wording. A claim that says "participants were less happy after a
partner's loss when they had chosen" covers a span reporting the interaction statistic for
exactly that effect, even though no digits are shared.

Do not stretch. If a claim is about a different condition, a different region, or a different
direction of effect, that is a gap, not a match.

**When genuinely torn between `covered` and `gap`, say `gap`.** The purpose of the exercise is
to find holes, and a false `covered` hides one where a false `gap` only costs a second look.

Captions and tables need a distinction the Results section does not. A caption sentence that
DESCRIBES what a panel displays is asserting a finding — judge it `covered` or `gap`. A caption
sentence that only explains the GRAPHICAL ENCODING asserts nothing about the world, and is
`not-an-assertion`. A raw table row is `not-an-assertion` — unless that row is the only place a
specific reported result appears, in which case it is a `gap`, and say which result it carries.

## Output

A JSON array, one object per span, no prose and no markdown fences:

    [{"uid": "results-026", "verdict": "gap", "claim": null, "why": "one short sentence"}]

`claim` is the slug for a `covered` verdict and null otherwise. `why` is one sentence
explaining the verdict — it becomes the body of the mark written into the paper, so write it
for a reader who will see it beside the sentence it judges.

The contract every role receives

Appended to each task above: the vocabulary of roles, claim types, relations and confidence, with one corpus example of each, and the exact shape to return. Generated from the code that defines those things, so the prompt and the checker cannot disagree.

contract/schema-candidate.md 33 lines · generated
# What a reader returns

Generated from `extract/claim_graphs/schema.py` by `claim-graphs contract --write`. Do not edit.

A JSON array of candidate claims and nothing else — no prose before or after, no code fence. Each element:

| Field | Value | Meaning |
|:--|:--|:--|
| `claim` | `string` | One declarative sentence in active voice, with the paper's own numbers and the paper's own epistemic verb. |
| `panel` (optional) | `string` or `null` | The panel that shows the result, lowercase, as the paper labels it: fig3a, fig5d-h, fig3s1c, table1. Several panels separated by commas when the claim spans them. null for a claim no panel shows. |
| `claim_type` | `empirical` / `interpretive` / `existence` / `synthesis` / `assessment` / `hypothesis` / `prediction` | The kind of proposition; see the vocabulary. |
| `role` | `hypothesis` / `prediction` / `empirical` / `control` / `scope` / `methodological` / `synthesis` / `interpretation` / `literature-context` | The work it does in the argument; see the vocabulary. |
| `addresses` (optional) | `string` or `null` | For a hypothesis or an alternative explanation, the research question it answers, as the paper states it or in one sentence; null for every other role. |
| `evidence` | `string` | A verbatim quote from the text you were given, at most two sentences, that grounds the claim. It is checked against the source. |
| `confidence` | `high` / `tentative` | high or tentative; see the vocabulary. |
| `span` (optional) | `string` or `null` | The id of the span the evidence quote comes from, exactly as it is bracketed before the sentence in your slice (results-026). null when the sentence shows no id. |
| `notes` (optional) | `string` or `null` | Hedges, alternative readings, or what made this tentative. null when there is nothing to say. |
| `evidence_verified` (optional) | `boolean` or `null` | Filled by the runner. Leave null. |
| `evidence_verified_against` (optional) | `span` / `slice` or `null` | Filled by the runner: whether the quote matched the cited span or only the wider slice. Leave null. |

```json
[
  {
    "claim": "Doubling distal dendritic inhibition reduces somatic firing from approximately 5.5 Hz to approximately 0.2 Hz.",
    "panel": "fig4a",
    "claim_type": "empirical",
    "role": "empirical",
    "evidence": "Doubling the strength of distal inhibition reduced the firing rate from 5.5 \u00b1 0.9 Hz to 0.2 \u00b1 0.2 Hz.",
    "confidence": "high",
    "notes": null
  }
]
```
contract/schema-draft.md 66 lines · generated
# What the reconciler and the reviewer return

Generated from `extract/claim_graphs/schema.py` by `claim-graphs contract --write`. Do not edit.

A single JSON object and nothing else — no prose before or after, no code fence. Top level:

| Field | Value | Meaning |
|:--|:--|:--|
| `paper_slug` | `string` | As given in the input. |
| `paper_doi` | `string` | As given in the input. |
| `paper_title` (optional) | `string` or `null` | As given in the input. |
| `extraction_path` (optional) | `jats` / `pdf` or `null` | Filled by the runner from prepared.json. Leave null. |
| `extraction_path_note` (optional) | `string` or `null` |  |
| `per_agent_counts` (optional) | object | How many candidates each reader proposed. |
| `model` (optional) | `string` or `null` |  |
| `claims` | list of ReconciledClaim | Every surviving claim. |
| `config_snapshot` (optional) | object | Filled by the runner: models and prompt variant. Leave empty. |

Each element of `claims`:

| Field | Value | Meaning |
|:--|:--|:--|
| `claim` | `string` | The canonical sentence. Where readers phrased one proposition differently, the most precise phrasing, grounded in their evidence. |
| `panel` (optional) | `string` or `null` | As for a reader. Where readers disagree, the caption reader's panel. |
| `claim_type` | `empirical` / `interpretive` / `existence` / `synthesis` / `assessment` / `hypothesis` / `prediction` | See the vocabulary. |
| `role` | `hypothesis` / `prediction` / `empirical` / `control` / `scope` / `methodological` / `synthesis` / `interpretation` / `literature-context` | See the vocabulary. |
| `addresses` (optional) | `string` or `null` | For a hypothesis or an alternative explanation, the research question it answers, as the paper states it or in one sentence; null for every other role. |
| `confidence` | `high` / `contested` / `single-source` | A fact about agreement: single-source for one reader, high for several who agree, contested for several who disagree. |
| `sources` | list of `results` / `caption` / `structure` / `reviewer` | The readers that surfaced this claim: results, caption, structure; reviewer for a claim the review pass added. |
| `evidence_by_agent` (optional) | object | For each reader in sources, the verbatim quote it gave. |
| `span_by_agent` (optional) | object | For each reader in sources that cited one, the span id its evidence quote came from. |
| `evidence_verified` (optional) | object | Filled by the runner. Leave null. |
| `evidence_verified_against` (optional) | object | Filled by the runner: per reader, `span` or `slice`. Leave null. |
| `notes` (optional) | `string` or `null` | What the readers disagreed about, or why a single-source claim deserves a second look. A review pass prefixes its notes with [reviewer]. |
| `part_of` (optional) | `string` or `null` | The exact `claim` sentence of another claim in this same table that this one is a component of — one comparison, condition, measure or study of a proposition that claim states whole. Keep both; the writer resolves the sentence to the whole's slug and writes `part-of`. null when this claim is not a part of another. |

```json
{
  "paper_slug": "headley-2026-inhibitory-rhythms",
  "paper_doi": "10.7554/eLife.95562",
  "paper_title": "Spatially targeted inhibitory rhythms differentially affect neuronal integration",
  "per_agent_counts": {
    "results": 18,
    "caption": 22,
    "structure": 7
  },
  "claims": [
    {
      "claim": "Doubling distal dendritic inhibition reduces somatic firing from approximately 5.5 Hz to approximately 0.2 Hz.",
      "panel": "fig4a",
      "claim_type": "empirical",
      "role": "empirical",
      "confidence": "high",
      "sources": [
        "results",
        "caption"
      ],
      "evidence_by_agent": {
        "results": "distal inhibition nearly silenced the cell (0.2 Hz)",
        "caption": "Doubling the strength of distal inhibition reduced the firing rate from 5.5 \u00b1 0.9 Hz to 0.2 \u00b1 0.2 Hz."
      },
      "notes": null
    }
  ]
}
```
contract/schema-review-patch.md 69 lines · generated
# What the external reviewer returns

Generated from `extract/claim_graphs/schema.py` by `claim-graphs contract --write`. Do not edit.

A single JSON object with two keys — `edits` and `additions` — and nothing else: no prose before or after, no code fence.

Top level:

| Field | Value | Meaning |
|:--|:--|:--|
| `edits` (optional) | list of ReviewEdit | Zero or more targeted changes to existing claims, keyed by their `claim` sentence. |
| `additions` (optional) | list of ReconciledClaim | Zero or more new claims to append to the draft; each must carry confidence=single-source and sources=[reviewer]. |

Each element of `edits` (targeted changes to existing claims):

| Field | Value | Meaning |
|:--|:--|:--|
| `claim` | `string` | The exact `claim` sentence of the claim to edit, as it appears in the draft. |
| `role` (optional) | `hypothesis` / `prediction` / `empirical` / `control` / `scope` / `methodological` / `synthesis` / `interpretation` / `literature-context` or `null` | New role to assign; omit or null to leave unchanged. |
| `claim_type` (optional) | `empirical` / `interpretive` / `existence` / `synthesis` / `assessment` / `hypothesis` / `prediction` or `null` | New claim_type to assign; omit or null to leave unchanged. |
| `panel` (optional) | `string` or `null` | New panel value; omit or null to leave unchanged. |
| `notes` (optional) | `string` or `null` | Replacement notes string (prefix with [reviewer]); omit to leave unchanged. |

Each element of `additions` (new claims appended to the draft):

| Field | Value | Meaning |
|:--|:--|:--|
| `claim` | `string` | The canonical sentence. Where readers phrased one proposition differently, the most precise phrasing, grounded in their evidence. |
| `panel` (optional) | `string` or `null` | As for a reader. Where readers disagree, the caption reader's panel. |
| `claim_type` | `empirical` / `interpretive` / `existence` / `synthesis` / `assessment` / `hypothesis` / `prediction` | See the vocabulary. |
| `role` | `hypothesis` / `prediction` / `empirical` / `control` / `scope` / `methodological` / `synthesis` / `interpretation` / `literature-context` | See the vocabulary. |
| `addresses` (optional) | `string` or `null` | For a hypothesis or an alternative explanation, the research question it answers, as the paper states it or in one sentence; null for every other role. |
| `confidence` | `high` / `contested` / `single-source` | A fact about agreement: single-source for one reader, high for several who agree, contested for several who disagree. |
| `sources` | list of `results` / `caption` / `structure` / `reviewer` | The readers that surfaced this claim: results, caption, structure; reviewer for a claim the review pass added. |
| `evidence_by_agent` (optional) | object | For each reader in sources, the verbatim quote it gave. |
| `span_by_agent` (optional) | object | For each reader in sources that cited one, the span id its evidence quote came from. |
| `evidence_verified` (optional) | object | Filled by the runner. Leave null. |
| `evidence_verified_against` (optional) | object | Filled by the runner: per reader, `span` or `slice`. Leave null. |
| `notes` (optional) | `string` or `null` | What the readers disagreed about, or why a single-source claim deserves a second look. A review pass prefixes its notes with [reviewer]. |
| `part_of` (optional) | `string` or `null` | The exact `claim` sentence of another claim in this same table that this one is a component of — one comparison, condition, measure or study of a proposition that claim states whole. Keep both; the writer resolves the sentence to the whole's slug and writes `part-of`. null when this claim is not a part of another. |

```json
{
  "edits": [
    {
      "claim": "Doubling distal dendritic inhibition reduces somatic firing from approximately 5.5 Hz to approximately 0.2 Hz.",
      "role": "empirical",
      "notes": "[reviewer] role: empirical \u2192 control. This result rules out that somatic firing is driven by distal input rather than proximal."
    }
  ],
  "additions": [
    {
      "claim": "Distal inhibition more strongly reduces somatic firing than proximal inhibition.",
      "panel": null,
      "claim_type": "synthesis",
      "role": "synthesis",
      "addresses": null,
      "confidence": "single-source",
      "sources": [
        "reviewer"
      ],
      "evidence_by_agent": {
        "reviewer": "The contrast across fig4a (distal) and fig4b (proximal) is stated in the Discussion: distal inhibition is uniquely effective."
      },
      "notes": "[reviewer] added: synthesis across panels."
    }
  ]
}
```
contract/vocabulary.md 255 lines · generated
# The claim vocabulary

Generated from `extract/claim_graphs/vocabulary.py` and `scripts/relations.py` by `claim-graphs contract --write`. Do not edit; edit the source and regenerate. Every example is a claim in the corpus, quoted from its file.

A claim is one declarative sentence in active voice, quantitative where the result is quantitative, carrying the paper's own epistemic verb. It has a **claim type** (what kind of proposition it is), a **role** (the work it does in this paper's argument), and **relations** to other claims. Type and role are independent axes: a measurement can play the role of a result, a control or a scope condition.

## Claim types

| `claim_type` | Meaning |
|:--|:--|
| `empirical` | a directly observed or computed result |
| `interpretive` | an inference drawn from one or more empirical claims |
| `existence` | an assertion that a phenomenon, entity or resource exists |
| `synthesis` | a claim integrating results across several analyses or papers |
| `assessment` | a methodological, scope or quality claim about how the work was done |
| `hypothesis` | a proposition bet on, not yet evidenced by this paper's results |
| `prediction` | a deduced expectation, to be tested by an empirical claim |

## Questions

A **question** is what the paper set out to answer. It is not a claim — a claim is a declarative sentence — so it is recorded on the paper rather than as a node in the graph. The paper states it, in the abstract or in the opening of the Introduction. Each `hypothesis` and each rejected alternative addresses one: the hypothesis is the answer the paper commits to, and the alternatives it rules out are the other answers to the same question. Never turn a question into a hollow hypothesis such as "X involves neural mechanisms" — return the question the paper actually asked, or return none.

## Roles

Nine roles. The signal phrases are what the prose says when it is doing that work; they are cues, not tests.

### `hypothesis`

An answer the paper commits to, to a question it states: the proposition it bets on, phrased as a claim about the world rather than as the question. It carries no empirical content of its own; it is what the predictions are deduced from and what the results are gathered for, and the alternatives the paper rules out are the other answers to the same question. Each hypothesis carries `addresses`, the question it answers. Most papers have one to three.

- Typical claim type: `hypothesis`
- Signals: “we hypothesize”, “we propose that”, “we asked whether”, “we sought to test”, “the central question is whether”
- Carries: `entails` to each of its predictions; usually `panel: null`
- Example: `hypothesis-distinct-compartmental-roles` (hypothesis, Headley): “Perisomatic and distal dendritic inhibition serve distinct computational roles in layer 5 pyramidal neurons: perisomatic inhibition principally regulates somatic action potential generation (gain and threshold of axonal output), while distal dendritic inhibition principally regulates dendritic spike incidence and the temporal coupling of dendritic spikes to somatic APs.”

### `prediction`

What should be observed if the hypothesis holds, under stated conditions. A prediction is deduced, not measured: it is anchored in the model or in principled reasoning, and an empirical claim then tests it. Write it as a conditional when the paper does not.

- Typical claim type: `prediction`
- Signals: “if X, then we should observe Y”, “this predicts that”, “the model predicts”, “should be maximally effective at”, “is predicted to”
- Carries: `derived-from` back to its hypothesis (written mechanically); is the target of `tests`
- Example: `prediction-beta-optimal-distal` (prediction, Headley): “If the optimal frequency of rhythmic inhibition at a compartment is set by matching the rhythm period to the local spike timescale, then distal inhibition — where apical Ca²⁺ and NMDA spikes lead the soma by ~20 ms and ~25 ms respectively — should be maximally effective at beta frequencies (~20 Hz), whose period (~50 ms) matches the dendritic-spike lead time.”

### `empirical`

A measured or computed result, anchored to the panel that shows it, carrying the paper's own numbers and the paper's own epistemic verb. The largest role.

- Typical claim type: `empirical`
- Signals: “we found that”, “we observed”, “we measured”, “increased”, “did not differ”
- Carries: `tests` to the prediction it checks; `supports` and `requires` as the argument needs
- Example: `distal-inhib-drops-firing-02hz` (empirical, Headley): “Doubling the strength of distal dendritic inhibition reduces somatic firing rate from a baseline of approximately 5.5 Hz to approximately 0.2 Hz, primarily by suppressing the occurrence of dendritic Ca²⁺ and NMDA spikes rather than by directly raising AP threshold.”

### `control`

An empirical result whose work in the argument is to eliminate an alternative explanation or to show a manipulation did what it should. Its content is a measurement; what makes it a control is what it rules out. A null result is usually a control.

- Typical claim type: `empirical`
- Signals: “rules out”, “excludes”, “is not due to”, “no significant effect of”, “control condition”, “manipulation check”, “regardless of”
- Carries: `rules-out` to the alternative it eliminates; `validates` to the claim it defends
- Example: `risk-premiums-not-differ-between` (control, Gadeke): “Risk premiums did not differ between Solo and Social conditions in either study (Study 1: t(39) = 1.53, p = 0.134, d = 0.24, BF10 = 0.49; Study 2: t(43) = –0.21, p = 0.84, d = –0.03, BF10 = 0.17).”

### `scope`

A boundary condition on what the results can mean: the model class, the preparation, the population, the design. Often global. It asserts nothing about the world; it says where the paper's assertions apply.

- Typical claim type: `assessment`
- Signals: “all results come from”, “restricted to”, “in this preparation”, “was not a physiological pattern”, “N = ”
- Carries: `scopes` to the claims it bounds, or `scopes: ["*"]`
- Example: `l5-model-single-cell-scope` (scope, Headley): “All results derive from a single-cell compartmental model of one layer 5 pyramidal neuron; no network dynamics, recurrent excitation, or population-level inhibitory effects are simulated, and all firing rate effects are for a single isolated neuron receiving naturalistic presynaptic drive.”

### `methodological`

A capability or analytical commitment that a downstream result depends on for its interpretation: the sorting pipeline, the null distribution, the model fit that licenses a model-based analysis. Not procedure for its own sake — which software ran the task is not a claim unless a result turns on it. A localizer — a contrast run only to define a region or a set of trials for a later analysis — is `methodological`, not `empirical`, because the paper does not argue from it.

- Typical claim type: `assessment`
- Signals: “analysis is on”, “nulls are”, “fit better than”, “validated against”
- Carries: `enables-method` to the results it warrants
- Example: `preregistered-design-validates-mvpa` (methodological, Kammer): “The preregistered analysis plan (osf.io/rxacd) specifies the MVPA decoding pipeline, ROI definitions, and statistical tests in advance, reducing the risk of analytic flexibility inflating the decoding accuracy results.”

#### What is not a claim

Procedure is a claim only when a result turns on it. The test is one question: would any result mean something different if this had been done differently? If not, it is not a claim, however carefully the methods state it.

Procedure a result does turn on, and so is methodological:

- `preregistered-design-validates-mvpa` (methodological, Kammer): “The preregistered analysis plan (osf.io/rxacd) specifies the MVPA decoding pipeline, ROI definitions, and statistical tests in advance, reducing the risk of analytic flexibility inflating the decoding accuracy results.”
- `model-based-glm-entered-best-fitting-computational` (methodological, Gadeke): “A model-based GLM (GLM2) entered the best-fitting computational (Responsibility) model's variables — certain rewards (CR), expected value (EV), participant RPE (sRPE), and partner RPE from participant choices (social_pRPE) and from partner choices (partner_pRPE) — as regressors to locate brain regions reflecting them.”
- `momentary-happiness-modelled-five-computational` (methodological, Gadeke): “Momentary happiness was modelled with five computational models (Basic, Inequality, Guilt-envy, Responsibility, and Responsibility Redux) sharing separate, exponentially decaying terms for certain rewards, expected value, and reward prediction errors.”
- `parameter-recovery-procedure-synthetic-data-generated` (methodological, Gadeke): “A parameter-recovery procedure on synthetic data generated from each participant's estimated parameters showed the happiness-model parameters could be reliably recovered, verifying their stability.”

Procedure no result turns on, so not a claim — each is only what the clause says:

- “Happiness ratings were Z-scored per participant to remove the influence of differing rating variability across participants.” — a normalisation — no result reads differently for it.
- “Risk attitude was quantified as a risk premium — the EVdiff value yielding 50% risky choices from a fitted logistic regression — and compared between Solo and Social conditions with paired t-tests in both studies.” — a definition of a measure — the finding is that the premium did not differ, not that it was defined this way.
- “Two gPPI seed-to-voxel connectivity analyses used functionally defined seeds: the left insula cluster more sensitive to Risky versus Safe outcomes (GLM3) and the left STS cluster responding more to social_pRPE than partner_pRPE (GLM4), with identical seeds across participants.” — a seed choice — the connectivity result is the claim, not which seeds produced it.
- “All reported clusters survive a whole-brain family-wise-error-corrected threshold of p < 0.05 with a cluster-forming voxel-wise threshold of p < 0.001 (or a smaller volume where explicitly mentioned).” — a threshold applied to every result alike — a scope condition folded into the paper's scope claim, not a finding.
- “The fMRI data of four Study 2 participants were excluded from the fMRI analysis for excessive head motion (>3 mm or >3°).” — an exclusion count — a component of the paper's scope claim, not a claim of its own.
- “The Study 2 sample size of 44 was fixed a priori by a G*Power analysis based on Study 1's effect size (Cohen's d = 0.56), with alpha = 0.05 and power = 0.95.” — a power analysis fixing the sample — a component of the paper's scope claim.
- “The experiment was implemented in MATLAB using Psychtoolbox.” — a software choice — no result would mean anything different in another toolbox.

### `synthesis`

A higher-order proposition that integrates several of the paper's own results into one claim, staying inside the paper's evidence: the dissociation, the reconciliation, the summary that several panels jointly establish.

- Typical claim type: `synthesis`
- Signals: “taken together”, “these results show”, “this dissociation establishes”, “in summary”
- Carries: is the target of `supports` from the results it integrates
- Example: `orthogonality-derived-from-additivity` (synthesis, Meijer): “In a 2 × 2 factorial choice × stimulation experimental design with linear (PCA) projection of the population response, additive 5-HT modulation entails geometric orthogonality between the stimulation-effect direction and the choice direction. Under additive modulation, every choice trajectory is displaced by the same vector (the stimulation effect), so the displacement that distinguishes the stimulated from the unstimulated trajectories — averaged across choice — is by definition orthogonal to the displacement that distinguishes the choice trajectories — averaged across stimulation. The orthogonality result in PCA space (`5ht-axis-orthogonal-to-choice-axis`) is therefore a geometric corollary of the GLM additivity finding, not an independent empirical claim.”

### `interpretation`

A reading of the results through a theoretical lens from outside the paper's own evidence: a mapping onto a framework, a proposed mechanism, a functional meaning. It is an act of mapping, not a derivation.

- Typical claim type: `interpretive`
- Signals: “may provide a functional interpretation”, “suggests a role for”, “points to a mechanism whereby”, “is consistent with the view that”
- Carries: `interprets` to the empirical claims it reframes
- Example: `pv-gamma-sst-beta-correspondence` (interpretation, Headley): “The model provides mechanistic grounding for the empirical association of parvalbumin-positive interneurons with gamma rhythms and somatostatin-positive interneurons with beta rhythms: PV+ neurons target perisomatic locations where gamma is optimal for AP threshold modulation, while SST+ neurons target distal dendrites where beta is optimal for dendritic spike entrainment.”

### `literature-context`

A finding from cited prior work that the paper's argument inherits as a premise, recorded as a claim of its own so the inheritance is auditable. The citation may be implicit: prose that paraphrases a prior empirical pattern as background is literature-context whether or not it names the paper.

- Typical claim type: `interpretive`
- Signals: “as shown by”, “previous work has established”, “the reported association of”, “(Author, Year) found”
- Carries: is the target of `requires` or `interprets` from the claims that lean on it
- Example: `interprets-pv-gamma-sst-beta-associations` (literature-context, Headley): “Empirical work across the cortical-interneuron literature has established two correlated associations: parvalbumin-positive (PV+) fast-spiking interneurons target perisomatic compartments and are implicated in the generation and entrainment of cortical gamma rhythms (40–80 Hz), while somatostatin-positive (SST+) interneurons target distal dendritic compartments and are preferentially associated with beta rhythms (12–35 Hz). These are literature claims about anatomical targeting patterns and their correlation with specific oscillatory bands, not results of the present paper.”

### Roles that are confused for each other

**`hypothesis` versus `prediction`.** A hypothesis says what is the case; a prediction says what will be observed if it is. "Compartments serve distinct roles" is the bet; "if so, doubling distal inhibition should suppress dendritic spikes more than doubling perisomatic inhibition" is what it commits the paper to seeing. Surface both, and keep them apart: adding predictions never reduces the number of hypotheses.

- `hypothesis`: `hypothesis-distinct-compartmental-roles` (hypothesis, Headley): “Perisomatic and distal dendritic inhibition serve distinct computational roles in layer 5 pyramidal neurons: perisomatic inhibition principally regulates somatic action potential generation (gain and threshold of axonal output), while distal dendritic inhibition principally regulates dendritic spike incidence and the temporal coupling of dendritic spikes to somatic APs.”
- `prediction`: `prediction-distal-dendritic-spike-mechanism` (prediction, Headley): “If perisomatic and distal dendritic inhibition serve dissociated computational roles, then doubling distal dendritic inhibition should reduce somatic firing primarily by suppressing apical Ca²⁺ and NMDA dendritic spikes, with little change in the somatic AP voltage threshold.”

**`control` versus `empirical`.** Both are measurements. Ask what the result is *for*: if it demonstrates the effect the paper is about, it is empirical; if it shows that something else does not explain that effect, or that the manipulation worked, it is a control.

- `control`: `risk-premiums-not-differ-between` (control, Gadeke): “Risk premiums did not differ between Solo and Social conditions in either study (Study 1: t(39) = 1.53, p = 0.134, d = 0.24, BF10 = 0.49; Study 2: t(43) = –0.21, p = 0.84, d = –0.03, BF10 = 0.17).”
- `empirical`: `when-partner-received-low-lottery` (empirical, Gadeke): “When the partner received the low lottery outcome, participant happiness was lower when the participant rather than the partner had chosen the lottery — a significant partner-outcome × decision-maker interaction (Study 1: t(1180) = 3.52, p = 0.0004, β = 0.37; Study 2: t(937) = 2.85, p = 0.0045, β = 0.33) — operationalizing interpersonal guilt.”

**`synthesis` versus `interpretation`.** Synthesis stays inside the paper's own evidence and says what several results jointly establish. Interpretation reaches outside it, to a framework, a mechanism or a literature, and says what the results mean there.

- `synthesis`: `orthogonality-derived-from-additivity` (synthesis, Meijer): “In a 2 × 2 factorial choice × stimulation experimental design with linear (PCA) projection of the population response, additive 5-HT modulation entails geometric orthogonality between the stimulation-effect direction and the choice direction. Under additive modulation, every choice trajectory is displaced by the same vector (the stimulation effect), so the displacement that distinguishes the stimulated from the unstimulated trajectories — averaged across choice — is by definition orthogonal to the displacement that distinguishes the choice trajectories — averaged across stimulation. The orthogonality result in PCA space (`5ht-axis-orthogonal-to-choice-axis`) is therefore a geometric corollary of the GLM additivity finding, not an independent empirical claim.”
- `interpretation`: `pv-gamma-sst-beta-correspondence` (interpretation, Headley): “The model provides mechanistic grounding for the empirical association of parvalbumin-positive interneurons with gamma rhythms and somatostatin-positive interneurons with beta rhythms: PV+ neurons target perisomatic locations where gamma is optimal for AP threshold modulation, while SST+ neurons target distal dendrites where beta is optimal for dendritic spike entrainment.”

**`scope` versus `methodological`.** Scope bounds where the results apply and usually qualifies every empirical claim at once. A methodological claim is a specific capability that specific results depend on for their meaning. "All results come from a single-cell model" is scope; "the preregistered plan fixed the decoding pipeline, the ROIs and the tests in advance" is methodological, because the decoding results count as confirmatory only if it did.

- `scope`: `l5-model-single-cell-scope` (scope, Headley): “All results derive from a single-cell compartmental model of one layer 5 pyramidal neuron; no network dynamics, recurrent excitation, or population-level inhibitory effects are simulated, and all firing rate effects are for a single isolated neuron receiving naturalistic presynaptic drive.”
- `methodological`: `preregistered-design-validates-mvpa` (methodological, Kammer): “The preregistered analysis plan (osf.io/rxacd) specifies the MVPA decoding pipeline, ROI definitions, and statistical tests in advance, reducing the risk of analytic flexibility inflating the decoding accuracy results.”

**`literature-context` versus `interpretation`.** Literature-context restates what a cited paper found; it is someone else's result, inherited. Interpretation is this paper's reading of its own results, even when that reading leans on the literature. The inherited premise and the reading that uses it are two claims.

- `literature-context`: `interprets-pv-gamma-sst-beta-associations` (literature-context, Headley): “Empirical work across the cortical-interneuron literature has established two correlated associations: parvalbumin-positive (PV+) fast-spiking interneurons target perisomatic compartments and are implicated in the generation and entrainment of cortical gamma rhythms (40–80 Hz), while somatostatin-positive (SST+) interneurons target distal dendritic compartments and are preferentially associated with beta rhythms (12–35 Hz). These are literature claims about anatomical targeting patterns and their correlation with specific oscillatory bands, not results of the present paper.”
- `interpretation`: `pv-gamma-sst-beta-correspondence` (interpretation, Headley): “The model provides mechanistic grounding for the empirical association of parvalbumin-positive interneurons with gamma rhythms and somatostatin-positive interneurons with beta rhythms: PV+ neurons target perisomatic locations where gamma is optimal for AP threshold modulation, while SST+ neurons target distal dendrites where beta is optimal for dendritic spike entrainment.”

## Relations

A relation is a proposition about logical structure between two claims, not a citation. Each is directed; the direction column says which end is which. `derived-from` is written mechanically as the reciprocal of `entails` and should not be emitted.

| Relation | Asserts | Direction |
|:--|:--|:--|
| `confirms` | an empirical result confirms the prediction it tested — the positive outcome of a test | from the result to the prediction it confirms — the positive outcome of a test, aimed at a prediction only |
| `contradicts` | the source and target cannot both hold | from either claim to the other; they cannot both hold |
| `derived-from` | a prediction derived from its hypothesis (inverse of entails) | from the prediction back to its hypothesis; written mechanically as the reciprocal of entails |
| `dissociates-with` | the source and target jointly establish a dissociation — two claims whose difference across a condition, region, population or measure is itself the finding, neither bearing on the other's truth (symmetric) | symmetric: between the two empirical claims that together establish the contrast |
| `enables-method` | a result makes a downstream method possible | from the methodological claim to the result whose interpretability it warrants |
| `entails` | a hypothesis entails its prediction — the deductive step | from the hypothesis to the prediction it deductively implies |
| `extends` | the source extends the target beyond its original conditions | from the later or broader result to the claim it extends |
| `in-tension-with` | two claims the paper asserts, both of which stand, that pull a shared implication in opposite directions or cannot be jointly explained without a further claim (symmetric) | symmetric: between two claims the paper asserts whose implications pull against each other |
| `interprets` | one claim interprets another | from the interpretation to the empirical claim it reframes |
| `opposes` | the source stands against the target | from the claim that stands against to the one it stands against |
| `part-of` | a component of another claim — one comparison, condition, measure or study of a proposition the target states whole; the target is weakened but not falsified by the source alone | from the component to the claim it is a part of: the source states one comparison, condition, measure or study of what the target states as a whole |
| `predicts` | the source predicts the target, typically model to experiment | from the model or hypothesis to the observation it predicts |
| `qualifies` | a claim narrows another's applicability | from the qualifying result to the claim whose applicability it narrows |
| `refutes` | an empirical result refutes the prediction it tested — the negative outcome of a test | from the result to the prediction it came out against — the negative outcome of a test, aimed at a prediction only |
| `replicates` | an independent finding of the same result as the target | from the independent finding to the claim it reproduces |
| `requires` | a claim depends on another holding | from the dependent claim to its prerequisite: the source would be invalid if the target were false |
| `rules-out` | the source's evidence eliminates the target as an explanation | from the control or evidence to the alternative explanation it eliminates — a claim the paper entertains or rejects, never one it asserts |
| `scopes` | a scope constraint governs another claim's validity | from the scope claim to the claims it bounds, or to `*` for every empirical claim in the paper |
| `supports` | the source provides evidence for the target | from the evidence to the claim it is evidence for |
| `tests` | an empirical result tests the target prediction, closing the loop | from the empirical result to the prediction it tests |
| `validates` | a control whose specific result strengthens the target's warrant | from the control to the claim whose warrant it strengthens |

### One example of each

- `confirms`: `5ht-axis-orthogonal-to-choice-axis` → `prediction-5ht-axis-orthogonal-to-choice` (Meijer). Source: “In the manifold analysis built from a pooled super-session of trial-averaged peri-event PETHs across all neurons, sessions, and mice, the 5-HT effect axis (defined as the displacement vector between mean stimulated and mean unstimulated trajectories in 3-D PCA space, averaged across choice) is significantly more orthogonal to the choice axis (the displacement between mean left and mean right trajectories, averaged across stimulation) than expected from a block-aware shuffle null distribution. The orthogonality measure (1 - |dot product| of unit-normalized direction vectors) increases over the pre-choice timecourse and is significant in the time window approaching the choice moment.” Target: “If brain-wide 5-HT modulation is structurally separable from the choice computation, then in a low-dimensional projection of the population response (here, the first three principal components), the vector that separates 5-HT-stimulated from unstimulated trials should be near-orthogonal to the vector that separates left-choice from right-choice trials, with orthogonality measured as 1 - |dot product| of unit-normalized direction vectors.”
- `contradicts`: no use in the corpus yet.
- `derived-from`: `prediction-distal-dendritic-spike-mechanism` → `hypothesis-distinct-compartmental-roles` (Headley). Source: “If perisomatic and distal dendritic inhibition serve dissociated computational roles, then doubling distal dendritic inhibition should reduce somatic firing primarily by suppressing apical Ca²⁺ and NMDA dendritic spikes, with little change in the somatic AP voltage threshold.” Target: “Perisomatic and distal dendritic inhibition serve distinct computational roles in layer 5 pyramidal neurons: perisomatic inhibition principally regulates somatic action potential generation (gain and threshold of axonal output), while distal dendritic inhibition principally regulates dendritic spike incidence and the temporal coupling of dendritic spikes to somatic APs.”
- `dissociates-with`: `distal-inhib-drops-firing-02hz` → `perisomatic-inhib-drops-firing-07hz` (Headley). Source: “Doubling the strength of distal dendritic inhibition reduces somatic firing rate from a baseline of approximately 5.5 Hz to approximately 0.2 Hz, primarily by suppressing the occurrence of dendritic Ca²⁺ and NMDA spikes rather than by directly raising AP threshold.” Target: “Doubling perisomatic inhibition reduces somatic firing from approximately 5.5 Hz to approximately 0.7 Hz by elevating action potential voltage threshold, while dendritic spike rates are relatively preserved.”
- `enables-method`: `preregistered-design-validates-mvpa` → `foveal-v1-decodes-peripheral-saccade-target` (Kammer). Source: “The preregistered analysis plan (osf.io/rxacd) specifies the MVPA decoding pipeline, ROI definitions, and statistical tests in advance, reducing the risk of analytic flexibility inflating the decoding accuracy results.” Target: “Saccade target identity (4 stimuli varying in shape and semantic category) can be decoded from foveal V1 BOLD signal above chance in the experimental condition where targets disappear before fixation: 57.43% accuracy (t(27)=8.81, p<0.001), significantly above 50% chance.”
- `entails`: `hypothesis-distinct-compartmental-roles` → `prediction-distal-dendritic-spike-mechanism` (Headley). Source: “Perisomatic and distal dendritic inhibition serve distinct computational roles in layer 5 pyramidal neurons: perisomatic inhibition principally regulates somatic action potential generation (gain and threshold of axonal output), while distal dendritic inhibition principally regulates dendritic spike incidence and the temporal coupling of dendritic spikes to somatic APs.” Target: “If perisomatic and distal dendritic inhibition serve dissociated computational roles, then doubling distal dendritic inhibition should reduce somatic firing primarily by suppressing apical Ca²⁺ and NMDA dendritic spikes, with little change in the somatic AP voltage threshold.”
- `extends`: `v2-v3-generalize-shape-not-category` → `decoding-shape-sensitive-not-semantic` (Kammer). Source: “The shape-sensitive, category-insensitive pattern of foveal feedback decoding found in V1 generalizes to foveal regions of V2 and V3, as shown in supplementary figure analyses, suggesting that low-level feature specificity is a property of early visual cortex as a whole rather than V1 specifically.” Target: “Foveal V1 feedback contains low-to-mid-level visual information (shape-sensitive) but not higher-level semantic category information (animal vs instrument not decodable), indicating that feedback signals represent low-level feature properties of the saccade target.”
- `in-tension-with`: `dot-products-between-individual-neural` → `individual-grbs-dot-product-values-not` (Gadeke). Source: “Dot products between individual neural guilt responses and the Yu et al. (2020) guilt-related brain signature (GRBS) were overall positive (mean = 5.22, median = 6.97, sign test p = 0.017, Cliff's Delta = 0.4), providing convergent validity with a previously published neural guilt signature.” Target: “Individual GRBS dot-product values did not correlate with the behavioural guilt responses (Spearman's Rho = –0.058, p = 0.725), indicating the neural signature does not track individual differences in behavioural guilt sensitivity.”
- `interprets`: `pv-gamma-sst-beta-correspondence` → `beta-optimal-distal-dendritic-entrainment` (Headley). Source: “The model provides mechanistic grounding for the empirical association of parvalbumin-positive interneurons with gamma rhythms and somatostatin-positive interneurons with beta rhythms: PV+ neurons target perisomatic locations where gamma is optimal for AP threshold modulation, while SST+ neurons target distal dendrites where beta is optimal for dendritic spike entrainment.” Target: “In a frequency sweep from 0.5 to 80 Hz, beta frequencies near 20 Hz produce the strongest phase-dependent entrainment of dendritic Ca²⁺ and NMDA spike onsets by distal rhythmic inhibition, as measured by Pairwise Phase Consistency.”
- `opposes`: no use in the corpus yet.
- `part-of`: `difference-response-between-low-high` → `insula-rois-responded-more-low` (Gadeke). Source: “The difference in response between low and high lottery outcomes was greater in the Social than the Partner condition in left insula (0.44***), right insula (0.19***), and right middle temporal cortex (0.67***).” Target: “The insula ROIs responded more to low lottery outcomes for the partner in the Social than the Partner condition — even after subtracting responses to high outcomes — mirroring the behavioural guilt effect.”
- `predicts`: `hypothesis-state-switch-by-5ht` → `5ht-stim-dilates-pupil` (Meijer). Source: “A phasic burst of serotonin release from the dorsal raphe nucleus drives a rapid switch in internal arousal/behavioral state in the awake quiescent animal — from an "offline" state characterized by low arousal, hippocampal sharp-wave ripples, and minimal exploratory movement, to an "online" state characterized by pupil dilation, ripple suppression, and active exploration (whisking, sniffing).” Target: “A 1-second optogenetic activation of dorsal-raphe serotonergic neurons in SERT-Cre mice produces a significant increase in pupil size relative to baseline, with the divergence from wild-type controls becoming significant approximately 2 seconds after stimulation offset.”
- `qualifies`: no use in the corpus yet.
- `refutes`: `5ht-stim-leaves-decision-behavior-intact` → `prediction-5ht-shifts-psychometric` (Meijer). Source: “In the IBL steering-wheel perceptual decision-making task, optogenetic activation of dorsal-raphe serotonergic neurons starting at visual stimulus onset and lasting until choice (or 1 s, whichever is shorter) produces no detectable change in any of: psychometric slope (visual acuity), bias (prior influence), lapse rate (motivation), percent-correct, median reaction time, the speed of choice-bias updating after a prior-probability block switch, or the per-predictor weights of a probabilistic choice model fitting visual evidence, prior, and past choice with separate stimulated / unstimulated coefficients.” Target: “If phasic 5-HT release alters how the animal weighs sensory evidence against prior expectation, or alters motivation, attention, or visual sensitivity, then 5-HT stimulation should produce detectable changes in psychometric slope, lapse rate, bias, reaction time, or any of the predictor weights of a probabilistic choice model fitting visual evidence, prior, and past choice with separate stimulated/unstimulated coefficients.”
- `replicates`: no use in the corpus yet.
- `requires`: `distal-inhib-drops-firing-02hz` → `l5-model-single-cell-scope` (Headley). Source: “Doubling the strength of distal dendritic inhibition reduces somatic firing rate from a baseline of approximately 5.5 Hz to approximately 0.2 Hz, primarily by suppressing the occurrence of dendritic Ca²⁺ and NMDA spikes rather than by directly raising AP threshold.” Target: “All results derive from a single-cell compartmental model of one layer 5 pyramidal neuron; no network dynamics, recurrent excitation, or population-level inhibitory effects are simulated, and all firing rate effects are for a single isolated neuron receiving naturalistic presynaptic drive.”
- `rules-out`: `participant-happiness-lower-when-participant` → `alt-agency-aversion-not-guilt` (Gadeke). Source: “Participant happiness was lower when the participant was the decision-maker (Social + Solo vs. Partner), independent of outcome (Study 1: t(3600) = –3.92, p < 0.0001, β = –0.14; Study 2: t(2870) = –6.07, p < 0.0001, β = –0.24).” Target: “The happiness cost observed in the Social condition is general agency aversion — the unpleasantness of being the decision-maker as such — and is not contingent on responsibility for a negative outcome befalling the partner.”
- `scopes`: `l5-model-single-cell-scope` → `*` (Headley). “All results derive from a single-cell compartmental model of one layer 5 pyramidal neuron; no network dynamics, recurrent excitation, or population-level inhibitory effects are simulated, and all firing rate effects are for a single isolated neuron receiving naturalistic presynaptic drive.” — bounds every empirical claim in the paper.
- `supports`: `distal-inhib-drops-firing-02hz` → `hypothesis-distinct-compartmental-roles` (Headley). Source: “Doubling the strength of distal dendritic inhibition reduces somatic firing rate from a baseline of approximately 5.5 Hz to approximately 0.2 Hz, primarily by suppressing the occurrence of dendritic Ca²⁺ and NMDA spikes rather than by directly raising AP threshold.” Target: “Perisomatic and distal dendritic inhibition serve distinct computational roles in layer 5 pyramidal neurons: perisomatic inhibition principally regulates somatic action potential generation (gain and threshold of axonal output), while distal dendritic inhibition principally regulates dendritic spike incidence and the temporal coupling of dendritic spikes to somatic APs.”
- `tests`: `distal-inhib-drops-firing-02hz` → `prediction-distal-dendritic-spike-mechanism` (Headley). Source: “Doubling the strength of distal dendritic inhibition reduces somatic firing rate from a baseline of approximately 5.5 Hz to approximately 0.2 Hz, primarily by suppressing the occurrence of dendritic Ca²⁺ and NMDA spikes rather than by directly raising AP threshold.” Target: “If perisomatic and distal dendritic inhibition serve dissociated computational roles, then doubling distal dendritic inhibition should reduce somatic firing primarily by suppressing apical Ca²⁺ and NMDA dendritic spikes, with little change in the somatic AP voltage threshold.”
- `validates`: `wt-controls-rule-out-light-artifact` → `5ht-stim-dilates-pupil` (Meijer). Source: “Wild-type control mice (n=6), which received the same surgical procedure, optical fiber implantation, and optogenetic stimulation protocol but lack Cre-dependent ChR2 expression in DRN serotonergic neurons, show baseline-level DRN/control fluorescence ratio (Fig. 1c) and chance-level (~5%) significantly modulated neurons (Fig. 2f) — ruling out light delivery, optical-fiber heating, viral injection, surgical artifact, or other non-specific effects of the procedure as drivers of the observed pupil, ripple, and neural modulation effects.” Target: “A 1-second optogenetic activation of dorsal-raphe serotonergic neurons in SERT-Cre mice produces a significant increase in pupil size relative to baseline, with the divergence from wild-type controls becoming significant approximately 2 seconds after stimulation offset.”

### Relations that are confused for each other

**`requires` versus `supports`.** `requires` is a dependency: if the target were false the source would be invalid. `supports` is evidence: the source makes the target more credible and would survive its falsity. One empirical claim commonly carries both — it *requires* the scope claim that bounds the model it was computed in, and *supports* the hypothesis it was run to test.

**`entails` versus `tests`.** Both connect a hypothesis's arc, in opposite directions and from different roles. `entails` runs *down* from the hypothesis to a prediction and is deductive: the prediction follows if the hypothesis holds. `tests` runs *up* from an empirical result to the prediction it checks. A result never `entails` anything; a hypothesis never `tests`.

**`rules-out` versus `refutes`.** `rules-out` eliminates an alternative explanation — a claim the paper raises in order to reject, which has a node of its own with stance `entertains` or `rejects`. `refutes` is the negative outcome of a test, aimed at one of the paper's own predictions that the evidence came out against; a paper refuting its own prediction is the hypothetico-deductive loop closing. When a result bears against a hypothesis, do not aim `refutes` at the hypothesis — write the prediction the hypothesis entails and refute that. Never aim `rules-out` at a claim the same paper asserts.

**`confirms` versus `supports`.** `confirms` is the outcome of a stated prediction: an empirical result came out the way the prediction said it would, and the edge runs from the result to that prediction — its negative counterpart is `refutes`. `supports` is evidence for a hypothesis or higher-order claim, making it more credible without being the settling of a prediction. A result that tests a prediction carries `tests` and then `confirms` or `refutes` it; the same result may `supports` the hypothesis that prediction was derived from. Aim `confirms` and `refutes` at predictions only — a result that bears on a hypothesis directly takes `supports`.

**`dissociates-with` versus `contradicts`.** `dissociates-with` joins two results that are both true and *differ*: the contrast between them is the finding, and neither undermines the other. `contradicts` says two claims cannot both hold. Two conditions producing different effects is a dissociation, not a contradiction. A dissociation is also not a tension: a contrast is marked by “whereas”, “in contrast”, “selectively”, and the two results simply differ; a tension (`in-tension-with`) is marked by “although”, “however”, “despite”, “no correlation with”, “at the cost of”, and the two results pull a shared implication in opposite directions.

**`in-tension-with` versus `contradicts`.** `contradicts` says two claims cannot both hold. A tension says both do: the paper asserts each, and they stand together while pulling a shared implication in opposite directions.

**`in-tension-with` versus `qualifies`.** `qualifies` is directional — one result narrows the applicability of another. A tension has no narrower side: neither claim bounds the other, they simply pull against each other. The wengert case (a general impairment, and a layer that mostly escapes it) is arguably a qualification, and the reading is left to the reader rather than fixed by rule.

**`in-tension-with` versus `rules-out`.** `rules-out` eliminates an alternative the paper raised in order to reject — a claim with stance `entertains` or `rejects`. A tension is between two claims the paper *asserts*, both of which stand; nothing is being eliminated.

**`scopes` versus `requires`.** A scope claim bounds what a result can mean and is written from the scope claim *to* the results it bounds (or to `*`). `requires` is written from the result *to* what it depends on. The same pair of claims can carry both, in opposite directions: the result requires the scope; the scope scopes the result.

**`validates` versus `supports`.** `validates` is a control's edge: a check whose specific outcome (a null where a confound would have produced an effect, a sign-flip, a manipulation check) strengthens the warrant for a target. `supports` is ordinary evidence for a proposition. A control `validates`; a main result `supports`.

**`part-of` versus `supports`.** `part-of` is composition: the source is one comparison, condition, measure or study *of* the proposition the target states whole, and dropping it weakens the target without falsifying it. `supports` is evidence: an independent finding that makes the target more credible and would survive being removed. The insula-ROI result stated beside the voxel-wise result is a *part of* the claim that the insula tracks the guilt effect; a distinct finding that happens to bear on that claim merely *supports* it. A filter on `supports` cannot tell a component from an independent finding, which is why composition needs its own relation.

## Confidence

A reader marks its own claim:

- `high` — asserted directly in the text the reader was given, with a quotable sentence
- `tentative` — read between the lines, summarised across sentences, or ambiguous in the source; say why in `notes`

After reconciliation a claim's confidence is a fact about agreement between readers, not about the world:

- `high` — more than one reader surfaced the same proposition and they agree on its panel and its direction
- `contested` — more than one reader surfaced it and they disagree — about the panel, the direction, or whether it is a hypothesis, a prediction or a result. Record what each said in `notes`
- `single-source` — one reader surfaced it. Expected for panel-level numerics (caption reader only), scope and methodological claims (structure reader only) and synthesis (results reader only); not a mark against the claim

`confidence` on an assertion is this paper's own confidence in the proposition — analyst judgement, which the writer leaves absent rather than inventing. It is not `readers`, the agreement between extraction readers recorded beside it.

## Same, part, or different

When two candidates might be one claim, ask what would verify each. **Same**: the same computation on the same data, however the two are worded and whichever study each cites — merge, keep the more precise wording and carry both readers' evidence. **Part**: one states one comparison, condition, measure or study of what the other states as a whole — keep both and set the part's `part_of` to the whole. **Different**: a different computation, direction, region or condition — keep both, unrelated. On real pairs:

- **same.** “In both studies, participants felt worse after low lottery outcomes for the partner when those outcomes followed their own choice rather than the partner's, which the authors interpret as interpersonal guilt.” beside “When the partner received the low lottery outcome, participant happiness was lower when the participant rather than the partner had chosen the lottery — a significant partner-outcome × decision-maker interaction (Study 1: t(1180) = 3.52, p = 0.0004, β = 0.37; Study 2: t(937) = 2.85, p = 0.0045, β = 0.33) — operationalizing interpersonal guilt.” — The same partner-outcome × decision-maker interaction on the same happiness data, cited once as a cross-study synthesis and once as the result that computes it: merge, keep the wording with the coefficients.

- **part.** “One cluster in the left STS responded more to partner reward prediction errors resulting from participant rather than partner choices (pFWE = 0.022, T = 4.70, d = 0.53, 100 voxels, peak MNI [−52 –32 0]).” beside “The left superior temporal sulcus cluster responded to model-based regressors coding participant reward prediction resulting from participant and partner choices across both sessions of the experiment.” — The first states one directional contrast — participant-caused above partner-caused — of the broader responsiveness the second states as a whole: keep both, the first `part_of` the second.

- **different.** “During receipt of lottery versus safe outcomes (across all conditions), clusters were more active in the bilateral anterior insula, dmPFC, right STS, bilateral ventral striatum, right dorsolateral prefrontal cortex, and bilateral inferior parietal lobe.” beside “The bilateral ventral striatum was more active when participants chose the risky rather than the safe option (Cohen's d = 0.72 left, 0.85 right), irrespective of Social or Solo condition, replicating previous findings.” — Both light up the ventral striatum, but by different computations — one the lottery-versus-safe outcome-receipt contrast, the other the risky-versus-safe choice contrast: keep both, no relation between them here.

What they produced, per paper

A run belongs to the paper it read. Each reader's output — the claims it found and the quote it found them in — is on that paper's cell for this layer, with the prompt hash and the model that answered.

gadeke-2026-guilt-insula

1 of 10 papers. The other nine were extracted before runs were recorded, so their reader-level history does not exist and would have to be regenerated.

What this record does not cover

The published claim tree for this paper did not come from this run. Its claim files were first committed five months before this trace was recorded, and no committed claim has even a 0.80 textual match anywhere in the draft above — so this is not those claims at an earlier stage, it is a separate extraction. How the committed tree was produced is not recorded anywhere and cannot be recovered.

1 of 10 papers has a trace at all. The rest were extracted before runs were being saved. Rather than show nine empty templates, this page covers the one paper that has been run and states the gap.