dsg-artifacts / FINDINGS.md
GOVINDFROM's picture
reading findings
abefce4 verified
|
Raw History Blame Contribute Delete
8.81 kB

Findings

Every number below is read from artifacts/report/*/results.md, which is generated from results.json by python -m dsg report. Comparisons are paired bootstraps over books (10,000 resamples, 95% intervals). No LLM judge appears in any metric; identity is scored against PDNC's human alias annotation and LitBank's gold coreference, and mention linking against PDNC's annotated referring expressions.

Status. PDNC at 3B and 1.5B and LitBank at 7B are complete. The PDNC 7B run is still extracting; this document is updated when it lands.

Every trend against discourse position is reported with a shuffled-window control, because state growth alone produces such trends (Β§2, F4).


1. The headline

Carrying narrative state forward is necessary and self-poisoning. It buys a large amount of accuracy and, at the same time, fills the state with contradictions that a stateless reader cannot have. Fact-level revision is the antidote, and it costs nothing in accuracy.

On 28 PDNC novels with a 3B extractor, moving one rung at a time:

rung identity CoNLL F1 mention linking self-contradictory slots rollbacks
persistent state (append-only βˆ’ window-only) +0.088 +0.310 +0.323 (worse) 0
+ identity merging βˆ’0.004 (ns) βˆ’0.016 (worse) +0.010 (ns) +2.36
+ fact revision 0.000 (ns) 0.000 (ns) βˆ’0.433 +33.1
+ deferred commitment βˆ’0.002 (ns) +0.014 0.000 βˆ’2.57
cost of causality (retrospective βˆ’ dsg-full) +0.005 (ns) +0.031 0.000 ns

Bold entries have a 95% interval excluding zero. Read down the fourth column: persistent state introduces contradiction into a third of the state's slots, and fact revision removes it β€” on every one of the 28 books (0–28 win split), with no measurable cost to either accuracy metric.

In absolute terms, the share of single-valued slots holding two or more live conflicting values:

corpus / model window-only append-only dsg-full
PDNC, Qwen2.5-3B 0.101 0.423 0.000
PDNC, Qwen2.5-1.5B 0.004 0.107 0.000
LitBank, Qwen2.5-7B 0.091 0.181 0.000

2. The pre-registered falsifiers

Stated in docs/RESEARCH_PLAN.md Β§5 before any run.

F1 β€” "dsg-full β‰ˆ append-only on gold identity β‡’ premature commitment is not the bottleneck." This falsifier FIRED. On identity CoNLL F1 the difference is indistinguishable from zero at every model size (3B: βˆ’0.005, CI [βˆ’0.011, +0.001]; 1.5B: βˆ’0.003; LitBank: βˆ’0.003, CI spans zero), and mention linking is likewise flat. The revision calculus does not make the state more accurate. What it does is make the state consistent β€” a different claim, and the one the evidence supports. The paper's framing follows the evidence rather than the other way round.

F2 β€” "dsg-full β‰ͺ retrospective with no closing of the gap β‡’ causality, not commitment policy, is the binding constraint." Partially fired. Reading forward genuinely costs accuracy, and we can now price it: the non-causal oracle is +0.031 better on mention linking at 3B (CI [+0.018, +0.045], 23–4) and +0.050 at 1.5B (CI [+0.038, +0.064], 28–0). Both intervals exclude zero. Causality is a real constraint, not an artefact β€” but it is a few points, not a collapse, and it does not touch consistency, where the causal system already matches the oracle exactly (both 0.000).

F3 β€” "invariant violations near zero for append-only β‡’ the breakage premise is wrong." Did not fire; the premise holds decisively. 42.3% of append-only's slots are self-contradictory at 3B on full novels.

F4 β€” "revision events distribute uniformly over discourse position β‡’ the instrument has no signal." Fired for revision; did not fire for elaboration.

In narrative order both operations trend strongly with position: elaboration falls (Spearman ρ = βˆ’0.59, p = 0.007) and revision rises steeply, from 4.6% of operations in the first twentieth of a book to 29.8% in the last (ρ = +0.82, p = 9Γ—10⁻⁢). Read alone that is a tidy narratological story β€” exposition first, reversal later.

It is mostly an artefact, and only a control shows it. As a book proceeds the state holds more assertions, so any conflict-driven operation becomes mechanically more likely. Permuting the reading order destroys narrative order while preserving state growth exactly:

ordering elaboration ρ revision ρ
narrative βˆ’0.59 (p = 0.007) +0.82 (p = 9Γ—10⁻⁢)
shuffled +0.08 (p = 0.75) +0.69 (p = 8Γ—10⁻⁴)

The revision trend largely survives shuffling, so most of it is state growth rather than a property of the text; the honest residual is the gap between ρ = 0.82 and ρ = 0.69, which this design cannot cleanly separate. The elaboration trend disappears under shuffling. That one is real: refinement of underspecified structure is genuinely front-loaded in narrative order, and not explained by how much state has accumulated.

So the instrument has signal, but less than the uncontrolled numbers suggest, and in the opposite operation from the one we expected to carry it.

3. What each mechanism actually buys

  • Persistent state is where nearly all the accuracy comes from: +0.310 mention linking over reading each window in isolation. It is also where all the inconsistency comes from.
  • Identity merging alone makes things worse. On its own it costs mention accuracy (βˆ’0.016, CI excludes zero) and buys 2.4 rollbacks per book. Merging two characters unifies their assertion slots, and without revision the merged slot simply holds both rival values. The two mechanisms are complementary: merging is only safe once the state can revise.
  • Fact revision is the load-bearing contribution: βˆ’0.433 self-contradictory slots, 0–28, at no accuracy cost.
  • Deferred commitment converts rollbacks into monotone updates. The size of the effect varies with how fast nodes become committed: βˆ’5.50 rollbacks per book at 1.5B (0–28, a 99% reduction), βˆ’0.60 on LitBank (92%), but only βˆ’2.57 of 35.5 at 3B (7%). It also slightly helps mention linking (+0.014 at 3B, +0.006 at 1.5B, both intervals excluding zero), which is the opposite of the cost we expected to pay for holding identity open.

4. Three defects found by looking at the output

Recorded because they materially changed the numbers, and because the third was found only by checking an assumption that seemed safe.

  1. Titles were stripped before comparing surfaces, so Mrs. Bennet and Miss Bennet scored 0.9 against each other. In this corpus a title is the primary distinguisher between people sharing a surname.
  2. Nothing prevented merging two separately named characters. A single bad link fuses two people; the fused node then matches both name sets and attracts more. On Pride and Prejudice this collapsed 74 gold characters into 3 nodes.
  3. The link bind path bypassed the merge guard and the policy check, so even append-only was receiving identity resolution it should not have had.

Fixing these moved conflation on Pride and Prejudice from 1.33 to 0.05.

The obvious fix for (2) was to require corroboration β€” accept a name-to-name link only if it is proposed more than once. Measured before adopting, and it does not work: across 205 links on one novel, Mr. Darcy -> Elizabeth was proposed nine times, exactly as often as the correct Mr. Bingley -> Bingley. Repetition does not separate good links from bad, so the constraint had to be structural rather than statistical.

5. Limitations

  • No identity-accuracy gain. Stated plainly: on the metrics that measure accuracy rather than consistency, the calculus is neutral.
  • One corpus of 28 English novels, plus 100 LitBank excerpts. One model family.
  • Alias-level identity gold. PDNC's alias sets are surface strings, not full mention-level coreference; the mention-linking plane partly compensates but covers only referring expressions inside quotations.
  • Extraction is the binding constraint on several metrics. A 1.5B model fails to produce parseable output on 21% of windows, so the scale comparison is partly a comparison of yield; the parse rate is reported alongside it.
  • Nicknames are unrecoverable under the merge constraint: Lizzy and Elizabeth share no surface material, so the guard that blocks Jane -> Elizabeth also blocks the correct Lizzy -> Elizabeth.
  • SUPERSEDE versus REVISE is decided by a lexicon, and its accuracy has not itself been evaluated against annotation.