|
Download FINDINGS.md from GOVINDFROM/dsg-artifacts: direct link, hf CLI and curl.
- Browser
- Download file 8.81 kB
-
https://huggingface.co/GOVINDFROM/dsg-artifacts/resolve/main/FINDINGS.md
- Command line
-
hf download hf://GOVINDFROM/dsg-artifacts/FINDINGS.md
-
curl -L -o FINDINGS.md https://huggingface.co/GOVINDFROM/dsg-artifacts/resolve/main/FINDINGS.md
8.81 kB
| # Findings | |
| Every number below is read from `artifacts/report/*/results.md`, which is | |
| generated from `results.json` by `python -m dsg report`. Comparisons are paired | |
| bootstraps over books (10,000 resamples, 95% intervals). No LLM judge appears in | |
| any metric; identity is scored against PDNC's human alias annotation and | |
| LitBank's gold coreference, and mention linking against PDNC's annotated | |
| referring expressions. | |
| **Status.** PDNC at 3B and 1.5B and LitBank at 7B are complete. The PDNC 7B run | |
| is still extracting; this document is updated when it lands. | |
| Every trend against discourse position is reported with a shuffled-window | |
| control, because state growth alone produces such trends (Β§2, F4). | |
| --- | |
| ## 1. The headline | |
| Carrying narrative state forward is **necessary and self-poisoning**. It buys a | |
| large amount of accuracy and, at the same time, fills the state with | |
| contradictions that a stateless reader cannot have. Fact-level revision is the | |
| antidote, and it costs nothing in accuracy. | |
| On 28 PDNC novels with a 3B extractor, moving one rung at a time: | |
| | rung | identity CoNLL F1 | mention linking | self-contradictory slots | rollbacks | | |
| | --- | --- | --- | --- | --- | | |
| | persistent state (`append-only` β `window-only`) | **+0.088** | **+0.310** | **+0.323** *(worse)* | 0 | | |
| | + identity merging | β0.004 (ns) | **β0.016** *(worse)* | +0.010 (ns) | **+2.36** | | |
| | + fact revision | 0.000 (ns) | 0.000 (ns) | **β0.433** | +33.1 | | |
| | + deferred commitment | β0.002 (ns) | **+0.014** | 0.000 | **β2.57** | | |
| | cost of causality (`retrospective` β `dsg-full`) | +0.005 (ns) | **+0.031** | 0.000 | ns | | |
| Bold entries have a 95% interval excluding zero. Read down the fourth column: | |
| persistent state introduces contradiction into a third of the state's slots, and | |
| fact revision removes it β **on every one of the 28 books** (0β28 win split), | |
| with no measurable cost to either accuracy metric. | |
| In absolute terms, the share of single-valued slots holding two or more live | |
| conflicting values: | |
| | corpus / model | `window-only` | `append-only` | `dsg-full` | | |
| | --- | --- | --- | --- | | |
| | PDNC, Qwen2.5-3B | 0.101 | 0.423 | **0.000** | | |
| | PDNC, Qwen2.5-1.5B | 0.004 | 0.107 | **0.000** | | |
| | LitBank, Qwen2.5-7B | 0.091 | 0.181 | **0.000** | | |
| ## 2. The pre-registered falsifiers | |
| Stated in `docs/RESEARCH_PLAN.md` Β§5 before any run. | |
| **F1 β "`dsg-full` β `append-only` on gold identity β premature commitment is | |
| not the bottleneck." This falsifier FIRED.** | |
| On identity CoNLL F1 the difference is indistinguishable from zero at every | |
| model size (3B: β0.005, CI [β0.011, +0.001]; 1.5B: β0.003; LitBank: β0.003, CI | |
| spans zero), and mention linking is likewise flat. **The revision calculus does | |
| not make the state more accurate.** What it does is make the state | |
| *consistent* β a different claim, and the one the evidence supports. The paper's | |
| framing follows the evidence rather than the other way round. | |
| **F2 β "`dsg-full` βͺ `retrospective` with no closing of the gap β causality, not | |
| commitment policy, is the binding constraint." Partially fired.** | |
| Reading forward genuinely costs accuracy, and we can now price it: the | |
| non-causal oracle is **+0.031** better on mention linking at 3B (CI [+0.018, | |
| +0.045], 23β4) and **+0.050** at 1.5B (CI [+0.038, +0.064], 28β0). Both | |
| intervals exclude zero. Causality is a real constraint, not an artefact β but | |
| it is a few points, not a collapse, and it does not touch consistency, where | |
| the causal system already matches the oracle exactly (both 0.000). | |
| **F3 β "invariant violations near zero for `append-only` β the breakage premise | |
| is wrong." Did not fire; the premise holds decisively.** | |
| 42.3% of `append-only`'s slots are self-contradictory at 3B on full novels. | |
| **F4 β "revision events distribute uniformly over discourse position β the | |
| instrument has no signal." Fired for revision; did not fire for elaboration.** | |
| In narrative order both operations trend strongly with position: elaboration | |
| falls (Spearman Ο = β0.59, p = 0.007) and revision rises steeply, from 4.6% of | |
| operations in the first twentieth of a book to 29.8% in the last | |
| (Ο = +0.82, p = 9Γ10β»βΆ). Read alone that is a tidy narratological story β | |
| exposition first, reversal later. | |
| It is mostly an artefact, and only a control shows it. As a book proceeds the | |
| state holds more assertions, so *any* conflict-driven operation becomes | |
| mechanically more likely. Permuting the reading order destroys narrative order | |
| while preserving state growth exactly: | |
| | ordering | elaboration Ο | revision Ο | | |
| | --- | --- | --- | | |
| | narrative | **β0.59** (p = 0.007) | **+0.82** (p = 9Γ10β»βΆ) | | |
| | shuffled | +0.08 (p = 0.75) | **+0.69** (p = 8Γ10β»β΄) | | |
| The revision trend largely **survives** shuffling, so most of it is state | |
| growth rather than a property of the text; the honest residual is the gap | |
| between Ο = 0.82 and Ο = 0.69, which this design cannot cleanly separate. The | |
| elaboration trend **disappears** under shuffling. That one is real: refinement | |
| of underspecified structure is genuinely front-loaded in narrative order, and | |
| not explained by how much state has accumulated. | |
| So the instrument has signal, but less than the uncontrolled numbers suggest, | |
| and in the opposite operation from the one we expected to carry it. | |
| ## 3. What each mechanism actually buys | |
| - **Persistent state** is where nearly all the accuracy comes from: +0.310 | |
| mention linking over reading each window in isolation. It is also where all | |
| the inconsistency comes from. | |
| - **Identity merging alone makes things worse.** On its own it costs mention | |
| accuracy (β0.016, CI excludes zero) and buys 2.4 rollbacks per book. Merging | |
| two characters unifies their assertion slots, and without revision the merged | |
| slot simply holds both rival values. The two mechanisms are complementary: | |
| merging is only safe once the state can revise. | |
| - **Fact revision** is the load-bearing contribution: β0.433 self-contradictory | |
| slots, 0β28, at no accuracy cost. | |
| - **Deferred commitment** converts rollbacks into monotone updates. The size of | |
| the effect varies with how fast nodes become committed: β5.50 rollbacks per | |
| book at 1.5B (0β28, a 99% reduction), β0.60 on LitBank (92%), but only β2.57 | |
| of 35.5 at 3B (7%). It also slightly *helps* mention linking (+0.014 at 3B, | |
| +0.006 at 1.5B, both intervals excluding zero), which is the opposite of the | |
| cost we expected to pay for holding identity open. | |
| ## 4. Three defects found by looking at the output | |
| Recorded because they materially changed the numbers, and because the third was | |
| found only by checking an assumption that seemed safe. | |
| 1. **Titles were stripped before comparing surfaces**, so `Mrs. Bennet` and | |
| `Miss Bennet` scored 0.9 against each other. In this corpus a title is the | |
| primary distinguisher between people sharing a surname. | |
| 2. **Nothing prevented merging two separately named characters.** A single bad | |
| link fuses two people; the fused node then matches both name sets and | |
| attracts more. On *Pride and Prejudice* this collapsed 74 gold characters | |
| into 3 nodes. | |
| 3. **The link *bind* path bypassed the merge guard and the policy check**, so | |
| even `append-only` was receiving identity resolution it should not have had. | |
| Fixing these moved conflation on *Pride and Prejudice* from 1.33 to 0.05. | |
| The obvious fix for (2) was to require corroboration β accept a name-to-name | |
| link only if it is proposed more than once. **Measured before adopting, and it | |
| does not work:** across 205 links on one novel, `Mr. Darcy -> Elizabeth` was | |
| proposed nine times, exactly as often as the correct `Mr. Bingley -> Bingley`. | |
| Repetition does not separate good links from bad, so the constraint had to be | |
| structural rather than statistical. | |
| ## 5. Limitations | |
| - **No identity-accuracy gain.** Stated plainly: on the metrics that measure | |
| accuracy rather than consistency, the calculus is neutral. | |
| - **One corpus of 28 English novels**, plus 100 LitBank excerpts. One model | |
| family. | |
| - **Alias-level identity gold.** PDNC's alias sets are surface strings, not | |
| full mention-level coreference; the mention-linking plane partly compensates | |
| but covers only referring expressions inside quotations. | |
| - **Extraction is the binding constraint on several metrics.** A 1.5B model | |
| fails to produce parseable output on 21% of windows, so the scale comparison | |
| is partly a comparison of yield; the parse rate is reported alongside it. | |
| - **Nicknames are unrecoverable** under the merge constraint: `Lizzy` and | |
| `Elizabeth` share no surface material, so the guard that blocks | |
| `Jane -> Elizabeth` also blocks the correct `Lizzy -> Elizabeth`. | |
| - **`SUPERSEDE` versus `REVISE` is decided by a lexicon**, and its accuracy has | |
| not itself been evaluated against annotation. | |