|
Download METHODOLOGY.md from GOVINDFROM/dsg-artifacts: direct link, hf CLI and curl.
- Browser
- Download file 9.12 kB
-
https://huggingface.co/GOVINDFROM/dsg-artifacts/resolve/main/METHODOLOGY.md
- Command line
-
hf download hf://GOVINDFROM/dsg-artifacts/METHODOLOGY.md
-
curl -L -o METHODOLOGY.md https://huggingface.co/GOVINDFROM/dsg-artifacts/resolve/main/METHODOLOGY.md
9.12 kB
| # Structured narrative state: what it can and cannot do | |
| A methodology built from what the evidence survived, not from what was hoped | |
| for. Everything below rests on 28 human-annotated novels, 215 public-domain | |
| novels, and 16,200 machine-generated chapters, with no LLM in any metric. | |
| --- | |
| ## 1. The thesis | |
| > A revision-aware narrative state is a **representational** result, not a | |
| > conditioning signal and not a text-level instrument. It keeps a maintained | |
| > state coherent; it does not help a model write, and its contradiction count | |
| > does not measure a text. | |
| Three claims, each with its scope stated, and three pre-registered negatives | |
| that bound them. The negatives are load-bearing: they are what make the | |
| positive claim narrow enough to be true. | |
| ## 2. C1 β The revision calculus (positive, well evidenced) | |
| **Claim.** Given a fixed stream of extracted assertions, a store equipped with | |
| an explicit update calculus holds **0.000** self-contradictory slots, against | |
| **0.370β0.423** for an append-only store β at 1.5B, 3B and 7B extractors, over | |
| 28 novels. | |
| **Method.** Three update types that append-only pipelines conflate: | |
| | operation | fires when | effect | | |
| | --- | --- | --- | | |
| | `ELABORATE` | later text specifies what earlier text left open | refined in place; nothing retracted | | |
| | `SUPERSEDE` | the **story world** changed | validity interval closed; kept as history | | |
| | `REVISE` | the **reader** was wrong; it was never true | retracted, and so is anything derived from it | | |
| The `SUPERSEDE`/`REVISE` split is the operational form of a narratological | |
| distinction β an event in the story versus a disclosure in the telling β and is | |
| decided by an inspectable lexicon (predicate mutability, evidential source, | |
| revelation markers), never by a model. Seven structural invariants are checked | |
| after every reading step, including prefix causality: no assertion may cite text | |
| the reader has not reached. | |
| **Scope, stated precisely.** This is a claim about *representations given | |
| identical input*, established by holding extraction constant and varying only | |
| the update policy. It is **not** a claim that the resulting state is a faithful | |
| model of the text (see Β§5). | |
| **Controls.** Every policy replays a byte-identical cached proposal stream, so | |
| differences are attributable to the representation rather than to extractor | |
| variance. Each rung of the ladder adds exactly one mechanism, so each is priced | |
| separately: identity merging, fact revision, deferred commitment, and the cost | |
| of causality against a non-causal oracle. | |
| ## 3. C2 β A benchmark needing no annotation and no judge (positive) | |
| **Problem.** Generated stories have no gold. Human judgement is expensive; an | |
| LLM judge is the thing being avoided. | |
| **Method.** Plant the gold. Each premise fixes canon facts drawn from closed | |
| vocabularies whose contradictions are enumerable β eye colour, metal, kinship, | |
| trade, birthplace. A violation becomes a deterministic string test near a | |
| mention of its subject, with a nearby correction ("not gold but silver") | |
| correctly not counted. Canon is stated in chapter 1 and withheld afterwards, so | |
| the benchmark tests memory rather than prompt-following. | |
| **The metric that makes it trustworthy.** `on_premise` β does the chapter write | |
| about the story it was asked for. Violations are only counted near a mention of | |
| their subject, so a model that stops writing about the premise scores a perfect | |
| violation rate while producing nothing usable. This is not hypothetical: a | |
| fine-tuned writer here scored **0.054** against **0.367** for the base model, an | |
| apparent 7Γ win, purely by degenerating into pastiche of its training corpus. | |
| `on_premise` exposed it (1.000 against 0.765β0.812, with canon restatement | |
| 0.84 against 0.09) and the arm was excluded. | |
| **Why this is the durable contribution.** It is gold by construction, so unlike | |
| Β§5 it does not depend on extraction quality at all. | |
| ## 4. C3 β State conditioning does not help generation (negative, well powered) | |
| Four designs, three runs, 12,600 chapters, 30 stories, paired within story. | |
| | memory | violations β | tokens | verdict | | |
| | --- | --- | --- | --- | | |
| | none | 0.487 | 112 | floor | | |
| | previous chapter | 0.312 | 1,017 | strong, cheap | | |
| | full transcript | 0.254 | 5,362 | best, 5Γ the context | | |
| | serialised state digest | 0.588 | 357 | **worse than nothing** | | |
| | digest + previous chapter | 0.379 | 1,194 | loses to previous alone | | |
| | beat-conditioned retrieval | 0.571 | 212 | **worse than nothing** | | |
| | retrieval + previous chapter | 0.317 | 1,098 | ties previous alone | | |
| | + graph guard | 0.317 | 1,095 | no effect (ns) | | |
| Serialised state is worse than no memory (+0.083, CI [+0.025, +0.142]). | |
| Beat-conditioned retrieval β the design the prior literature predicts should | |
| win β is also worse than no memory (+0.083, CI [+0.038, +0.125]) and | |
| indistinguishable from the serialisation it was meant to fix. Retrieval plus | |
| recent text ties recent text alone (+0.004, ns). The guard never moves a number. | |
| **This is not an extraction ceiling.** The state held **0.492** of the planted | |
| canon at the moment of writing and conditioning on it still added nothing. | |
| **It replicates a published negative.** The Narrative World Model paper's own | |
| ablation reports serialised current state at 0.358 against query-conditioned | |
| retrieval at 0.898. Our `state-digest` is structurally their `State Memory`. We | |
| reproduce their failure with a 3B open model β and, unlike them, find that | |
| their success condition does not transfer. | |
| ## 5. C4 β The contradiction count does not measure a text (negative, and it | |
| constrains C1) | |
| We tested our own instrument's external validity and it failed. That test is | |
| reported because it bounds what C1 may claim. | |
| **Design.** 215 published novels and 30 generated stories, first 20 chapters | |
| each, read with the *same* extraction prompt, reconciler and policy. If the | |
| contradiction rate measures textual consistency, professionally edited novels | |
| should score far below machine-generated ones. | |
| **Result.** They do not separate: AUC **0.394**, 95% CI [0.259, 0.534] β the | |
| interval spans chance. Published novels score **0.238**, which is not credible | |
| as a rate of genuine self-contradiction in edited prose. | |
| **Diagnosis.** The flagged "contradictions" in real novels are extraction | |
| artefacts. In *Peter Pan*: `Mrs. Darling.occupation: [wife, mother]` β both | |
| true; `Nana.material: [Newfoundland dog, dog]` β the same thing at different | |
| precision; `Wendy.resides_in: [14, nursery]` β a house number and a room. | |
| **Correction applied.** `occupation` is genuinely multi-valued and was | |
| reclassified; refinement detection was broadened so a more precise restatement | |
| is not counted as a rival value. This lowered both populations (human | |
| 0.267β0.238) and still did not separate them. **Residual extraction noise | |
| dominates the absolute level.** | |
| **What this licenses and forbids.** C1 stands: it is a controlled comparison of | |
| policies on identical input, and that comparison is unaffected by a shared noise | |
| floor. What is *forbidden* is reporting the contradiction rate as a measure of | |
| how consistent a text is. We do not. | |
| ## 6. Design principles worth stating | |
| 1. **Hold extraction constant.** Every policy comparison replays a | |
| byte-identical cached stream, so representation and extractor never confound. | |
| 2. **Gold by construction beats gold by judgement.** The planted-canon benchmark | |
| needs no annotation and no model, which is why it survived when the | |
| extraction-dependent metric did not. | |
| 3. **Every headline metric needs a denominator that can expose its artefact.** | |
| `on_premise` for violation rate; the human-vs-machine test for the | |
| contradiction rate. One caught a 7Γ false positive; the other invalidated a | |
| claim we wanted to make. | |
| 4. **Pre-register the falsifiers.** Four were stated before the generation runs; | |
| three fired, and are reported as such. | |
| 5. **Price causality.** A non-causal oracle given the same extraction bounds | |
| what reading forward costs, separately from what the method buys. | |
| ## 7. Honest positioning | |
| This is a paper about the limits of a popular idea, with one solid positive | |
| result and a reusable benchmark. It is not a system paper claiming better story | |
| generation, because the evidence does not support one. | |
| The field is actively building state-conditioned writers. Evidence with | |
| intervals that four such designs fail β plus a metric that catches the | |
| degenerate-model artefact which makes them look like they work, plus a | |
| demonstration that the obvious consistency metric measures its own extractor β | |
| is a contribution of the kind that saves other people months. | |
| ## 8. What would extend it | |
| 1. **Mention-level identity gold** (BookCoref) to replace the alias-set proxy. | |
| 2. **Validate the `SUPERSEDE`/`REVISE` decision itself** against a few hundred | |
| hand-annotated conflict pairs. It is the conceptual centre and the least | |
| evaluated part. | |
| 3. **A second model family**, to show the negatives are not Qwen-specific. | |
| 4. **Human validation of the planted-canon metric** β do readers agree a flagged | |
| violation reads as a continuity error? | |