Download METHODOLOGY.md from GOVINDFROM/dsg-artifacts: direct link, hf CLI and curl.
- Browser
- Download file 9.12 kB
-
https://huggingface.co/GOVINDFROM/dsg-artifacts/resolve/main/METHODOLOGY.md
- Command line
-
hf download hf://GOVINDFROM/dsg-artifacts/METHODOLOGY.md
-
curl -L -o METHODOLOGY.md https://huggingface.co/GOVINDFROM/dsg-artifacts/resolve/main/METHODOLOGY.md
Structured narrative state: what it can and cannot do
A methodology built from what the evidence survived, not from what was hoped for. Everything below rests on 28 human-annotated novels, 215 public-domain novels, and 16,200 machine-generated chapters, with no LLM in any metric.
1. The thesis
A revision-aware narrative state is a representational result, not a conditioning signal and not a text-level instrument. It keeps a maintained state coherent; it does not help a model write, and its contradiction count does not measure a text.
Three claims, each with its scope stated, and three pre-registered negatives that bound them. The negatives are load-bearing: they are what make the positive claim narrow enough to be true.
2. C1 β The revision calculus (positive, well evidenced)
Claim. Given a fixed stream of extracted assertions, a store equipped with an explicit update calculus holds 0.000 self-contradictory slots, against 0.370β0.423 for an append-only store β at 1.5B, 3B and 7B extractors, over 28 novels.
Method. Three update types that append-only pipelines conflate:
| operation | fires when | effect |
|---|---|---|
ELABORATE |
later text specifies what earlier text left open | refined in place; nothing retracted |
SUPERSEDE |
the story world changed | validity interval closed; kept as history |
REVISE |
the reader was wrong; it was never true | retracted, and so is anything derived from it |
The SUPERSEDE/REVISE split is the operational form of a narratological
distinction β an event in the story versus a disclosure in the telling β and is
decided by an inspectable lexicon (predicate mutability, evidential source,
revelation markers), never by a model. Seven structural invariants are checked
after every reading step, including prefix causality: no assertion may cite text
the reader has not reached.
Scope, stated precisely. This is a claim about representations given identical input, established by holding extraction constant and varying only the update policy. It is not a claim that the resulting state is a faithful model of the text (see Β§5).
Controls. Every policy replays a byte-identical cached proposal stream, so differences are attributable to the representation rather than to extractor variance. Each rung of the ladder adds exactly one mechanism, so each is priced separately: identity merging, fact revision, deferred commitment, and the cost of causality against a non-causal oracle.
3. C2 β A benchmark needing no annotation and no judge (positive)
Problem. Generated stories have no gold. Human judgement is expensive; an LLM judge is the thing being avoided.
Method. Plant the gold. Each premise fixes canon facts drawn from closed vocabularies whose contradictions are enumerable β eye colour, metal, kinship, trade, birthplace. A violation becomes a deterministic string test near a mention of its subject, with a nearby correction ("not gold but silver") correctly not counted. Canon is stated in chapter 1 and withheld afterwards, so the benchmark tests memory rather than prompt-following.
The metric that makes it trustworthy. on_premise β does the chapter write
about the story it was asked for. Violations are only counted near a mention of
their subject, so a model that stops writing about the premise scores a perfect
violation rate while producing nothing usable. This is not hypothetical: a
fine-tuned writer here scored 0.054 against 0.367 for the base model, an
apparent 7Γ win, purely by degenerating into pastiche of its training corpus.
on_premise exposed it (1.000 against 0.765β0.812, with canon restatement
0.84 against 0.09) and the arm was excluded.
Why this is the durable contribution. It is gold by construction, so unlike Β§5 it does not depend on extraction quality at all.
4. C3 β State conditioning does not help generation (negative, well powered)
Four designs, three runs, 12,600 chapters, 30 stories, paired within story.
| memory | violations β | tokens | verdict |
|---|---|---|---|
| none | 0.487 | 112 | floor |
| previous chapter | 0.312 | 1,017 | strong, cheap |
| full transcript | 0.254 | 5,362 | best, 5Γ the context |
| serialised state digest | 0.588 | 357 | worse than nothing |
| digest + previous chapter | 0.379 | 1,194 | loses to previous alone |
| beat-conditioned retrieval | 0.571 | 212 | worse than nothing |
| retrieval + previous chapter | 0.317 | 1,098 | ties previous alone |
| + graph guard | 0.317 | 1,095 | no effect (ns) |
Serialised state is worse than no memory (+0.083, CI [+0.025, +0.142]). Beat-conditioned retrieval β the design the prior literature predicts should win β is also worse than no memory (+0.083, CI [+0.038, +0.125]) and indistinguishable from the serialisation it was meant to fix. Retrieval plus recent text ties recent text alone (+0.004, ns). The guard never moves a number.
This is not an extraction ceiling. The state held 0.492 of the planted canon at the moment of writing and conditioning on it still added nothing.
It replicates a published negative. The Narrative World Model paper's own
ablation reports serialised current state at 0.358 against query-conditioned
retrieval at 0.898. Our state-digest is structurally their State Memory. We
reproduce their failure with a 3B open model β and, unlike them, find that
their success condition does not transfer.
5. C4 β The contradiction count does not measure a text (negative, and it
constrains C1)
We tested our own instrument's external validity and it failed. That test is reported because it bounds what C1 may claim.
Design. 215 published novels and 30 generated stories, first 20 chapters each, read with the same extraction prompt, reconciler and policy. If the contradiction rate measures textual consistency, professionally edited novels should score far below machine-generated ones.
Result. They do not separate: AUC 0.394, 95% CI [0.259, 0.534] β the interval spans chance. Published novels score 0.238, which is not credible as a rate of genuine self-contradiction in edited prose.
Diagnosis. The flagged "contradictions" in real novels are extraction
artefacts. In Peter Pan: Mrs. Darling.occupation: [wife, mother] β both
true; Nana.material: [Newfoundland dog, dog] β the same thing at different
precision; Wendy.resides_in: [14, nursery] β a house number and a room.
Correction applied. occupation is genuinely multi-valued and was
reclassified; refinement detection was broadened so a more precise restatement
is not counted as a rival value. This lowered both populations (human
0.267β0.238) and still did not separate them. Residual extraction noise
dominates the absolute level.
What this licenses and forbids. C1 stands: it is a controlled comparison of policies on identical input, and that comparison is unaffected by a shared noise floor. What is forbidden is reporting the contradiction rate as a measure of how consistent a text is. We do not.
6. Design principles worth stating
- Hold extraction constant. Every policy comparison replays a byte-identical cached stream, so representation and extractor never confound.
- Gold by construction beats gold by judgement. The planted-canon benchmark needs no annotation and no model, which is why it survived when the extraction-dependent metric did not.
- Every headline metric needs a denominator that can expose its artefact.
on_premisefor violation rate; the human-vs-machine test for the contradiction rate. One caught a 7Γ false positive; the other invalidated a claim we wanted to make. - Pre-register the falsifiers. Four were stated before the generation runs; three fired, and are reported as such.
- Price causality. A non-causal oracle given the same extraction bounds what reading forward costs, separately from what the method buys.
7. Honest positioning
This is a paper about the limits of a popular idea, with one solid positive result and a reusable benchmark. It is not a system paper claiming better story generation, because the evidence does not support one.
The field is actively building state-conditioned writers. Evidence with intervals that four such designs fail β plus a metric that catches the degenerate-model artefact which makes them look like they work, plus a demonstration that the obvious consistency metric measures its own extractor β is a contribution of the kind that saves other people months.
8. What would extend it
- Mention-level identity gold (BookCoref) to replace the alias-set proxy.
- Validate the
SUPERSEDE/REVISEdecision itself against a few hundred hand-annotated conflict pairs. It is the conceptual centre and the least evaluated part. - A second model family, to show the negatives are not Qwen-specific.
- Human validation of the planted-canon metric β do readers agree a flagged violation reads as a continuity error?