dsg-artifacts / METHODOLOGY.md
GOVINDFROM's picture
final methodology
59387b7 verified
|
Raw History Blame Contribute Delete
9.12 kB
# Structured narrative state: what it can and cannot do
A methodology built from what the evidence survived, not from what was hoped
for. Everything below rests on 28 human-annotated novels, 215 public-domain
novels, and 16,200 machine-generated chapters, with no LLM in any metric.
---
## 1. The thesis
> A revision-aware narrative state is a **representational** result, not a
> conditioning signal and not a text-level instrument. It keeps a maintained
> state coherent; it does not help a model write, and its contradiction count
> does not measure a text.
Three claims, each with its scope stated, and three pre-registered negatives
that bound them. The negatives are load-bearing: they are what make the
positive claim narrow enough to be true.
## 2. C1 β€” The revision calculus (positive, well evidenced)
**Claim.** Given a fixed stream of extracted assertions, a store equipped with
an explicit update calculus holds **0.000** self-contradictory slots, against
**0.370–0.423** for an append-only store β€” at 1.5B, 3B and 7B extractors, over
28 novels.
**Method.** Three update types that append-only pipelines conflate:
| operation | fires when | effect |
| --- | --- | --- |
| `ELABORATE` | later text specifies what earlier text left open | refined in place; nothing retracted |
| `SUPERSEDE` | the **story world** changed | validity interval closed; kept as history |
| `REVISE` | the **reader** was wrong; it was never true | retracted, and so is anything derived from it |
The `SUPERSEDE`/`REVISE` split is the operational form of a narratological
distinction β€” an event in the story versus a disclosure in the telling β€” and is
decided by an inspectable lexicon (predicate mutability, evidential source,
revelation markers), never by a model. Seven structural invariants are checked
after every reading step, including prefix causality: no assertion may cite text
the reader has not reached.
**Scope, stated precisely.** This is a claim about *representations given
identical input*, established by holding extraction constant and varying only
the update policy. It is **not** a claim that the resulting state is a faithful
model of the text (see Β§5).
**Controls.** Every policy replays a byte-identical cached proposal stream, so
differences are attributable to the representation rather than to extractor
variance. Each rung of the ladder adds exactly one mechanism, so each is priced
separately: identity merging, fact revision, deferred commitment, and the cost
of causality against a non-causal oracle.
## 3. C2 β€” A benchmark needing no annotation and no judge (positive)
**Problem.** Generated stories have no gold. Human judgement is expensive; an
LLM judge is the thing being avoided.
**Method.** Plant the gold. Each premise fixes canon facts drawn from closed
vocabularies whose contradictions are enumerable β€” eye colour, metal, kinship,
trade, birthplace. A violation becomes a deterministic string test near a
mention of its subject, with a nearby correction ("not gold but silver")
correctly not counted. Canon is stated in chapter 1 and withheld afterwards, so
the benchmark tests memory rather than prompt-following.
**The metric that makes it trustworthy.** `on_premise` β€” does the chapter write
about the story it was asked for. Violations are only counted near a mention of
their subject, so a model that stops writing about the premise scores a perfect
violation rate while producing nothing usable. This is not hypothetical: a
fine-tuned writer here scored **0.054** against **0.367** for the base model, an
apparent 7Γ— win, purely by degenerating into pastiche of its training corpus.
`on_premise` exposed it (1.000 against 0.765–0.812, with canon restatement
0.84 against 0.09) and the arm was excluded.
**Why this is the durable contribution.** It is gold by construction, so unlike
Β§5 it does not depend on extraction quality at all.
## 4. C3 β€” State conditioning does not help generation (negative, well powered)
Four designs, three runs, 12,600 chapters, 30 stories, paired within story.
| memory | violations ↓ | tokens | verdict |
| --- | --- | --- | --- |
| none | 0.487 | 112 | floor |
| previous chapter | 0.312 | 1,017 | strong, cheap |
| full transcript | 0.254 | 5,362 | best, 5Γ— the context |
| serialised state digest | 0.588 | 357 | **worse than nothing** |
| digest + previous chapter | 0.379 | 1,194 | loses to previous alone |
| beat-conditioned retrieval | 0.571 | 212 | **worse than nothing** |
| retrieval + previous chapter | 0.317 | 1,098 | ties previous alone |
| + graph guard | 0.317 | 1,095 | no effect (ns) |
Serialised state is worse than no memory (+0.083, CI [+0.025, +0.142]).
Beat-conditioned retrieval β€” the design the prior literature predicts should
win β€” is also worse than no memory (+0.083, CI [+0.038, +0.125]) and
indistinguishable from the serialisation it was meant to fix. Retrieval plus
recent text ties recent text alone (+0.004, ns). The guard never moves a number.
**This is not an extraction ceiling.** The state held **0.492** of the planted
canon at the moment of writing and conditioning on it still added nothing.
**It replicates a published negative.** The Narrative World Model paper's own
ablation reports serialised current state at 0.358 against query-conditioned
retrieval at 0.898. Our `state-digest` is structurally their `State Memory`. We
reproduce their failure with a 3B open model β€” and, unlike them, find that
their success condition does not transfer.
## 5. C4 β€” The contradiction count does not measure a text (negative, and it
constrains C1)
We tested our own instrument's external validity and it failed. That test is
reported because it bounds what C1 may claim.
**Design.** 215 published novels and 30 generated stories, first 20 chapters
each, read with the *same* extraction prompt, reconciler and policy. If the
contradiction rate measures textual consistency, professionally edited novels
should score far below machine-generated ones.
**Result.** They do not separate: AUC **0.394**, 95% CI [0.259, 0.534] β€” the
interval spans chance. Published novels score **0.238**, which is not credible
as a rate of genuine self-contradiction in edited prose.
**Diagnosis.** The flagged "contradictions" in real novels are extraction
artefacts. In *Peter Pan*: `Mrs. Darling.occupation: [wife, mother]` β€” both
true; `Nana.material: [Newfoundland dog, dog]` β€” the same thing at different
precision; `Wendy.resides_in: [14, nursery]` β€” a house number and a room.
**Correction applied.** `occupation` is genuinely multi-valued and was
reclassified; refinement detection was broadened so a more precise restatement
is not counted as a rival value. This lowered both populations (human
0.267β†’0.238) and still did not separate them. **Residual extraction noise
dominates the absolute level.**
**What this licenses and forbids.** C1 stands: it is a controlled comparison of
policies on identical input, and that comparison is unaffected by a shared noise
floor. What is *forbidden* is reporting the contradiction rate as a measure of
how consistent a text is. We do not.
## 6. Design principles worth stating
1. **Hold extraction constant.** Every policy comparison replays a
byte-identical cached stream, so representation and extractor never confound.
2. **Gold by construction beats gold by judgement.** The planted-canon benchmark
needs no annotation and no model, which is why it survived when the
extraction-dependent metric did not.
3. **Every headline metric needs a denominator that can expose its artefact.**
`on_premise` for violation rate; the human-vs-machine test for the
contradiction rate. One caught a 7Γ— false positive; the other invalidated a
claim we wanted to make.
4. **Pre-register the falsifiers.** Four were stated before the generation runs;
three fired, and are reported as such.
5. **Price causality.** A non-causal oracle given the same extraction bounds
what reading forward costs, separately from what the method buys.
## 7. Honest positioning
This is a paper about the limits of a popular idea, with one solid positive
result and a reusable benchmark. It is not a system paper claiming better story
generation, because the evidence does not support one.
The field is actively building state-conditioned writers. Evidence with
intervals that four such designs fail β€” plus a metric that catches the
degenerate-model artefact which makes them look like they work, plus a
demonstration that the obvious consistency metric measures its own extractor β€”
is a contribution of the kind that saves other people months.
## 8. What would extend it
1. **Mention-level identity gold** (BookCoref) to replace the alias-set proxy.
2. **Validate the `SUPERSEDE`/`REVISE` decision itself** against a few hundred
hand-annotated conflict pairs. It is the conceptual centre and the least
evaluated part.
3. **A second model family**, to show the negatives are not Qwen-specific.
4. **Human validation of the planted-canon metric** β€” do readers agree a flagged
violation reads as a continuity error?