File size: 9,124 Bytes
59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 42709ea 59387b7 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 | # Structured narrative state: what it can and cannot do
A methodology built from what the evidence survived, not from what was hoped
for. Everything below rests on 28 human-annotated novels, 215 public-domain
novels, and 16,200 machine-generated chapters, with no LLM in any metric.
---
## 1. The thesis
> A revision-aware narrative state is a **representational** result, not a
> conditioning signal and not a text-level instrument. It keeps a maintained
> state coherent; it does not help a model write, and its contradiction count
> does not measure a text.
Three claims, each with its scope stated, and three pre-registered negatives
that bound them. The negatives are load-bearing: they are what make the
positive claim narrow enough to be true.
## 2. C1 β The revision calculus (positive, well evidenced)
**Claim.** Given a fixed stream of extracted assertions, a store equipped with
an explicit update calculus holds **0.000** self-contradictory slots, against
**0.370β0.423** for an append-only store β at 1.5B, 3B and 7B extractors, over
28 novels.
**Method.** Three update types that append-only pipelines conflate:
| operation | fires when | effect |
| --- | --- | --- |
| `ELABORATE` | later text specifies what earlier text left open | refined in place; nothing retracted |
| `SUPERSEDE` | the **story world** changed | validity interval closed; kept as history |
| `REVISE` | the **reader** was wrong; it was never true | retracted, and so is anything derived from it |
The `SUPERSEDE`/`REVISE` split is the operational form of a narratological
distinction β an event in the story versus a disclosure in the telling β and is
decided by an inspectable lexicon (predicate mutability, evidential source,
revelation markers), never by a model. Seven structural invariants are checked
after every reading step, including prefix causality: no assertion may cite text
the reader has not reached.
**Scope, stated precisely.** This is a claim about *representations given
identical input*, established by holding extraction constant and varying only
the update policy. It is **not** a claim that the resulting state is a faithful
model of the text (see Β§5).
**Controls.** Every policy replays a byte-identical cached proposal stream, so
differences are attributable to the representation rather than to extractor
variance. Each rung of the ladder adds exactly one mechanism, so each is priced
separately: identity merging, fact revision, deferred commitment, and the cost
of causality against a non-causal oracle.
## 3. C2 β A benchmark needing no annotation and no judge (positive)
**Problem.** Generated stories have no gold. Human judgement is expensive; an
LLM judge is the thing being avoided.
**Method.** Plant the gold. Each premise fixes canon facts drawn from closed
vocabularies whose contradictions are enumerable β eye colour, metal, kinship,
trade, birthplace. A violation becomes a deterministic string test near a
mention of its subject, with a nearby correction ("not gold but silver")
correctly not counted. Canon is stated in chapter 1 and withheld afterwards, so
the benchmark tests memory rather than prompt-following.
**The metric that makes it trustworthy.** `on_premise` β does the chapter write
about the story it was asked for. Violations are only counted near a mention of
their subject, so a model that stops writing about the premise scores a perfect
violation rate while producing nothing usable. This is not hypothetical: a
fine-tuned writer here scored **0.054** against **0.367** for the base model, an
apparent 7Γ win, purely by degenerating into pastiche of its training corpus.
`on_premise` exposed it (1.000 against 0.765β0.812, with canon restatement
0.84 against 0.09) and the arm was excluded.
**Why this is the durable contribution.** It is gold by construction, so unlike
Β§5 it does not depend on extraction quality at all.
## 4. C3 β State conditioning does not help generation (negative, well powered)
Four designs, three runs, 12,600 chapters, 30 stories, paired within story.
| memory | violations β | tokens | verdict |
| --- | --- | --- | --- |
| none | 0.487 | 112 | floor |
| previous chapter | 0.312 | 1,017 | strong, cheap |
| full transcript | 0.254 | 5,362 | best, 5Γ the context |
| serialised state digest | 0.588 | 357 | **worse than nothing** |
| digest + previous chapter | 0.379 | 1,194 | loses to previous alone |
| beat-conditioned retrieval | 0.571 | 212 | **worse than nothing** |
| retrieval + previous chapter | 0.317 | 1,098 | ties previous alone |
| + graph guard | 0.317 | 1,095 | no effect (ns) |
Serialised state is worse than no memory (+0.083, CI [+0.025, +0.142]).
Beat-conditioned retrieval β the design the prior literature predicts should
win β is also worse than no memory (+0.083, CI [+0.038, +0.125]) and
indistinguishable from the serialisation it was meant to fix. Retrieval plus
recent text ties recent text alone (+0.004, ns). The guard never moves a number.
**This is not an extraction ceiling.** The state held **0.492** of the planted
canon at the moment of writing and conditioning on it still added nothing.
**It replicates a published negative.** The Narrative World Model paper's own
ablation reports serialised current state at 0.358 against query-conditioned
retrieval at 0.898. Our `state-digest` is structurally their `State Memory`. We
reproduce their failure with a 3B open model β and, unlike them, find that
their success condition does not transfer.
## 5. C4 β The contradiction count does not measure a text (negative, and it
constrains C1)
We tested our own instrument's external validity and it failed. That test is
reported because it bounds what C1 may claim.
**Design.** 215 published novels and 30 generated stories, first 20 chapters
each, read with the *same* extraction prompt, reconciler and policy. If the
contradiction rate measures textual consistency, professionally edited novels
should score far below machine-generated ones.
**Result.** They do not separate: AUC **0.394**, 95% CI [0.259, 0.534] β the
interval spans chance. Published novels score **0.238**, which is not credible
as a rate of genuine self-contradiction in edited prose.
**Diagnosis.** The flagged "contradictions" in real novels are extraction
artefacts. In *Peter Pan*: `Mrs. Darling.occupation: [wife, mother]` β both
true; `Nana.material: [Newfoundland dog, dog]` β the same thing at different
precision; `Wendy.resides_in: [14, nursery]` β a house number and a room.
**Correction applied.** `occupation` is genuinely multi-valued and was
reclassified; refinement detection was broadened so a more precise restatement
is not counted as a rival value. This lowered both populations (human
0.267β0.238) and still did not separate them. **Residual extraction noise
dominates the absolute level.**
**What this licenses and forbids.** C1 stands: it is a controlled comparison of
policies on identical input, and that comparison is unaffected by a shared noise
floor. What is *forbidden* is reporting the contradiction rate as a measure of
how consistent a text is. We do not.
## 6. Design principles worth stating
1. **Hold extraction constant.** Every policy comparison replays a
byte-identical cached stream, so representation and extractor never confound.
2. **Gold by construction beats gold by judgement.** The planted-canon benchmark
needs no annotation and no model, which is why it survived when the
extraction-dependent metric did not.
3. **Every headline metric needs a denominator that can expose its artefact.**
`on_premise` for violation rate; the human-vs-machine test for the
contradiction rate. One caught a 7Γ false positive; the other invalidated a
claim we wanted to make.
4. **Pre-register the falsifiers.** Four were stated before the generation runs;
three fired, and are reported as such.
5. **Price causality.** A non-causal oracle given the same extraction bounds
what reading forward costs, separately from what the method buys.
## 7. Honest positioning
This is a paper about the limits of a popular idea, with one solid positive
result and a reusable benchmark. It is not a system paper claiming better story
generation, because the evidence does not support one.
The field is actively building state-conditioned writers. Evidence with
intervals that four such designs fail β plus a metric that catches the
degenerate-model artefact which makes them look like they work, plus a
demonstration that the obvious consistency metric measures its own extractor β
is a contribution of the kind that saves other people months.
## 8. What would extend it
1. **Mention-level identity gold** (BookCoref) to replace the alias-set proxy.
2. **Validate the `SUPERSEDE`/`REVISE` decision itself** against a few hundred
hand-annotated conflict pairs. It is the conceptual centre and the least
evaluated part.
3. **A second model family**, to show the negatives are not Qwen-specific.
4. **Human validation of the planted-canon metric** β do readers agree a flagged
violation reads as a continuity error?
|