File size: 9,124 Bytes
59387b7
42709ea
59387b7
 
 
42709ea
 
 
59387b7
42709ea
59387b7
 
 
 
42709ea
59387b7
 
 
42709ea
59387b7
42709ea
59387b7
 
 
 
42709ea
59387b7
42709ea
59387b7
 
 
 
 
42709ea
59387b7
 
 
 
 
 
42709ea
59387b7
 
 
 
42709ea
59387b7
 
 
 
 
42709ea
59387b7
42709ea
59387b7
 
42709ea
59387b7
 
 
 
 
 
42709ea
59387b7
 
 
 
 
 
 
 
42709ea
59387b7
 
42709ea
59387b7
42709ea
59387b7
42709ea
 
 
 
59387b7
 
 
 
 
 
42709ea
 
59387b7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
# Structured narrative state: what it can and cannot do

A methodology built from what the evidence survived, not from what was hoped
for. Everything below rests on 28 human-annotated novels, 215 public-domain
novels, and 16,200 machine-generated chapters, with no LLM in any metric.

---

## 1. The thesis

> A revision-aware narrative state is a **representational** result, not a
> conditioning signal and not a text-level instrument. It keeps a maintained
> state coherent; it does not help a model write, and its contradiction count
> does not measure a text.

Three claims, each with its scope stated, and three pre-registered negatives
that bound them. The negatives are load-bearing: they are what make the
positive claim narrow enough to be true.

## 2. C1 β€” The revision calculus (positive, well evidenced)

**Claim.** Given a fixed stream of extracted assertions, a store equipped with
an explicit update calculus holds **0.000** self-contradictory slots, against
**0.370–0.423** for an append-only store β€” at 1.5B, 3B and 7B extractors, over
28 novels.

**Method.** Three update types that append-only pipelines conflate:

| operation | fires when | effect |
| --- | --- | --- |
| `ELABORATE` | later text specifies what earlier text left open | refined in place; nothing retracted |
| `SUPERSEDE` | the **story world** changed | validity interval closed; kept as history |
| `REVISE` | the **reader** was wrong; it was never true | retracted, and so is anything derived from it |

The `SUPERSEDE`/`REVISE` split is the operational form of a narratological
distinction β€” an event in the story versus a disclosure in the telling β€” and is
decided by an inspectable lexicon (predicate mutability, evidential source,
revelation markers), never by a model. Seven structural invariants are checked
after every reading step, including prefix causality: no assertion may cite text
the reader has not reached.

**Scope, stated precisely.** This is a claim about *representations given
identical input*, established by holding extraction constant and varying only
the update policy. It is **not** a claim that the resulting state is a faithful
model of the text (see Β§5).

**Controls.** Every policy replays a byte-identical cached proposal stream, so
differences are attributable to the representation rather than to extractor
variance. Each rung of the ladder adds exactly one mechanism, so each is priced
separately: identity merging, fact revision, deferred commitment, and the cost
of causality against a non-causal oracle.

## 3. C2 β€” A benchmark needing no annotation and no judge (positive)

**Problem.** Generated stories have no gold. Human judgement is expensive; an
LLM judge is the thing being avoided.

**Method.** Plant the gold. Each premise fixes canon facts drawn from closed
vocabularies whose contradictions are enumerable β€” eye colour, metal, kinship,
trade, birthplace. A violation becomes a deterministic string test near a
mention of its subject, with a nearby correction ("not gold but silver")
correctly not counted. Canon is stated in chapter 1 and withheld afterwards, so
the benchmark tests memory rather than prompt-following.

**The metric that makes it trustworthy.** `on_premise` β€” does the chapter write
about the story it was asked for. Violations are only counted near a mention of
their subject, so a model that stops writing about the premise scores a perfect
violation rate while producing nothing usable. This is not hypothetical: a
fine-tuned writer here scored **0.054** against **0.367** for the base model, an
apparent 7Γ— win, purely by degenerating into pastiche of its training corpus.
`on_premise` exposed it (1.000 against 0.765–0.812, with canon restatement
0.84 against 0.09) and the arm was excluded.

**Why this is the durable contribution.** It is gold by construction, so unlike
Β§5 it does not depend on extraction quality at all.

## 4. C3 β€” State conditioning does not help generation (negative, well powered)

Four designs, three runs, 12,600 chapters, 30 stories, paired within story.

| memory | violations ↓ | tokens | verdict |
| --- | --- | --- | --- |
| none | 0.487 | 112 | floor |
| previous chapter | 0.312 | 1,017 | strong, cheap |
| full transcript | 0.254 | 5,362 | best, 5Γ— the context |
| serialised state digest | 0.588 | 357 | **worse than nothing** |
| digest + previous chapter | 0.379 | 1,194 | loses to previous alone |
| beat-conditioned retrieval | 0.571 | 212 | **worse than nothing** |
| retrieval + previous chapter | 0.317 | 1,098 | ties previous alone |
| + graph guard | 0.317 | 1,095 | no effect (ns) |

Serialised state is worse than no memory (+0.083, CI [+0.025, +0.142]).
Beat-conditioned retrieval β€” the design the prior literature predicts should
win β€” is also worse than no memory (+0.083, CI [+0.038, +0.125]) and
indistinguishable from the serialisation it was meant to fix. Retrieval plus
recent text ties recent text alone (+0.004, ns). The guard never moves a number.

**This is not an extraction ceiling.** The state held **0.492** of the planted
canon at the moment of writing and conditioning on it still added nothing.

**It replicates a published negative.** The Narrative World Model paper's own
ablation reports serialised current state at 0.358 against query-conditioned
retrieval at 0.898. Our `state-digest` is structurally their `State Memory`. We
reproduce their failure with a 3B open model β€” and, unlike them, find that
their success condition does not transfer.

## 5. C4 β€” The contradiction count does not measure a text (negative, and it
constrains C1)

We tested our own instrument's external validity and it failed. That test is
reported because it bounds what C1 may claim.

**Design.** 215 published novels and 30 generated stories, first 20 chapters
each, read with the *same* extraction prompt, reconciler and policy. If the
contradiction rate measures textual consistency, professionally edited novels
should score far below machine-generated ones.

**Result.** They do not separate: AUC **0.394**, 95% CI [0.259, 0.534] β€” the
interval spans chance. Published novels score **0.238**, which is not credible
as a rate of genuine self-contradiction in edited prose.

**Diagnosis.** The flagged "contradictions" in real novels are extraction
artefacts. In *Peter Pan*: `Mrs. Darling.occupation: [wife, mother]` β€” both
true; `Nana.material: [Newfoundland dog, dog]` β€” the same thing at different
precision; `Wendy.resides_in: [14, nursery]` β€” a house number and a room.

**Correction applied.** `occupation` is genuinely multi-valued and was
reclassified; refinement detection was broadened so a more precise restatement
is not counted as a rival value. This lowered both populations (human
0.267β†’0.238) and still did not separate them. **Residual extraction noise
dominates the absolute level.**

**What this licenses and forbids.** C1 stands: it is a controlled comparison of
policies on identical input, and that comparison is unaffected by a shared noise
floor. What is *forbidden* is reporting the contradiction rate as a measure of
how consistent a text is. We do not.

## 6. Design principles worth stating

1. **Hold extraction constant.** Every policy comparison replays a
   byte-identical cached stream, so representation and extractor never confound.
2. **Gold by construction beats gold by judgement.** The planted-canon benchmark
   needs no annotation and no model, which is why it survived when the
   extraction-dependent metric did not.
3. **Every headline metric needs a denominator that can expose its artefact.**
   `on_premise` for violation rate; the human-vs-machine test for the
   contradiction rate. One caught a 7Γ— false positive; the other invalidated a
   claim we wanted to make.
4. **Pre-register the falsifiers.** Four were stated before the generation runs;
   three fired, and are reported as such.
5. **Price causality.** A non-causal oracle given the same extraction bounds
   what reading forward costs, separately from what the method buys.

## 7. Honest positioning

This is a paper about the limits of a popular idea, with one solid positive
result and a reusable benchmark. It is not a system paper claiming better story
generation, because the evidence does not support one.

The field is actively building state-conditioned writers. Evidence with
intervals that four such designs fail β€” plus a metric that catches the
degenerate-model artefact which makes them look like they work, plus a
demonstration that the obvious consistency metric measures its own extractor β€”
is a contribution of the kind that saves other people months.

## 8. What would extend it

1. **Mention-level identity gold** (BookCoref) to replace the alias-set proxy.
2. **Validate the `SUPERSEDE`/`REVISE` decision itself** against a few hundred
   hand-annotated conflict pairs. It is the conceptual centre and the least
   evaluated part.
3. **A second model family**, to show the negatives are not Qwen-specific.
4. **Human validation of the planted-canon metric** β€” do readers agree a flagged
   violation reads as a continuity error?