File size: 38,366 Bytes
b68816f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f07443e
b68816f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f07443e
 
b68816f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f07443e
b68816f
 
 
f07443e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b68816f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f07443e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b68816f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f07443e
b68816f
 
 
 
 
 
 
 
 
 
 
 
 
 
f07443e
b68816f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f07443e
b68816f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f07443e
b68816f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f07443e
b68816f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f07443e
 
b68816f
 
 
 
 
 
 
 
f07443e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b68816f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
# Knowledge Pipeline — Output Contract **v2**

> **Implemented 2026-08-26.** `src/knowledge_extraction/` emits this shape.
> The schema review of 2026-08-26 landed with it: ids on every entity, formulas referenced rather
> than restated, rules linked to formulas and terms, and 16 fields cut.
>
> **Not yet re-measured.** The prompt changes this required have not been scored — no source
> document, parsed artifact or MinerU install exists on the dev box (all three `data/` directories
> are gitignored and absent), so the free stages have nothing to read. The last measured figures in
> §12 predate these edits. See §13.

**Supersedes:** v1 of this document (2026-08-24)
**Reference run:** `v2_full_document_2026-08-24_093051` · BUMA `STD_2026_006_MNO` · 9 pages · `gpt-5.4-nano`
**Producer:** `python -m src.knowledge_extraction.cli <artifact.json> --extract`
**Status:** offline artifacts only — no endpoint, no table.

---

## 0 · What changed from v1

Harry's review of v1 asked for entity ids, a formula reference instead of a restated formula, and a
link between rules and formulas. All are adopted. The schema was also cut back — v1 carried fields
nothing consumes.

| Change | Detail |
|---|---|
| **Ids on every entity** | `term_id`, `formula_id`, `brief_id`; `rule_id` re-specified. All deterministic and content-derived |
| **Glossary no longer restates formulas** | `formula_latex` removed → `defining_formula_id` reference |
| **Rules link to formulas and terms** | `formula_ids[]`, `term_ids[]` |
| **`rule_type` added** | `interpretation` \| `calculation` — see §2 for why |
| **16 fields cut, 8 added** | 50 → 42 fields total |
| **`chunk_id` separator fixed** | `DOC#0007` → **`DOC::0007`**, adopting the parser's format |
| **Reference numbers corrected** | v1 quoted a superseded run (13 chunks / 66 clusters). Current is **31 chunks / 58 clusters** |

### Why the ids had to be re-specified

Neither id in v1 survived a re-run. `cluster_id` was a **frequency-rank ordinal**
(`f"{doc_id}#c{i:03d}"` over a list sorted by mention count), and `rule_id` was **written by the
model** (`rule.txt` asks for SCREAMING_SNAKE_CASE). Either way, re-running with a tuned prompt
reshuffles the ids — and if an expert has approved 40 entries, we have lost which 40.

**Approval continuity, not linking, is the real reason these exist.**

---

## 1 · Id scheme

All ids are deterministic, content-derived, and stable across re-runs of the same document.
None is model-generated; none is positional.

```
term_id     = "t_" + sha256(doc_id + "|" + canonical)[:10]
formula_id  = "f_" + sha256(doc_id + "|" + name_or_latex_or_chunk_id)[:10]
rule_id     = "r_" + sha256(doc_id + "|" + chunk_id + "|" + char_start)[:10]
brief_id    = "b_" + sha256(doc_id)[:10]    # per-document card; `d_` belongs to the
                                            # scope-level DomainContext in src/knowledge_domain/
```

Implemented in `src/knowledge_extraction/ids.py`. `term_id` is keyed on the **cluster canonical**,
not on the `term` the model returned — the cluster is the stable thing across runs, and the model is
free to answer "PA" or "Physical Availability" for the same cluster. `formula_id` falls back through
`name` → `formula_latex` → `chunk_id`, because both of the first two are Optional and without the
fallback every unnamed formula in a document would collide.

**Scope is per document.** The same term appearing in two documents receives two different
`term_id`s. Deciding that BUMA's "PA" and a textbook's "PA" are the same concept is an **expert
judgement at review time**, not something the pipeline asserts — the same reasoning that makes the
pipeline record "Physical *of* Availability" instead of normalising it, and the reason MTTR at BUMA
(breakdown duration ÷ breakdown frequency) must not silently merge with MTTR in IT.

**Known limit.** Ids derive from content, so if content changes the id changes — retune clustering
such that a term canonicalises differently and its `term_id` moves. The alternative is a persisted
registry, which belongs with persistence (DEV_PLAN §0.8 D2).

---

## 2 · The four artifacts

| File | Entity | Holds |
|---|---|---|
| `glossary.json` | `GlossaryEntry[]` | Terms and their definitions |
| `rules.json` | `RuleEntry[]` | Rules of thumb — how to interpret data or behaviour. **Renamed from `interpretation_pack.json` 2026-09-08** — the file, the entity and the DB `kind` now share one name. *The Interpretation Pack* remains the deliverable's name in prose |
| `formulas.json` | `FormulaEntry[]` | Formulas, transcribed from the document |
| `document_brief.json` | `DocumentBrief` | One per-document card. Prefix history, not churn: `b_` (BriefContext) -> `d_` (DomainContext, 2026-09-01) -> **`b_` again** (DocumentBrief, 2026-09-08), once v4 gave `d_` to the real scope-level `DomainContext` in `src/knowledge_domain/` |

Plus three derived/audit payloads: `review_queue.json`, `rejected.json`, `usage.json`.

### Rule vs domain knowledge — where the line is

*Added 2026-09-14 (o2). The two blurred on screen during the 2026-09-08 review,
and a boundary nobody has written down is one the prompts cannot hold.*

| | **`rule`** | **`domain`** |
|---|---|---|
| Grain | One statement, from one document | One per company, composed across all of them |
| Shape | A condition and a consequence | Identity, boundary, conventions, measures |
| Evidence | Span-checked against its source | Derived from entries, or expert-declared |
| Answers | *"…and how do I compute this correctly?"* | *"…what does this company measure, and what does it call it?"* |

**The test to apply.** If removing it would change **how a number is computed
or sourced**, it is a rule — *"if joint-survey data is unavailable, use truck
count"* changes the figure. If removing it would only leave an agent **less
oriented** — not knowing that this company works in overburden and BCM — it is
domain.

Two consequences worth stating, because they are what the distinction is *for*:

- **A rule is actionable alone; domain knowledge is not.** A planner can apply
  a rule to a specific calculation. It cannot apply "this is a mining company"
  to anything — that shapes which questions make sense, not which column to sum.
- **Domain is high-level and rules are specific, but "high-level" is not the
  test.** A long, general rule is still a rule. The grain question — *does this
  come from one document's sentence, or from the corpus as a whole?* — settles
  it faster and does not drift.

Where it is genuinely both, prefer the **rule**: it carries provenance and a
span, so an expert can check it. Domain fields are the ones the pipeline can
least easily prove.

### Why `RuleEntry` carries a `rule_type`

The Interpretation Pack exists to **improve analytics insight from expert interpretation rules** —
interpretation logic, action benchmarks, and rules for when to conclude or not conclude from data.

Measured against the reference document, **0 of the 15 gold rules are interpretation rules.** They
are calculation conventions (`PA_COMPOSITE_WEIGHTED`), data-sourcing rules
(`PTY_PRODUCTION_SOURCE`), and unit conventions (`GAINLOSS_UNITS`). This is a property of the
document type: a *Standard Parameter* document defines how to **calculate** parameters, not how to
**read** them. Interpretation rules live in other document types, or in the expert's head.

`rule_type` keeps both in one artifact rather than discarding the conventions we do extract:

- `interpretation` — how to read a value or behaviour. **The target.**
- `calculation` — how a value is computed or sourced. What this corpus actually yields.

A consumer wanting only interpretation logic filters on `rule_type`.

---

## 2c · Evidence now carries figures and tables *(parsing 0.4.0, 2026-09-10)*

The artifact this half consumes changed shape: a figure is no longer its own chunk. It is folded into
the prose chunk at its reading-order position as `![](asset://<id>)`, and travels as an `Asset` on that
chunk. `adapter.py` carries `assets`, `referenced_by` and `table_html` across; `evidence_block` renders
them into the prompt under `ASSETS REFERENCED BY THIS CHUNK`.

**What that means for an extraction branch**, and the distinction is enforced in the prompts:

| | may be quoted? | may be a provenance span? |
|---|---|---|
| `Asset.caption` — printed in the document | **yes** | **yes** |
| `Asset.description` — written by a vision model | no | **no** — it is not in the document, so the span check rejects it and the field is discarded |

Why it mattered: before 0.4.0 a figure chunk had an EMPTY `text`, so it produced no mentions, joined no
cluster, and was never selected as evidence. 13 figures across three documents, each already described
by a paid vision call, reached **zero** prompts. Full detail:
`knowledge_pipeline_runs/CHANGELOG_v2.md` §2.

> ✅ **RESOLVED 2026-09-11 — §1, §2 and §8 now match the code.** The note is kept for the record.
>
> ⚠️ **Correction to §8 of this document (noted 2026-09-10).** The code writes
> `document_brief.json` with **`brief_id`** and a `b_` prefix; §8 below documents
> `domain_context.json` with `domain_id` and `d_`. The v3 rename was partly reverted for the
> per-document object once the scope-level `DomainContext` became a separate thing in
> `src/knowledge_domain/`. **§8 does not describe what the code emits.** Worth one edit, and it matters
> because `entity_id` prefixes are the key an expert's approvals hang on.

---

## 3 · Transport reality

Every payload is written by the CLI as a file into `--out-dir` (default `out/knowledge/`), as a
**bare JSON array** with no envelope. There is no endpoint and no table.

Three conventions hold across every payload:

- **`page` is 0-based**, exactly as the parser reports; **`page_no` is the 1-based number a human
  reads** and is what a review UI binds to. Both travel together on every entity, derived from one
  value so they cannot disagree (S6b, 2026-09-02).
- **All content fields are nullable by design.** `null` is a valid, correct answer — the model
  abstaining, not failing. On the reference run **48 of 58** entries carry no definition.
- **`provenance` is mandatory and `provenance.span` is verbatim-checked.** A field whose span cannot
  be located in the source is set to `null` and logged to `rejected.json` — **never repaired.**

---

## 4 · Common types

### `Provenance` — required on every entity

```json
{
  "doc_id": "STD_2026_006_MNO",
  "span": "Adalah jumlah equipment/unit yang digunakan untuk menghasilkan produksi dalam suatu periode.",
  "page": 3,
  "page_no": 4,
  "section_no": "2.1.2",
  "chunk_id": "STD_2026_006_MNO::0007"
}
```

| Field | Type | Req | Notes |
|---|---|---|---|
| `doc_id` | string | ✅ | |
| `span` | string | ✅ | **Verbatim** from the source. The anti-hallucination control — the reviewer checks a quote against a page, not a claim against memory |
| `page` | int \| null | — | **0-based**, exactly as the parser reports |
| `page_no` | int \| null | — | **1-based. Added 2026-09-02 (S6b).** DERIVED from `page`, never stored, so it cannot drift. **A review UI must render this one** — an off-by-one is invisible until an expert opens the wrong page and concludes the provenance is wrong. `null` when `page` is unknown |
| `section_no` | string \| null | — | e.g. `"2.1.2"`; null when the document is unnumbered |
| `chunk_id` | string \| null | — | Back-reference into the parsed artifact. **Format `<doc_id>::<seq>`**, matching `KNOWLEDGE_PARSING_OUTPUT_CONTRACT.md`. Treat as opaque — do not split it |

### Enums

| Enum | Values |
|---|---|
| `subdomain_tags[]` | `production` · `maintenance` · `hauling` · `loading` · `drilling_blasting` · `equipment` · `safety` · `quality` · `planning` · `cost` · `geology` · `other` |
| `extraction_status` | `ok` · `no_definition_found` · `escalated` |
| `diff_status` | `new` · `duplicate` · `conflicting` |
| `rule_type` | `interpretation` · `calculation` |
| `latex_verification` | `verified` · `unverified_no_markup` · `unverified_operator` |

`subdomain_tags` is a **closed enum — classification, not generation.** Adding a member changes what
the model is allowed to answer, which makes it a prompt change, not a data change.

---

## 5 · `glossary.json` → `GlossaryEntry[]` — 13 fields

```json
[
  {
    "term_id": "t_9f2a41c0b7",
    "term": "Qty",
    "full_name": "Quantity",
    "source_wording": "2.1.2. Quantity (Qty)",
    "definition": "Adalah jumlah equipment/unit yang digunakan untuk menghasilkan produksi dalam suatu periode.",
    "defining_formula_id": "f_31de08aa95",
    "subdomain_tags": ["production", "equipment"],
    "mention_count": 20,
    "provenance": {
      "doc_id": "STD_2026_006_MNO",
      "span": "Adalah jumlah equipment/unit yang digunakan untuk menghasilkan produksi dalam suatu periode.",
      "page": 3,
      "section_no": "2.1.2",
      "chunk_id": "STD_2026_006_MNO::0007"
    },
    "extraction_status": "ok",
    "diff_status": "new",
    "definition_conflict": false,
    "conflict_variants": []
  },
  {
    "term_id": "t_5b71ce0d34",
    "term": "Gain/Loss",
    "full_name": null,
    "source_wording": null,
    "definition": null,
    "defining_formula_id": null,
    "subdomain_tags": ["production"],
    "mention_count": 4,
    "provenance": {
      "doc_id": "STD_2026_006_MNO",
      "span": "Gain/Loss",
      "page": 6,
      "section_no": null,
      "chunk_id": "STD_2026_006_MNO::0011"
    },
    "extraction_status": "no_definition_found",
    "diff_status": "new",
    "definition_conflict": false,
    "conflict_variants": []
  }
]
```

| Field | Type | Req | Notes |
|---|---|---|---|
| `term_id` | string | ✅ | **New in v2.** Stable across re-runs — the key an approval hangs on |
| `term` | string | ✅ | The only required content field |
| `full_name` | string \| null | — | Expanded name. Span-guarded as a literal transcription |
| `source_wording` | string \| null | — | **The document's literal wording**, un-normalised. Read deterministically from the section heading, *not* asked of the model — asked directly, it returned the tidied form |
| `definition` | string \| null | — | `null` = correct abstention (**48 of 58** on the reference run) |
| `defining_formula_id` | string \| null | — | **New in v2**, replaces `formula_latex`. The formula that *defines* this term. Null when the formula branch did not extract one |
| `subdomain_tags` | enum[] | ✅ | Defaults `[]` |
| `mention_count` | int | ✅ | Default `0`. **Drives review-queue ordering** |
| `provenance` | Provenance | ✅ | |
| `extraction_status` | enum | ✅ | Default `ok` |
| `diff_status` | enum \| null | — | Set by the diff stage against the active glossary version |
| `definition_conflict` | bool | ✅ | Default `false` |
| `conflict_variants` | string[] | ✅ | Default `[]`. The competing definitions — the pipeline **never picks a winner** |

**Removed in v2:** `formula_latex` (→ `defining_formula_id`) · `interpretation` (belongs in the
Interpretation Pack, reached via `term_id`) · `domain` · `company` · `language`.

> `domain` and `company` were never in the extraction prompt's field guide — the model filled them
> by copying worked examples — and neither was span-guarded, so a wrong value was undetectable.
> `language` is a genuine three-value field (`id`/`en`/`mixed`) but nothing consumes it and it is
> recoverable from the document; it is cut for simplicity and can return if a consumer needs it.

> **Writer split.** The LLM fills only `term`, `full_name`, `definition` and `subdomain_tags`.
> `term_id`, `defining_formula_id`, `source_wording`, `mention_count`, `extraction_status`,
> `diff_status`, `definition_conflict` and `conflict_variants` are all set by deterministic code —
> **the model never writes its own ids, links or audit fields**, so none of them can be hallucinated.

---

## 6 · `rules.json` → `RuleEntry[]` — 8 fields

```json
[
  {
    "rule_id": "r_c07be41f22",
    "rule_type": "calculation",
    "statement": "Production yang digunakan dalam perhitungan adalah produksi hasil joint survey.",
    "condition": "Apabila data joint survey belum tersedia",
    "consequence": "maka digunakan data produksi berdasarkan truck count sebagai dasar perhitungan",
    "formula_ids": ["f_77c1ab9e40"],
    "term_ids": ["t_a4e2f81b06"],
    "provenance": {
      "doc_id": "STD_2026_006_MNO",
      "span": "Apabila data joint survey belum tersedia maka digunakan data produksi berdasarkan truck count sebagai dasar perhitungan",
      "page": 4,
      "section_no": "2.1.5",
      "chunk_id": "STD_2026_006_MNO::0009"
    }
  }
]
```

| Field | Type | Req | Notes |
|---|---|---|---|
| `rule_id` | string | ✅ | **Re-specified in v2.** Was written by the model, therefore unstable; now deterministic |
| `rule_type` | enum | ✅ | `interpretation` \| `calculation` — see §2 |
| `statement` | string \| null | — | The rule as prose. **Prose, not an equation** — evidence carrying both yields the prose, and the equation goes to the formula branch |
| `condition` | string \| null | — | The triggering condition, split out — not left buried in prose |
| `consequence` | string \| null | — | What follows when the condition holds |
| `formula_ids` | string[] | ✅ | **New in v2.** The formulas this rule constrains. Default `[]` |
| `term_ids` | string[] | ✅ | **New in v2.** The terms this rule governs. Default `[]` |
| `provenance` | Provenance | ✅ | |

**Removed in v2:** `applies_to` (free text → `term_ids[]`) · `subdomain_tags` · `language` ·
`extraction_status`.

`condition` + `consequence` is the point of this artifact, and the seed of the interpretation logic
tree: a rule stored as one paragraph cannot be attached to a skill later; a trigger and its
consequence can.

> **Deliberately not modelled yet:** the interpretation logic tree, action benchmarks, and explicit
> conclude / do-not-conclude verdicts. Those are the direction of travel, but designing their schema
> against zero real examples would be guessing. They land when a document containing them does.

---

## 7 · `formulas.json` → `FormulaEntry[]` — 7 fields

```json
[
  {
    "formula_id": "f_77c1ab9e40",
    "name": "Production",
    "formula_latex": "Production = MOHH \\times Qty \\times PA \\times UA \\times Pty",
    "variables": [
      { "symbol": "MOHH", "term_id": "t_1c9d5e7a83", "meaning": "Machine on Hand Hours" },
      { "symbol": "Qty",  "term_id": "t_9f2a41c0b7", "meaning": "Quantity" },
      { "symbol": "PA",   "term_id": "t_a4e2f81b06", "meaning": "Physical Availability" },
      { "symbol": "UA",   "term_id": "t_6d40b2fc17", "meaning": "Utilization of Availability" },
      { "symbol": "Pty",  "term_id": "t_e8b3092d55", "meaning": "Productivity" }
    ],
    "unit": "BCM",
    "latex_verification": "unverified_operator",
    "provenance": {
      "doc_id": "STD_2026_006_MNO",
      "span": "Production = MOHH x Qty x PA x UA x Pty",
      "page": 1,
      "section_no": "2.1",
      "chunk_id": "STD_2026_006_MNO::0003"
    }
  }
]
```

| Field | Type | Req | Notes |
|---|---|---|---|
| `formula_id` | string | ✅ | **New in v2** |
| `name` | string \| null | — | What the formula computes, as named in the source |
| `formula_latex` | string \| null | — | **Transcribed, never derived.** The pipeline copies what the document states |
| `variables[]` | object[] | ✅ | `{ symbol (required), term_id \| null, meaning \| null }`. `term_id` is **new in v2** |
| `unit` | string \| null | — | Result unit if stated |
| `latex_verification` | enum \| null | — | **Added 2026-08-26.** Which guarantee `formula_latex` actually carries. Set by the span check, never by the model. `null` when there is no `formula_latex` to characterise |
| `provenance` | Provenance | ✅ | |

**Removed in v2:** `extraction_status`.

> **Read `latex_verification` before trusting `formula_latex`.** Three paths leave the field populated and
> they are not equally strong. `verified` means the transcription was located in the source markup
> (`Chunk.latex`), compared in a notation-aware canonical form. The two `unverified_*` values mean the
> claim could not be checked either way and the entry is guarded by its provenance span alone — the same
> guarantee everything carried before 2026-08-26.
>
> `unverified_operator` is not rare. MinerU's **`pipeline`** backend transcribes multiplication as a bare
> letter `x`, so every product formula from such an artifact lands here — including the sample above.
> An artifact parsed with `hybrid/high` carries `\times` and verifies normally.

> `variables[].symbol` is not always a legend abbreviation — the extraction prompt's own worked
> example emits `"Total Hours"` and `"Breakdown"` as symbols. So `term_id` resolution is
> **best-effort and nulls are expected**; the link stage reports its dangle rate rather than hiding it.

---

## 8 · `document_brief.json` → `DocumentBrief` — 6 fields *(single object)*

> ✅ **The v3 reshape is complete, corrected 2026-09-08.** This banner previously said the
> LLM-facing parts — `purpose` as a span-guarded verbatim quote, and the `subdomains` aggregate —
> were *"still proposed and not built"*. They shipped on 2026-09-02 and the banner was never
> updated, so it contradicted the field table directly below it, which dates both changes. Verified
> live: a `domain` entry persisted on 2026-09-07 carries `purpose_verbatim` and `subdomains`.
>
> The delivery really was split, and that part was deliberate: the deterministic half (`outline`
> read off the artifact, `key_parameters` corroborated against extracted terms) shipped first
> because it costs nothing to prove, and the LLM-facing half followed once an eval run could pay
> for itself.

```json
{
  "brief_id": "d_2ef60a8c19",
  "title": "Standard Parameter Produksi & ECA",
  "purpose_verbatim": "Standard Parameter ini merupakan pedoman dalam perhitungan parameter produksi yang meliputi Quantity (Qty) Physical Availability (PA), Utilization of Availability (UA), Productivity (PTY), serta Equipment Capacity Analysis (ECA).",
  "outline": [
    "1. TUJUAN PARAMETER",
    "2.1.2. Quantity (Qty)",
    "2.1.3. Physical of Availability (PA)",
    "2.2. Equipment Capacity Analysis (ECA)"
  ],
  "key_parameters": [
    {"surface": "Physical Availability (PA)", "term_id": "t_8c7359e27f"},
    {"surface": "Equipment Capacity Analysis (ECA)", "term_id": null}
  ],
  "subdomains": ["equipment", "production"],
  "n_terms": 5,
  "n_formulas": 6,
  "n_rules": 0,
  "provenance": {
    "doc_id": "STD_2026_006_MNO",
    "span": "Standard Parameter ini merupakan pedoman dalam perhitungan parameter produksi",
    "page": 0,
    "section_no": "1",
    "chunk_id": "STD_2026_006_MNO::0000"
  }
}
```

| Field | Type | Req |
|---|---|---|
| `brief_id` | string | ✅ **New in v2.** `b_` prefix — reconciled with the code 2026-09-11 |
| `title` | string \| null | — |
| `purpose_verbatim` | string \| null | — **Renamed + guarded 2026-09-02** (was `purpose`) |
| `outline` | string[] | ✅ Default `[]`. **Added 2026-09-02** |
| `subdomains` | SubdomainEnum[] | ✅ Default `[]`. **Added 2026-09-02** |
| `n_terms` · `n_formulas` · `n_rules` | int | ✅ Default `0`. **Added 2026-09-02** |
| `key_parameters` | KeyParameter[] | ✅ Default `[]`. **Shape changed 2026-09-02** — was `string[]` |
| `provenance` | Provenance | ✅ |

**`KeyParameter`** — `{surface: string, term_id: string | null}`. `surface` is what the model named;
`term_id` is the glossary term that corroborates it, resolved deterministically by `link.py`. A
**null `term_id` means nothing in the extracted glossary supports that name** — reported, never
guessed, and it earns its own `review_queue.json` row (§9). Dangle count is reported as
`key_parameters_unresolved`.

**`subdomains`** is AGGREGATED from the glossary entries' own `subdomain_tags`, never asked of the
model — each tag was already chosen once, per term, with that term's evidence in front of the model,
and a second document-level classification would be the same judgement made with less context and no
way to check it. Ordered by tag frequency descending, alphabetical tiebreak, so two runs of one
document agree. **`n_terms` / `n_formulas` / `n_rules`** are counts of the other three artifacts,
computed after every branch and the diff have run.

**`outline`** is DERIVED from the artifact's `heading_path`, in reading order, de-duplicated. No LLM
and no spend. It carries the strong guarantee the rest of this entry cannot: it is verbatim source
structure, so there is nothing to hallucinate — including wording a reader would be tempted to fix
(the reference document heads a section *"Physical of Availability (PA)"*). Empty is normal for a
document with no headings.

**Removed in v2:** `scope` (overlaps `purpose`, frequently null) · `summary_md`.

> ✅ **This branch is span-checked as of 2026-09-02.** It used to be the one exception: it asked the
> model to SUMMARISE, and a summary is not verbatim by construction, so the primary
> anti-hallucination control could not apply. It now asks the model to **locate** — `purpose_verbatim`
> is the document's own statement of purpose, copied, not a sentence composed about it — so `title`
> and `purpose_verbatim` are guarded like any other transcription. A value that is not verbatim is
> nulled and recorded in `rejected.json`, never repaired.
>
> Measured on the reference document: **0 fields rejected**, and a deliberately fabricated purpose is
> refused with `purpose_verbatim is not a verbatim transcription of the source`.
>
> Two consequences worth knowing. **`title` will look wrong sometimes and that is correct** — the
> reference document yields `"BUMA | STANDARD PARAMETERProduction Parameter andEquipment Capacity
> Analysis (ECA)"`, missing spaces and all, because that is what the PDF says; tidying it is the
> UI's job at display time, never the pipeline's. And this weakens the case that the branch needs a
> larger model tier (DEV_PLAN §0.8 D3, risk R5): locating a sentence is a much smaller ask than
> composing one.

---

## 9 · `review_queue.json` → `ReviewQueueRow[]` — 12 fields

**This is the consumption surface** — the payload a review UI binds to. Deliberately denormalised so
a row renders with no joins.

> ⚠️ **Changed 2026-09-02: the queue is no longer glossary-only.** Every row now carries
> **`row_kind`**, and a consumer must switch on it rather than assume every row is a term. A
> `key_parameter` row has a surface and **no `term_id`, no definition and no mention count** — a
> client that binds `term_id` unconditionally will break on it.

```json
[
  {
    "rank": 1,
    "row_kind": "term",
    "term_id": "t_a4e2f81b06",
    "term": "PA",
    "definition": "Adalah ketersediaan fisik suatu equipment/unit yang menunjukkan proporsi waktu equipment/unit tersebut berada pada kondisi available (siap pakai) selama suatu periode tertentu.",
    "source_wording": "2.1.3. Physical of Availability (PA)",
    "mention_count": 43,
    "page": 3,
    "section_no": "2.1.3",
    "span": "Adalah ketersediaan fisik suatu equipment/unit",
    "review_reason": "source wording differs from the expanded name — confirm which is correct"
  },
  {
    "rank": 2,
    "row_kind": "key_parameter",
    "term_id": null,
    "brief_id": "d_2ef60a8c19",
    "term": "Grouping (Composite)",
    "definition": null,
    "source_wording": null,
    "mention_count": 0,
    "page": 0,
    "section_no": "1",
    "span": "Standard Parameter ini merupakan pedoman dalam perhitungan parameter produksi",
    "review_reason": "named as a key parameter but no extracted term corroborates it — confirm it is real"
  }
]
```

| Field | Type | Req | Notes |
|---|---|---|---|
| `rank` | int | ✅ | 1-based display order — see sorting below |
| `row_kind` | `term` \| `key_parameter` | ✅ | **New 2026-09-02.** Switch on this before reading anything else |
| `term_id` | string \| null | ✅ | **New in v2.** What an approve / edit / reject decision attaches to, so it survives a re-run. **Null on a `key_parameter` row** — that it resolves to no term is the reason the row exists |
| `brief_id` | string | — | `key_parameter` rows only — what the decision attaches to instead. **Named `brief_id` on the wire**, not `domain_id`; see the correction under Sort order |
| `surface_full` | string \| null | — | `key_parameter` rows only. **New 2026-09-11.** The untruncated surface when `term` had to be shortened to a head; `null` when it did not |
| `term` · `definition` · `source_wording` · `mention_count` | | | Copied from the entry |
| `page` · `page_no` · `section_no` · `span` | | | **Lifted out of `provenance`** so a row is self-contained. Render **`page_no`** (1-based), not `page` |
| `review_reason` | string | ✅ | Human-readable, one of five |

**Removed in v2:** `extraction_status` · `diff_status` · `definition_conflict` ·
`conflict_variants` — all four are already expressed by `review_reason`, and remain available on the
entry itself.

**Sort order** — **changed 2026-09-11.** `(conflicts, then mention_count descending, with unresolved
key parameters spliced in at most one per five term rows)`.

Conflicts still outrank everything, because a contradiction is a decision only the expert can make.
What changed is the tier beneath them. Unresolved key parameters used to occupy one, which meant
**every** unresolved parameter outranked **every** routine term regardless of how many there were —
and the branch produces them in bulk on exactly the documents it understands least. On the Komatsu
shop manual that was **24 junk rows standing in front of `Engine`, a term with 108 mentions**, which
sat at rank 25.

The original reason for promoting them is preserved and is still sound: that branch has **no span
check at all**, so an unresolvable surface is the one hallucination signal it can raise, and it must
not sink beneath sixty routine rows where nobody reaches it. A **cap** expresses that without letting
the tail take the queue over. Measured by replaying the three persisted v2 runs:

| Document | unresolved key parameters | first one at rank | top term, rank |
|---|---|---|---|
| BUMA | 0 | — | `Qty` 1 → 1 (unchanged) |
| Komatsu | 24 | 1 → **6** | `Engine` **25 → 1** |
| Open Pit | 7 | 1 → **6** | `Resource` **8 → 1** |

Row counts are identical before and after on all three — **nothing is dropped**; parameters the
splice cannot place are appended rather than discarded.

**Also new on `key_parameter` rows: `surface_full`.** The branch is supposed to return a parameter
*name* and on a textbook it returns the right concept wrapped in its whole defining sentence (185
characters, in one measured case). `term` now carries a head — split at the first `.`, `:` or newline
— and `surface_full` carries the untruncated text, or `null` when no shortening was needed. The head
is only taken past 60 characters, which is what keeps a legitimate short name containing a period
("No. of units") intact.

> ⚠️ **Field-name correction, 2026-09-11.** This document described the `key_parameter` row's owning
> id as **`domain_id`**. The code emits **`brief_id`** and always has: the K12 rename changed the
> entity from `BriefContext` to `DomainContext` but left the id field's name alone, deliberately —
> renaming it is a migration that invalidates every expert approval recorded against the old prefix.
> **The wire field is `brief_id`.** Corrected here rather than in code, per the trust order.

**`review_reason` values:**

| Value | Meaning |
|---|---|
| `conflicting definitions — expert decision required` | Contradiction found; the pipeline picked no winner |
| `source wording differs from the expanded name — confirm which is correct` | e.g. "Physical **of** Availability" vs "Physical Availability" |
| `term found but no definition in document` | Correct abstention — **48 of 58** on the reference run |
| `definition rejected by span check or absent` | Control fired; see `rejected.json` |
| `routine confirmation` | Nothing anomalous |
| `named as a key parameter but no extracted term corroborates it — confirm it is real` | `key_parameter` rows only. Added 2026-09-02 |

---

## 10 · Linking model

`GlossaryEntry` is the hub. Every link is computed **after** both branches run, by a deterministic
never-throw link stage. **No additional LLM calls, and nothing a model could fabricate.**

```
RuleEntry ──term_ids[]────────► GlossaryEntry ◄──variables[].term_id── FormulaEntry
     └────formula_ids[]──────────────────────────────────────────────────────► ▲
                          GlossaryEntry ──defining_formula_id───────────────────┘
```

Two properties that must hold:

1. **A dangling link is `null` or an empty array — never a fabrication.** Same rule as spans.
2. **The "appears in" edge is derived, not stored.** That PA appears inside
   `Production = MOHH × Qty × PA × UA × Pty` is recoverable by scanning `formulas[].variables[]`.
   Do not add a field for it.

---

## 11 · Audit payloads — unchanged from v1

**`rejected.json` → `RejectedField[]`** — what the span check *caught*, kept so a reviewer sees the
control working, not only what it let through. The reference run rejected **2 fields**.

```json
[
  {
    "entry_term": "<term>",
    "field": "full_name",
    "offending_value": "<value the model returned>",
    "reason": "full_name is not a verbatim transcription of the source",
    "branch": "glossary"
  }
]
```

| Field | Type | Req |
|---|---|---|
| `entry_term` | string | ✅ |
| `field` | string | ✅ |
| `offending_value` | string | ✅ |
| `reason` | string | ✅ |
| `branch` | `glossary` \| `rule` \| `formula` \| `summary` | ✅ |

**`usage.json` → `CallUsage[]`** — per-call accounting. `cached_tokens` is read from the API, never
modelled: caching does not engage below 1024 prompt tokens, so assuming it would understate input
cost by roughly 10×.

```json
[
  {
    "branch": "glossary",
    "deployment": "gpt-5.4-nano",
    "tier": "nano",
    "prompt_tokens": 2676,
    "cached_tokens": 2389,
    "completion_tokens": 214,
    "latency_s": 3.4,
    "retries": 0,
    "structured_output_mode": "json_schema",
    "simulated": false
  }
]
```

`simulated: true` marks a `--mock` run — **never a quality measurement.**

---

## 12 · Reference run — real figures

From `eval/knowledge/results/v2_full_document_2026-08-24_093051.json`. The parsing module built the
artifact; extraction ran all four branches.

| | |
|---|---|
| Artifact | `schema_version` 0.2.0 · **31 chunks** · 9 pages · `mineru/pipeline 3.4.4` |
| Funnel | 188 raw mentions → 163 after noise → **58 clusters** (compression 2.81×) |
| Entries | **58 glossary · 8 rules · 6 formulas** |
| LLM | 77 calls · 139,283 prompt tokens · **79.6% cached** · 168.4 s |
| Quality | **48 abstained**, 10 with a definition · **2 fields rejected** by the span check |
| E1 term-filter recall | **0.8049** (kill line 0.70, PASS) |
| E3 schema-fill precision | **0.90** (kill line 0.80, **PASS** — the prototype failed this at 0.75) |

> v1 of this contract quoted 13 chunks / 66 clusters / 66 entries. That run predates the
> section-aware chunker. **Anything calibrated against those figures should be re-checked.**
>
> These figures also predate the v2 schema itself: they were measured on the v1 prompts, which asked
> for fields that no longer exist. The funnel counts (chunks, clusters, calls) should carry over
> unchanged — the filter, cluster and rank stages were not touched — but **E3 must be re-scored**,
> because trimming a prompt can move definition quality even when the scored fields are unchanged.

---

## 13 · Open items

| # | Item | Owner |
|---|---|---|
| 1 | **Unevaluated.** Cutting `interpretation` / `formula_latex` and adding `rule_type` are prompt changes, and no eval has run against them — there is no document or parsed artifact on the dev box to run one. Same position X20 was committed in, and it needs the same first paid run to clear | Rifqi |
| 2 | **Rule-branch quality is a separate track.** The branch measures 3/15, an unevaluated prompt fix is already in flight (X20), and a prompt/gold contradiction caps it at 14/15 (X23). `rule_type` and the X23 de-contamination make this the *third and fourth* uncommitted-to-measurement changes on that one prompt — the first paid run measures the prompt as a whole, and isolating any single edit would cost one full run each. **Expect the raw score to move in both directions:** X20 should add rules back, X23 removes a free hit (example 1 was gold `PTY_PRODUCTION_SOURCE` verbatim) and lifts the 14/15 ceiling | Rifqi |
| 2b | **`rule.txt` grew again.** X20 already pushed it past the 1024-token cache floor; example 4 adds more. Only the API's `cached_tokens` proves a hit — verify on the first paid run | Rifqi |
| 3 | **Persistence.** Stage output is JSON on disk; parsed artifacts, candidate entries, glossary versions and the approval audit trail still need one consolidated DDL handoff. Go owns the schema — Python never executes DDL | Rifqi → Harry |
| 4 | **Page indexing at the API boundary** — stays 0-based to the UI, or converts once at the boundary? Currently 0-based everywhere | Harry + Rifqi |
| 5 | **Endpoint shape** — four endpoints, or one with an `?artifact=` parameter? Not started; the offline CLI is the honest first milestone | Rifqi |

---

## Caveat on the sample values

Field shapes, types and defaults are the **agreed v2 target**. Run-level figures in §12 are literal,
from the cited result file. The per-entity **sample values are illustrative** — the underlying
`out/*.json` from that run is not in version control, and all `*_id` values shown are placeholders
that demonstrate format, not real hashes. Replace them with literal output once v2 is implemented
and a run exists.