# Knowledge Pipeline — Output Contract **v2** > **Implemented 2026-08-26.** `src/knowledge_extraction/` emits this shape. > The schema review of 2026-08-26 landed with it: ids on every entity, formulas referenced rather > than restated, rules linked to formulas and terms, and 16 fields cut. > > **Not yet re-measured.** The prompt changes this required have not been scored — no source > document, parsed artifact or MinerU install exists on the dev box (all three `data/` directories > are gitignored and absent), so the free stages have nothing to read. The last measured figures in > §12 predate these edits. See §13. **Supersedes:** v1 of this document (2026-08-24) **Reference run:** `v2_full_document_2026-08-24_093051` · BUMA `STD_2026_006_MNO` · 9 pages · `gpt-5.4-nano` **Producer:** `python -m src.knowledge_extraction.cli --extract` **Status:** offline artifacts only — no endpoint, no table. --- ## 0 · What changed from v1 Harry's review of v1 asked for entity ids, a formula reference instead of a restated formula, and a link between rules and formulas. All are adopted. The schema was also cut back — v1 carried fields nothing consumes. | Change | Detail | |---|---| | **Ids on every entity** | `term_id`, `formula_id`, `brief_id`; `rule_id` re-specified. All deterministic and content-derived | | **Glossary no longer restates formulas** | `formula_latex` removed → `defining_formula_id` reference | | **Rules link to formulas and terms** | `formula_ids[]`, `term_ids[]` | | **`rule_type` added** | `interpretation` \| `calculation` — see §2 for why | | **16 fields cut, 8 added** | 50 → 42 fields total | | **`chunk_id` separator fixed** | `DOC#0007` → **`DOC::0007`**, adopting the parser's format | | **Reference numbers corrected** | v1 quoted a superseded run (13 chunks / 66 clusters). Current is **31 chunks / 58 clusters** | ### Why the ids had to be re-specified Neither id in v1 survived a re-run. `cluster_id` was a **frequency-rank ordinal** (`f"{doc_id}#c{i:03d}"` over a list sorted by mention count), and `rule_id` was **written by the model** (`rule.txt` asks for SCREAMING_SNAKE_CASE). Either way, re-running with a tuned prompt reshuffles the ids — and if an expert has approved 40 entries, we have lost which 40. **Approval continuity, not linking, is the real reason these exist.** --- ## 1 · Id scheme All ids are deterministic, content-derived, and stable across re-runs of the same document. None is model-generated; none is positional. ``` term_id = "t_" + sha256(doc_id + "|" + canonical)[:10] formula_id = "f_" + sha256(doc_id + "|" + name_or_latex_or_chunk_id)[:10] rule_id = "r_" + sha256(doc_id + "|" + chunk_id + "|" + char_start)[:10] brief_id = "b_" + sha256(doc_id)[:10] # per-document card; `d_` belongs to the # scope-level DomainContext in src/knowledge_domain/ ``` Implemented in `src/knowledge_extraction/ids.py`. `term_id` is keyed on the **cluster canonical**, not on the `term` the model returned — the cluster is the stable thing across runs, and the model is free to answer "PA" or "Physical Availability" for the same cluster. `formula_id` falls back through `name` → `formula_latex` → `chunk_id`, because both of the first two are Optional and without the fallback every unnamed formula in a document would collide. **Scope is per document.** The same term appearing in two documents receives two different `term_id`s. Deciding that BUMA's "PA" and a textbook's "PA" are the same concept is an **expert judgement at review time**, not something the pipeline asserts — the same reasoning that makes the pipeline record "Physical *of* Availability" instead of normalising it, and the reason MTTR at BUMA (breakdown duration ÷ breakdown frequency) must not silently merge with MTTR in IT. **Known limit.** Ids derive from content, so if content changes the id changes — retune clustering such that a term canonicalises differently and its `term_id` moves. The alternative is a persisted registry, which belongs with persistence (DEV_PLAN §0.8 D2). --- ## 2 · The four artifacts | File | Entity | Holds | |---|---|---| | `glossary.json` | `GlossaryEntry[]` | Terms and their definitions | | `rules.json` | `RuleEntry[]` | Rules of thumb — how to interpret data or behaviour. **Renamed from `interpretation_pack.json` 2026-09-08** — the file, the entity and the DB `kind` now share one name. *The Interpretation Pack* remains the deliverable's name in prose | | `formulas.json` | `FormulaEntry[]` | Formulas, transcribed from the document | | `document_brief.json` | `DocumentBrief` | One per-document card. Prefix history, not churn: `b_` (BriefContext) -> `d_` (DomainContext, 2026-09-01) -> **`b_` again** (DocumentBrief, 2026-09-08), once v4 gave `d_` to the real scope-level `DomainContext` in `src/knowledge_domain/` | Plus three derived/audit payloads: `review_queue.json`, `rejected.json`, `usage.json`. ### Rule vs domain knowledge — where the line is *Added 2026-09-14 (o2). The two blurred on screen during the 2026-09-08 review, and a boundary nobody has written down is one the prompts cannot hold.* | | **`rule`** | **`domain`** | |---|---|---| | Grain | One statement, from one document | One per company, composed across all of them | | Shape | A condition and a consequence | Identity, boundary, conventions, measures | | Evidence | Span-checked against its source | Derived from entries, or expert-declared | | Answers | *"…and how do I compute this correctly?"* | *"…what does this company measure, and what does it call it?"* | **The test to apply.** If removing it would change **how a number is computed or sourced**, it is a rule — *"if joint-survey data is unavailable, use truck count"* changes the figure. If removing it would only leave an agent **less oriented** — not knowing that this company works in overburden and BCM — it is domain. Two consequences worth stating, because they are what the distinction is *for*: - **A rule is actionable alone; domain knowledge is not.** A planner can apply a rule to a specific calculation. It cannot apply "this is a mining company" to anything — that shapes which questions make sense, not which column to sum. - **Domain is high-level and rules are specific, but "high-level" is not the test.** A long, general rule is still a rule. The grain question — *does this come from one document's sentence, or from the corpus as a whole?* — settles it faster and does not drift. Where it is genuinely both, prefer the **rule**: it carries provenance and a span, so an expert can check it. Domain fields are the ones the pipeline can least easily prove. ### Why `RuleEntry` carries a `rule_type` The Interpretation Pack exists to **improve analytics insight from expert interpretation rules** — interpretation logic, action benchmarks, and rules for when to conclude or not conclude from data. Measured against the reference document, **0 of the 15 gold rules are interpretation rules.** They are calculation conventions (`PA_COMPOSITE_WEIGHTED`), data-sourcing rules (`PTY_PRODUCTION_SOURCE`), and unit conventions (`GAINLOSS_UNITS`). This is a property of the document type: a *Standard Parameter* document defines how to **calculate** parameters, not how to **read** them. Interpretation rules live in other document types, or in the expert's head. `rule_type` keeps both in one artifact rather than discarding the conventions we do extract: - `interpretation` — how to read a value or behaviour. **The target.** - `calculation` — how a value is computed or sourced. What this corpus actually yields. A consumer wanting only interpretation logic filters on `rule_type`. --- ## 2c · Evidence now carries figures and tables *(parsing 0.4.0, 2026-09-10)* The artifact this half consumes changed shape: a figure is no longer its own chunk. It is folded into the prose chunk at its reading-order position as `![](asset://)`, and travels as an `Asset` on that chunk. `adapter.py` carries `assets`, `referenced_by` and `table_html` across; `evidence_block` renders them into the prompt under `ASSETS REFERENCED BY THIS CHUNK`. **What that means for an extraction branch**, and the distinction is enforced in the prompts: | | may be quoted? | may be a provenance span? | |---|---|---| | `Asset.caption` — printed in the document | **yes** | **yes** | | `Asset.description` — written by a vision model | no | **no** — it is not in the document, so the span check rejects it and the field is discarded | Why it mattered: before 0.4.0 a figure chunk had an EMPTY `text`, so it produced no mentions, joined no cluster, and was never selected as evidence. 13 figures across three documents, each already described by a paid vision call, reached **zero** prompts. Full detail: `knowledge_pipeline_runs/CHANGELOG_v2.md` §2. > ✅ **RESOLVED 2026-09-11 — §1, §2 and §8 now match the code.** The note is kept for the record. > > ⚠️ **Correction to §8 of this document (noted 2026-09-10).** The code writes > `document_brief.json` with **`brief_id`** and a `b_` prefix; §8 below documents > `domain_context.json` with `domain_id` and `d_`. The v3 rename was partly reverted for the > per-document object once the scope-level `DomainContext` became a separate thing in > `src/knowledge_domain/`. **§8 does not describe what the code emits.** Worth one edit, and it matters > because `entity_id` prefixes are the key an expert's approvals hang on. --- ## 3 · Transport reality Every payload is written by the CLI as a file into `--out-dir` (default `out/knowledge/`), as a **bare JSON array** with no envelope. There is no endpoint and no table. Three conventions hold across every payload: - **`page` is 0-based**, exactly as the parser reports; **`page_no` is the 1-based number a human reads** and is what a review UI binds to. Both travel together on every entity, derived from one value so they cannot disagree (S6b, 2026-09-02). - **All content fields are nullable by design.** `null` is a valid, correct answer — the model abstaining, not failing. On the reference run **48 of 58** entries carry no definition. - **`provenance` is mandatory and `provenance.span` is verbatim-checked.** A field whose span cannot be located in the source is set to `null` and logged to `rejected.json` — **never repaired.** --- ## 4 · Common types ### `Provenance` — required on every entity ```json { "doc_id": "STD_2026_006_MNO", "span": "Adalah jumlah equipment/unit yang digunakan untuk menghasilkan produksi dalam suatu periode.", "page": 3, "page_no": 4, "section_no": "2.1.2", "chunk_id": "STD_2026_006_MNO::0007" } ``` | Field | Type | Req | Notes | |---|---|---|---| | `doc_id` | string | ✅ | | | `span` | string | ✅ | **Verbatim** from the source. The anti-hallucination control — the reviewer checks a quote against a page, not a claim against memory | | `page` | int \| null | — | **0-based**, exactly as the parser reports | | `page_no` | int \| null | — | **1-based. Added 2026-09-02 (S6b).** DERIVED from `page`, never stored, so it cannot drift. **A review UI must render this one** — an off-by-one is invisible until an expert opens the wrong page and concludes the provenance is wrong. `null` when `page` is unknown | | `section_no` | string \| null | — | e.g. `"2.1.2"`; null when the document is unnumbered | | `chunk_id` | string \| null | — | Back-reference into the parsed artifact. **Format `::`**, matching `KNOWLEDGE_PARSING_OUTPUT_CONTRACT.md`. Treat as opaque — do not split it | ### Enums | Enum | Values | |---|---| | `subdomain_tags[]` | `production` · `maintenance` · `hauling` · `loading` · `drilling_blasting` · `equipment` · `safety` · `quality` · `planning` · `cost` · `geology` · `other` | | `extraction_status` | `ok` · `no_definition_found` · `escalated` | | `diff_status` | `new` · `duplicate` · `conflicting` | | `rule_type` | `interpretation` · `calculation` | | `latex_verification` | `verified` · `unverified_no_markup` · `unverified_operator` | `subdomain_tags` is a **closed enum — classification, not generation.** Adding a member changes what the model is allowed to answer, which makes it a prompt change, not a data change. --- ## 5 · `glossary.json` → `GlossaryEntry[]` — 13 fields ```json [ { "term_id": "t_9f2a41c0b7", "term": "Qty", "full_name": "Quantity", "source_wording": "2.1.2. Quantity (Qty)", "definition": "Adalah jumlah equipment/unit yang digunakan untuk menghasilkan produksi dalam suatu periode.", "defining_formula_id": "f_31de08aa95", "subdomain_tags": ["production", "equipment"], "mention_count": 20, "provenance": { "doc_id": "STD_2026_006_MNO", "span": "Adalah jumlah equipment/unit yang digunakan untuk menghasilkan produksi dalam suatu periode.", "page": 3, "section_no": "2.1.2", "chunk_id": "STD_2026_006_MNO::0007" }, "extraction_status": "ok", "diff_status": "new", "definition_conflict": false, "conflict_variants": [] }, { "term_id": "t_5b71ce0d34", "term": "Gain/Loss", "full_name": null, "source_wording": null, "definition": null, "defining_formula_id": null, "subdomain_tags": ["production"], "mention_count": 4, "provenance": { "doc_id": "STD_2026_006_MNO", "span": "Gain/Loss", "page": 6, "section_no": null, "chunk_id": "STD_2026_006_MNO::0011" }, "extraction_status": "no_definition_found", "diff_status": "new", "definition_conflict": false, "conflict_variants": [] } ] ``` | Field | Type | Req | Notes | |---|---|---|---| | `term_id` | string | ✅ | **New in v2.** Stable across re-runs — the key an approval hangs on | | `term` | string | ✅ | The only required content field | | `full_name` | string \| null | — | Expanded name. Span-guarded as a literal transcription | | `source_wording` | string \| null | — | **The document's literal wording**, un-normalised. Read deterministically from the section heading, *not* asked of the model — asked directly, it returned the tidied form | | `definition` | string \| null | — | `null` = correct abstention (**48 of 58** on the reference run) | | `defining_formula_id` | string \| null | — | **New in v2**, replaces `formula_latex`. The formula that *defines* this term. Null when the formula branch did not extract one | | `subdomain_tags` | enum[] | ✅ | Defaults `[]` | | `mention_count` | int | ✅ | Default `0`. **Drives review-queue ordering** | | `provenance` | Provenance | ✅ | | | `extraction_status` | enum | ✅ | Default `ok` | | `diff_status` | enum \| null | — | Set by the diff stage against the active glossary version | | `definition_conflict` | bool | ✅ | Default `false` | | `conflict_variants` | string[] | ✅ | Default `[]`. The competing definitions — the pipeline **never picks a winner** | **Removed in v2:** `formula_latex` (→ `defining_formula_id`) · `interpretation` (belongs in the Interpretation Pack, reached via `term_id`) · `domain` · `company` · `language`. > `domain` and `company` were never in the extraction prompt's field guide — the model filled them > by copying worked examples — and neither was span-guarded, so a wrong value was undetectable. > `language` is a genuine three-value field (`id`/`en`/`mixed`) but nothing consumes it and it is > recoverable from the document; it is cut for simplicity and can return if a consumer needs it. > **Writer split.** The LLM fills only `term`, `full_name`, `definition` and `subdomain_tags`. > `term_id`, `defining_formula_id`, `source_wording`, `mention_count`, `extraction_status`, > `diff_status`, `definition_conflict` and `conflict_variants` are all set by deterministic code — > **the model never writes its own ids, links or audit fields**, so none of them can be hallucinated. --- ## 6 · `rules.json` → `RuleEntry[]` — 8 fields ```json [ { "rule_id": "r_c07be41f22", "rule_type": "calculation", "statement": "Production yang digunakan dalam perhitungan adalah produksi hasil joint survey.", "condition": "Apabila data joint survey belum tersedia", "consequence": "maka digunakan data produksi berdasarkan truck count sebagai dasar perhitungan", "formula_ids": ["f_77c1ab9e40"], "term_ids": ["t_a4e2f81b06"], "provenance": { "doc_id": "STD_2026_006_MNO", "span": "Apabila data joint survey belum tersedia maka digunakan data produksi berdasarkan truck count sebagai dasar perhitungan", "page": 4, "section_no": "2.1.5", "chunk_id": "STD_2026_006_MNO::0009" } } ] ``` | Field | Type | Req | Notes | |---|---|---|---| | `rule_id` | string | ✅ | **Re-specified in v2.** Was written by the model, therefore unstable; now deterministic | | `rule_type` | enum | ✅ | `interpretation` \| `calculation` — see §2 | | `statement` | string \| null | — | The rule as prose. **Prose, not an equation** — evidence carrying both yields the prose, and the equation goes to the formula branch | | `condition` | string \| null | — | The triggering condition, split out — not left buried in prose | | `consequence` | string \| null | — | What follows when the condition holds | | `formula_ids` | string[] | ✅ | **New in v2.** The formulas this rule constrains. Default `[]` | | `term_ids` | string[] | ✅ | **New in v2.** The terms this rule governs. Default `[]` | | `provenance` | Provenance | ✅ | | **Removed in v2:** `applies_to` (free text → `term_ids[]`) · `subdomain_tags` · `language` · `extraction_status`. `condition` + `consequence` is the point of this artifact, and the seed of the interpretation logic tree: a rule stored as one paragraph cannot be attached to a skill later; a trigger and its consequence can. > **Deliberately not modelled yet:** the interpretation logic tree, action benchmarks, and explicit > conclude / do-not-conclude verdicts. Those are the direction of travel, but designing their schema > against zero real examples would be guessing. They land when a document containing them does. --- ## 7 · `formulas.json` → `FormulaEntry[]` — 7 fields ```json [ { "formula_id": "f_77c1ab9e40", "name": "Production", "formula_latex": "Production = MOHH \\times Qty \\times PA \\times UA \\times Pty", "variables": [ { "symbol": "MOHH", "term_id": "t_1c9d5e7a83", "meaning": "Machine on Hand Hours" }, { "symbol": "Qty", "term_id": "t_9f2a41c0b7", "meaning": "Quantity" }, { "symbol": "PA", "term_id": "t_a4e2f81b06", "meaning": "Physical Availability" }, { "symbol": "UA", "term_id": "t_6d40b2fc17", "meaning": "Utilization of Availability" }, { "symbol": "Pty", "term_id": "t_e8b3092d55", "meaning": "Productivity" } ], "unit": "BCM", "latex_verification": "unverified_operator", "provenance": { "doc_id": "STD_2026_006_MNO", "span": "Production = MOHH x Qty x PA x UA x Pty", "page": 1, "section_no": "2.1", "chunk_id": "STD_2026_006_MNO::0003" } } ] ``` | Field | Type | Req | Notes | |---|---|---|---| | `formula_id` | string | ✅ | **New in v2** | | `name` | string \| null | — | What the formula computes, as named in the source | | `formula_latex` | string \| null | — | **Transcribed, never derived.** The pipeline copies what the document states | | `variables[]` | object[] | ✅ | `{ symbol (required), term_id \| null, meaning \| null }`. `term_id` is **new in v2** | | `unit` | string \| null | — | Result unit if stated | | `latex_verification` | enum \| null | — | **Added 2026-08-26.** Which guarantee `formula_latex` actually carries. Set by the span check, never by the model. `null` when there is no `formula_latex` to characterise | | `provenance` | Provenance | ✅ | | **Removed in v2:** `extraction_status`. > **Read `latex_verification` before trusting `formula_latex`.** Three paths leave the field populated and > they are not equally strong. `verified` means the transcription was located in the source markup > (`Chunk.latex`), compared in a notation-aware canonical form. The two `unverified_*` values mean the > claim could not be checked either way and the entry is guarded by its provenance span alone — the same > guarantee everything carried before 2026-08-26. > > `unverified_operator` is not rare. MinerU's **`pipeline`** backend transcribes multiplication as a bare > letter `x`, so every product formula from such an artifact lands here — including the sample above. > An artifact parsed with `hybrid/high` carries `\times` and verifies normally. > `variables[].symbol` is not always a legend abbreviation — the extraction prompt's own worked > example emits `"Total Hours"` and `"Breakdown"` as symbols. So `term_id` resolution is > **best-effort and nulls are expected**; the link stage reports its dangle rate rather than hiding it. --- ## 8 · `document_brief.json` → `DocumentBrief` — 6 fields *(single object)* > ✅ **The v3 reshape is complete, corrected 2026-09-08.** This banner previously said the > LLM-facing parts — `purpose` as a span-guarded verbatim quote, and the `subdomains` aggregate — > were *"still proposed and not built"*. They shipped on 2026-09-02 and the banner was never > updated, so it contradicted the field table directly below it, which dates both changes. Verified > live: a `domain` entry persisted on 2026-09-07 carries `purpose_verbatim` and `subdomains`. > > The delivery really was split, and that part was deliberate: the deterministic half (`outline` > read off the artifact, `key_parameters` corroborated against extracted terms) shipped first > because it costs nothing to prove, and the LLM-facing half followed once an eval run could pay > for itself. ```json { "brief_id": "d_2ef60a8c19", "title": "Standard Parameter Produksi & ECA", "purpose_verbatim": "Standard Parameter ini merupakan pedoman dalam perhitungan parameter produksi yang meliputi Quantity (Qty) Physical Availability (PA), Utilization of Availability (UA), Productivity (PTY), serta Equipment Capacity Analysis (ECA).", "outline": [ "1. TUJUAN PARAMETER", "2.1.2. Quantity (Qty)", "2.1.3. Physical of Availability (PA)", "2.2. Equipment Capacity Analysis (ECA)" ], "key_parameters": [ {"surface": "Physical Availability (PA)", "term_id": "t_8c7359e27f"}, {"surface": "Equipment Capacity Analysis (ECA)", "term_id": null} ], "subdomains": ["equipment", "production"], "n_terms": 5, "n_formulas": 6, "n_rules": 0, "provenance": { "doc_id": "STD_2026_006_MNO", "span": "Standard Parameter ini merupakan pedoman dalam perhitungan parameter produksi", "page": 0, "section_no": "1", "chunk_id": "STD_2026_006_MNO::0000" } } ``` | Field | Type | Req | |---|---|---| | `brief_id` | string | ✅ **New in v2.** `b_` prefix — reconciled with the code 2026-09-11 | | `title` | string \| null | — | | `purpose_verbatim` | string \| null | — **Renamed + guarded 2026-09-02** (was `purpose`) | | `outline` | string[] | ✅ Default `[]`. **Added 2026-09-02** | | `subdomains` | SubdomainEnum[] | ✅ Default `[]`. **Added 2026-09-02** | | `n_terms` · `n_formulas` · `n_rules` | int | ✅ Default `0`. **Added 2026-09-02** | | `key_parameters` | KeyParameter[] | ✅ Default `[]`. **Shape changed 2026-09-02** — was `string[]` | | `provenance` | Provenance | ✅ | **`KeyParameter`** — `{surface: string, term_id: string | null}`. `surface` is what the model named; `term_id` is the glossary term that corroborates it, resolved deterministically by `link.py`. A **null `term_id` means nothing in the extracted glossary supports that name** — reported, never guessed, and it earns its own `review_queue.json` row (§9). Dangle count is reported as `key_parameters_unresolved`. **`subdomains`** is AGGREGATED from the glossary entries' own `subdomain_tags`, never asked of the model — each tag was already chosen once, per term, with that term's evidence in front of the model, and a second document-level classification would be the same judgement made with less context and no way to check it. Ordered by tag frequency descending, alphabetical tiebreak, so two runs of one document agree. **`n_terms` / `n_formulas` / `n_rules`** are counts of the other three artifacts, computed after every branch and the diff have run. **`outline`** is DERIVED from the artifact's `heading_path`, in reading order, de-duplicated. No LLM and no spend. It carries the strong guarantee the rest of this entry cannot: it is verbatim source structure, so there is nothing to hallucinate — including wording a reader would be tempted to fix (the reference document heads a section *"Physical of Availability (PA)"*). Empty is normal for a document with no headings. **Removed in v2:** `scope` (overlaps `purpose`, frequently null) · `summary_md`. > ✅ **This branch is span-checked as of 2026-09-02.** It used to be the one exception: it asked the > model to SUMMARISE, and a summary is not verbatim by construction, so the primary > anti-hallucination control could not apply. It now asks the model to **locate** — `purpose_verbatim` > is the document's own statement of purpose, copied, not a sentence composed about it — so `title` > and `purpose_verbatim` are guarded like any other transcription. A value that is not verbatim is > nulled and recorded in `rejected.json`, never repaired. > > Measured on the reference document: **0 fields rejected**, and a deliberately fabricated purpose is > refused with `purpose_verbatim is not a verbatim transcription of the source`. > > Two consequences worth knowing. **`title` will look wrong sometimes and that is correct** — the > reference document yields `"BUMA | STANDARD PARAMETERProduction Parameter andEquipment Capacity > Analysis (ECA)"`, missing spaces and all, because that is what the PDF says; tidying it is the > UI's job at display time, never the pipeline's. And this weakens the case that the branch needs a > larger model tier (DEV_PLAN §0.8 D3, risk R5): locating a sentence is a much smaller ask than > composing one. --- ## 9 · `review_queue.json` → `ReviewQueueRow[]` — 12 fields **This is the consumption surface** — the payload a review UI binds to. Deliberately denormalised so a row renders with no joins. > ⚠️ **Changed 2026-09-02: the queue is no longer glossary-only.** Every row now carries > **`row_kind`**, and a consumer must switch on it rather than assume every row is a term. A > `key_parameter` row has a surface and **no `term_id`, no definition and no mention count** — a > client that binds `term_id` unconditionally will break on it. ```json [ { "rank": 1, "row_kind": "term", "term_id": "t_a4e2f81b06", "term": "PA", "definition": "Adalah ketersediaan fisik suatu equipment/unit yang menunjukkan proporsi waktu equipment/unit tersebut berada pada kondisi available (siap pakai) selama suatu periode tertentu.", "source_wording": "2.1.3. Physical of Availability (PA)", "mention_count": 43, "page": 3, "section_no": "2.1.3", "span": "Adalah ketersediaan fisik suatu equipment/unit", "review_reason": "source wording differs from the expanded name — confirm which is correct" }, { "rank": 2, "row_kind": "key_parameter", "term_id": null, "brief_id": "d_2ef60a8c19", "term": "Grouping (Composite)", "definition": null, "source_wording": null, "mention_count": 0, "page": 0, "section_no": "1", "span": "Standard Parameter ini merupakan pedoman dalam perhitungan parameter produksi", "review_reason": "named as a key parameter but no extracted term corroborates it — confirm it is real" } ] ``` | Field | Type | Req | Notes | |---|---|---|---| | `rank` | int | ✅ | 1-based display order — see sorting below | | `row_kind` | `term` \| `key_parameter` | ✅ | **New 2026-09-02.** Switch on this before reading anything else | | `term_id` | string \| null | ✅ | **New in v2.** What an approve / edit / reject decision attaches to, so it survives a re-run. **Null on a `key_parameter` row** — that it resolves to no term is the reason the row exists | | `brief_id` | string | — | `key_parameter` rows only — what the decision attaches to instead. **Named `brief_id` on the wire**, not `domain_id`; see the correction under Sort order | | `surface_full` | string \| null | — | `key_parameter` rows only. **New 2026-09-11.** The untruncated surface when `term` had to be shortened to a head; `null` when it did not | | `term` · `definition` · `source_wording` · `mention_count` | | | Copied from the entry | | `page` · `page_no` · `section_no` · `span` | | | **Lifted out of `provenance`** so a row is self-contained. Render **`page_no`** (1-based), not `page` | | `review_reason` | string | ✅ | Human-readable, one of five | **Removed in v2:** `extraction_status` · `diff_status` · `definition_conflict` · `conflict_variants` — all four are already expressed by `review_reason`, and remain available on the entry itself. **Sort order** — **changed 2026-09-11.** `(conflicts, then mention_count descending, with unresolved key parameters spliced in at most one per five term rows)`. Conflicts still outrank everything, because a contradiction is a decision only the expert can make. What changed is the tier beneath them. Unresolved key parameters used to occupy one, which meant **every** unresolved parameter outranked **every** routine term regardless of how many there were — and the branch produces them in bulk on exactly the documents it understands least. On the Komatsu shop manual that was **24 junk rows standing in front of `Engine`, a term with 108 mentions**, which sat at rank 25. The original reason for promoting them is preserved and is still sound: that branch has **no span check at all**, so an unresolvable surface is the one hallucination signal it can raise, and it must not sink beneath sixty routine rows where nobody reaches it. A **cap** expresses that without letting the tail take the queue over. Measured by replaying the three persisted v2 runs: | Document | unresolved key parameters | first one at rank | top term, rank | |---|---|---|---| | BUMA | 0 | — | `Qty` 1 → 1 (unchanged) | | Komatsu | 24 | 1 → **6** | `Engine` **25 → 1** | | Open Pit | 7 | 1 → **6** | `Resource` **8 → 1** | Row counts are identical before and after on all three — **nothing is dropped**; parameters the splice cannot place are appended rather than discarded. **Also new on `key_parameter` rows: `surface_full`.** The branch is supposed to return a parameter *name* and on a textbook it returns the right concept wrapped in its whole defining sentence (185 characters, in one measured case). `term` now carries a head — split at the first `.`, `:` or newline — and `surface_full` carries the untruncated text, or `null` when no shortening was needed. The head is only taken past 60 characters, which is what keeps a legitimate short name containing a period ("No. of units") intact. > ⚠️ **Field-name correction, 2026-09-11.** This document described the `key_parameter` row's owning > id as **`domain_id`**. The code emits **`brief_id`** and always has: the K12 rename changed the > entity from `BriefContext` to `DomainContext` but left the id field's name alone, deliberately — > renaming it is a migration that invalidates every expert approval recorded against the old prefix. > **The wire field is `brief_id`.** Corrected here rather than in code, per the trust order. **`review_reason` values:** | Value | Meaning | |---|---| | `conflicting definitions — expert decision required` | Contradiction found; the pipeline picked no winner | | `source wording differs from the expanded name — confirm which is correct` | e.g. "Physical **of** Availability" vs "Physical Availability" | | `term found but no definition in document` | Correct abstention — **48 of 58** on the reference run | | `definition rejected by span check or absent` | Control fired; see `rejected.json` | | `routine confirmation` | Nothing anomalous | | `named as a key parameter but no extracted term corroborates it — confirm it is real` | `key_parameter` rows only. Added 2026-09-02 | --- ## 10 · Linking model `GlossaryEntry` is the hub. Every link is computed **after** both branches run, by a deterministic never-throw link stage. **No additional LLM calls, and nothing a model could fabricate.** ``` RuleEntry ──term_ids[]────────► GlossaryEntry ◄──variables[].term_id── FormulaEntry └────formula_ids[]──────────────────────────────────────────────────────► ▲ GlossaryEntry ──defining_formula_id───────────────────┘ ``` Two properties that must hold: 1. **A dangling link is `null` or an empty array — never a fabrication.** Same rule as spans. 2. **The "appears in" edge is derived, not stored.** That PA appears inside `Production = MOHH × Qty × PA × UA × Pty` is recoverable by scanning `formulas[].variables[]`. Do not add a field for it. --- ## 11 · Audit payloads — unchanged from v1 **`rejected.json` → `RejectedField[]`** — what the span check *caught*, kept so a reviewer sees the control working, not only what it let through. The reference run rejected **2 fields**. ```json [ { "entry_term": "", "field": "full_name", "offending_value": "", "reason": "full_name is not a verbatim transcription of the source", "branch": "glossary" } ] ``` | Field | Type | Req | |---|---|---| | `entry_term` | string | ✅ | | `field` | string | ✅ | | `offending_value` | string | ✅ | | `reason` | string | ✅ | | `branch` | `glossary` \| `rule` \| `formula` \| `summary` | ✅ | **`usage.json` → `CallUsage[]`** — per-call accounting. `cached_tokens` is read from the API, never modelled: caching does not engage below 1024 prompt tokens, so assuming it would understate input cost by roughly 10×. ```json [ { "branch": "glossary", "deployment": "gpt-5.4-nano", "tier": "nano", "prompt_tokens": 2676, "cached_tokens": 2389, "completion_tokens": 214, "latency_s": 3.4, "retries": 0, "structured_output_mode": "json_schema", "simulated": false } ] ``` `simulated: true` marks a `--mock` run — **never a quality measurement.** --- ## 12 · Reference run — real figures From `eval/knowledge/results/v2_full_document_2026-08-24_093051.json`. The parsing module built the artifact; extraction ran all four branches. | | | |---|---| | Artifact | `schema_version` 0.2.0 · **31 chunks** · 9 pages · `mineru/pipeline 3.4.4` | | Funnel | 188 raw mentions → 163 after noise → **58 clusters** (compression 2.81×) | | Entries | **58 glossary · 8 rules · 6 formulas** | | LLM | 77 calls · 139,283 prompt tokens · **79.6% cached** · 168.4 s | | Quality | **48 abstained**, 10 with a definition · **2 fields rejected** by the span check | | E1 term-filter recall | **0.8049** (kill line 0.70, PASS) | | E3 schema-fill precision | **0.90** (kill line 0.80, **PASS** — the prototype failed this at 0.75) | > v1 of this contract quoted 13 chunks / 66 clusters / 66 entries. That run predates the > section-aware chunker. **Anything calibrated against those figures should be re-checked.** > > These figures also predate the v2 schema itself: they were measured on the v1 prompts, which asked > for fields that no longer exist. The funnel counts (chunks, clusters, calls) should carry over > unchanged — the filter, cluster and rank stages were not touched — but **E3 must be re-scored**, > because trimming a prompt can move definition quality even when the scored fields are unchanged. --- ## 13 · Open items | # | Item | Owner | |---|---|---| | 1 | **Unevaluated.** Cutting `interpretation` / `formula_latex` and adding `rule_type` are prompt changes, and no eval has run against them — there is no document or parsed artifact on the dev box to run one. Same position X20 was committed in, and it needs the same first paid run to clear | Rifqi | | 2 | **Rule-branch quality is a separate track.** The branch measures 3/15, an unevaluated prompt fix is already in flight (X20), and a prompt/gold contradiction caps it at 14/15 (X23). `rule_type` and the X23 de-contamination make this the *third and fourth* uncommitted-to-measurement changes on that one prompt — the first paid run measures the prompt as a whole, and isolating any single edit would cost one full run each. **Expect the raw score to move in both directions:** X20 should add rules back, X23 removes a free hit (example 1 was gold `PTY_PRODUCTION_SOURCE` verbatim) and lifts the 14/15 ceiling | Rifqi | | 2b | **`rule.txt` grew again.** X20 already pushed it past the 1024-token cache floor; example 4 adds more. Only the API's `cached_tokens` proves a hit — verify on the first paid run | Rifqi | | 3 | **Persistence.** Stage output is JSON on disk; parsed artifacts, candidate entries, glossary versions and the approval audit trail still need one consolidated DDL handoff. Go owns the schema — Python never executes DDL | Rifqi → Harry | | 4 | **Page indexing at the API boundary** — stays 0-based to the UI, or converts once at the boundary? Currently 0-based everywhere | Harry + Rifqi | | 5 | **Endpoint shape** — four endpoints, or one with an `?artifact=` parameter? Not started; the offline CLI is the honest first milestone | Rifqi | --- ## Caveat on the sample values Field shapes, types and defaults are the **agreed v2 target**. Run-level figures in §12 are literal, from the cited result file. The per-entity **sample values are illustrative** — the underlying `out/*.json` from that run is not in version control, and all `*_id` values shown are placeholders that demonstrate format, not real hashes. Replace them with literal output once v2 is implemented and a run exists.