|
Download docs/knowledge/KNOWLEDGE_OUTPUT_CONTRACT.md from DataEyond/Agentic-Service-Data-Eyond-Catalog: direct link, hf CLI and curl.
- Browser
- Download file 38.4 kB
-
https://huggingface.co/spaces/DataEyond/Agentic-Service-Data-Eyond-Catalog/resolve/main/docs/knowledge/KNOWLEDGE_OUTPUT_CONTRACT.md
- Command line
-
hf download hf://spaces/DataEyond/Agentic-Service-Data-Eyond-Catalog/docs/knowledge/KNOWLEDGE_OUTPUT_CONTRACT.md
-
curl -L -o KNOWLEDGE_OUTPUT_CONTRACT.md https://huggingface.co/spaces/DataEyond/Agentic-Service-Data-Eyond-Catalog/resolve/main/docs/knowledge/KNOWLEDGE_OUTPUT_CONTRACT.md
38.4 kB
| # Knowledge Pipeline β Output Contract **v2** | |
| > **Implemented 2026-08-26.** `src/knowledge_extraction/` emits this shape. | |
| > The schema review of 2026-08-26 landed with it: ids on every entity, formulas referenced rather | |
| > than restated, rules linked to formulas and terms, and 16 fields cut. | |
| > | |
| > **Not yet re-measured.** The prompt changes this required have not been scored β no source | |
| > document, parsed artifact or MinerU install exists on the dev box (all three `data/` directories | |
| > are gitignored and absent), so the free stages have nothing to read. The last measured figures in | |
| > Β§12 predate these edits. See Β§13. | |
| **Supersedes:** v1 of this document (2026-08-24) | |
| **Reference run:** `v2_full_document_2026-08-24_093051` Β· BUMA `STD_2026_006_MNO` Β· 9 pages Β· `gpt-5.4-nano` | |
| **Producer:** `python -m src.knowledge_extraction.cli <artifact.json> --extract` | |
| **Status:** offline artifacts only β no endpoint, no table. | |
| --- | |
| ## 0 Β· What changed from v1 | |
| Harry's review of v1 asked for entity ids, a formula reference instead of a restated formula, and a | |
| link between rules and formulas. All are adopted. The schema was also cut back β v1 carried fields | |
| nothing consumes. | |
| | Change | Detail | | |
| |---|---| | |
| | **Ids on every entity** | `term_id`, `formula_id`, `brief_id`; `rule_id` re-specified. All deterministic and content-derived | | |
| | **Glossary no longer restates formulas** | `formula_latex` removed β `defining_formula_id` reference | | |
| | **Rules link to formulas and terms** | `formula_ids[]`, `term_ids[]` | | |
| | **`rule_type` added** | `interpretation` \| `calculation` β see Β§2 for why | | |
| | **16 fields cut, 8 added** | 50 β 42 fields total | | |
| | **`chunk_id` separator fixed** | `DOC#0007` β **`DOC::0007`**, adopting the parser's format | | |
| | **Reference numbers corrected** | v1 quoted a superseded run (13 chunks / 66 clusters). Current is **31 chunks / 58 clusters** | | |
| ### Why the ids had to be re-specified | |
| Neither id in v1 survived a re-run. `cluster_id` was a **frequency-rank ordinal** | |
| (`f"{doc_id}#c{i:03d}"` over a list sorted by mention count), and `rule_id` was **written by the | |
| model** (`rule.txt` asks for SCREAMING_SNAKE_CASE). Either way, re-running with a tuned prompt | |
| reshuffles the ids β and if an expert has approved 40 entries, we have lost which 40. | |
| **Approval continuity, not linking, is the real reason these exist.** | |
| --- | |
| ## 1 Β· Id scheme | |
| All ids are deterministic, content-derived, and stable across re-runs of the same document. | |
| None is model-generated; none is positional. | |
| ``` | |
| term_id = "t_" + sha256(doc_id + "|" + canonical)[:10] | |
| formula_id = "f_" + sha256(doc_id + "|" + name_or_latex_or_chunk_id)[:10] | |
| rule_id = "r_" + sha256(doc_id + "|" + chunk_id + "|" + char_start)[:10] | |
| brief_id = "b_" + sha256(doc_id)[:10] # per-document card; `d_` belongs to the | |
| # scope-level DomainContext in src/knowledge_domain/ | |
| ``` | |
| Implemented in `src/knowledge_extraction/ids.py`. `term_id` is keyed on the **cluster canonical**, | |
| not on the `term` the model returned β the cluster is the stable thing across runs, and the model is | |
| free to answer "PA" or "Physical Availability" for the same cluster. `formula_id` falls back through | |
| `name` β `formula_latex` β `chunk_id`, because both of the first two are Optional and without the | |
| fallback every unnamed formula in a document would collide. | |
| **Scope is per document.** The same term appearing in two documents receives two different | |
| `term_id`s. Deciding that BUMA's "PA" and a textbook's "PA" are the same concept is an **expert | |
| judgement at review time**, not something the pipeline asserts β the same reasoning that makes the | |
| pipeline record "Physical *of* Availability" instead of normalising it, and the reason MTTR at BUMA | |
| (breakdown duration Γ· breakdown frequency) must not silently merge with MTTR in IT. | |
| **Known limit.** Ids derive from content, so if content changes the id changes β retune clustering | |
| such that a term canonicalises differently and its `term_id` moves. The alternative is a persisted | |
| registry, which belongs with persistence (DEV_PLAN Β§0.8 D2). | |
| --- | |
| ## 2 Β· The four artifacts | |
| | File | Entity | Holds | | |
| |---|---|---| | |
| | `glossary.json` | `GlossaryEntry[]` | Terms and their definitions | | |
| | `rules.json` | `RuleEntry[]` | Rules of thumb β how to interpret data or behaviour. **Renamed from `interpretation_pack.json` 2026-09-08** β the file, the entity and the DB `kind` now share one name. *The Interpretation Pack* remains the deliverable's name in prose | | |
| | `formulas.json` | `FormulaEntry[]` | Formulas, transcribed from the document | | |
| | `document_brief.json` | `DocumentBrief` | One per-document card. Prefix history, not churn: `b_` (BriefContext) -> `d_` (DomainContext, 2026-09-01) -> **`b_` again** (DocumentBrief, 2026-09-08), once v4 gave `d_` to the real scope-level `DomainContext` in `src/knowledge_domain/` | | |
| Plus three derived/audit payloads: `review_queue.json`, `rejected.json`, `usage.json`. | |
| ### Rule vs domain knowledge β where the line is | |
| *Added 2026-09-14 (o2). The two blurred on screen during the 2026-09-08 review, | |
| and a boundary nobody has written down is one the prompts cannot hold.* | |
| | | **`rule`** | **`domain`** | | |
| |---|---|---| | |
| | Grain | One statement, from one document | One per company, composed across all of them | | |
| | Shape | A condition and a consequence | Identity, boundary, conventions, measures | | |
| | Evidence | Span-checked against its source | Derived from entries, or expert-declared | | |
| | Answers | *"β¦and how do I compute this correctly?"* | *"β¦what does this company measure, and what does it call it?"* | | |
| **The test to apply.** If removing it would change **how a number is computed | |
| or sourced**, it is a rule β *"if joint-survey data is unavailable, use truck | |
| count"* changes the figure. If removing it would only leave an agent **less | |
| oriented** β not knowing that this company works in overburden and BCM β it is | |
| domain. | |
| Two consequences worth stating, because they are what the distinction is *for*: | |
| - **A rule is actionable alone; domain knowledge is not.** A planner can apply | |
| a rule to a specific calculation. It cannot apply "this is a mining company" | |
| to anything β that shapes which questions make sense, not which column to sum. | |
| - **Domain is high-level and rules are specific, but "high-level" is not the | |
| test.** A long, general rule is still a rule. The grain question β *does this | |
| come from one document's sentence, or from the corpus as a whole?* β settles | |
| it faster and does not drift. | |
| Where it is genuinely both, prefer the **rule**: it carries provenance and a | |
| span, so an expert can check it. Domain fields are the ones the pipeline can | |
| least easily prove. | |
| ### Why `RuleEntry` carries a `rule_type` | |
| The Interpretation Pack exists to **improve analytics insight from expert interpretation rules** β | |
| interpretation logic, action benchmarks, and rules for when to conclude or not conclude from data. | |
| Measured against the reference document, **0 of the 15 gold rules are interpretation rules.** They | |
| are calculation conventions (`PA_COMPOSITE_WEIGHTED`), data-sourcing rules | |
| (`PTY_PRODUCTION_SOURCE`), and unit conventions (`GAINLOSS_UNITS`). This is a property of the | |
| document type: a *Standard Parameter* document defines how to **calculate** parameters, not how to | |
| **read** them. Interpretation rules live in other document types, or in the expert's head. | |
| `rule_type` keeps both in one artifact rather than discarding the conventions we do extract: | |
| - `interpretation` β how to read a value or behaviour. **The target.** | |
| - `calculation` β how a value is computed or sourced. What this corpus actually yields. | |
| A consumer wanting only interpretation logic filters on `rule_type`. | |
| --- | |
| ## 2c Β· Evidence now carries figures and tables *(parsing 0.4.0, 2026-09-10)* | |
| The artifact this half consumes changed shape: a figure is no longer its own chunk. It is folded into | |
| the prose chunk at its reading-order position as ``, and travels as an `Asset` on that | |
| chunk. `adapter.py` carries `assets`, `referenced_by` and `table_html` across; `evidence_block` renders | |
| them into the prompt under `ASSETS REFERENCED BY THIS CHUNK`. | |
| **What that means for an extraction branch**, and the distinction is enforced in the prompts: | |
| | | may be quoted? | may be a provenance span? | | |
| |---|---|---| | |
| | `Asset.caption` β printed in the document | **yes** | **yes** | | |
| | `Asset.description` β written by a vision model | no | **no** β it is not in the document, so the span check rejects it and the field is discarded | | |
| Why it mattered: before 0.4.0 a figure chunk had an EMPTY `text`, so it produced no mentions, joined no | |
| cluster, and was never selected as evidence. 13 figures across three documents, each already described | |
| by a paid vision call, reached **zero** prompts. Full detail: | |
| `knowledge_pipeline_runs/CHANGELOG_v2.md` Β§2. | |
| > β **RESOLVED 2026-09-11 β Β§1, Β§2 and Β§8 now match the code.** The note is kept for the record. | |
| > | |
| > β οΈ **Correction to Β§8 of this document (noted 2026-09-10).** The code writes | |
| > `document_brief.json` with **`brief_id`** and a `b_` prefix; Β§8 below documents | |
| > `domain_context.json` with `domain_id` and `d_`. The v3 rename was partly reverted for the | |
| > per-document object once the scope-level `DomainContext` became a separate thing in | |
| > `src/knowledge_domain/`. **Β§8 does not describe what the code emits.** Worth one edit, and it matters | |
| > because `entity_id` prefixes are the key an expert's approvals hang on. | |
| --- | |
| ## 3 Β· Transport reality | |
| Every payload is written by the CLI as a file into `--out-dir` (default `out/knowledge/`), as a | |
| **bare JSON array** with no envelope. There is no endpoint and no table. | |
| Three conventions hold across every payload: | |
| - **`page` is 0-based**, exactly as the parser reports; **`page_no` is the 1-based number a human | |
| reads** and is what a review UI binds to. Both travel together on every entity, derived from one | |
| value so they cannot disagree (S6b, 2026-09-02). | |
| - **All content fields are nullable by design.** `null` is a valid, correct answer β the model | |
| abstaining, not failing. On the reference run **48 of 58** entries carry no definition. | |
| - **`provenance` is mandatory and `provenance.span` is verbatim-checked.** A field whose span cannot | |
| be located in the source is set to `null` and logged to `rejected.json` β **never repaired.** | |
| --- | |
| ## 4 Β· Common types | |
| ### `Provenance` β required on every entity | |
| ```json | |
| { | |
| "doc_id": "STD_2026_006_MNO", | |
| "span": "Adalah jumlah equipment/unit yang digunakan untuk menghasilkan produksi dalam suatu periode.", | |
| "page": 3, | |
| "page_no": 4, | |
| "section_no": "2.1.2", | |
| "chunk_id": "STD_2026_006_MNO::0007" | |
| } | |
| ``` | |
| | Field | Type | Req | Notes | | |
| |---|---|---|---| | |
| | `doc_id` | string | β | | | |
| | `span` | string | β | **Verbatim** from the source. The anti-hallucination control β the reviewer checks a quote against a page, not a claim against memory | | |
| | `page` | int \| null | β | **0-based**, exactly as the parser reports | | |
| | `page_no` | int \| null | β | **1-based. Added 2026-09-02 (S6b).** DERIVED from `page`, never stored, so it cannot drift. **A review UI must render this one** β an off-by-one is invisible until an expert opens the wrong page and concludes the provenance is wrong. `null` when `page` is unknown | | |
| | `section_no` | string \| null | β | e.g. `"2.1.2"`; null when the document is unnumbered | | |
| | `chunk_id` | string \| null | β | Back-reference into the parsed artifact. **Format `<doc_id>::<seq>`**, matching `KNOWLEDGE_PARSING_OUTPUT_CONTRACT.md`. Treat as opaque β do not split it | | |
| ### Enums | |
| | Enum | Values | | |
| |---|---| | |
| | `subdomain_tags[]` | `production` Β· `maintenance` Β· `hauling` Β· `loading` Β· `drilling_blasting` Β· `equipment` Β· `safety` Β· `quality` Β· `planning` Β· `cost` Β· `geology` Β· `other` | | |
| | `extraction_status` | `ok` Β· `no_definition_found` Β· `escalated` | | |
| | `diff_status` | `new` Β· `duplicate` Β· `conflicting` | | |
| | `rule_type` | `interpretation` Β· `calculation` | | |
| | `latex_verification` | `verified` Β· `unverified_no_markup` Β· `unverified_operator` | | |
| `subdomain_tags` is a **closed enum β classification, not generation.** Adding a member changes what | |
| the model is allowed to answer, which makes it a prompt change, not a data change. | |
| --- | |
| ## 5 Β· `glossary.json` β `GlossaryEntry[]` β 13 fields | |
| ```json | |
| [ | |
| { | |
| "term_id": "t_9f2a41c0b7", | |
| "term": "Qty", | |
| "full_name": "Quantity", | |
| "source_wording": "2.1.2. Quantity (Qty)", | |
| "definition": "Adalah jumlah equipment/unit yang digunakan untuk menghasilkan produksi dalam suatu periode.", | |
| "defining_formula_id": "f_31de08aa95", | |
| "subdomain_tags": ["production", "equipment"], | |
| "mention_count": 20, | |
| "provenance": { | |
| "doc_id": "STD_2026_006_MNO", | |
| "span": "Adalah jumlah equipment/unit yang digunakan untuk menghasilkan produksi dalam suatu periode.", | |
| "page": 3, | |
| "section_no": "2.1.2", | |
| "chunk_id": "STD_2026_006_MNO::0007" | |
| }, | |
| "extraction_status": "ok", | |
| "diff_status": "new", | |
| "definition_conflict": false, | |
| "conflict_variants": [] | |
| }, | |
| { | |
| "term_id": "t_5b71ce0d34", | |
| "term": "Gain/Loss", | |
| "full_name": null, | |
| "source_wording": null, | |
| "definition": null, | |
| "defining_formula_id": null, | |
| "subdomain_tags": ["production"], | |
| "mention_count": 4, | |
| "provenance": { | |
| "doc_id": "STD_2026_006_MNO", | |
| "span": "Gain/Loss", | |
| "page": 6, | |
| "section_no": null, | |
| "chunk_id": "STD_2026_006_MNO::0011" | |
| }, | |
| "extraction_status": "no_definition_found", | |
| "diff_status": "new", | |
| "definition_conflict": false, | |
| "conflict_variants": [] | |
| } | |
| ] | |
| ``` | |
| | Field | Type | Req | Notes | | |
| |---|---|---|---| | |
| | `term_id` | string | β | **New in v2.** Stable across re-runs β the key an approval hangs on | | |
| | `term` | string | β | The only required content field | | |
| | `full_name` | string \| null | β | Expanded name. Span-guarded as a literal transcription | | |
| | `source_wording` | string \| null | β | **The document's literal wording**, un-normalised. Read deterministically from the section heading, *not* asked of the model β asked directly, it returned the tidied form | | |
| | `definition` | string \| null | β | `null` = correct abstention (**48 of 58** on the reference run) | | |
| | `defining_formula_id` | string \| null | β | **New in v2**, replaces `formula_latex`. The formula that *defines* this term. Null when the formula branch did not extract one | | |
| | `subdomain_tags` | enum[] | β | Defaults `[]` | | |
| | `mention_count` | int | β | Default `0`. **Drives review-queue ordering** | | |
| | `provenance` | Provenance | β | | | |
| | `extraction_status` | enum | β | Default `ok` | | |
| | `diff_status` | enum \| null | β | Set by the diff stage against the active glossary version | | |
| | `definition_conflict` | bool | β | Default `false` | | |
| | `conflict_variants` | string[] | β | Default `[]`. The competing definitions β the pipeline **never picks a winner** | | |
| **Removed in v2:** `formula_latex` (β `defining_formula_id`) Β· `interpretation` (belongs in the | |
| Interpretation Pack, reached via `term_id`) Β· `domain` Β· `company` Β· `language`. | |
| > `domain` and `company` were never in the extraction prompt's field guide β the model filled them | |
| > by copying worked examples β and neither was span-guarded, so a wrong value was undetectable. | |
| > `language` is a genuine three-value field (`id`/`en`/`mixed`) but nothing consumes it and it is | |
| > recoverable from the document; it is cut for simplicity and can return if a consumer needs it. | |
| > **Writer split.** The LLM fills only `term`, `full_name`, `definition` and `subdomain_tags`. | |
| > `term_id`, `defining_formula_id`, `source_wording`, `mention_count`, `extraction_status`, | |
| > `diff_status`, `definition_conflict` and `conflict_variants` are all set by deterministic code β | |
| > **the model never writes its own ids, links or audit fields**, so none of them can be hallucinated. | |
| --- | |
| ## 6 Β· `rules.json` β `RuleEntry[]` β 8 fields | |
| ```json | |
| [ | |
| { | |
| "rule_id": "r_c07be41f22", | |
| "rule_type": "calculation", | |
| "statement": "Production yang digunakan dalam perhitungan adalah produksi hasil joint survey.", | |
| "condition": "Apabila data joint survey belum tersedia", | |
| "consequence": "maka digunakan data produksi berdasarkan truck count sebagai dasar perhitungan", | |
| "formula_ids": ["f_77c1ab9e40"], | |
| "term_ids": ["t_a4e2f81b06"], | |
| "provenance": { | |
| "doc_id": "STD_2026_006_MNO", | |
| "span": "Apabila data joint survey belum tersedia maka digunakan data produksi berdasarkan truck count sebagai dasar perhitungan", | |
| "page": 4, | |
| "section_no": "2.1.5", | |
| "chunk_id": "STD_2026_006_MNO::0009" | |
| } | |
| } | |
| ] | |
| ``` | |
| | Field | Type | Req | Notes | | |
| |---|---|---|---| | |
| | `rule_id` | string | β | **Re-specified in v2.** Was written by the model, therefore unstable; now deterministic | | |
| | `rule_type` | enum | β | `interpretation` \| `calculation` β see Β§2 | | |
| | `statement` | string \| null | β | The rule as prose. **Prose, not an equation** β evidence carrying both yields the prose, and the equation goes to the formula branch | | |
| | `condition` | string \| null | β | The triggering condition, split out β not left buried in prose | | |
| | `consequence` | string \| null | β | What follows when the condition holds | | |
| | `formula_ids` | string[] | β | **New in v2.** The formulas this rule constrains. Default `[]` | | |
| | `term_ids` | string[] | β | **New in v2.** The terms this rule governs. Default `[]` | | |
| | `provenance` | Provenance | β | | | |
| **Removed in v2:** `applies_to` (free text β `term_ids[]`) Β· `subdomain_tags` Β· `language` Β· | |
| `extraction_status`. | |
| `condition` + `consequence` is the point of this artifact, and the seed of the interpretation logic | |
| tree: a rule stored as one paragraph cannot be attached to a skill later; a trigger and its | |
| consequence can. | |
| > **Deliberately not modelled yet:** the interpretation logic tree, action benchmarks, and explicit | |
| > conclude / do-not-conclude verdicts. Those are the direction of travel, but designing their schema | |
| > against zero real examples would be guessing. They land when a document containing them does. | |
| --- | |
| ## 7 Β· `formulas.json` β `FormulaEntry[]` β 7 fields | |
| ```json | |
| [ | |
| { | |
| "formula_id": "f_77c1ab9e40", | |
| "name": "Production", | |
| "formula_latex": "Production = MOHH \\times Qty \\times PA \\times UA \\times Pty", | |
| "variables": [ | |
| { "symbol": "MOHH", "term_id": "t_1c9d5e7a83", "meaning": "Machine on Hand Hours" }, | |
| { "symbol": "Qty", "term_id": "t_9f2a41c0b7", "meaning": "Quantity" }, | |
| { "symbol": "PA", "term_id": "t_a4e2f81b06", "meaning": "Physical Availability" }, | |
| { "symbol": "UA", "term_id": "t_6d40b2fc17", "meaning": "Utilization of Availability" }, | |
| { "symbol": "Pty", "term_id": "t_e8b3092d55", "meaning": "Productivity" } | |
| ], | |
| "unit": "BCM", | |
| "latex_verification": "unverified_operator", | |
| "provenance": { | |
| "doc_id": "STD_2026_006_MNO", | |
| "span": "Production = MOHH x Qty x PA x UA x Pty", | |
| "page": 1, | |
| "section_no": "2.1", | |
| "chunk_id": "STD_2026_006_MNO::0003" | |
| } | |
| } | |
| ] | |
| ``` | |
| | Field | Type | Req | Notes | | |
| |---|---|---|---| | |
| | `formula_id` | string | β | **New in v2** | | |
| | `name` | string \| null | β | What the formula computes, as named in the source | | |
| | `formula_latex` | string \| null | β | **Transcribed, never derived.** The pipeline copies what the document states | | |
| | `variables[]` | object[] | β | `{ symbol (required), term_id \| null, meaning \| null }`. `term_id` is **new in v2** | | |
| | `unit` | string \| null | β | Result unit if stated | | |
| | `latex_verification` | enum \| null | β | **Added 2026-08-26.** Which guarantee `formula_latex` actually carries. Set by the span check, never by the model. `null` when there is no `formula_latex` to characterise | | |
| | `provenance` | Provenance | β | | | |
| **Removed in v2:** `extraction_status`. | |
| > **Read `latex_verification` before trusting `formula_latex`.** Three paths leave the field populated and | |
| > they are not equally strong. `verified` means the transcription was located in the source markup | |
| > (`Chunk.latex`), compared in a notation-aware canonical form. The two `unverified_*` values mean the | |
| > claim could not be checked either way and the entry is guarded by its provenance span alone β the same | |
| > guarantee everything carried before 2026-08-26. | |
| > | |
| > `unverified_operator` is not rare. MinerU's **`pipeline`** backend transcribes multiplication as a bare | |
| > letter `x`, so every product formula from such an artifact lands here β including the sample above. | |
| > An artifact parsed with `hybrid/high` carries `\times` and verifies normally. | |
| > `variables[].symbol` is not always a legend abbreviation β the extraction prompt's own worked | |
| > example emits `"Total Hours"` and `"Breakdown"` as symbols. So `term_id` resolution is | |
| > **best-effort and nulls are expected**; the link stage reports its dangle rate rather than hiding it. | |
| --- | |
| ## 8 Β· `document_brief.json` β `DocumentBrief` β 6 fields *(single object)* | |
| > β **The v3 reshape is complete, corrected 2026-09-08.** This banner previously said the | |
| > LLM-facing parts β `purpose` as a span-guarded verbatim quote, and the `subdomains` aggregate β | |
| > were *"still proposed and not built"*. They shipped on 2026-09-02 and the banner was never | |
| > updated, so it contradicted the field table directly below it, which dates both changes. Verified | |
| > live: a `domain` entry persisted on 2026-09-07 carries `purpose_verbatim` and `subdomains`. | |
| > | |
| > The delivery really was split, and that part was deliberate: the deterministic half (`outline` | |
| > read off the artifact, `key_parameters` corroborated against extracted terms) shipped first | |
| > because it costs nothing to prove, and the LLM-facing half followed once an eval run could pay | |
| > for itself. | |
| ```json | |
| { | |
| "brief_id": "d_2ef60a8c19", | |
| "title": "Standard Parameter Produksi & ECA", | |
| "purpose_verbatim": "Standard Parameter ini merupakan pedoman dalam perhitungan parameter produksi yang meliputi Quantity (Qty) Physical Availability (PA), Utilization of Availability (UA), Productivity (PTY), serta Equipment Capacity Analysis (ECA).", | |
| "outline": [ | |
| "1. TUJUAN PARAMETER", | |
| "2.1.2. Quantity (Qty)", | |
| "2.1.3. Physical of Availability (PA)", | |
| "2.2. Equipment Capacity Analysis (ECA)" | |
| ], | |
| "key_parameters": [ | |
| {"surface": "Physical Availability (PA)", "term_id": "t_8c7359e27f"}, | |
| {"surface": "Equipment Capacity Analysis (ECA)", "term_id": null} | |
| ], | |
| "subdomains": ["equipment", "production"], | |
| "n_terms": 5, | |
| "n_formulas": 6, | |
| "n_rules": 0, | |
| "provenance": { | |
| "doc_id": "STD_2026_006_MNO", | |
| "span": "Standard Parameter ini merupakan pedoman dalam perhitungan parameter produksi", | |
| "page": 0, | |
| "section_no": "1", | |
| "chunk_id": "STD_2026_006_MNO::0000" | |
| } | |
| } | |
| ``` | |
| | Field | Type | Req | | |
| |---|---|---| | |
| | `brief_id` | string | β **New in v2.** `b_` prefix β reconciled with the code 2026-09-11 | | |
| | `title` | string \| null | β | | |
| | `purpose_verbatim` | string \| null | β **Renamed + guarded 2026-09-02** (was `purpose`) | | |
| | `outline` | string[] | β Default `[]`. **Added 2026-09-02** | | |
| | `subdomains` | SubdomainEnum[] | β Default `[]`. **Added 2026-09-02** | | |
| | `n_terms` Β· `n_formulas` Β· `n_rules` | int | β Default `0`. **Added 2026-09-02** | | |
| | `key_parameters` | KeyParameter[] | β Default `[]`. **Shape changed 2026-09-02** β was `string[]` | | |
| | `provenance` | Provenance | β | | |
| **`KeyParameter`** β `{surface: string, term_id: string | null}`. `surface` is what the model named; | |
| `term_id` is the glossary term that corroborates it, resolved deterministically by `link.py`. A | |
| **null `term_id` means nothing in the extracted glossary supports that name** β reported, never | |
| guessed, and it earns its own `review_queue.json` row (Β§9). Dangle count is reported as | |
| `key_parameters_unresolved`. | |
| **`subdomains`** is AGGREGATED from the glossary entries' own `subdomain_tags`, never asked of the | |
| model β each tag was already chosen once, per term, with that term's evidence in front of the model, | |
| and a second document-level classification would be the same judgement made with less context and no | |
| way to check it. Ordered by tag frequency descending, alphabetical tiebreak, so two runs of one | |
| document agree. **`n_terms` / `n_formulas` / `n_rules`** are counts of the other three artifacts, | |
| computed after every branch and the diff have run. | |
| **`outline`** is DERIVED from the artifact's `heading_path`, in reading order, de-duplicated. No LLM | |
| and no spend. It carries the strong guarantee the rest of this entry cannot: it is verbatim source | |
| structure, so there is nothing to hallucinate β including wording a reader would be tempted to fix | |
| (the reference document heads a section *"Physical of Availability (PA)"*). Empty is normal for a | |
| document with no headings. | |
| **Removed in v2:** `scope` (overlaps `purpose`, frequently null) Β· `summary_md`. | |
| > β **This branch is span-checked as of 2026-09-02.** It used to be the one exception: it asked the | |
| > model to SUMMARISE, and a summary is not verbatim by construction, so the primary | |
| > anti-hallucination control could not apply. It now asks the model to **locate** β `purpose_verbatim` | |
| > is the document's own statement of purpose, copied, not a sentence composed about it β so `title` | |
| > and `purpose_verbatim` are guarded like any other transcription. A value that is not verbatim is | |
| > nulled and recorded in `rejected.json`, never repaired. | |
| > | |
| > Measured on the reference document: **0 fields rejected**, and a deliberately fabricated purpose is | |
| > refused with `purpose_verbatim is not a verbatim transcription of the source`. | |
| > | |
| > Two consequences worth knowing. **`title` will look wrong sometimes and that is correct** β the | |
| > reference document yields `"BUMA | STANDARD PARAMETERProduction Parameter andEquipment Capacity | |
| > Analysis (ECA)"`, missing spaces and all, because that is what the PDF says; tidying it is the | |
| > UI's job at display time, never the pipeline's. And this weakens the case that the branch needs a | |
| > larger model tier (DEV_PLAN Β§0.8 D3, risk R5): locating a sentence is a much smaller ask than | |
| > composing one. | |
| --- | |
| ## 9 Β· `review_queue.json` β `ReviewQueueRow[]` β 12 fields | |
| **This is the consumption surface** β the payload a review UI binds to. Deliberately denormalised so | |
| a row renders with no joins. | |
| > β οΈ **Changed 2026-09-02: the queue is no longer glossary-only.** Every row now carries | |
| > **`row_kind`**, and a consumer must switch on it rather than assume every row is a term. A | |
| > `key_parameter` row has a surface and **no `term_id`, no definition and no mention count** β a | |
| > client that binds `term_id` unconditionally will break on it. | |
| ```json | |
| [ | |
| { | |
| "rank": 1, | |
| "row_kind": "term", | |
| "term_id": "t_a4e2f81b06", | |
| "term": "PA", | |
| "definition": "Adalah ketersediaan fisik suatu equipment/unit yang menunjukkan proporsi waktu equipment/unit tersebut berada pada kondisi available (siap pakai) selama suatu periode tertentu.", | |
| "source_wording": "2.1.3. Physical of Availability (PA)", | |
| "mention_count": 43, | |
| "page": 3, | |
| "section_no": "2.1.3", | |
| "span": "Adalah ketersediaan fisik suatu equipment/unit", | |
| "review_reason": "source wording differs from the expanded name β confirm which is correct" | |
| }, | |
| { | |
| "rank": 2, | |
| "row_kind": "key_parameter", | |
| "term_id": null, | |
| "brief_id": "d_2ef60a8c19", | |
| "term": "Grouping (Composite)", | |
| "definition": null, | |
| "source_wording": null, | |
| "mention_count": 0, | |
| "page": 0, | |
| "section_no": "1", | |
| "span": "Standard Parameter ini merupakan pedoman dalam perhitungan parameter produksi", | |
| "review_reason": "named as a key parameter but no extracted term corroborates it β confirm it is real" | |
| } | |
| ] | |
| ``` | |
| | Field | Type | Req | Notes | | |
| |---|---|---|---| | |
| | `rank` | int | β | 1-based display order β see sorting below | | |
| | `row_kind` | `term` \| `key_parameter` | β | **New 2026-09-02.** Switch on this before reading anything else | | |
| | `term_id` | string \| null | β | **New in v2.** What an approve / edit / reject decision attaches to, so it survives a re-run. **Null on a `key_parameter` row** β that it resolves to no term is the reason the row exists | | |
| | `brief_id` | string | β | `key_parameter` rows only β what the decision attaches to instead. **Named `brief_id` on the wire**, not `domain_id`; see the correction under Sort order | | |
| | `surface_full` | string \| null | β | `key_parameter` rows only. **New 2026-09-11.** The untruncated surface when `term` had to be shortened to a head; `null` when it did not | | |
| | `term` Β· `definition` Β· `source_wording` Β· `mention_count` | | | Copied from the entry | | |
| | `page` Β· `page_no` Β· `section_no` Β· `span` | | | **Lifted out of `provenance`** so a row is self-contained. Render **`page_no`** (1-based), not `page` | | |
| | `review_reason` | string | β | Human-readable, one of five | | |
| **Removed in v2:** `extraction_status` Β· `diff_status` Β· `definition_conflict` Β· | |
| `conflict_variants` β all four are already expressed by `review_reason`, and remain available on the | |
| entry itself. | |
| **Sort order** β **changed 2026-09-11.** `(conflicts, then mention_count descending, with unresolved | |
| key parameters spliced in at most one per five term rows)`. | |
| Conflicts still outrank everything, because a contradiction is a decision only the expert can make. | |
| What changed is the tier beneath them. Unresolved key parameters used to occupy one, which meant | |
| **every** unresolved parameter outranked **every** routine term regardless of how many there were β | |
| and the branch produces them in bulk on exactly the documents it understands least. On the Komatsu | |
| shop manual that was **24 junk rows standing in front of `Engine`, a term with 108 mentions**, which | |
| sat at rank 25. | |
| The original reason for promoting them is preserved and is still sound: that branch has **no span | |
| check at all**, so an unresolvable surface is the one hallucination signal it can raise, and it must | |
| not sink beneath sixty routine rows where nobody reaches it. A **cap** expresses that without letting | |
| the tail take the queue over. Measured by replaying the three persisted v2 runs: | |
| | Document | unresolved key parameters | first one at rank | top term, rank | | |
| |---|---|---|---| | |
| | BUMA | 0 | β | `Qty` 1 β 1 (unchanged) | | |
| | Komatsu | 24 | 1 β **6** | `Engine` **25 β 1** | | |
| | Open Pit | 7 | 1 β **6** | `Resource` **8 β 1** | | |
| Row counts are identical before and after on all three β **nothing is dropped**; parameters the | |
| splice cannot place are appended rather than discarded. | |
| **Also new on `key_parameter` rows: `surface_full`.** The branch is supposed to return a parameter | |
| *name* and on a textbook it returns the right concept wrapped in its whole defining sentence (185 | |
| characters, in one measured case). `term` now carries a head β split at the first `.`, `:` or newline | |
| β and `surface_full` carries the untruncated text, or `null` when no shortening was needed. The head | |
| is only taken past 60 characters, which is what keeps a legitimate short name containing a period | |
| ("No. of units") intact. | |
| > β οΈ **Field-name correction, 2026-09-11.** This document described the `key_parameter` row's owning | |
| > id as **`domain_id`**. The code emits **`brief_id`** and always has: the K12 rename changed the | |
| > entity from `BriefContext` to `DomainContext` but left the id field's name alone, deliberately β | |
| > renaming it is a migration that invalidates every expert approval recorded against the old prefix. | |
| > **The wire field is `brief_id`.** Corrected here rather than in code, per the trust order. | |
| **`review_reason` values:** | |
| | Value | Meaning | | |
| |---|---| | |
| | `conflicting definitions β expert decision required` | Contradiction found; the pipeline picked no winner | | |
| | `source wording differs from the expanded name β confirm which is correct` | e.g. "Physical **of** Availability" vs "Physical Availability" | | |
| | `term found but no definition in document` | Correct abstention β **48 of 58** on the reference run | | |
| | `definition rejected by span check or absent` | Control fired; see `rejected.json` | | |
| | `routine confirmation` | Nothing anomalous | | |
| | `named as a key parameter but no extracted term corroborates it β confirm it is real` | `key_parameter` rows only. Added 2026-09-02 | | |
| --- | |
| ## 10 Β· Linking model | |
| `GlossaryEntry` is the hub. Every link is computed **after** both branches run, by a deterministic | |
| never-throw link stage. **No additional LLM calls, and nothing a model could fabricate.** | |
| ``` | |
| RuleEntry ββterm_ids[]βββββββββΊ GlossaryEntry βββvariables[].term_idββ FormulaEntry | |
| βββββformula_ids[]βββββββββββββββββββββββββββββββββββββββββββββββββββββββΊ β² | |
| GlossaryEntry ββdefining_formula_idββββββββββββββββββββ | |
| ``` | |
| Two properties that must hold: | |
| 1. **A dangling link is `null` or an empty array β never a fabrication.** Same rule as spans. | |
| 2. **The "appears in" edge is derived, not stored.** That PA appears inside | |
| `Production = MOHH Γ Qty Γ PA Γ UA Γ Pty` is recoverable by scanning `formulas[].variables[]`. | |
| Do not add a field for it. | |
| --- | |
| ## 11 Β· Audit payloads β unchanged from v1 | |
| **`rejected.json` β `RejectedField[]`** β what the span check *caught*, kept so a reviewer sees the | |
| control working, not only what it let through. The reference run rejected **2 fields**. | |
| ```json | |
| [ | |
| { | |
| "entry_term": "<term>", | |
| "field": "full_name", | |
| "offending_value": "<value the model returned>", | |
| "reason": "full_name is not a verbatim transcription of the source", | |
| "branch": "glossary" | |
| } | |
| ] | |
| ``` | |
| | Field | Type | Req | | |
| |---|---|---| | |
| | `entry_term` | string | β | | |
| | `field` | string | β | | |
| | `offending_value` | string | β | | |
| | `reason` | string | β | | |
| | `branch` | `glossary` \| `rule` \| `formula` \| `summary` | β | | |
| **`usage.json` β `CallUsage[]`** β per-call accounting. `cached_tokens` is read from the API, never | |
| modelled: caching does not engage below 1024 prompt tokens, so assuming it would understate input | |
| cost by roughly 10Γ. | |
| ```json | |
| [ | |
| { | |
| "branch": "glossary", | |
| "deployment": "gpt-5.4-nano", | |
| "tier": "nano", | |
| "prompt_tokens": 2676, | |
| "cached_tokens": 2389, | |
| "completion_tokens": 214, | |
| "latency_s": 3.4, | |
| "retries": 0, | |
| "structured_output_mode": "json_schema", | |
| "simulated": false | |
| } | |
| ] | |
| ``` | |
| `simulated: true` marks a `--mock` run β **never a quality measurement.** | |
| --- | |
| ## 12 Β· Reference run β real figures | |
| From `eval/knowledge/results/v2_full_document_2026-08-24_093051.json`. The parsing module built the | |
| artifact; extraction ran all four branches. | |
| | | | | |
| |---|---| | |
| | Artifact | `schema_version` 0.2.0 Β· **31 chunks** Β· 9 pages Β· `mineru/pipeline 3.4.4` | | |
| | Funnel | 188 raw mentions β 163 after noise β **58 clusters** (compression 2.81Γ) | | |
| | Entries | **58 glossary Β· 8 rules Β· 6 formulas** | | |
| | LLM | 77 calls Β· 139,283 prompt tokens Β· **79.6% cached** Β· 168.4 s | | |
| | Quality | **48 abstained**, 10 with a definition Β· **2 fields rejected** by the span check | | |
| | E1 term-filter recall | **0.8049** (kill line 0.70, PASS) | | |
| | E3 schema-fill precision | **0.90** (kill line 0.80, **PASS** β the prototype failed this at 0.75) | | |
| > v1 of this contract quoted 13 chunks / 66 clusters / 66 entries. That run predates the | |
| > section-aware chunker. **Anything calibrated against those figures should be re-checked.** | |
| > | |
| > These figures also predate the v2 schema itself: they were measured on the v1 prompts, which asked | |
| > for fields that no longer exist. The funnel counts (chunks, clusters, calls) should carry over | |
| > unchanged β the filter, cluster and rank stages were not touched β but **E3 must be re-scored**, | |
| > because trimming a prompt can move definition quality even when the scored fields are unchanged. | |
| --- | |
| ## 13 Β· Open items | |
| | # | Item | Owner | | |
| |---|---|---| | |
| | 1 | **Unevaluated.** Cutting `interpretation` / `formula_latex` and adding `rule_type` are prompt changes, and no eval has run against them β there is no document or parsed artifact on the dev box to run one. Same position X20 was committed in, and it needs the same first paid run to clear | Rifqi | | |
| | 2 | **Rule-branch quality is a separate track.** The branch measures 3/15, an unevaluated prompt fix is already in flight (X20), and a prompt/gold contradiction caps it at 14/15 (X23). `rule_type` and the X23 de-contamination make this the *third and fourth* uncommitted-to-measurement changes on that one prompt β the first paid run measures the prompt as a whole, and isolating any single edit would cost one full run each. **Expect the raw score to move in both directions:** X20 should add rules back, X23 removes a free hit (example 1 was gold `PTY_PRODUCTION_SOURCE` verbatim) and lifts the 14/15 ceiling | Rifqi | | |
| | 2b | **`rule.txt` grew again.** X20 already pushed it past the 1024-token cache floor; example 4 adds more. Only the API's `cached_tokens` proves a hit β verify on the first paid run | Rifqi | | |
| | 3 | **Persistence.** Stage output is JSON on disk; parsed artifacts, candidate entries, glossary versions and the approval audit trail still need one consolidated DDL handoff. Go owns the schema β Python never executes DDL | Rifqi β Harry | | |
| | 4 | **Page indexing at the API boundary** β stays 0-based to the UI, or converts once at the boundary? Currently 0-based everywhere | Harry + Rifqi | | |
| | 5 | **Endpoint shape** β four endpoints, or one with an `?artifact=` parameter? Not started; the offline CLI is the honest first milestone | Rifqi | | |
| --- | |
| ## Caveat on the sample values | |
| Field shapes, types and defaults are the **agreed v2 target**. Run-level figures in Β§12 are literal, | |
| from the cited result file. The per-entity **sample values are illustrative** β the underlying | |
| `out/*.json` from that run is not in version control, and all `*_id` values shown are placeholders | |
| that demonstrate format, not real hashes. Replace them with literal output once v2 is implemented | |
| and a run exists. | |