Agentic-Service-Data-Eyond-Catalog / docs /knowledge /KNOWLEDGE_OUTPUT_CONTRACT.md
ishaq101's picture
/fix parsing and term extract (#21)
f07443e
|
Raw History Blame Contribute Delete
38.4 kB

Knowledge Pipeline β€” Output Contract v2

Implemented 2026-08-26. src/knowledge_extraction/ emits this shape. The schema review of 2026-08-26 landed with it: ids on every entity, formulas referenced rather than restated, rules linked to formulas and terms, and 16 fields cut.

Not yet re-measured. The prompt changes this required have not been scored β€” no source document, parsed artifact or MinerU install exists on the dev box (all three data/ directories are gitignored and absent), so the free stages have nothing to read. The last measured figures in Β§12 predate these edits. See Β§13.

Supersedes: v1 of this document (2026-08-24) Reference run: v2_full_document_2026-08-24_093051 Β· BUMA STD_2026_006_MNO Β· 9 pages Β· gpt-5.4-nano Producer: python -m src.knowledge_extraction.cli <artifact.json> --extract Status: offline artifacts only β€” no endpoint, no table.


0 Β· What changed from v1

Harry's review of v1 asked for entity ids, a formula reference instead of a restated formula, and a link between rules and formulas. All are adopted. The schema was also cut back β€” v1 carried fields nothing consumes.

Change Detail
Ids on every entity term_id, formula_id, brief_id; rule_id re-specified. All deterministic and content-derived
Glossary no longer restates formulas formula_latex removed β†’ defining_formula_id reference
Rules link to formulas and terms formula_ids[], term_ids[]
rule_type added interpretation | calculation β€” see Β§2 for why
16 fields cut, 8 added 50 β†’ 42 fields total
chunk_id separator fixed DOC#0007 β†’ DOC::0007, adopting the parser's format
Reference numbers corrected v1 quoted a superseded run (13 chunks / 66 clusters). Current is 31 chunks / 58 clusters

Why the ids had to be re-specified

Neither id in v1 survived a re-run. cluster_id was a frequency-rank ordinal (f"{doc_id}#c{i:03d}" over a list sorted by mention count), and rule_id was written by the model (rule.txt asks for SCREAMING_SNAKE_CASE). Either way, re-running with a tuned prompt reshuffles the ids β€” and if an expert has approved 40 entries, we have lost which 40.

Approval continuity, not linking, is the real reason these exist.


1 Β· Id scheme

All ids are deterministic, content-derived, and stable across re-runs of the same document. None is model-generated; none is positional.

term_id     = "t_" + sha256(doc_id + "|" + canonical)[:10]
formula_id  = "f_" + sha256(doc_id + "|" + name_or_latex_or_chunk_id)[:10]
rule_id     = "r_" + sha256(doc_id + "|" + chunk_id + "|" + char_start)[:10]
brief_id    = "b_" + sha256(doc_id)[:10]    # per-document card; `d_` belongs to the
                                            # scope-level DomainContext in src/knowledge_domain/

Implemented in src/knowledge_extraction/ids.py. term_id is keyed on the cluster canonical, not on the term the model returned β€” the cluster is the stable thing across runs, and the model is free to answer "PA" or "Physical Availability" for the same cluster. formula_id falls back through name β†’ formula_latex β†’ chunk_id, because both of the first two are Optional and without the fallback every unnamed formula in a document would collide.

Scope is per document. The same term appearing in two documents receives two different term_ids. Deciding that BUMA's "PA" and a textbook's "PA" are the same concept is an expert judgement at review time, not something the pipeline asserts β€” the same reasoning that makes the pipeline record "Physical of Availability" instead of normalising it, and the reason MTTR at BUMA (breakdown duration Γ· breakdown frequency) must not silently merge with MTTR in IT.

Known limit. Ids derive from content, so if content changes the id changes β€” retune clustering such that a term canonicalises differently and its term_id moves. The alternative is a persisted registry, which belongs with persistence (DEV_PLAN Β§0.8 D2).


2 Β· The four artifacts

File Entity Holds
glossary.json GlossaryEntry[] Terms and their definitions
rules.json RuleEntry[] Rules of thumb β€” how to interpret data or behaviour. Renamed from interpretation_pack.json 2026-09-08 β€” the file, the entity and the DB kind now share one name. The Interpretation Pack remains the deliverable's name in prose
formulas.json FormulaEntry[] Formulas, transcribed from the document
document_brief.json DocumentBrief One per-document card. Prefix history, not churn: b_ (BriefContext) -> d_ (DomainContext, 2026-09-01) -> b_ again (DocumentBrief, 2026-09-08), once v4 gave d_ to the real scope-level DomainContext in src/knowledge_domain/

Plus three derived/audit payloads: review_queue.json, rejected.json, usage.json.

Rule vs domain knowledge β€” where the line is

Added 2026-09-14 (o2). The two blurred on screen during the 2026-09-08 review, and a boundary nobody has written down is one the prompts cannot hold.

rule domain
Grain One statement, from one document One per company, composed across all of them
Shape A condition and a consequence Identity, boundary, conventions, measures
Evidence Span-checked against its source Derived from entries, or expert-declared
Answers "…and how do I compute this correctly?" "…what does this company measure, and what does it call it?"

The test to apply. If removing it would change how a number is computed or sourced, it is a rule β€” "if joint-survey data is unavailable, use truck count" changes the figure. If removing it would only leave an agent less oriented β€” not knowing that this company works in overburden and BCM β€” it is domain.

Two consequences worth stating, because they are what the distinction is for:

  • A rule is actionable alone; domain knowledge is not. A planner can apply a rule to a specific calculation. It cannot apply "this is a mining company" to anything β€” that shapes which questions make sense, not which column to sum.
  • Domain is high-level and rules are specific, but "high-level" is not the test. A long, general rule is still a rule. The grain question β€” does this come from one document's sentence, or from the corpus as a whole? β€” settles it faster and does not drift.

Where it is genuinely both, prefer the rule: it carries provenance and a span, so an expert can check it. Domain fields are the ones the pipeline can least easily prove.

Why RuleEntry carries a rule_type

The Interpretation Pack exists to improve analytics insight from expert interpretation rules β€” interpretation logic, action benchmarks, and rules for when to conclude or not conclude from data.

Measured against the reference document, 0 of the 15 gold rules are interpretation rules. They are calculation conventions (PA_COMPOSITE_WEIGHTED), data-sourcing rules (PTY_PRODUCTION_SOURCE), and unit conventions (GAINLOSS_UNITS). This is a property of the document type: a Standard Parameter document defines how to calculate parameters, not how to read them. Interpretation rules live in other document types, or in the expert's head.

rule_type keeps both in one artifact rather than discarding the conventions we do extract:

  • interpretation β€” how to read a value or behaviour. The target.
  • calculation β€” how a value is computed or sourced. What this corpus actually yields.

A consumer wanting only interpretation logic filters on rule_type.


2c Β· Evidence now carries figures and tables (parsing 0.4.0, 2026-09-10)

The artifact this half consumes changed shape: a figure is no longer its own chunk. It is folded into the prose chunk at its reading-order position as ![](asset://<id>), and travels as an Asset on that chunk. adapter.py carries assets, referenced_by and table_html across; evidence_block renders them into the prompt under ASSETS REFERENCED BY THIS CHUNK.

What that means for an extraction branch, and the distinction is enforced in the prompts:

may be quoted? may be a provenance span?
Asset.caption β€” printed in the document yes yes
Asset.description β€” written by a vision model no no β€” it is not in the document, so the span check rejects it and the field is discarded

Why it mattered: before 0.4.0 a figure chunk had an EMPTY text, so it produced no mentions, joined no cluster, and was never selected as evidence. 13 figures across three documents, each already described by a paid vision call, reached zero prompts. Full detail: knowledge_pipeline_runs/CHANGELOG_v2.md Β§2.

βœ… RESOLVED 2026-09-11 β€” Β§1, Β§2 and Β§8 now match the code. The note is kept for the record.

⚠️ Correction to §8 of this document (noted 2026-09-10). The code writes document_brief.json with brief_id and a b_ prefix; §8 below documents domain_context.json with domain_id and d_. The v3 rename was partly reverted for the per-document object once the scope-level DomainContext became a separate thing in src/knowledge_domain/. §8 does not describe what the code emits. Worth one edit, and it matters because entity_id prefixes are the key an expert's approvals hang on.


3 Β· Transport reality

Every payload is written by the CLI as a file into --out-dir (default out/knowledge/), as a bare JSON array with no envelope. There is no endpoint and no table.

Three conventions hold across every payload:

  • page is 0-based, exactly as the parser reports; page_no is the 1-based number a human reads and is what a review UI binds to. Both travel together on every entity, derived from one value so they cannot disagree (S6b, 2026-09-02).
  • All content fields are nullable by design. null is a valid, correct answer β€” the model abstaining, not failing. On the reference run 48 of 58 entries carry no definition.
  • provenance is mandatory and provenance.span is verbatim-checked. A field whose span cannot be located in the source is set to null and logged to rejected.json β€” never repaired.

4 Β· Common types

Provenance β€” required on every entity

{
  "doc_id": "STD_2026_006_MNO",
  "span": "Adalah jumlah equipment/unit yang digunakan untuk menghasilkan produksi dalam suatu periode.",
  "page": 3,
  "page_no": 4,
  "section_no": "2.1.2",
  "chunk_id": "STD_2026_006_MNO::0007"
}
Field Type Req Notes
doc_id string βœ…
span string βœ… Verbatim from the source. The anti-hallucination control β€” the reviewer checks a quote against a page, not a claim against memory
page int | null β€” 0-based, exactly as the parser reports
page_no int | null β€” 1-based. Added 2026-09-02 (S6b). DERIVED from page, never stored, so it cannot drift. A review UI must render this one β€” an off-by-one is invisible until an expert opens the wrong page and concludes the provenance is wrong. null when page is unknown
section_no string | null β€” e.g. "2.1.2"; null when the document is unnumbered
chunk_id string | null β€” Back-reference into the parsed artifact. Format <doc_id>::<seq>, matching KNOWLEDGE_PARSING_OUTPUT_CONTRACT.md. Treat as opaque β€” do not split it

Enums

Enum Values
subdomain_tags[] production Β· maintenance Β· hauling Β· loading Β· drilling_blasting Β· equipment Β· safety Β· quality Β· planning Β· cost Β· geology Β· other
extraction_status ok Β· no_definition_found Β· escalated
diff_status new Β· duplicate Β· conflicting
rule_type interpretation Β· calculation
latex_verification verified Β· unverified_no_markup Β· unverified_operator

subdomain_tags is a closed enum β€” classification, not generation. Adding a member changes what the model is allowed to answer, which makes it a prompt change, not a data change.


5 Β· glossary.json β†’ GlossaryEntry[] β€” 13 fields

[
  {
    "term_id": "t_9f2a41c0b7",
    "term": "Qty",
    "full_name": "Quantity",
    "source_wording": "2.1.2. Quantity (Qty)",
    "definition": "Adalah jumlah equipment/unit yang digunakan untuk menghasilkan produksi dalam suatu periode.",
    "defining_formula_id": "f_31de08aa95",
    "subdomain_tags": ["production", "equipment"],
    "mention_count": 20,
    "provenance": {
      "doc_id": "STD_2026_006_MNO",
      "span": "Adalah jumlah equipment/unit yang digunakan untuk menghasilkan produksi dalam suatu periode.",
      "page": 3,
      "section_no": "2.1.2",
      "chunk_id": "STD_2026_006_MNO::0007"
    },
    "extraction_status": "ok",
    "diff_status": "new",
    "definition_conflict": false,
    "conflict_variants": []
  },
  {
    "term_id": "t_5b71ce0d34",
    "term": "Gain/Loss",
    "full_name": null,
    "source_wording": null,
    "definition": null,
    "defining_formula_id": null,
    "subdomain_tags": ["production"],
    "mention_count": 4,
    "provenance": {
      "doc_id": "STD_2026_006_MNO",
      "span": "Gain/Loss",
      "page": 6,
      "section_no": null,
      "chunk_id": "STD_2026_006_MNO::0011"
    },
    "extraction_status": "no_definition_found",
    "diff_status": "new",
    "definition_conflict": false,
    "conflict_variants": []
  }
]
Field Type Req Notes
term_id string βœ… New in v2. Stable across re-runs β€” the key an approval hangs on
term string βœ… The only required content field
full_name string | null β€” Expanded name. Span-guarded as a literal transcription
source_wording string | null β€” The document's literal wording, un-normalised. Read deterministically from the section heading, not asked of the model β€” asked directly, it returned the tidied form
definition string | null β€” null = correct abstention (48 of 58 on the reference run)
defining_formula_id string | null β€” New in v2, replaces formula_latex. The formula that defines this term. Null when the formula branch did not extract one
subdomain_tags enum[] βœ… Defaults []
mention_count int βœ… Default 0. Drives review-queue ordering
provenance Provenance βœ…
extraction_status enum βœ… Default ok
diff_status enum | null β€” Set by the diff stage against the active glossary version
definition_conflict bool βœ… Default false
conflict_variants string[] βœ… Default []. The competing definitions β€” the pipeline never picks a winner

Removed in v2: formula_latex (β†’ defining_formula_id) Β· interpretation (belongs in the Interpretation Pack, reached via term_id) Β· domain Β· company Β· language.

domain and company were never in the extraction prompt's field guide β€” the model filled them by copying worked examples β€” and neither was span-guarded, so a wrong value was undetectable. language is a genuine three-value field (id/en/mixed) but nothing consumes it and it is recoverable from the document; it is cut for simplicity and can return if a consumer needs it.

Writer split. The LLM fills only term, full_name, definition and subdomain_tags. term_id, defining_formula_id, source_wording, mention_count, extraction_status, diff_status, definition_conflict and conflict_variants are all set by deterministic code β€” the model never writes its own ids, links or audit fields, so none of them can be hallucinated.


6 Β· rules.json β†’ RuleEntry[] β€” 8 fields

[
  {
    "rule_id": "r_c07be41f22",
    "rule_type": "calculation",
    "statement": "Production yang digunakan dalam perhitungan adalah produksi hasil joint survey.",
    "condition": "Apabila data joint survey belum tersedia",
    "consequence": "maka digunakan data produksi berdasarkan truck count sebagai dasar perhitungan",
    "formula_ids": ["f_77c1ab9e40"],
    "term_ids": ["t_a4e2f81b06"],
    "provenance": {
      "doc_id": "STD_2026_006_MNO",
      "span": "Apabila data joint survey belum tersedia maka digunakan data produksi berdasarkan truck count sebagai dasar perhitungan",
      "page": 4,
      "section_no": "2.1.5",
      "chunk_id": "STD_2026_006_MNO::0009"
    }
  }
]
Field Type Req Notes
rule_id string βœ… Re-specified in v2. Was written by the model, therefore unstable; now deterministic
rule_type enum βœ… interpretation | calculation β€” see Β§2
statement string | null β€” The rule as prose. Prose, not an equation β€” evidence carrying both yields the prose, and the equation goes to the formula branch
condition string | null β€” The triggering condition, split out β€” not left buried in prose
consequence string | null β€” What follows when the condition holds
formula_ids string[] βœ… New in v2. The formulas this rule constrains. Default []
term_ids string[] βœ… New in v2. The terms this rule governs. Default []
provenance Provenance βœ…

Removed in v2: applies_to (free text β†’ term_ids[]) Β· subdomain_tags Β· language Β· extraction_status.

condition + consequence is the point of this artifact, and the seed of the interpretation logic tree: a rule stored as one paragraph cannot be attached to a skill later; a trigger and its consequence can.

Deliberately not modelled yet: the interpretation logic tree, action benchmarks, and explicit conclude / do-not-conclude verdicts. Those are the direction of travel, but designing their schema against zero real examples would be guessing. They land when a document containing them does.


7 Β· formulas.json β†’ FormulaEntry[] β€” 7 fields

[
  {
    "formula_id": "f_77c1ab9e40",
    "name": "Production",
    "formula_latex": "Production = MOHH \\times Qty \\times PA \\times UA \\times Pty",
    "variables": [
      { "symbol": "MOHH", "term_id": "t_1c9d5e7a83", "meaning": "Machine on Hand Hours" },
      { "symbol": "Qty",  "term_id": "t_9f2a41c0b7", "meaning": "Quantity" },
      { "symbol": "PA",   "term_id": "t_a4e2f81b06", "meaning": "Physical Availability" },
      { "symbol": "UA",   "term_id": "t_6d40b2fc17", "meaning": "Utilization of Availability" },
      { "symbol": "Pty",  "term_id": "t_e8b3092d55", "meaning": "Productivity" }
    ],
    "unit": "BCM",
    "latex_verification": "unverified_operator",
    "provenance": {
      "doc_id": "STD_2026_006_MNO",
      "span": "Production = MOHH x Qty x PA x UA x Pty",
      "page": 1,
      "section_no": "2.1",
      "chunk_id": "STD_2026_006_MNO::0003"
    }
  }
]
Field Type Req Notes
formula_id string βœ… New in v2
name string | null β€” What the formula computes, as named in the source
formula_latex string | null β€” Transcribed, never derived. The pipeline copies what the document states
variables[] object[] βœ… { symbol (required), term_id | null, meaning | null }. term_id is new in v2
unit string | null β€” Result unit if stated
latex_verification enum | null β€” Added 2026-08-26. Which guarantee formula_latex actually carries. Set by the span check, never by the model. null when there is no formula_latex to characterise
provenance Provenance βœ…

Removed in v2: extraction_status.

Read latex_verification before trusting formula_latex. Three paths leave the field populated and they are not equally strong. verified means the transcription was located in the source markup (Chunk.latex), compared in a notation-aware canonical form. The two unverified_* values mean the claim could not be checked either way and the entry is guarded by its provenance span alone β€” the same guarantee everything carried before 2026-08-26.

unverified_operator is not rare. MinerU's pipeline backend transcribes multiplication as a bare letter x, so every product formula from such an artifact lands here β€” including the sample above. An artifact parsed with hybrid/high carries \times and verifies normally.

variables[].symbol is not always a legend abbreviation β€” the extraction prompt's own worked example emits "Total Hours" and "Breakdown" as symbols. So term_id resolution is best-effort and nulls are expected; the link stage reports its dangle rate rather than hiding it.


8 Β· document_brief.json β†’ DocumentBrief β€” 6 fields (single object)

βœ… The v3 reshape is complete, corrected 2026-09-08. This banner previously said the LLM-facing parts β€” purpose as a span-guarded verbatim quote, and the subdomains aggregate β€” were "still proposed and not built". They shipped on 2026-09-02 and the banner was never updated, so it contradicted the field table directly below it, which dates both changes. Verified live: a domain entry persisted on 2026-09-07 carries purpose_verbatim and subdomains.

The delivery really was split, and that part was deliberate: the deterministic half (outline read off the artifact, key_parameters corroborated against extracted terms) shipped first because it costs nothing to prove, and the LLM-facing half followed once an eval run could pay for itself.

{
  "brief_id": "d_2ef60a8c19",
  "title": "Standard Parameter Produksi & ECA",
  "purpose_verbatim": "Standard Parameter ini merupakan pedoman dalam perhitungan parameter produksi yang meliputi Quantity (Qty) Physical Availability (PA), Utilization of Availability (UA), Productivity (PTY), serta Equipment Capacity Analysis (ECA).",
  "outline": [
    "1. TUJUAN PARAMETER",
    "2.1.2. Quantity (Qty)",
    "2.1.3. Physical of Availability (PA)",
    "2.2. Equipment Capacity Analysis (ECA)"
  ],
  "key_parameters": [
    {"surface": "Physical Availability (PA)", "term_id": "t_8c7359e27f"},
    {"surface": "Equipment Capacity Analysis (ECA)", "term_id": null}
  ],
  "subdomains": ["equipment", "production"],
  "n_terms": 5,
  "n_formulas": 6,
  "n_rules": 0,
  "provenance": {
    "doc_id": "STD_2026_006_MNO",
    "span": "Standard Parameter ini merupakan pedoman dalam perhitungan parameter produksi",
    "page": 0,
    "section_no": "1",
    "chunk_id": "STD_2026_006_MNO::0000"
  }
}
Field Type Req
brief_id string βœ… New in v2. b_ prefix β€” reconciled with the code 2026-09-11
title string | null β€”
purpose_verbatim string | null β€” Renamed + guarded 2026-09-02 (was purpose)
outline string[] βœ… Default []. Added 2026-09-02
subdomains SubdomainEnum[] βœ… Default []. Added 2026-09-02
n_terms Β· n_formulas Β· n_rules int βœ… Default 0. Added 2026-09-02
key_parameters KeyParameter[] βœ… Default []. Shape changed 2026-09-02 β€” was string[]
provenance Provenance βœ…

KeyParameter β€” {surface: string, term_id: string | null}. surface is what the model named; term_id is the glossary term that corroborates it, resolved deterministically by link.py. A null term_id means nothing in the extracted glossary supports that name β€” reported, never guessed, and it earns its own review_queue.json row (Β§9). Dangle count is reported as key_parameters_unresolved.

subdomains is AGGREGATED from the glossary entries' own subdomain_tags, never asked of the model β€” each tag was already chosen once, per term, with that term's evidence in front of the model, and a second document-level classification would be the same judgement made with less context and no way to check it. Ordered by tag frequency descending, alphabetical tiebreak, so two runs of one document agree. n_terms / n_formulas / n_rules are counts of the other three artifacts, computed after every branch and the diff have run.

outline is DERIVED from the artifact's heading_path, in reading order, de-duplicated. No LLM and no spend. It carries the strong guarantee the rest of this entry cannot: it is verbatim source structure, so there is nothing to hallucinate β€” including wording a reader would be tempted to fix (the reference document heads a section "Physical of Availability (PA)"). Empty is normal for a document with no headings.

Removed in v2: scope (overlaps purpose, frequently null) Β· summary_md.

βœ… This branch is span-checked as of 2026-09-02. It used to be the one exception: it asked the model to SUMMARISE, and a summary is not verbatim by construction, so the primary anti-hallucination control could not apply. It now asks the model to locate β€” purpose_verbatim is the document's own statement of purpose, copied, not a sentence composed about it β€” so title and purpose_verbatim are guarded like any other transcription. A value that is not verbatim is nulled and recorded in rejected.json, never repaired.

Measured on the reference document: 0 fields rejected, and a deliberately fabricated purpose is refused with purpose_verbatim is not a verbatim transcription of the source.

Two consequences worth knowing. title will look wrong sometimes and that is correct β€” the reference document yields "BUMA | STANDARD PARAMETERProduction Parameter andEquipment Capacity Analysis (ECA)", missing spaces and all, because that is what the PDF says; tidying it is the UI's job at display time, never the pipeline's. And this weakens the case that the branch needs a larger model tier (DEV_PLAN Β§0.8 D3, risk R5): locating a sentence is a much smaller ask than composing one.


9 Β· review_queue.json β†’ ReviewQueueRow[] β€” 12 fields

This is the consumption surface β€” the payload a review UI binds to. Deliberately denormalised so a row renders with no joins.

⚠️ Changed 2026-09-02: the queue is no longer glossary-only. Every row now carries row_kind, and a consumer must switch on it rather than assume every row is a term. A key_parameter row has a surface and no term_id, no definition and no mention count β€” a client that binds term_id unconditionally will break on it.

[
  {
    "rank": 1,
    "row_kind": "term",
    "term_id": "t_a4e2f81b06",
    "term": "PA",
    "definition": "Adalah ketersediaan fisik suatu equipment/unit yang menunjukkan proporsi waktu equipment/unit tersebut berada pada kondisi available (siap pakai) selama suatu periode tertentu.",
    "source_wording": "2.1.3. Physical of Availability (PA)",
    "mention_count": 43,
    "page": 3,
    "section_no": "2.1.3",
    "span": "Adalah ketersediaan fisik suatu equipment/unit",
    "review_reason": "source wording differs from the expanded name β€” confirm which is correct"
  },
  {
    "rank": 2,
    "row_kind": "key_parameter",
    "term_id": null,
    "brief_id": "d_2ef60a8c19",
    "term": "Grouping (Composite)",
    "definition": null,
    "source_wording": null,
    "mention_count": 0,
    "page": 0,
    "section_no": "1",
    "span": "Standard Parameter ini merupakan pedoman dalam perhitungan parameter produksi",
    "review_reason": "named as a key parameter but no extracted term corroborates it β€” confirm it is real"
  }
]
Field Type Req Notes
rank int βœ… 1-based display order β€” see sorting below
row_kind term | key_parameter βœ… New 2026-09-02. Switch on this before reading anything else
term_id string | null βœ… New in v2. What an approve / edit / reject decision attaches to, so it survives a re-run. Null on a key_parameter row β€” that it resolves to no term is the reason the row exists
brief_id string β€” key_parameter rows only β€” what the decision attaches to instead. Named brief_id on the wire, not domain_id; see the correction under Sort order
surface_full string | null β€” key_parameter rows only. New 2026-09-11. The untruncated surface when term had to be shortened to a head; null when it did not
term Β· definition Β· source_wording Β· mention_count Copied from the entry
page Β· page_no Β· section_no Β· span Lifted out of provenance so a row is self-contained. Render page_no (1-based), not page
review_reason string βœ… Human-readable, one of five

Removed in v2: extraction_status Β· diff_status Β· definition_conflict Β· conflict_variants β€” all four are already expressed by review_reason, and remain available on the entry itself.

Sort order β€” changed 2026-09-11. (conflicts, then mention_count descending, with unresolved key parameters spliced in at most one per five term rows).

Conflicts still outrank everything, because a contradiction is a decision only the expert can make. What changed is the tier beneath them. Unresolved key parameters used to occupy one, which meant every unresolved parameter outranked every routine term regardless of how many there were β€” and the branch produces them in bulk on exactly the documents it understands least. On the Komatsu shop manual that was 24 junk rows standing in front of Engine, a term with 108 mentions, which sat at rank 25.

The original reason for promoting them is preserved and is still sound: that branch has no span check at all, so an unresolvable surface is the one hallucination signal it can raise, and it must not sink beneath sixty routine rows where nobody reaches it. A cap expresses that without letting the tail take the queue over. Measured by replaying the three persisted v2 runs:

Document unresolved key parameters first one at rank top term, rank
BUMA 0 β€” Qty 1 β†’ 1 (unchanged)
Komatsu 24 1 β†’ 6 Engine 25 β†’ 1
Open Pit 7 1 β†’ 6 Resource 8 β†’ 1

Row counts are identical before and after on all three β€” nothing is dropped; parameters the splice cannot place are appended rather than discarded.

Also new on key_parameter rows: surface_full. The branch is supposed to return a parameter name and on a textbook it returns the right concept wrapped in its whole defining sentence (185 characters, in one measured case). term now carries a head β€” split at the first ., : or newline β€” and surface_full carries the untruncated text, or null when no shortening was needed. The head is only taken past 60 characters, which is what keeps a legitimate short name containing a period ("No. of units") intact.

⚠️ Field-name correction, 2026-09-11. This document described the key_parameter row's owning id as domain_id. The code emits brief_id and always has: the K12 rename changed the entity from BriefContext to DomainContext but left the id field's name alone, deliberately β€” renaming it is a migration that invalidates every expert approval recorded against the old prefix. The wire field is brief_id. Corrected here rather than in code, per the trust order.

review_reason values:

Value Meaning
conflicting definitions β€” expert decision required Contradiction found; the pipeline picked no winner
source wording differs from the expanded name β€” confirm which is correct e.g. "Physical of Availability" vs "Physical Availability"
term found but no definition in document Correct abstention β€” 48 of 58 on the reference run
definition rejected by span check or absent Control fired; see rejected.json
routine confirmation Nothing anomalous
named as a key parameter but no extracted term corroborates it β€” confirm it is real key_parameter rows only. Added 2026-09-02

10 Β· Linking model

GlossaryEntry is the hub. Every link is computed after both branches run, by a deterministic never-throw link stage. No additional LLM calls, and nothing a model could fabricate.

RuleEntry ──term_ids[]────────► GlossaryEntry ◄──variables[].term_id── FormulaEntry
     └────formula_ids[]──────────────────────────────────────────────────────► β–²
                          GlossaryEntry ──defining_formula_idβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Two properties that must hold:

  1. A dangling link is null or an empty array β€” never a fabrication. Same rule as spans.
  2. The "appears in" edge is derived, not stored. That PA appears inside Production = MOHH Γ— Qty Γ— PA Γ— UA Γ— Pty is recoverable by scanning formulas[].variables[]. Do not add a field for it.

11 Β· Audit payloads β€” unchanged from v1

rejected.json β†’ RejectedField[] β€” what the span check caught, kept so a reviewer sees the control working, not only what it let through. The reference run rejected 2 fields.

[
  {
    "entry_term": "<term>",
    "field": "full_name",
    "offending_value": "<value the model returned>",
    "reason": "full_name is not a verbatim transcription of the source",
    "branch": "glossary"
  }
]
Field Type Req
entry_term string βœ…
field string βœ…
offending_value string βœ…
reason string βœ…
branch glossary | rule | formula | summary βœ…

usage.json β†’ CallUsage[] β€” per-call accounting. cached_tokens is read from the API, never modelled: caching does not engage below 1024 prompt tokens, so assuming it would understate input cost by roughly 10Γ—.

[
  {
    "branch": "glossary",
    "deployment": "gpt-5.4-nano",
    "tier": "nano",
    "prompt_tokens": 2676,
    "cached_tokens": 2389,
    "completion_tokens": 214,
    "latency_s": 3.4,
    "retries": 0,
    "structured_output_mode": "json_schema",
    "simulated": false
  }
]

simulated: true marks a --mock run β€” never a quality measurement.


12 Β· Reference run β€” real figures

From eval/knowledge/results/v2_full_document_2026-08-24_093051.json. The parsing module built the artifact; extraction ran all four branches.

Artifact schema_version 0.2.0 Β· 31 chunks Β· 9 pages Β· mineru/pipeline 3.4.4
Funnel 188 raw mentions β†’ 163 after noise β†’ 58 clusters (compression 2.81Γ—)
Entries 58 glossary Β· 8 rules Β· 6 formulas
LLM 77 calls Β· 139,283 prompt tokens Β· 79.6% cached Β· 168.4 s
Quality 48 abstained, 10 with a definition Β· 2 fields rejected by the span check
E1 term-filter recall 0.8049 (kill line 0.70, PASS)
E3 schema-fill precision 0.90 (kill line 0.80, PASS β€” the prototype failed this at 0.75)

v1 of this contract quoted 13 chunks / 66 clusters / 66 entries. That run predates the section-aware chunker. Anything calibrated against those figures should be re-checked.

These figures also predate the v2 schema itself: they were measured on the v1 prompts, which asked for fields that no longer exist. The funnel counts (chunks, clusters, calls) should carry over unchanged β€” the filter, cluster and rank stages were not touched β€” but E3 must be re-scored, because trimming a prompt can move definition quality even when the scored fields are unchanged.


13 Β· Open items

# Item Owner
1 Unevaluated. Cutting interpretation / formula_latex and adding rule_type are prompt changes, and no eval has run against them β€” there is no document or parsed artifact on the dev box to run one. Same position X20 was committed in, and it needs the same first paid run to clear Rifqi
2 Rule-branch quality is a separate track. The branch measures 3/15, an unevaluated prompt fix is already in flight (X20), and a prompt/gold contradiction caps it at 14/15 (X23). rule_type and the X23 de-contamination make this the third and fourth uncommitted-to-measurement changes on that one prompt β€” the first paid run measures the prompt as a whole, and isolating any single edit would cost one full run each. Expect the raw score to move in both directions: X20 should add rules back, X23 removes a free hit (example 1 was gold PTY_PRODUCTION_SOURCE verbatim) and lifts the 14/15 ceiling Rifqi
2b rule.txt grew again. X20 already pushed it past the 1024-token cache floor; example 4 adds more. Only the API's cached_tokens proves a hit β€” verify on the first paid run Rifqi
3 Persistence. Stage output is JSON on disk; parsed artifacts, candidate entries, glossary versions and the approval audit trail still need one consolidated DDL handoff. Go owns the schema β€” Python never executes DDL Rifqi β†’ Harry
4 Page indexing at the API boundary β€” stays 0-based to the UI, or converts once at the boundary? Currently 0-based everywhere Harry + Rifqi
5 Endpoint shape β€” four endpoints, or one with an ?artifact= parameter? Not started; the offline CLI is the honest first milestone Rifqi

Caveat on the sample values

Field shapes, types and defaults are the agreed v2 target. Run-level figures in Β§12 are literal, from the cited result file. The per-entity sample values are illustrative β€” the underlying out/*.json from that run is not in version control, and all *_id values shown are placeholders that demonstrate format, not real hashes. Replace them with literal output once v2 is implemented and a run exists.