Download docs/knowledge/KNOWLEDGE_OUTPUT_CONTRACT.md from DataEyond/Agentic-Service-Data-Eyond-Catalog: direct link, hf CLI and curl.
- Browser
- Download file 38.4 kB
-
https://huggingface.co/spaces/DataEyond/Agentic-Service-Data-Eyond-Catalog/resolve/main/docs/knowledge/KNOWLEDGE_OUTPUT_CONTRACT.md
- Command line
-
hf download hf://spaces/DataEyond/Agentic-Service-Data-Eyond-Catalog/docs/knowledge/KNOWLEDGE_OUTPUT_CONTRACT.md
-
curl -L -o KNOWLEDGE_OUTPUT_CONTRACT.md https://huggingface.co/spaces/DataEyond/Agentic-Service-Data-Eyond-Catalog/resolve/main/docs/knowledge/KNOWLEDGE_OUTPUT_CONTRACT.md
Knowledge Pipeline β Output Contract v2
Implemented 2026-08-26.
src/knowledge_extraction/emits this shape. The schema review of 2026-08-26 landed with it: ids on every entity, formulas referenced rather than restated, rules linked to formulas and terms, and 16 fields cut.Not yet re-measured. The prompt changes this required have not been scored β no source document, parsed artifact or MinerU install exists on the dev box (all three
data/directories are gitignored and absent), so the free stages have nothing to read. The last measured figures in Β§12 predate these edits. See Β§13.
Supersedes: v1 of this document (2026-08-24)
Reference run: v2_full_document_2026-08-24_093051 Β· BUMA STD_2026_006_MNO Β· 9 pages Β· gpt-5.4-nano
Producer: python -m src.knowledge_extraction.cli <artifact.json> --extract
Status: offline artifacts only β no endpoint, no table.
0 Β· What changed from v1
Harry's review of v1 asked for entity ids, a formula reference instead of a restated formula, and a link between rules and formulas. All are adopted. The schema was also cut back β v1 carried fields nothing consumes.
| Change | Detail |
|---|---|
| Ids on every entity | term_id, formula_id, brief_id; rule_id re-specified. All deterministic and content-derived |
| Glossary no longer restates formulas | formula_latex removed β defining_formula_id reference |
| Rules link to formulas and terms | formula_ids[], term_ids[] |
rule_type added |
interpretation | calculation β see Β§2 for why |
| 16 fields cut, 8 added | 50 β 42 fields total |
chunk_id separator fixed |
DOC#0007 β DOC::0007, adopting the parser's format |
| Reference numbers corrected | v1 quoted a superseded run (13 chunks / 66 clusters). Current is 31 chunks / 58 clusters |
Why the ids had to be re-specified
Neither id in v1 survived a re-run. cluster_id was a frequency-rank ordinal
(f"{doc_id}#c{i:03d}" over a list sorted by mention count), and rule_id was written by the
model (rule.txt asks for SCREAMING_SNAKE_CASE). Either way, re-running with a tuned prompt
reshuffles the ids β and if an expert has approved 40 entries, we have lost which 40.
Approval continuity, not linking, is the real reason these exist.
1 Β· Id scheme
All ids are deterministic, content-derived, and stable across re-runs of the same document. None is model-generated; none is positional.
term_id = "t_" + sha256(doc_id + "|" + canonical)[:10]
formula_id = "f_" + sha256(doc_id + "|" + name_or_latex_or_chunk_id)[:10]
rule_id = "r_" + sha256(doc_id + "|" + chunk_id + "|" + char_start)[:10]
brief_id = "b_" + sha256(doc_id)[:10] # per-document card; `d_` belongs to the
# scope-level DomainContext in src/knowledge_domain/
Implemented in src/knowledge_extraction/ids.py. term_id is keyed on the cluster canonical,
not on the term the model returned β the cluster is the stable thing across runs, and the model is
free to answer "PA" or "Physical Availability" for the same cluster. formula_id falls back through
name β formula_latex β chunk_id, because both of the first two are Optional and without the
fallback every unnamed formula in a document would collide.
Scope is per document. The same term appearing in two documents receives two different
term_ids. Deciding that BUMA's "PA" and a textbook's "PA" are the same concept is an expert
judgement at review time, not something the pipeline asserts β the same reasoning that makes the
pipeline record "Physical of Availability" instead of normalising it, and the reason MTTR at BUMA
(breakdown duration Γ· breakdown frequency) must not silently merge with MTTR in IT.
Known limit. Ids derive from content, so if content changes the id changes β retune clustering
such that a term canonicalises differently and its term_id moves. The alternative is a persisted
registry, which belongs with persistence (DEV_PLAN Β§0.8 D2).
2 Β· The four artifacts
| File | Entity | Holds |
|---|---|---|
glossary.json |
GlossaryEntry[] |
Terms and their definitions |
rules.json |
RuleEntry[] |
Rules of thumb β how to interpret data or behaviour. Renamed from interpretation_pack.json 2026-09-08 β the file, the entity and the DB kind now share one name. The Interpretation Pack remains the deliverable's name in prose |
formulas.json |
FormulaEntry[] |
Formulas, transcribed from the document |
document_brief.json |
DocumentBrief |
One per-document card. Prefix history, not churn: b_ (BriefContext) -> d_ (DomainContext, 2026-09-01) -> b_ again (DocumentBrief, 2026-09-08), once v4 gave d_ to the real scope-level DomainContext in src/knowledge_domain/ |
Plus three derived/audit payloads: review_queue.json, rejected.json, usage.json.
Rule vs domain knowledge β where the line is
Added 2026-09-14 (o2). The two blurred on screen during the 2026-09-08 review, and a boundary nobody has written down is one the prompts cannot hold.
rule |
domain |
|
|---|---|---|
| Grain | One statement, from one document | One per company, composed across all of them |
| Shape | A condition and a consequence | Identity, boundary, conventions, measures |
| Evidence | Span-checked against its source | Derived from entries, or expert-declared |
| Answers | "β¦and how do I compute this correctly?" | "β¦what does this company measure, and what does it call it?" |
The test to apply. If removing it would change how a number is computed or sourced, it is a rule β "if joint-survey data is unavailable, use truck count" changes the figure. If removing it would only leave an agent less oriented β not knowing that this company works in overburden and BCM β it is domain.
Two consequences worth stating, because they are what the distinction is for:
- A rule is actionable alone; domain knowledge is not. A planner can apply a rule to a specific calculation. It cannot apply "this is a mining company" to anything β that shapes which questions make sense, not which column to sum.
- Domain is high-level and rules are specific, but "high-level" is not the test. A long, general rule is still a rule. The grain question β does this come from one document's sentence, or from the corpus as a whole? β settles it faster and does not drift.
Where it is genuinely both, prefer the rule: it carries provenance and a span, so an expert can check it. Domain fields are the ones the pipeline can least easily prove.
Why RuleEntry carries a rule_type
The Interpretation Pack exists to improve analytics insight from expert interpretation rules β interpretation logic, action benchmarks, and rules for when to conclude or not conclude from data.
Measured against the reference document, 0 of the 15 gold rules are interpretation rules. They
are calculation conventions (PA_COMPOSITE_WEIGHTED), data-sourcing rules
(PTY_PRODUCTION_SOURCE), and unit conventions (GAINLOSS_UNITS). This is a property of the
document type: a Standard Parameter document defines how to calculate parameters, not how to
read them. Interpretation rules live in other document types, or in the expert's head.
rule_type keeps both in one artifact rather than discarding the conventions we do extract:
interpretationβ how to read a value or behaviour. The target.calculationβ how a value is computed or sourced. What this corpus actually yields.
A consumer wanting only interpretation logic filters on rule_type.
2c Β· Evidence now carries figures and tables (parsing 0.4.0, 2026-09-10)
The artifact this half consumes changed shape: a figure is no longer its own chunk. It is folded into
the prose chunk at its reading-order position as , and travels as an Asset on that
chunk. adapter.py carries assets, referenced_by and table_html across; evidence_block renders
them into the prompt under ASSETS REFERENCED BY THIS CHUNK.
What that means for an extraction branch, and the distinction is enforced in the prompts:
| may be quoted? | may be a provenance span? | |
|---|---|---|
Asset.caption β printed in the document |
yes | yes |
Asset.description β written by a vision model |
no | no β it is not in the document, so the span check rejects it and the field is discarded |
Why it mattered: before 0.4.0 a figure chunk had an EMPTY text, so it produced no mentions, joined no
cluster, and was never selected as evidence. 13 figures across three documents, each already described
by a paid vision call, reached zero prompts. Full detail:
knowledge_pipeline_runs/CHANGELOG_v2.md Β§2.
β RESOLVED 2026-09-11 β Β§1, Β§2 and Β§8 now match the code. The note is kept for the record.
β οΈ Correction to Β§8 of this document (noted 2026-09-10). The code writes
document_brief.jsonwithbrief_idand ab_prefix; Β§8 below documentsdomain_context.jsonwithdomain_idandd_. The v3 rename was partly reverted for the per-document object once the scope-levelDomainContextbecame a separate thing insrc/knowledge_domain/. Β§8 does not describe what the code emits. Worth one edit, and it matters becauseentity_idprefixes are the key an expert's approvals hang on.
3 Β· Transport reality
Every payload is written by the CLI as a file into --out-dir (default out/knowledge/), as a
bare JSON array with no envelope. There is no endpoint and no table.
Three conventions hold across every payload:
pageis 0-based, exactly as the parser reports;page_nois the 1-based number a human reads and is what a review UI binds to. Both travel together on every entity, derived from one value so they cannot disagree (S6b, 2026-09-02).- All content fields are nullable by design.
nullis a valid, correct answer β the model abstaining, not failing. On the reference run 48 of 58 entries carry no definition. provenanceis mandatory andprovenance.spanis verbatim-checked. A field whose span cannot be located in the source is set tonulland logged torejected.jsonβ never repaired.
4 Β· Common types
Provenance β required on every entity
{
"doc_id": "STD_2026_006_MNO",
"span": "Adalah jumlah equipment/unit yang digunakan untuk menghasilkan produksi dalam suatu periode.",
"page": 3,
"page_no": 4,
"section_no": "2.1.2",
"chunk_id": "STD_2026_006_MNO::0007"
}
| Field | Type | Req | Notes |
|---|---|---|---|
doc_id |
string | β | |
span |
string | β | Verbatim from the source. The anti-hallucination control β the reviewer checks a quote against a page, not a claim against memory |
page |
int | null | β | 0-based, exactly as the parser reports |
page_no |
int | null | β | 1-based. Added 2026-09-02 (S6b). DERIVED from page, never stored, so it cannot drift. A review UI must render this one β an off-by-one is invisible until an expert opens the wrong page and concludes the provenance is wrong. null when page is unknown |
section_no |
string | null | β | e.g. "2.1.2"; null when the document is unnumbered |
chunk_id |
string | null | β | Back-reference into the parsed artifact. Format <doc_id>::<seq>, matching KNOWLEDGE_PARSING_OUTPUT_CONTRACT.md. Treat as opaque β do not split it |
Enums
| Enum | Values |
|---|---|
subdomain_tags[] |
production Β· maintenance Β· hauling Β· loading Β· drilling_blasting Β· equipment Β· safety Β· quality Β· planning Β· cost Β· geology Β· other |
extraction_status |
ok Β· no_definition_found Β· escalated |
diff_status |
new Β· duplicate Β· conflicting |
rule_type |
interpretation Β· calculation |
latex_verification |
verified Β· unverified_no_markup Β· unverified_operator |
subdomain_tags is a closed enum β classification, not generation. Adding a member changes what
the model is allowed to answer, which makes it a prompt change, not a data change.
5 Β· glossary.json β GlossaryEntry[] β 13 fields
[
{
"term_id": "t_9f2a41c0b7",
"term": "Qty",
"full_name": "Quantity",
"source_wording": "2.1.2. Quantity (Qty)",
"definition": "Adalah jumlah equipment/unit yang digunakan untuk menghasilkan produksi dalam suatu periode.",
"defining_formula_id": "f_31de08aa95",
"subdomain_tags": ["production", "equipment"],
"mention_count": 20,
"provenance": {
"doc_id": "STD_2026_006_MNO",
"span": "Adalah jumlah equipment/unit yang digunakan untuk menghasilkan produksi dalam suatu periode.",
"page": 3,
"section_no": "2.1.2",
"chunk_id": "STD_2026_006_MNO::0007"
},
"extraction_status": "ok",
"diff_status": "new",
"definition_conflict": false,
"conflict_variants": []
},
{
"term_id": "t_5b71ce0d34",
"term": "Gain/Loss",
"full_name": null,
"source_wording": null,
"definition": null,
"defining_formula_id": null,
"subdomain_tags": ["production"],
"mention_count": 4,
"provenance": {
"doc_id": "STD_2026_006_MNO",
"span": "Gain/Loss",
"page": 6,
"section_no": null,
"chunk_id": "STD_2026_006_MNO::0011"
},
"extraction_status": "no_definition_found",
"diff_status": "new",
"definition_conflict": false,
"conflict_variants": []
}
]
| Field | Type | Req | Notes |
|---|---|---|---|
term_id |
string | β | New in v2. Stable across re-runs β the key an approval hangs on |
term |
string | β | The only required content field |
full_name |
string | null | β | Expanded name. Span-guarded as a literal transcription |
source_wording |
string | null | β | The document's literal wording, un-normalised. Read deterministically from the section heading, not asked of the model β asked directly, it returned the tidied form |
definition |
string | null | β | null = correct abstention (48 of 58 on the reference run) |
defining_formula_id |
string | null | β | New in v2, replaces formula_latex. The formula that defines this term. Null when the formula branch did not extract one |
subdomain_tags |
enum[] | β | Defaults [] |
mention_count |
int | β | Default 0. Drives review-queue ordering |
provenance |
Provenance | β | |
extraction_status |
enum | β | Default ok |
diff_status |
enum | null | β | Set by the diff stage against the active glossary version |
definition_conflict |
bool | β | Default false |
conflict_variants |
string[] | β | Default []. The competing definitions β the pipeline never picks a winner |
Removed in v2: formula_latex (β defining_formula_id) Β· interpretation (belongs in the
Interpretation Pack, reached via term_id) Β· domain Β· company Β· language.
domainandcompanywere never in the extraction prompt's field guide β the model filled them by copying worked examples β and neither was span-guarded, so a wrong value was undetectable.languageis a genuine three-value field (id/en/mixed) but nothing consumes it and it is recoverable from the document; it is cut for simplicity and can return if a consumer needs it.
Writer split. The LLM fills only
term,full_name,definitionandsubdomain_tags.term_id,defining_formula_id,source_wording,mention_count,extraction_status,diff_status,definition_conflictandconflict_variantsare all set by deterministic code β the model never writes its own ids, links or audit fields, so none of them can be hallucinated.
6 Β· rules.json β RuleEntry[] β 8 fields
[
{
"rule_id": "r_c07be41f22",
"rule_type": "calculation",
"statement": "Production yang digunakan dalam perhitungan adalah produksi hasil joint survey.",
"condition": "Apabila data joint survey belum tersedia",
"consequence": "maka digunakan data produksi berdasarkan truck count sebagai dasar perhitungan",
"formula_ids": ["f_77c1ab9e40"],
"term_ids": ["t_a4e2f81b06"],
"provenance": {
"doc_id": "STD_2026_006_MNO",
"span": "Apabila data joint survey belum tersedia maka digunakan data produksi berdasarkan truck count sebagai dasar perhitungan",
"page": 4,
"section_no": "2.1.5",
"chunk_id": "STD_2026_006_MNO::0009"
}
}
]
| Field | Type | Req | Notes |
|---|---|---|---|
rule_id |
string | β | Re-specified in v2. Was written by the model, therefore unstable; now deterministic |
rule_type |
enum | β | interpretation | calculation β see Β§2 |
statement |
string | null | β | The rule as prose. Prose, not an equation β evidence carrying both yields the prose, and the equation goes to the formula branch |
condition |
string | null | β | The triggering condition, split out β not left buried in prose |
consequence |
string | null | β | What follows when the condition holds |
formula_ids |
string[] | β | New in v2. The formulas this rule constrains. Default [] |
term_ids |
string[] | β | New in v2. The terms this rule governs. Default [] |
provenance |
Provenance | β |
Removed in v2: applies_to (free text β term_ids[]) Β· subdomain_tags Β· language Β·
extraction_status.
condition + consequence is the point of this artifact, and the seed of the interpretation logic
tree: a rule stored as one paragraph cannot be attached to a skill later; a trigger and its
consequence can.
Deliberately not modelled yet: the interpretation logic tree, action benchmarks, and explicit conclude / do-not-conclude verdicts. Those are the direction of travel, but designing their schema against zero real examples would be guessing. They land when a document containing them does.
7 Β· formulas.json β FormulaEntry[] β 7 fields
[
{
"formula_id": "f_77c1ab9e40",
"name": "Production",
"formula_latex": "Production = MOHH \\times Qty \\times PA \\times UA \\times Pty",
"variables": [
{ "symbol": "MOHH", "term_id": "t_1c9d5e7a83", "meaning": "Machine on Hand Hours" },
{ "symbol": "Qty", "term_id": "t_9f2a41c0b7", "meaning": "Quantity" },
{ "symbol": "PA", "term_id": "t_a4e2f81b06", "meaning": "Physical Availability" },
{ "symbol": "UA", "term_id": "t_6d40b2fc17", "meaning": "Utilization of Availability" },
{ "symbol": "Pty", "term_id": "t_e8b3092d55", "meaning": "Productivity" }
],
"unit": "BCM",
"latex_verification": "unverified_operator",
"provenance": {
"doc_id": "STD_2026_006_MNO",
"span": "Production = MOHH x Qty x PA x UA x Pty",
"page": 1,
"section_no": "2.1",
"chunk_id": "STD_2026_006_MNO::0003"
}
}
]
| Field | Type | Req | Notes |
|---|---|---|---|
formula_id |
string | β | New in v2 |
name |
string | null | β | What the formula computes, as named in the source |
formula_latex |
string | null | β | Transcribed, never derived. The pipeline copies what the document states |
variables[] |
object[] | β | { symbol (required), term_id | null, meaning | null }. term_id is new in v2 |
unit |
string | null | β | Result unit if stated |
latex_verification |
enum | null | β | Added 2026-08-26. Which guarantee formula_latex actually carries. Set by the span check, never by the model. null when there is no formula_latex to characterise |
provenance |
Provenance | β |
Removed in v2: extraction_status.
Read
latex_verificationbefore trustingformula_latex. Three paths leave the field populated and they are not equally strong.verifiedmeans the transcription was located in the source markup (Chunk.latex), compared in a notation-aware canonical form. The twounverified_*values mean the claim could not be checked either way and the entry is guarded by its provenance span alone β the same guarantee everything carried before 2026-08-26.
unverified_operatoris not rare. MinerU'spipelinebackend transcribes multiplication as a bare letterx, so every product formula from such an artifact lands here β including the sample above. An artifact parsed withhybrid/highcarries\timesand verifies normally.
variables[].symbolis not always a legend abbreviation β the extraction prompt's own worked example emits"Total Hours"and"Breakdown"as symbols. Soterm_idresolution is best-effort and nulls are expected; the link stage reports its dangle rate rather than hiding it.
8 Β· document_brief.json β DocumentBrief β 6 fields (single object)
β The v3 reshape is complete, corrected 2026-09-08. This banner previously said the LLM-facing parts β
purposeas a span-guarded verbatim quote, and thesubdomainsaggregate β were "still proposed and not built". They shipped on 2026-09-02 and the banner was never updated, so it contradicted the field table directly below it, which dates both changes. Verified live: adomainentry persisted on 2026-09-07 carriespurpose_verbatimandsubdomains.The delivery really was split, and that part was deliberate: the deterministic half (
outlineread off the artifact,key_parameterscorroborated against extracted terms) shipped first because it costs nothing to prove, and the LLM-facing half followed once an eval run could pay for itself.
{
"brief_id": "d_2ef60a8c19",
"title": "Standard Parameter Produksi & ECA",
"purpose_verbatim": "Standard Parameter ini merupakan pedoman dalam perhitungan parameter produksi yang meliputi Quantity (Qty) Physical Availability (PA), Utilization of Availability (UA), Productivity (PTY), serta Equipment Capacity Analysis (ECA).",
"outline": [
"1. TUJUAN PARAMETER",
"2.1.2. Quantity (Qty)",
"2.1.3. Physical of Availability (PA)",
"2.2. Equipment Capacity Analysis (ECA)"
],
"key_parameters": [
{"surface": "Physical Availability (PA)", "term_id": "t_8c7359e27f"},
{"surface": "Equipment Capacity Analysis (ECA)", "term_id": null}
],
"subdomains": ["equipment", "production"],
"n_terms": 5,
"n_formulas": 6,
"n_rules": 0,
"provenance": {
"doc_id": "STD_2026_006_MNO",
"span": "Standard Parameter ini merupakan pedoman dalam perhitungan parameter produksi",
"page": 0,
"section_no": "1",
"chunk_id": "STD_2026_006_MNO::0000"
}
}
| Field | Type | Req |
|---|---|---|
brief_id |
string | β
New in v2. b_ prefix β reconciled with the code 2026-09-11 |
title |
string | null | β |
purpose_verbatim |
string | null | β Renamed + guarded 2026-09-02 (was purpose) |
outline |
string[] | β
Default []. Added 2026-09-02 |
subdomains |
SubdomainEnum[] | β
Default []. Added 2026-09-02 |
n_terms Β· n_formulas Β· n_rules |
int | β
Default 0. Added 2026-09-02 |
key_parameters |
KeyParameter[] | β
Default []. Shape changed 2026-09-02 β was string[] |
provenance |
Provenance | β |
KeyParameter β {surface: string, term_id: string | null}. surface is what the model named;
term_id is the glossary term that corroborates it, resolved deterministically by link.py. A
null term_id means nothing in the extracted glossary supports that name β reported, never
guessed, and it earns its own review_queue.json row (Β§9). Dangle count is reported as
key_parameters_unresolved.
subdomains is AGGREGATED from the glossary entries' own subdomain_tags, never asked of the
model β each tag was already chosen once, per term, with that term's evidence in front of the model,
and a second document-level classification would be the same judgement made with less context and no
way to check it. Ordered by tag frequency descending, alphabetical tiebreak, so two runs of one
document agree. n_terms / n_formulas / n_rules are counts of the other three artifacts,
computed after every branch and the diff have run.
outline is DERIVED from the artifact's heading_path, in reading order, de-duplicated. No LLM
and no spend. It carries the strong guarantee the rest of this entry cannot: it is verbatim source
structure, so there is nothing to hallucinate β including wording a reader would be tempted to fix
(the reference document heads a section "Physical of Availability (PA)"). Empty is normal for a
document with no headings.
Removed in v2: scope (overlaps purpose, frequently null) Β· summary_md.
β This branch is span-checked as of 2026-09-02. It used to be the one exception: it asked the model to SUMMARISE, and a summary is not verbatim by construction, so the primary anti-hallucination control could not apply. It now asks the model to locate β
purpose_verbatimis the document's own statement of purpose, copied, not a sentence composed about it β sotitleandpurpose_verbatimare guarded like any other transcription. A value that is not verbatim is nulled and recorded inrejected.json, never repaired.Measured on the reference document: 0 fields rejected, and a deliberately fabricated purpose is refused with
purpose_verbatim is not a verbatim transcription of the source.Two consequences worth knowing.
titlewill look wrong sometimes and that is correct β the reference document yields"BUMA | STANDARD PARAMETERProduction Parameter andEquipment Capacity Analysis (ECA)", missing spaces and all, because that is what the PDF says; tidying it is the UI's job at display time, never the pipeline's. And this weakens the case that the branch needs a larger model tier (DEV_PLAN Β§0.8 D3, risk R5): locating a sentence is a much smaller ask than composing one.
9 Β· review_queue.json β ReviewQueueRow[] β 12 fields
This is the consumption surface β the payload a review UI binds to. Deliberately denormalised so a row renders with no joins.
β οΈ Changed 2026-09-02: the queue is no longer glossary-only. Every row now carries
row_kind, and a consumer must switch on it rather than assume every row is a term. Akey_parameterrow has a surface and noterm_id, no definition and no mention count β a client that bindsterm_idunconditionally will break on it.
[
{
"rank": 1,
"row_kind": "term",
"term_id": "t_a4e2f81b06",
"term": "PA",
"definition": "Adalah ketersediaan fisik suatu equipment/unit yang menunjukkan proporsi waktu equipment/unit tersebut berada pada kondisi available (siap pakai) selama suatu periode tertentu.",
"source_wording": "2.1.3. Physical of Availability (PA)",
"mention_count": 43,
"page": 3,
"section_no": "2.1.3",
"span": "Adalah ketersediaan fisik suatu equipment/unit",
"review_reason": "source wording differs from the expanded name β confirm which is correct"
},
{
"rank": 2,
"row_kind": "key_parameter",
"term_id": null,
"brief_id": "d_2ef60a8c19",
"term": "Grouping (Composite)",
"definition": null,
"source_wording": null,
"mention_count": 0,
"page": 0,
"section_no": "1",
"span": "Standard Parameter ini merupakan pedoman dalam perhitungan parameter produksi",
"review_reason": "named as a key parameter but no extracted term corroborates it β confirm it is real"
}
]
| Field | Type | Req | Notes |
|---|---|---|---|
rank |
int | β | 1-based display order β see sorting below |
row_kind |
term | key_parameter |
β | New 2026-09-02. Switch on this before reading anything else |
term_id |
string | null | β | New in v2. What an approve / edit / reject decision attaches to, so it survives a re-run. Null on a key_parameter row β that it resolves to no term is the reason the row exists |
brief_id |
string | β | key_parameter rows only β what the decision attaches to instead. Named brief_id on the wire, not domain_id; see the correction under Sort order |
surface_full |
string | null | β | key_parameter rows only. New 2026-09-11. The untruncated surface when term had to be shortened to a head; null when it did not |
term Β· definition Β· source_wording Β· mention_count |
Copied from the entry | ||
page Β· page_no Β· section_no Β· span |
Lifted out of provenance so a row is self-contained. Render page_no (1-based), not page |
||
review_reason |
string | β | Human-readable, one of five |
Removed in v2: extraction_status Β· diff_status Β· definition_conflict Β·
conflict_variants β all four are already expressed by review_reason, and remain available on the
entry itself.
Sort order β changed 2026-09-11. (conflicts, then mention_count descending, with unresolved key parameters spliced in at most one per five term rows).
Conflicts still outrank everything, because a contradiction is a decision only the expert can make.
What changed is the tier beneath them. Unresolved key parameters used to occupy one, which meant
every unresolved parameter outranked every routine term regardless of how many there were β
and the branch produces them in bulk on exactly the documents it understands least. On the Komatsu
shop manual that was 24 junk rows standing in front of Engine, a term with 108 mentions, which
sat at rank 25.
The original reason for promoting them is preserved and is still sound: that branch has no span check at all, so an unresolvable surface is the one hallucination signal it can raise, and it must not sink beneath sixty routine rows where nobody reaches it. A cap expresses that without letting the tail take the queue over. Measured by replaying the three persisted v2 runs:
| Document | unresolved key parameters | first one at rank | top term, rank |
|---|---|---|---|
| BUMA | 0 | β | Qty 1 β 1 (unchanged) |
| Komatsu | 24 | 1 β 6 | Engine 25 β 1 |
| Open Pit | 7 | 1 β 6 | Resource 8 β 1 |
Row counts are identical before and after on all three β nothing is dropped; parameters the splice cannot place are appended rather than discarded.
Also new on key_parameter rows: surface_full. The branch is supposed to return a parameter
name and on a textbook it returns the right concept wrapped in its whole defining sentence (185
characters, in one measured case). term now carries a head β split at the first ., : or newline
β and surface_full carries the untruncated text, or null when no shortening was needed. The head
is only taken past 60 characters, which is what keeps a legitimate short name containing a period
("No. of units") intact.
β οΈ Field-name correction, 2026-09-11. This document described the
key_parameterrow's owning id asdomain_id. The code emitsbrief_idand always has: the K12 rename changed the entity fromBriefContexttoDomainContextbut left the id field's name alone, deliberately β renaming it is a migration that invalidates every expert approval recorded against the old prefix. The wire field isbrief_id. Corrected here rather than in code, per the trust order.
review_reason values:
| Value | Meaning |
|---|---|
conflicting definitions β expert decision required |
Contradiction found; the pipeline picked no winner |
source wording differs from the expanded name β confirm which is correct |
e.g. "Physical of Availability" vs "Physical Availability" |
term found but no definition in document |
Correct abstention β 48 of 58 on the reference run |
definition rejected by span check or absent |
Control fired; see rejected.json |
routine confirmation |
Nothing anomalous |
named as a key parameter but no extracted term corroborates it β confirm it is real |
key_parameter rows only. Added 2026-09-02 |
10 Β· Linking model
GlossaryEntry is the hub. Every link is computed after both branches run, by a deterministic
never-throw link stage. No additional LLM calls, and nothing a model could fabricate.
RuleEntry ββterm_ids[]βββββββββΊ GlossaryEntry βββvariables[].term_idββ FormulaEntry
βββββformula_ids[]βββββββββββββββββββββββββββββββββββββββββββββββββββββββΊ β²
GlossaryEntry ββdefining_formula_idββββββββββββββββββββ
Two properties that must hold:
- A dangling link is
nullor an empty array β never a fabrication. Same rule as spans. - The "appears in" edge is derived, not stored. That PA appears inside
Production = MOHH Γ Qty Γ PA Γ UA Γ Ptyis recoverable by scanningformulas[].variables[]. Do not add a field for it.
11 Β· Audit payloads β unchanged from v1
rejected.json β RejectedField[] β what the span check caught, kept so a reviewer sees the
control working, not only what it let through. The reference run rejected 2 fields.
[
{
"entry_term": "<term>",
"field": "full_name",
"offending_value": "<value the model returned>",
"reason": "full_name is not a verbatim transcription of the source",
"branch": "glossary"
}
]
| Field | Type | Req |
|---|---|---|
entry_term |
string | β |
field |
string | β |
offending_value |
string | β |
reason |
string | β |
branch |
glossary | rule | formula | summary |
β |
usage.json β CallUsage[] β per-call accounting. cached_tokens is read from the API, never
modelled: caching does not engage below 1024 prompt tokens, so assuming it would understate input
cost by roughly 10Γ.
[
{
"branch": "glossary",
"deployment": "gpt-5.4-nano",
"tier": "nano",
"prompt_tokens": 2676,
"cached_tokens": 2389,
"completion_tokens": 214,
"latency_s": 3.4,
"retries": 0,
"structured_output_mode": "json_schema",
"simulated": false
}
]
simulated: true marks a --mock run β never a quality measurement.
12 Β· Reference run β real figures
From eval/knowledge/results/v2_full_document_2026-08-24_093051.json. The parsing module built the
artifact; extraction ran all four branches.
| Artifact | schema_version 0.2.0 Β· 31 chunks Β· 9 pages Β· mineru/pipeline 3.4.4 |
| Funnel | 188 raw mentions β 163 after noise β 58 clusters (compression 2.81Γ) |
| Entries | 58 glossary Β· 8 rules Β· 6 formulas |
| LLM | 77 calls Β· 139,283 prompt tokens Β· 79.6% cached Β· 168.4 s |
| Quality | 48 abstained, 10 with a definition Β· 2 fields rejected by the span check |
| E1 term-filter recall | 0.8049 (kill line 0.70, PASS) |
| E3 schema-fill precision | 0.90 (kill line 0.80, PASS β the prototype failed this at 0.75) |
v1 of this contract quoted 13 chunks / 66 clusters / 66 entries. That run predates the section-aware chunker. Anything calibrated against those figures should be re-checked.
These figures also predate the v2 schema itself: they were measured on the v1 prompts, which asked for fields that no longer exist. The funnel counts (chunks, clusters, calls) should carry over unchanged β the filter, cluster and rank stages were not touched β but E3 must be re-scored, because trimming a prompt can move definition quality even when the scored fields are unchanged.
13 Β· Open items
| # | Item | Owner |
|---|---|---|
| 1 | Unevaluated. Cutting interpretation / formula_latex and adding rule_type are prompt changes, and no eval has run against them β there is no document or parsed artifact on the dev box to run one. Same position X20 was committed in, and it needs the same first paid run to clear |
Rifqi |
| 2 | Rule-branch quality is a separate track. The branch measures 3/15, an unevaluated prompt fix is already in flight (X20), and a prompt/gold contradiction caps it at 14/15 (X23). rule_type and the X23 de-contamination make this the third and fourth uncommitted-to-measurement changes on that one prompt β the first paid run measures the prompt as a whole, and isolating any single edit would cost one full run each. Expect the raw score to move in both directions: X20 should add rules back, X23 removes a free hit (example 1 was gold PTY_PRODUCTION_SOURCE verbatim) and lifts the 14/15 ceiling |
Rifqi |
| 2b | rule.txt grew again. X20 already pushed it past the 1024-token cache floor; example 4 adds more. Only the API's cached_tokens proves a hit β verify on the first paid run |
Rifqi |
| 3 | Persistence. Stage output is JSON on disk; parsed artifacts, candidate entries, glossary versions and the approval audit trail still need one consolidated DDL handoff. Go owns the schema β Python never executes DDL | Rifqi β Harry |
| 4 | Page indexing at the API boundary β stays 0-based to the UI, or converts once at the boundary? Currently 0-based everywhere | Harry + Rifqi |
| 5 | Endpoint shape β four endpoints, or one with an ?artifact= parameter? Not started; the offline CLI is the honest first milestone |
Rifqi |
Caveat on the sample values
Field shapes, types and defaults are the agreed v2 target. Run-level figures in Β§12 are literal,
from the cited result file. The per-entity sample values are illustrative β the underlying
out/*.json from that run is not in version control, and all *_id values shown are placeholders
that demonstrate format, not real hashes. Replace them with literal output once v2 is implemented
and a run exists.