ishaq101's picture sofhiaazzhr's picture
/feat knowledge management (#20)
b68816f
Raw History Blame Contribute Delete
3.81 kB
"""Deterministic entity ids.
Every id here is **content-derived and stable across re-runs** of the same
document. That property is the whole point, and it is not cosmetic:
- The product is an expert working through a review queue. If Mas Beta approves
40 entries, we re-run with a tuned prompt, and the ids have moved, we have
lost which 40 he approved. Approval continuity β€” not linking β€” is the reason
these exist.
- The two ids v1 carried both failed that test. `cluster_id` was a
**frequency-rank ordinal** (`#c000`, `#c001`, … over a list sorted by mention
count), so one changed mention count reshuffled every id after it. `rule_id`
was **written by the model**, so consecutive runs at temperature 0 could
return different slugs for the same rule.
Two rules for anything added here:
1. **Never positional.** An id must not depend on where an entity landed in a
list, because ordering is a function of measurements that legitimately move.
2. **Never model-supplied.** A model cannot be asked for a stable identifier;
it has no way to know what it emitted last time.
**Scope is per document.** The same term in two documents gets two ids. Deciding
that BUMA's "PA" and a textbook's "PA" are the same concept is an expert
judgement at review time β€” the same reasoning that makes the pipeline record
"Physical *of* Availability" rather than normalising it, and the reason MTTR at
BUMA must not silently merge with MTTR in IT.
**Known limit:** ids derive from content, so if the content changes the id
changes β€” retune clustering such that a term canonicalises differently and its
`term_id` moves. The durable fix is a persisted registry, which belongs with
persistence (DEV_PLAN Β§0.8 D2), not here.
"""
from __future__ import annotations
import hashlib
# sha256 rather than sha1: no security requirement either way, but sha1 trips
# the repo's bandit lint rules and buys nothing here.
_DIGEST_LEN = 10
def _digest(*parts: str) -> str:
payload = "|".join(p or "" for p in parts).encode("utf-8")
return hashlib.sha256(payload).hexdigest()[:_DIGEST_LEN]
def term_id(doc_id: str, canonical: str) -> str:
"""Identity of one term as observed in one document."""
return f"t_{_digest(doc_id, canonical)}"
def formula_id(doc_id: str, name: str | None, latex: str | None, chunk_id: str) -> str:
"""Identity of one formula.
`name` and `formula_latex` are both Optional β€” a formula entry may carry
neither β€” so the fallback chain ends at `chunk_id`, which always exists.
Without the fallback an unnamed formula would collide with every other
unnamed formula in the document.
"""
return f"f_{_digest(doc_id, name or latex or chunk_id)}"
def rule_id(doc_id: str, chunk_id: str, char_start: int) -> str:
"""Identity of one rule, keyed to where its cue was found.
Position **within a chunk** is stable in a way position within a list is
not: it is a property of the document, not of a ranking.
"""
return f"r_{_digest(doc_id, chunk_id, str(char_start))}"
def brief_id(doc_id: str) -> str:
"""Identity of one document's brief, so the document identifies it.
Prefix history, and it is not churn: `b_` (BriefContext) -> `d_`
(DomainContext, 2026-09-01) -> `b_` again (DocumentBrief, 2026-09-08). The
middle move renamed the object without changing what it held - a per-document
card. v4 builds the real scope-level `DomainContext` and gives `d_` to it, so
this one goes back to the name and prefix that describe it.
Safe to move only because nothing is approved against it: an id prefix change
after an expert has ruled is a migration that loses their approvals, which is
the exact failure this whole scheme exists to prevent.
"""
return f"b_{_digest(doc_id)}"