"""Deterministic entity ids. Every id here is **content-derived and stable across re-runs** of the same document. That property is the whole point, and it is not cosmetic: - The product is an expert working through a review queue. If Mas Beta approves 40 entries, we re-run with a tuned prompt, and the ids have moved, we have lost which 40 he approved. Approval continuity — not linking — is the reason these exist. - The two ids v1 carried both failed that test. `cluster_id` was a **frequency-rank ordinal** (`#c000`, `#c001`, … over a list sorted by mention count), so one changed mention count reshuffled every id after it. `rule_id` was **written by the model**, so consecutive runs at temperature 0 could return different slugs for the same rule. Two rules for anything added here: 1. **Never positional.** An id must not depend on where an entity landed in a list, because ordering is a function of measurements that legitimately move. 2. **Never model-supplied.** A model cannot be asked for a stable identifier; it has no way to know what it emitted last time. **Scope is per document.** The same term in two documents gets two ids. Deciding that BUMA's "PA" and a textbook's "PA" are the same concept is an expert judgement at review time — the same reasoning that makes the pipeline record "Physical *of* Availability" rather than normalising it, and the reason MTTR at BUMA must not silently merge with MTTR in IT. **Known limit:** ids derive from content, so if the content changes the id changes — retune clustering such that a term canonicalises differently and its `term_id` moves. The durable fix is a persisted registry, which belongs with persistence (DEV_PLAN §0.8 D2), not here. """ from __future__ import annotations import hashlib # sha256 rather than sha1: no security requirement either way, but sha1 trips # the repo's bandit lint rules and buys nothing here. _DIGEST_LEN = 10 def _digest(*parts: str) -> str: payload = "|".join(p or "" for p in parts).encode("utf-8") return hashlib.sha256(payload).hexdigest()[:_DIGEST_LEN] def term_id(doc_id: str, canonical: str) -> str: """Identity of one term as observed in one document.""" return f"t_{_digest(doc_id, canonical)}" def formula_id(doc_id: str, name: str | None, latex: str | None, chunk_id: str) -> str: """Identity of one formula. `name` and `formula_latex` are both Optional — a formula entry may carry neither — so the fallback chain ends at `chunk_id`, which always exists. Without the fallback an unnamed formula would collide with every other unnamed formula in the document. """ return f"f_{_digest(doc_id, name or latex or chunk_id)}" def rule_id(doc_id: str, chunk_id: str, char_start: int) -> str: """Identity of one rule, keyed to where its cue was found. Position **within a chunk** is stable in a way position within a list is not: it is a property of the document, not of a ranking. """ return f"r_{_digest(doc_id, chunk_id, str(char_start))}" def brief_id(doc_id: str) -> str: """Identity of one document's brief, so the document identifies it. Prefix history, and it is not churn: `b_` (BriefContext) -> `d_` (DomainContext, 2026-09-01) -> `b_` again (DocumentBrief, 2026-09-08). The middle move renamed the object without changing what it held - a per-document card. v4 builds the real scope-level `DomainContext` and gives `d_` to it, so this one goes back to the name and prefix that describe it. Safe to move only because nothing is approved against it: an id prefix change after an expert has ruled is a migration that loses their approvals, which is the exact failure this whole scheme exists to prevent. """ return f"b_{_digest(doc_id)}"