Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download docs/architecture/06-evidence-and-confidence.md from thundercode/SatQuery: direct link, hf CLI and curl.
- Browser
- Download file 108 kB
-
https://huggingface.co/thundercode/SatQuery/resolve/main/docs/architecture/06-evidence-and-confidence.md
- Command line
-
hf download hf://thundercode/SatQuery/docs/architecture/06-evidence-and-confidence.md
-
curl -L -o 06-evidence-and-confidence.md https://huggingface.co/thundercode/SatQuery/resolve/main/docs/architecture/06-evidence-and-confidence.md
108 kB
| # 06 β Evidence and Confidence | |
| **Parent:** [Architecture hub](README.md) Β· **Sibling chapters:** | |
| [01 System overview](01-system-overview.md) Β· [05 Specialists](05-specialists.md) Β· | |
| [07 Configuration freeze](07-configuration-freeze.md) | |
| **Primary sources read for this chapter (all under `C:/Users/anish/satquery-ai/`):** | |
| | Source | What it establishes here | | |
| |---|---| | |
| | `evidence/engine.py` (714 lines) | the aggregation pipeline, purity contract, identity/sort/claim keys, `_deduplicate`, `_record_agreement`, `_renumber`, `evidence_type_for`, `evidence_from_box/region/geospatial`, `confidence_for`, `evidence_digest` | | |
| | `evidence/confidence.py` (434 lines) | the honesty rule, `TemperatureCalibration`, `_is_effective`, `_logit`/`_sigmoid`, `_EPS`, `calibrate`, `calibrate_result`, `load_calibration` and its ordered candidate search | | |
| | `core/schemas.py` (462 lines) | `Evidence`, `EvidenceType`, `ConfidenceBreakdown`, `ExecutionTrace`, `TraceStep`, `ModelRef`, `SpecialistResult`, `CoordinateSystem`, and every validator | | |
| | `configs/base.yaml` (Β§`evidence`, Β§`confidence`) | `evidence.max_items: 32`, `confidence.temperature_scaling: true`, `confidence.calibration_file: calibration_v001.json` | | |
| | `artifacts/calibration_v001.json` | the measured fitted temperature and the reliability diagram | | |
| | `frontend/assets/js/core.js` | `SQ.EVENT_NAMES` β the eight execution events | | |
| | `frontend/assets/js/mission.js` | `markState()` and the `.trace__fill` width formula | | |
| | `docs/ARCHITECTURE_FREEZE.md` Β§1, Β§3, Β§5 | the layer verbs and the frozen non-negotiables | | |
| | `docs/PHASE13_EVIDENCE_ENGINE.md` | the phase record for this package (77 tests) | | |
| | `docs/API_CONTRACT.md` Β§2.4, Β§3.3, Β§4 | the client-facing shape of evidence and confidence | | |
| | `docs/DEPLOYMENT_ARCHITECTURE.md` Β§5.5 | the F-16 owner ruling on artifact refs | | |
| | `docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§7, Β§8 | the measured calibration and evidence integration results | | |
| --- | |
| ## 1. Where this subsystem sits | |
| `docs/ARCHITECTURE_FREEZE.md` Β§5 gives every layer exactly one verb: | |
| > Router *understands*; policy engine *decides*; specialists *compute*; VLM *explains*; | |
| > **evidence engine *proves*.** | |
| Two consecutive stages implement that last verb: | |
| ```mermaid | |
| flowchart LR | |
| SPEC["Specialists<br/>compute"] -->|"SpecialistResult.evidence"| EV["Evidence engine<br/>aggregate()<br/>collect β dedup β sort β<br/>annotate β renumber β cap"] | |
| EV -->|"EvidenceCollection"| CONF["Confidence<br/>confidence_for() / calibrate()"] | |
| CONF -->|"ConfidenceBreakdown"| RES["ResultEnvelope<br/>+ ExecutionTrace"] | |
| style EV fill:#1f6feb22,stroke:#1f6feb | |
| style CONF fill:#1f6feb22,stroke:#1f6feb | |
| ``` | |
| The freeze Β§1 pipeline diagram places them adjacently: | |
| ``` | |
| +---------+---------+---------+ | |
| | | | | | |
| VQA/CAP GROUNDING CHANGE OPTICAL-SAR | |
| | | | | | |
| +---------+---------+---------+ | |
| | | |
| v | |
| Evidence engine <- this chapter, part A | |
| | | |
| v | |
| Confidence calibration <- this chapter, part B | |
| | | |
| v | |
| Result normaliser -> JSON + trace + PDF | |
| ``` | |
| **Division of labour, stated precisely.** `evidence/engine.py` does **not** re-derive any | |
| specialist's claim. The module docstring is explicit (`evidence/engine.py:10-14`): | |
| > This module is the "proves" step. Specialists each emit `Evidence` for what *they* | |
| > computed (see `GroundingSpecialist._build_evidence`). This engine does not re-derive any | |
| > of that. Its job is aggregation: | |
| > | |
| > collect across specialists -> order -> deduplicate -> renumber -> bound | |
| **Status of this subsystem.** | |
| | Component | Status | Basis | | |
| |---|---|---| | |
| | Evidence aggregation | `IMPLEMENTED` + `VERIFIED` | `evidence/engine.py`; 77 tests in `tests/unit/test_evidence_engine.py` (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§7) | | |
| | Dedup / ordering / renumber / cap | `IMPLEMENTED` + `VERIFIED` | pinned by named tests (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§2.2βΒ§2.5) | | |
| | Confidence honesty rule (uncalibrated pass-through) | `IMPLEMENTED` + `VERIFIED` | `evidence/confidence.py:296-311` | | |
| | Temperature-scaling transform | `IMPLEMENTED` + `VERIFIED` | monotonicity + endpoint-safety tests (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.3) | | |
| | Calibration artifact fitted | `MEASURED` | `artifacts/calibration_v001.json` | | |
| | Calibration **improves** ECE | **`REJECTED` β it does not** | ECE 0.013755 β 0.014929 (worse); see Β§7 | | |
| | Evidence engine wired into the controller | `IMPLEMENTED` (Phase 14) | `docs/PHASE13_EVIDENCE_ENGINE.md` Β§8 recorded it as "not yet wired"; the live run path emits the eight events (Β§9) | | |
| | Artifact rendering / retrieval | `OPEN` β deliberately `null` in v1 | F-16 ruling, Β§3.2 | | |
| --- | |
| # Part A β The evidence schema | |
| ## 2. `Evidence` β the canonical record | |
| `Evidence` is defined once, in `core/schemas.py:216-253`. Every specialist emits it; nothing | |
| invents a second shape (`docs/ARCHITECTURE_FREEZE.md` Β§3: *"Single schema for every | |
| specialist. No specialist invents its own shape."*). | |
| ```python | |
| class Evidence(BaseModel): | |
| model_config = ConfigDict(extra="forbid") | |
| evidence_id: str = Field(default_factory=lambda: _new_id("ev")) | |
| type: EvidenceType | |
| source_specialist: str | |
| coordinate_system: CoordinateSystem | None = None | |
| coordinates: list[float] | None = None | |
| score: float | None = Field(default=None, ge=0.0, le=1.0) | |
| artifact_ref: str | None = Field( | |
| default=None, | |
| description=( | |
| "Reference to an externally retrievable artifact. NEVER a " | |
| "filesystem path (F-16, owner ruling 2026-09-23): v1 exposes no " | |
| "artifact-serving endpoint, so this is null unless a deployment " | |
| "supplies a client-fetchable reference. An artifact may still be " | |
| "written server-side where configured; being written is not the " | |
| "same as being retrievable." | |
| ), | |
| ) | |
| payload: dict[str, Any] = Field(default_factory=dict) | |
| ``` | |
| ### 2.1 Every field, exhaustively | |
| | Field | Type | Required | Default | Meaning | Note | | |
| |---|---|---|---|---|---| | |
| | `evidence_id` | `str` | no | `_new_id("ev")` β `ev_<12 hex>` | the item's identity | **run-local**, engine-assigned after aggregation (Β§5.8); a uuid before that | | |
| | `type` | `EvidenceType` | **yes** | β | the *kind* of proof | closed 11-member enum, Β§3 | | |
| | `source_specialist` | `str` | **yes** | β | which specialist made this claim | a **single string**, not a set β this is load-bearing, Β§5.7 | | |
| | `coordinate_system` | `CoordinateSystem \| None` | no | `None` | the frame the coordinates live in | **mandatory in practice** for spatial types β see the validator, Β§4 | | |
| | `coordinates` | `list[float] \| None` | no | `None` | the geometry | a flat list; a box is `[x1, y1, x2, y2]` | | |
| | `score` | `float \| None` | no | `None` | reliability, bounded | `ge=0.0, le=1.0` β Pydantic rejects out-of-range | | |
| | `artifact_ref` | `str \| None` | no | `None` | a *retrievable* artifact reference | **always `null` in v1** β Β§2.3 | | |
| | `payload` | `dict[str, Any]` | no | `{}` | observations that *support* the claim without being it | where `corroborated_by` is written, Β§5.7 | | |
| ### 2.2 What `Evidence` deliberately does **not** have | |
| These absences are load-bearing; each one is why a downstream design decision exists. | |
| | Absent field | Why it is absent | Consequence | | |
| |---|---|---| | |
| | any ordering field (`index`, `rank`, `order`) | the engine owns order; a specialist must not pre-empt it | order is defined by `_sort_key`, Β§5.6 | | |
| | `value` | the field is called **`score`** | `docs/PHASE19_FINAL_HARDENING.md` Β§3.6 records that the docs once said `value` and that was one of nine validated defects | | |
| | `source` | the field is called **`source_specialist`** | same defect class; a frontend using `source` renders nothing | | |
| | `summary` | not in the schema | same defect class | | |
| | `contributing_specialists: list[str]` | **proposed and NOT adopted** | see Β§5.7 and `docs/PHASE13_EVIDENCE_ENGINE.md` Β§5 | | |
| | `label` at top level | labels live in `payload` | `evidence_from_box` writes `payload["label"]` | | |
| | a `retrievable` boolean | retrieval is not a v1 capability | `artifact_ref` being `null` *is* the signal, Β§2.3 | | |
| `model_config = ConfigDict(extra="forbid")` is what makes all of the above enforceable: a | |
| response carrying `value` instead of `score` is a **422**, not a silently-ignored key. The | |
| conformance test asserts this directly β `docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§D records | |
| "`extra="forbid"` on all eight client-facing models; `GeoMetadata` is the one documented | |
| `extra="allow"` exception." | |
| ### 2.3 `artifact_ref` is permanently `null` in v1 β owner ruling F-16 | |
| This is the single most mis-documented field in the system, and the ruling is recorded in | |
| three places: the field's own `description` (`core/schemas.py:225-235`), the client contract | |
| (`docs/API_CONTRACT.md` Β§2.4 "Artifact refs β `null` in v1, and why"), and the audit | |
| (`docs/DEPLOYMENT_ARCHITECTURE.md` Β§5.5). | |
| **The ruling (F-16, owner, 2026-09-23): never expose filesystem paths.** | |
| `docs/API_CONTRACT.md:586-602` states the contract in full: | |
| > **Every `artifact_ref` and `change_map` in a v1 response is `null`.** This is a deliberate | |
| > contract, not a missing value. | |
| > | |
| > F-16 (owner ruling 2026-09-23): **never expose filesystem paths.** The specialists *do* | |
| > render their artifacts β the change map and the optical/SAR views are written | |
| > server-side β but their location is an operator fact, not a client-facing one. A response | |
| > that carried the server's path would disclose the deployment's directory layout to an | |
| > unauthenticated caller, and nothing the frontend can do requires it. | |
| > | |
| > **No `artifact://` URI is fabricated in its place.** v1 has **no artifact-serving | |
| > endpoint**, so a URI would be a promise the service cannot keep β strictly worse than | |
| > `null`, because the frontend would build a link that 404s. | |
| The docstring's own phrasing is the crispest statement of the principle: | |
| > **being written is not the same as being retrievable.** | |
| ```mermaid | |
| flowchart TB | |
| subgraph Server["Server side (operator facts)"] | |
| R["specialist renders change map / views"] | |
| W["file written where configured<br/>change.artifact_dir"] | |
| R --> W | |
| end | |
| subgraph Client["Client-facing contract (v1)"] | |
| N["artifact_ref = null"] | |
| P["payload statistics<br/>total_change_pixels<br/>n_components_kept<br/>threshold"] | |
| WARN["warnings[] entry:<br/>artifact NOT retrievable"] | |
| end | |
| W -.->|"deliberately NOT exposed"| X["filesystem path<br/>(never sent)"] | |
| R --> N | |
| R --> P | |
| R --> WARN | |
| style X fill:#f8514922,stroke:#f85149 | |
| style N fill:#3fb95022,stroke:#3fb950 | |
| ``` | |
| **What replaces the ref** (`docs/API_CONTRACT.md:604-610`): | |
| | Removed | Replaced by | | |
| |---|---| | |
| | `change_map` path | `null`, plus the change statistics in the CHANGE_MAP evidence's `payload` (`total_change_pixels`, `n_components_kept`, `threshold`) | | |
| | view `artifact_ref` path | `null`, plus `payload.rendered` / `payload.retrievable` / `payload.retrieval` | | |
| | β | an explicit `warnings[]` entry saying the artifact is **NOT retrievable** | | |
| **Two live carriers, and the half-implementation that followed.** The audit records that the | |
| ruling was initially applied to only one carrier | |
| (`docs/DEPLOYMENT_ARCHITECTURE.md:361`): | |
| > **Two live carriers** β `change` and `croma` β and **both** had to be brought into line: | |
| > fixing only the `croma` carrier left the ruling half implemented. | |
| `docs/STATUS.md:21` names the lesson: *"The finding worth carrying forward: the F-16 ruling | |
| was HALF IMPLEMENTED, and the suite was green."* | |
| **F-16c β the consequence the fix introduced, measured and left open.** The ruling removed a | |
| proxy and thereby changed a `degraded` signal. The `SpecialistResult` validator used to read | |
| (`core/schemas.py:362-372`, retained as a comment): | |
| ```python | |
| if self.task is Task.CHANGE and self.change_map is None and not self.regions: | |
| ... "change analysis produced no spatial output" ... degraded = True | |
| ``` | |
| `change_map` was a *proxy* for "a map was produced". With the ref permanently `null`, the | |
| clause collapsed to `not regions`, and a **successful no-change analysis began reporting | |
| `degraded: true`**. The clause was therefore **removed** under ruling F-16c | |
| (`core/schemas.py:362-395`). The comment left in its place is worth quoting because it states | |
| the reasoning better than a summary could: | |
| > That conflated two different things, and the conflation WAS the defect. `degraded` means | |
| > "the analysis could not be fully performed": no trained detector, or co-registration too | |
| > poor to support a spatial claim. Both are set by the specialist itself | |
| > (`specialists/change/specialist.py:345` and `:365`) and neither is this validator's to | |
| > invent. **"No change was detected" is a NORMAL, successful outcome.** | |
| And, on why no replacement clause was added: | |
| > No replacement clause is added, deliberately. A narrower "empty answer => degraded" rule | |
| > was tried and backed out: it is not what the ruling asked for, it invented a semantic the | |
| > specialist already owns, and it made a pre-existing, unrelated fixture | |
| > (`test_change_with_regions_is_not_degraded`, a result with regions and no answer) fail. **A | |
| > fix that forces edits to tests it has nothing to do with is signalling over-reach, not | |
| > diligence.** CHANGE is the one task whose `degraded` flag is now set entirely by its | |
| > specialist. | |
| **Status:** F-16 `CLOSED` (both carriers conform). F-16c `RESOLVED` by clause removal β the | |
| narrower replacement rule is `REJECTED`. `mask_ref` is *not* an exemption: the audit | |
| corrected an earlier working note that had listed it alongside `artifact_ref` | |
| (`docs/DEPLOYMENT_ARCHITECTURE.md:868-871`). | |
| ## 3. `EvidenceType` β the closed 11-member vocabulary | |
| ```python | |
| class EvidenceType(str, Enum): | |
| IMAGE_CROP = "image_crop" | |
| TILE = "tile" | |
| BOUNDING_BOX = "bounding_box" | |
| MASK = "mask" | |
| CHANGE_MAP = "change_map" | |
| OPTICAL_VIEW = "optical_view" | |
| SAR_VIEW = "sar_view" | |
| JOINT_FEATURE_REGION = "joint_feature_region" | |
| STATISTIC = "statistic" | |
| GEOLOCATION = "geolocation" | |
| AVAILABILITY_MASK = "availability_mask" # C-1: modality trust evidence | |
| ``` | |
| `core/schemas.py:65-76`. **All eleven members, with what each one asserts:** | |
| | # | Member | Wire value | Asserts | Typically emitted by | | |
| |---|---|---|---|---| | |
| | 1 | `IMAGE_CROP` | `image_crop` | a rectangular raster crop exists | preprocessing / tiling | | |
| | 2 | `TILE` | `tile` | one tile of the tiling policy was examined | `preprocessing/tiling.py` | | |
| | 3 | `BOUNDING_BOX` | `bounding_box` | a rectangle locates a referent | grounding, change | | |
| | 4 | `MASK` | `mask` | a per-pixel region exists | grounding, change | | |
| | 5 | `CHANGE_MAP` | `change_map` | a bitemporal difference map exists | change specialist | | |
| | 6 | `OPTICAL_VIEW` | `optical_view` | the optical rendering of an optical/SAR pair | optical-SAR specialist | | |
| | 7 | `SAR_VIEW` | `sar_view` | the SAR rendering of an optical/SAR pair | optical-SAR specialist | | |
| | 8 | `JOINT_FEATURE_REGION` | `joint_feature_region` | a region defined in CROMA's **joint** embedding space | optical-SAR specialist | | |
| | 9 | `STATISTIC` | `statistic` | a measured scalar or count, with no geometry | VQA, caption, unsupported, and any geometry-less region | | |
| | 10 | `GEOLOCATION` | `geolocation` | the raster is georeferenced and *where* it sits | geospatial layer | | |
| | 11 | `AVAILABILITY_MASK` | `availability_mask` | which modality channels were actually available | sensor adapter (C-1) | | |
| ### 3.1 `availability_mask` β the member the freeze prose omits | |
| `docs/ARCHITECTURE_FREEZE.md` Β§3's prose list of evidence types stops at ten members and does | |
| **not** name `availability_mask`. The enum member is nevertheless real, and the engine says | |
| so explicitly (`evidence/engine.py:441-447`): | |
| > The members are exactly those of `core.schemas.EvidenceType`. The freeze's section 3 prose | |
| > list omits `availability_mask`; that member is real (C-1: modality trust evidence) and is | |
| > included here because **a missing member would otherwise force a wrong fallback.** | |
| Its existence follows from finding C-1: the channel-availability mask is a first-class fusion | |
| input and is **never** a CROMA input (`core/schemas.py:7`; `docs/PHASE0_CONTRACT_VALIDATION.md` | |
| Β§1.2). Because the mask is a first-class *input*, it is also a first-class *observation* β | |
| "these twelve optical channels were present and these two SAR channels were not" is a fact a | |
| consumer can audit. `docs/API_CONTRACT.md:579` lists all eleven values for the frontend. | |
| ### 3.2 Which types are spatial | |
| The engine's own validator names the spatial set (`core/schemas.py:240-247`) β this is the | |
| authoritative list, not a prose paraphrase: | |
| ```python | |
| spatial = { | |
| EvidenceType.BOUNDING_BOX, | |
| EvidenceType.MASK, | |
| EvidenceType.CHANGE_MAP, | |
| EvidenceType.TILE, | |
| EvidenceType.IMAGE_CROP, | |
| EvidenceType.JOINT_FEATURE_REGION, | |
| } | |
| ``` | |
| | Type | Spatial? | Why | | |
| |---|---|---| | |
| | `BOUNDING_BOX` | **yes** | a box is meaningless without a frame | | |
| | `MASK` | **yes** | pixels need a frame | | |
| | `CHANGE_MAP` | **yes** | a map is a raster | | |
| | `TILE` | **yes** | a tile is a rectangle of a raster | | |
| | `IMAGE_CROP` | **yes** | a crop is a rectangle of a raster | | |
| | `JOINT_FEATURE_REGION` | **yes** | a region in a feature grid still needs a frame | | |
| | `STATISTIC` | no | a scalar has no geometry | | |
| | `GEOLOCATION` | no *(by this validator)* | geo bounds are already self-describing via `payload["crs"]`; `evidence_from_geospatial` sets `CoordinateSystem.GEO` when bounds are given | | |
| | `OPTICAL_VIEW` | no *(by this validator)* | a whole-frame view is not a sub-region claim | | |
| | `SAR_VIEW` | no *(by this validator)* | as above | | |
| | `AVAILABILITY_MASK` | no *(by this validator)* | a channel-presence vector is not spatial | | |
| ## 4. The validator that refuses coordinates without a frame | |
| ```python | |
| @model_validator(mode="after") | |
| def _spatial_needs_crs(self) -> "Evidence": | |
| spatial = { | |
| EvidenceType.BOUNDING_BOX, | |
| EvidenceType.MASK, | |
| EvidenceType.CHANGE_MAP, | |
| EvidenceType.TILE, | |
| EvidenceType.IMAGE_CROP, | |
| EvidenceType.JOINT_FEATURE_REGION, | |
| } | |
| if self.type in spatial and self.coordinates and self.coordinate_system is None: | |
| raise ValueError( | |
| f"evidence type '{self.type.value}' carries coordinates " | |
| "but no coordinate_system" | |
| ) | |
| return self | |
| ``` | |
| `core/schemas.py:238-253`. Three things about this are deliberate: | |
| 1. **It is a `model_validator(mode="after")`, not a `field_validator`.** The rule is | |
| *cross-field* β it depends on `type`, `coordinates` and `coordinate_system` together, so | |
| it cannot be expressed on any single field. | |
| 2. **The trigger is `self.coordinates`, not `is not None`.** An empty list is falsy and does | |
| not trip the rule; a non-empty list does. A spatial item with **no** coordinates at all is | |
| legal (a `MASK` whose geometry lives only in a `mask_ref`). | |
| 3. **It raises rather than defaulting.** Defaulting to `normalized_0_1` would silently | |
| mislabel pixel or geo coordinates as normalized, which is the failure mode the finding | |
| exists to prevent. `docs/ARCHITECTURE_FREEZE.md` Β§3 puts the requirement bluntly: | |
| > Every spatial object carries `coordinate_system` β `{normalized_0_1, pixel, geo}`. | |
| `docs/ARCHITECTURE_FREEZE.md` Β§2.3 states the origin: internal box coordinates are | |
| **normalised 0β1**, **always with an explicit `coordinate_system` field** (finding C-5). | |
| ### 4.1 `CoordinateSystem` β three values, exact spellings | |
| ```python | |
| class CoordinateSystem(str, Enum): | |
| """Never omit this. A bare box is meaningless without it. (C-5)""" | |
| NORMALIZED_0_1 = "normalized_0_1" | |
| PIXEL = "pixel" | |
| GEO = "geo" | |
| ``` | |
| `core/schemas.py:57-62`. | |
| | Member | Wire value | Frame | | |
| |---|---|---| | |
| | `NORMALIZED_0_1` | `normalized_0_1` | fractions of the frame, `[0, 1]` | | |
| | `PIXEL` | `pixel` | integer-ish pixel indices | | |
| | `GEO` | `geo` | a projected/geographic CRS named in `payload["crs"]` | | |
| > **Documentation hazard, measured.** `docs/PHASE19_FINAL_HARDENING.md` Β§3.6 records that the | |
| > docs once spelled these `normalized` / `geographic`. The real values are | |
| > `normalized_0_1` / `geo`. *"A frontend using the documented names would send values the | |
| > server rejects with 422."* `docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§D confirms the live enum is | |
| > exactly `{normalized_0_1, pixel, geo}`. | |
| `Evidence`'s `coordinate_system` is `CoordinateSystem | None` β the field is optional *in | |
| type* but effectively mandatory *in practice*, because the validator above enforces it for | |
| every spatial type that carries coordinates. | |
| --- | |
| # Part B β The aggregation pipeline | |
| ## 5. `EvidenceEngine.aggregate()` β collect β dedup β sort β annotate β renumber β cap | |
| ### 5.1 Why the module exists: three problems a specialist cannot solve | |
| The module docstring enumerates them (`evidence/engine.py:18-36`); each is real and none is | |
| fixable inside a single specialist: | |
| | # | Problem | Why a specialist cannot fix it | | |
| |---|---|---| | |
| | 1 | **Stable identity** | `Evidence.evidence_id` defaults to a random uuid β fine within one result, useless once four specialists' evidence is merged into one trace. *"Nothing can be cited."* | | |
| | 2 | **Order** | `Evidence` has no ordering field, so the collection's order follows **specialist completion order** β the same inputs produce different JSON on different runs. *"Pure functions must not do that, and this whole system's reproducibility argument rests on it."* | | |
| | 3 | **Duplication** | the VQA specialist emits a `STATISTIC` carrying its answer, *and* that answer travels in `SpecialistResult.answer`; two specialists that georeference the same asset emit the same `GEOLOCATION`. *"Neither is wrong, and neither knows about the other."* | | |
| ### 5.2 The pipeline, in one place | |
| ```mermaid | |
| flowchart TB | |
| A["results: SpecialistResult<br/>or Sequence[SpecialistResult]<br/>or evidence: Iterable[Evidence]"] --> B["collect<br/>extend(result.evidence)"] | |
| B --> C["sources = sorted({source_specialist})<br/><i>recorded BEFORE the cap</i>"] | |
| C --> D["_deduplicate<br/>identity = _identity_key<br/>payloads MERGED on collision"] | |
| D --> E["sorted(key=_sort_key)<br/>type β Β· specialist β Β·<br/>score β Β· coordinates β"] | |
| E --> F["_record_agreement<br/>_claim_key groups β<br/>payload['corroborated_by']"] | |
| F --> G["annotated[:max_items]<br/>dropped_over_limit = total β len(capped)"] | |
| G --> H["_renumber β evidence_001β¦"] | |
| H --> I["EvidenceCollection"] | |
| style D fill:#1f6feb22,stroke:#1f6feb | |
| style F fill:#1f6feb22,stroke:#1f6feb | |
| ``` | |
| The implementation, verbatim (`evidence/engine.py:324-350`): | |
| ```python | |
| def _canonicalise(self, raw: Iterable[Evidence]) -> EvidenceCollection: | |
| """Dedup -> sort -> renumber -> cap. The whole pipeline, in one place.""" | |
| items = list(raw) | |
| sources = sorted({item.source_specialist for item in items}) | |
| deduped, dropped_duplicates = self._deduplicate(items) | |
| ordered = sorted(deduped, key=_sort_key) | |
| # Two specialists can make the same claim, and `deduplicate` keeps both | |
| # (different `source_specialist` -> different key). Record the agreement | |
| # now that ordering is fixed, so the annotation is part of the pure | |
| # pipeline rather than a post-hoc edit a caller might forget. | |
| annotated = self._record_agreement(ordered) | |
| total = len(annotated) | |
| capped = annotated[: self.max_items] | |
| dropped_over_limit = total - len(capped) | |
| return EvidenceCollection( | |
| items=[ | |
| self._renumber(item, index) for index, item in enumerate(capped, start=1) | |
| ], | |
| sources=sources, | |
| dropped_duplicates=dropped_duplicates, | |
| dropped_over_limit=dropped_over_limit, | |
| total_before_limit=total, | |
| truncated=dropped_over_limit > 0, | |
| ) | |
| ``` | |
| Note the ordering choice: **annotate comes after sort but before cap**. Annotating before | |
| sorting would be equivalent (the annotation is a payload key, which is not in `_sort_key`), | |
| but the docstring's stated reason is that the annotation must be *inside* the pure pipeline | |
| rather than "a post-hoc edit a caller might forget". | |
| ### 5.3 The entry point and its argument rules | |
| ```python | |
| def aggregate( | |
| self, | |
| results: SpecialistResult | Sequence[SpecialistResult] | None = None, | |
| *, | |
| evidence: Iterable[Evidence] | None = None, | |
| ) -> EvidenceCollection: | |
| ``` | |
| `evidence/engine.py:280-320`. Two argument invariants, both raising: | |
| | Condition | Behaviour | Reason given in the docstring | | |
| |---|---|---| | |
| | both `results` and `evidence` are `None` | `ValueError("aggregate() requires results= or evidence=")` | there is nothing to aggregate | | |
| | both are supplied | `ValueError("aggregate() accepts results= or evidence=, not both")` | ambiguity about precedence | | |
| | empty list | returns an empty `EvidenceCollection` | *"Never raises on empty input."* | | |
| **Only `result.evidence` is read.** The docstring explains why reading both forms would | |
| double-count (`evidence/engine.py:292-297`): | |
| > Specialists state their evidence twice β in `SpecialistResult.evidence` and via | |
| > `Specialist.produce_evidence`, which currently delegates to the same list β so only | |
| > `result.evidence` is read. Reading both would double-count every item and inflate the | |
| > duplicate count. | |
| ### 5.4 The purity contract | |
| ```python | |
| """THE PURITY CONTRACT | |
| ------------------- | |
| `aggregate` is pure and deterministic: | |
| * source results are never mutated -- `model_copy` is used to rebuild the | |
| collection rather than editing `Evidence.evidence_id` in place; | |
| * ids are assigned from the sorted position, not from input order; | |
| * dedup is order-insensitive by construction (the key is built from the | |
| sorted specialist list); | |
| * no clock, no RNG, no I/O. | |
| Same inputs -> byte-identical output. `tests/unit/test_evidence_engine.py` | |
| pins this. | |
| """ | |
| ``` | |
| `evidence/engine.py:38-50`. | |
| | Property | Mechanism | Pinned by | | |
| |---|---|---| | |
| | source non-mutation | `model_copy(update={...})` at every write β `_renumber` (`:430`), payload merge (`:385`), agreement (`:424`) | *"the original keeps its uuid and the aggregated copy is a different object"* (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§2.1) | | |
| | ids from sorted position | `enumerate(capped, start=1)` after `sorted(...)` | Β§5.8 | | |
| | order-insensitive dedup | key content-only; `_record_agreement` sorts `names` | Β§5.5, Β§5.7 | | |
| | no clock / RNG / I/O | the module imports only `hashlib`, `collections.abc`, `dataclasses`, `typing` | `evidence/engine.py:80-102` | | |
| The dedup merge rebuilds rather than mutates, and the comment says why | |
| (`evidence/engine.py:379-382`): | |
| > Survivor's own keys win; the loser only fills gaps. Mutating the survivor in place would be | |
| > fine here because `merged` holds the same object, but rebuilding keeps the "never mutate a | |
| > source result" contract true even if the caller kept a reference. | |
| ### 5.5 Deduplication β identity is *defined*, not assumed | |
| The module docstring (`evidence/engine.py:52-68`) states the principle: | |
| > `Evidence.evidence_id` is a uuid default, so it cannot be part of an identity key β two | |
| > structurally identical items from two runs would never dedup. The key is therefore the | |
| > *content* of the observation. | |
| ```python | |
| def _identity_key(item: Evidence) -> tuple[Any, ...]: | |
| coords = ( | |
| tuple(_round(c) for c in item.coordinates) | |
| if item.coordinates is not None | |
| else None | |
| ) | |
| return ( | |
| item.type.value, | |
| item.source_specialist, | |
| item.coordinate_system.value if item.coordinate_system else None, | |
| coords, | |
| _round(item.score), | |
| ) | |
| ``` | |
| `evidence/engine.py:125-142`. | |
| | Key component | Included? | Rationale | | |
| |---|---|---| | |
| | `type` | **yes** | a box and a mask are different claims even at the same coordinates | | |
| | `source_specialist` | **yes** | keeps two specialists' identical claims as two items (Β§5.7) | | |
| | `coordinate_system` | **yes** | the same four numbers in `pixel` and in `geo` are different claims | | |
| | rounded `coordinates` | **yes** | disagreeing geometry β two claims; suppressing it would be a silent contradiction | | |
| | rounded `score` | **yes** | different reliability is a different observation | | |
| | `evidence_id` | **NO** | *"including it would mean dedup never fires in production"* | | |
| | `payload` | **NO** | *"a detail of the claim, not a different claim"* | | |
| | `artifact_ref` | **NO** | *"two renderings of one claim are still one claim"* | | |
| The docstring on the exclusions is the clearest statement of intent | |
| (`evidence/engine.py:60-68`): | |
| > `payload` and `artifact_ref` are deliberately EXCLUDED. Two items that agree on the same | |
| > geolocation, one carrying a `crs` in its payload and one not, have made the same claim | |
| > about the world; deduplicating them is correct, and the surviving item's payload is merged | |
| > with the discarded one's so the `crs` is not lost. Conversely two items that share a type | |
| > and score but disagree on coordinates are two different claims and are both kept β which | |
| > matters, because **suppressing a spatial disagreement would be exactly the silent | |
| > contradiction the freeze forbids.** | |
| **The rounding constant.** | |
| ```python | |
| #: Rounding applied before an evidence item is turned into a dedup key. Six | |
| #: decimals on normalized coordinates is ~1e-6 of the frame, far below any | |
| #: meaningful spatial difference, but enough to absorb float noise from two | |
| #: specialists rounding the same value two different ways. | |
| _KEY_PRECISION = 6 | |
| def _round(value: float | None) -> float | None: | |
| return None if value is None else round(float(value), _KEY_PRECISION) | |
| ``` | |
| `evidence/engine.py:114-122`. `_round` is `None`-safe: a `None` coordinate stays `None` and | |
| does not become `0.0` in the *identity* key (unlike the sort key, Β§5.6, where a missing | |
| coordinate becomes `0.0` because a total order needs a value). | |
| **Payloads are MERGED on collision.** | |
| ```python | |
| dropped += 1 | |
| # Survivor's own keys win; the loser only fills gaps. | |
| payload = {**item.payload, **existing.payload} | |
| if payload != existing.payload: | |
| updated = existing.model_copy(update={"payload": payload}) | |
| seen[key] = updated | |
| merged[merged.index(existing)] = updated | |
| ``` | |
| `evidence/engine.py:378-387`. The spread order is `{**loser, **survivor}` β because Python | |
| dict unpacking is last-wins, `existing.payload` (the survivor's) overrides the loser's on | |
| key collision. This is the "survivor's own keys win; the loser only fills gaps" rule | |
| expressed directly in the merge order. | |
| | Behaviour | Result | | |
| |---|---| | |
| | survivor had `crs`, loser did not | survivor's `crs` survives | | |
| | loser had `crs`, survivor did not | `crs` is **filled in** from the loser β the information is not lost | | |
| | both had `crs` with different values | survivor's wins; the loser's value is dropped | | |
| | no payload difference | no `model_copy` is performed (the `if payload != existing.payload` guard) | | |
| `docs/PHASE13_EVIDENCE_ENGINE.md` Β§2.3 names the test that pins this: | |
| `test_deduplication_ignores_payload_differences_and_merges_them`. | |
| **What dedup does *not* merge.** Because the key includes `source_specialist`, *"only | |
| same-specialist repeats are merged here; two specialists agreeing is handled by | |
| `_record_agreement`, which preserves the second specialist's identity rather than discarding | |
| it"* (`evidence/engine.py:356-361`). | |
| ### 5.6 The total sort order β and why each level exists | |
| ```python | |
| def _sort_key(item: Evidence) -> tuple[Any, ...]: | |
| if item.score is None: | |
| score_key: tuple[int, float] = (1, 0.0) | |
| else: | |
| score_key = (0, -float(item.score)) | |
| coords = ( | |
| tuple(_round(c) or 0.0 for c in item.coordinates) | |
| if item.coordinates is not None | |
| else () | |
| ) | |
| return (item.type.value, item.source_specialist, score_key, coords) | |
| ``` | |
| `evidence/engine.py:165-190`. | |
| | # | Key | Direction | Why it exists (verbatim from the docstring) | | |
| |---|---|---|---| | |
| | 1 | `type` | ascending | *"groups like with like, so a reader sees all the boxes together"* | | |
| | 2 | `source_specialist` | ascending | *"stable and meaningful, unlike the uuid"* | | |
| | 3 | `score` | **descending** | *"within a type, the strongest claim leads"* β implemented as `-score` so the tuple stays uniformly ascending | | |
| | 4 | `coordinates` | ascending | *"the tie-breaker that makes the order total. Without it two items of the same type, specialist and score would sort by Python's stable-sort insertion order, which reintroduces input-order dependence."* | | |
| `docs/PHASE13_EVIDENCE_ENGINE.md` Β§2.2 restates key 4 and names the pinning test: | |
| > Key 4 is not decoration. Without it, two items of the same type, specialist and score would | |
| > sort by Python's stable-sort insertion order β which reintroduces exactly the input-order | |
| > dependence the sort exists to remove. Pinned by | |
| > `test_equal_scores_order_deterministically_by_coordinates`. | |
| **Unscored items sort last.** *"Items with no score sort AFTER items with a score: a missing | |
| score is not evidence of strength."* (`evidence/engine.py:178-179`). The mechanism is a | |
| `(0, β¦)` / `(1, β¦)` discriminator in `score_key`: a scored item leads with `0`, an unscored | |
| item leads with `1`, so every scored item precedes every unscored item **within the same | |
| `(type, specialist)` group**. | |
| ```mermaid | |
| flowchart LR | |
| subgraph G1["type=bounding_box, specialist=grounding"] | |
| direction TB | |
| A1["score 0.91"] --> A2["score 0.62"] --> A3["score None<br/><i>unscored last</i>"] | |
| end | |
| subgraph G2["type=statistic, specialist=vqa"] | |
| direction TB | |
| B1["score 0.74"] | |
| end | |
| G1 --> G2 | |
| style A3 fill:#d2992222,stroke:#d29922 | |
| ``` | |
| ### 5.7 `_record_agreement` β corroboration, and why nothing is ever removed | |
| **The problem.** `source_specialist` is a single `str`, not a set (the schema is frozen), so | |
| two specialists making the same claim are two *different* items under `_identity_key` and | |
| **both survive**. The docstring calls this *"the correct conservative behaviour"* | |
| (`evidence/engine.py:70-77`): | |
| > That is the correct conservative behaviour: collapsing them would require the engine to pick | |
| > a winner, and the engine has no basis for that. What it does instead is keep both and record | |
| > the disagreement-free agreement in the collection's `sources` and in each item's `payload` | |
| > via `_record_agreement`. | |
| **The claim key β deliberately distinct from the identity key.** | |
| ```python | |
| def _claim_key(item: Evidence) -> tuple[Any, ...]: | |
| """Identity of the *claim*, ignoring which specialist made it. | |
| Used only to detect corroboration. Deliberately distinct from | |
| `_identity_key`: that one answers "is this the same item", this one answers | |
| "are these two specialists saying the same thing about the world". | |
| """ | |
| coords = ( | |
| tuple(_round(c) for c in item.coordinates) | |
| if item.coordinates is not None | |
| else None | |
| ) | |
| return ( | |
| item.type.value, | |
| item.coordinate_system.value if item.coordinate_system else None, | |
| coords, | |
| _round(item.score), | |
| ) | |
| ``` | |
| `evidence/engine.py:145-162`. The **only** difference from `_identity_key` is the absent | |
| `source_specialist` component. | |
| ```mermaid | |
| flowchart TB | |
| I["_identity_key =<br/>(type, specialist, crs, coords, score)"] --> IQ{"same item?"} | |
| C["_claim_key =<br/>(type, crs, coords, score)"] --> CQ{"same claim about the world?"} | |
| IQ -->|"yes β merge payloads"| DEDUP["_deduplicate"] | |
| CQ -->|"yes, β₯2 distinct specialists β annotate"| AGR["_record_agreement"] | |
| style I fill:#1f6feb22,stroke:#1f6feb | |
| style C fill:#3fb95022,stroke:#3fb950 | |
| ``` | |
| **The annotation.** | |
| ```python | |
| AGREEMENT_KEY = "corroborated_by" | |
| ... | |
| out = list(items) | |
| for indexes in groups.values(): | |
| if len(indexes) < 2: | |
| continue | |
| names = sorted({items[i].source_specialist for i in indexes}) | |
| if len(names) < 2: | |
| # Same specialist restating itself across the collection. The | |
| # identity key would normally have merged these, so reaching | |
| # here means the payloads differed; it is not corroboration. | |
| continue | |
| for i in indexes: | |
| item = out[i] | |
| others = [n for n in names if n != item.source_specialist] | |
| payload = {**item.payload, EvidenceCollection.AGREEMENT_KEY: others} | |
| out[i] = item.model_copy(update={"payload": payload}) | |
| return out | |
| ``` | |
| `evidence/engine.py:391-425`. | |
| **Why the payload, and not a schema field.** The constant's own comment states it | |
| (`evidence/engine.py:234-238`): | |
| > Payload key under which `_record_agreement` stashes the specialists that made the identical | |
| > claim. **A single string field cannot hold a set, so the agreement is recorded in the | |
| > payload** β which is exactly what the payload is for: observations that support the evidence | |
| > but are not the claim. | |
| The `contributing_specialists: list[str]` field that would have made this first-class was | |
| **proposed and rejected** β `docs/PHASE13_EVIDENCE_ENGINE.md` Β§5 records the full | |
| ARCHITECTURE CHANGE entry. The three reasons: | |
| | # | Reason | | |
| |---|---| | |
| | 1 | Adding a field is a frozen-contract change, and the freeze rule requires evidence that the current design is insufficient. No such evidence exists: identity comes from `evidence_id`; provenance is already served by `EvidenceCollection.sources` + `Evidence.source_specialist`; corroboration is recorded in `payload["corroborated_by"]`. | | |
| | 2 | Because `source_specialist` is a single string, two specialists making the identical claim are two items and **both survive**. Collapsing them would require the engine to pick a winner, and the engine has no basis for that. | | |
| | 3 | Reversing this later is cheap: the engine already computes `_claim_key()`, which groups items by claim ignoring the specialist. Adopting the field would mean emitting one item per group instead of N. | | |
| **Decision: proposal `REJECTED`; `core/schemas.py` unchanged.** Adopted interim = | |
| `source_specialist` (single) + `EvidenceCollection.sources` (run-level provenance) + | |
| `payload["corroborated_by"]` (per-claim corroboration). *"Revisit only if a downstream | |
| consumer needs corroboration as a first-class, queryable field."* | |
| **Nothing is ever removed.** The docstring is unambiguous (`evidence/engine.py:401-403`): | |
| > **Nothing is removed:** dropping a member would throw away a specialist's attribution, and | |
| > the freeze does not permit silently discarding a specialist's evidence. | |
| **The two guards inside the loop** are worth naming because they are the two ways a | |
| non-corroboration could masquerade as one: | |
| | Guard | Condition | Why | | |
| |---|---|---| | |
| | group size | `len(indexes) < 2` β skip | a lone item is not agreement | | |
| | distinct specialists | `len(names) < 2` β skip | the same specialist restating itself is *not* corroboration; if this is reached, the payloads differed (otherwise `_identity_key` would have merged them) | | |
| `corroborated_by` is always the **other** specialists (`others = [n for n in names if n != | |
| item.source_specialist]`), so an item never lists itself. | |
| **Excluded from the digest.** *"The corroboration annotation in the payload IS excluded, | |
| because it is a derived observation about the collection, not part of the claim."* | |
| (`evidence/engine.py:692-693`) β consistent with `payload` being excluded from | |
| `_identity_key`. | |
| ### 5.8 Renumbering to `evidence_001β¦` | |
| ```python | |
| ID_PREFIX = "evidence" | |
| @staticmethod | |
| def _renumber(item: Evidence, index: int) -> Evidence: | |
| """Give one item its canonical id, leaving the source untouched.""" | |
| return item.model_copy(update={"evidence_id": f"{ID_PREFIX}_{index:03d}"}) | |
| ``` | |
| `evidence/engine.py:104-107`, `:427-430`. | |
| | Property | Value | Reason | | |
| |---|---|---| | |
| | format | `evidence_001` β¦ | zero-padded to three digits | | |
| | why three digits | *"so lexical sort matches numeric sort up to 999 items -- well past the `evidence.max_items` bound of 32"* (`:104-107`) | | |
| | assignment basis | **sorted position**, not input order | `enumerate(capped, start=1)` after `sorted` | | |
| | scope | **run-local** | *"Ids restart at `001` for each `aggregate` call β they are run-local citations, not global identities."* (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§2.4) | | |
| | source untouched | `model_copy` | the source keeps its `ev_<uuid>` | | |
| **Why this matters at all** (`evidence/engine.py:21-25`): | |
| > `Evidence.evidence_id` defaults to a random uuid, which is fine within one result but | |
| > useless when the controller merges four specialists' evidence into the trace of a single | |
| > run. Nothing can be cited. The engine renumbers to `evidence_001`, `evidence_002`, β¦ so a | |
| > downstream artefact (report, UI, audit) can reference one item deterministically. | |
| ### 5.9 The cap, and loss accounting that is never silent | |
| ```python | |
| DEFAULT_MAX_ITEMS = 32 | |
| ``` | |
| `evidence/engine.py:109-112`, mirroring `configs/base.yaml`: | |
| ```yaml | |
| evidence: | |
| max_items: 32 | |
| coordinate_system_default: normalized_0_1 | |
| ``` | |
| `EvidenceEngine.__init__` **refuses** a non-positive cap (`evidence/engine.py:267-276`): | |
| ```python | |
| if max_items < 1: | |
| raise ValueError(f"max_items must be >= 1, got {max_items}") | |
| ``` | |
| > a cap of zero would silently discard every piece of evidence, which is a policy decision | |
| > this engine has no business making on its own. (`evidence/engine.py:258-261`) | |
| **The accounting fields, all five:** | |
| | Field | Type | Meaning | Non-zero means | | |
| |---|---|---|---| | |
| | `sources` | `list[str]` | sorted specialist names that contributed anything, **BEFORE** the cap | provenance β *"Provenance is a property of the run, not of the surviving items."* | | |
| | `dropped_duplicates` | `int` | items merged into an existing claim | **normal and healthy** β two specialists agreed | | |
| | `dropped_over_limit` | `int` | items the cap discarded | **a warning** β a specialist's evidence did not survive | | |
| | `total_before_limit` | `int` | deduplicated count prior to capping | the denominator for the two drop counts | | |
| | `truncated` | `bool` | `dropped_over_limit > 0` | a boolean summary so a trace can branch without recomputing | | |
| `evidence/engine.py:196-217`; restated in `docs/PHASE13_EVIDENCE_ENGINE.md` Β§2.5. | |
| > Losses are **counted, never silent**. | |
| And on why `sources` is recorded pre-cap: | |
| > `sources` is recorded pre-cap deliberately: provenance is a property of the *run*, not of | |
| > the surviving items. A specialist whose evidence did not fit under the cap still ran, and | |
| > the trace must still say so. (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§2.5) | |
| This is the same principle as `core/config.py`'s refusal to silently truncate and | |
| `evidence/confidence.py`'s refusal to fabricate a number: **a bounded output must declare | |
| what the bound cost.** | |
| ### 5.10 `EvidenceCollection` β the returned container | |
| ```python | |
| @dataclass(frozen=True) | |
| class EvidenceCollection: | |
| items: list[Evidence] = field(default_factory=list) | |
| sources: list[str] = field(default_factory=list) | |
| dropped_duplicates: int = 0 | |
| dropped_over_limit: int = 0 | |
| total_before_limit: int = 0 | |
| truncated: bool = False | |
| ``` | |
| `evidence/engine.py:193-250`. It is `frozen=True` β a dataclass decorator that blocks | |
| attribute *rebinding* on the collection, complementing the `model_copy` discipline that keeps | |
| the individual `Evidence` items unmutated. | |
| | Method | Returns | Purpose | | |
| |---|---|---| | |
| | `__len__` | `int` | `len(items)` | | |
| | `__iter__` | iterator | iterate the items directly | | |
| | `ids()` | `list[str]` | `[item.evidence_id for item in self.items]` | | |
| | `by_type(t)` | `list[Evidence]` | filter by `EvidenceType` (identity comparison: `item.type is evidence_type`) | | |
| | `by_specialist(name)` | `list[Evidence]` | filter by `source_specialist` | | |
| | `summary()` | `dict[str, Any]` | the observable trace facts | | |
| | `AGREEMENT_KEY` | `"corroborated_by"` | the payload key for corroboration | | |
| **`summary()` β the trace-facing view.** | |
| ```python | |
| def summary(self) -> dict[str, Any]: | |
| """Observable trace facts. No chain-of-thought, no interpretation.""" | |
| return { | |
| "returned": len(self.items), | |
| "total_before_limit": self.total_before_limit, | |
| "dropped_duplicates": self.dropped_duplicates, | |
| "dropped_over_limit": self.dropped_over_limit, | |
| "truncated": self.truncated, | |
| "sources": list(self.sources), | |
| "types": sorted({item.type.value for item in self.items}), | |
| } | |
| ``` | |
| `evidence/engine.py:240-250`. Note `types` is a **sorted set** of the type *values* actually | |
| present in the surviving items β so a consumer can see which evidence kinds survived the cap | |
| without walking the list. The docstring's first line is a constraint, not a description: *"No | |
| chain-of-thought, no interpretation."* This mirrors `docs/ARCHITECTURE_FREEZE.md` Β§5's *"Every | |
| result carries an observable execution trace. No chain-of-thought."* | |
| ### 5.11 `evidence_digest()` β reproducibility as an assertion | |
| ```python | |
| def evidence_digest(collection: EvidenceCollection | Sequence[Evidence]) -> str: | |
| items = ( | |
| collection.items | |
| if isinstance(collection, EvidenceCollection) | |
| else list(collection) | |
| ) | |
| hasher = hashlib.sha256() | |
| for item in items: | |
| hasher.update(repr(_identity_key(item)).encode("utf-8")) | |
| hasher.update(b"\n") | |
| return hasher.hexdigest() | |
| ``` | |
| `evidence/engine.py:680-704`. | |
| | Design choice | Consequence | | |
| |---|---| | |
| | built from `_identity_key`, **in order** | changes when the *claims* change | | |
| | **not** from `evidence_id` | does **not** change when a specialist restates the same claim under a new uuid | | |
| | **not** from `payload` | payload is a detail of the claim, not the claim | | |
| | `repr(...)` + `b"\n"` separator | unambiguous framing; no two distinct sequences can hash the same | | |
| | corroboration annotation excluded | *"it is a derived observation about the collection, not part of the claim"* | | |
| The docstring states the contract it enables (`evidence/engine.py:687-690`): | |
| > That is what makes it useful as a reproducibility assertion: | |
| > | |
| > aggregate(inputs_a) is reproducible iff digest(a) == digest(b) | |
| > | |
| > against the same input, **even across processes where the uuid defaults differ.** | |
| ### 5.12 The canonical vocabulary: `evidence_type_for` | |
| ```python | |
| @staticmethod | |
| def evidence_type_for(result: SpecialistResult) -> EvidenceType: | |
| from core.schemas import Task | |
| mapping = { | |
| Task.GROUNDING: EvidenceType.BOUNDING_BOX, | |
| Task.CHANGE: EvidenceType.CHANGE_MAP, | |
| Task.OPTICAL_SAR: EvidenceType.JOINT_FEATURE_REGION, | |
| Task.VQA: EvidenceType.STATISTIC, | |
| Task.CAPTION: EvidenceType.STATISTIC, | |
| Task.UNSUPPORTED: EvidenceType.STATISTIC, | |
| } | |
| return mapping.get(result.task, EvidenceType.STATISTIC) | |
| ``` | |
| `evidence/engine.py:434-458`. | |
| > This is the ONE place that knows the mapping from specialist task to evidence vocabulary, so | |
| > no specialist needs to branch on it. Note it returns the *primary* type; a specialist's own | |
| > `produce_evidence` emits its real per-artefact types, and those are preserved by `aggregate`. | |
| | Task | Primary `EvidenceType` | Why | | |
| |---|---|---| | |
| | `grounding` | `BOUNDING_BOX` | the task *is* localisation | | |
| | `change` | `CHANGE_MAP` | the task *is* a change map | | |
| | `optical_sar` | `JOINT_FEATURE_REGION` | the claim lives in CROMA's joint embedding space | | |
| | `vqa` | `STATISTIC` | an answer is a measured scalar, not geometry | | |
| | `caption` | `STATISTIC` | as above | | |
| | `unsupported` | `STATISTIC` | there is no claim to localise | | |
| | *(fallback)* | `STATISTIC` | `.get(..., STATISTIC)` β an unknown task degrades to a non-spatial claim rather than a wrong spatial one | | |
| **`change_vqa` is absent from the mapping.** `Task.CHANGE_VQA` exists in the enum | |
| (`core/schemas.py:46`) but has no row, so it falls through to `STATISTIC` β which is correct: | |
| change-VQA returns a short *answer*, not a map. The docstring's note that this returns the | |
| *primary* type matters here: a change-VQA result may still carry `CHANGE_MAP` evidence | |
| emitted by the detector it shares, and `aggregate` preserves that. | |
| ### 5.13 The three `evidence_from_*` constructors | |
| These exist so the controller does not hand-build `Evidence` objects β the module docstring | |
| calls them *"Convenience for the controller when it needs evidence for a box that a specialist | |
| emitted but did not itself describe."* | |
| #### `evidence_from_box` | |
| ```python | |
| def evidence_from_box(self, box, *, source_specialist, payload=None, asset_ref=None) -> Evidence: | |
| return Evidence( | |
| type=EvidenceType.BOUNDING_BOX, | |
| source_specialist=source_specialist, | |
| coordinate_system=box.coordinate_system, | |
| coordinates=[box.x1, box.y1, box.x2, box.y2], | |
| score=box.score, | |
| artifact_ref=asset_ref, | |
| payload={"label": box.label, **(payload or {})}, | |
| ) | |
| ``` | |
| `evidence/engine.py:462-485`. | |
| > The `coordinate_system` is copied from the box, **never assumed** β `core.schemas` requires | |
| > it and `Evidence` rejects spatial coordinates without one. | |
| The coordinates are flattened to `[x1, y1, x2, y2]` β matching the flat-`Box` shape that | |
| `docs/PHASE19_FINAL_HARDENING.md` Β§3.6 records as a corrected documentation defect. | |
| #### `evidence_from_region` β a three-way branch | |
| `evidence/engine.py:487-540`. A region may carry a mask, a box, both, or neither, and each | |
| case maps to a different evidence type: | |
| ```mermaid | |
| flowchart TB | |
| R["Region"] --> M{"region.mask_ref?"} | |
| M -->|yes| MASK["EvidenceType.MASK<br/>artifact_ref = asset_ref or mask_ref<br/>coords = box if box else None"] | |
| M -->|no| B{"region.box?"} | |
| B -->|yes| BOX["EvidenceType.BOUNDING_BOX<br/>coords from box<br/>score = region.score or box.score"] | |
| B -->|no| STAT["EvidenceType.STATISTIC<br/>no coordinates<br/>score = region.score"] | |
| style MASK fill:#1f6feb22,stroke:#1f6feb | |
| style BOX fill:#1f6feb22,stroke:#1f6feb | |
| style STAT fill:#d2992222,stroke:#d29922 | |
| ``` | |
| > A region carrying a mask reference becomes `MASK` evidence; one with only a box becomes | |
| > `BOUNDING_BOX` evidence. A region with neither is reported as a `STATISTIC` **rather than a | |
| > spatial claim, because there is no geometry to prove.** (`evidence/engine.py:496-500`) | |
| Two details worth naming: | |
| - The `MASK` branch's `coordinates` are `None` when `region.box is None` β a mask without a | |
| bounding box is legal, and the validator (Β§4) only fires when coordinates are *present*. | |
| - The `BOUNDING_BOX` branch reads `coordinate_system` from **`region.box.coordinate_system`** | |
| (the box's own frame), not from `region.coordinate_system`. The frame of the geometry is | |
| the frame of the geometry. | |
| #### `evidence_from_geospatial` β may return `None` | |
| ```python | |
| def evidence_from_geospatial(self, geo, *, source_specialist, bounds=None, score=None) -> Evidence | None: | |
| if not geo.has_crs and not geo.crs: | |
| return None | |
| return Evidence( | |
| type=EvidenceType.GEOLOCATION, | |
| source_specialist=source_specialist, | |
| coordinate_system=CoordinateSystem.GEO if bounds else None, | |
| coordinates=list(bounds) if bounds else None, | |
| score=score, | |
| payload={ | |
| "crs": geo.crs, | |
| "bounds": list(geo.bounds) if geo.bounds else None, | |
| "width": geo.width, | |
| "height": geo.height, | |
| "is_georeferenced": geo.is_georeferenced, | |
| }, | |
| ) | |
| ``` | |
| `evidence/engine.py:542-572`. | |
| > Returns `None` when there is no CRS to report. Emitting a `GEOLOCATION` item without a CRS | |
| > would assert a placement the data does not support, which is the same failure mode the | |
| > grounding degenerate-box guard exists to prevent. | |
| This is the honesty rule applied to geometry: **a claim that cannot be supported is not | |
| emitted, rather than emitted with a placeholder.** Note the guard is `not geo.has_crs and not | |
| geo.crs` β either signal is sufficient, so a raster that reports a CRS string without the | |
| `has_crs` flag still produces evidence. | |
| ## 6. `confidence_for()` β calibrating through the engine's artifact | |
| ```python | |
| def confidence_for( | |
| self, | |
| result: SpecialistResult | ConfidenceBreakdown, | |
| *, | |
| extra_components: dict[str, float] | None = None, | |
| degraded: bool | None = None, | |
| degradation_reason: str | None = None, | |
| ) -> ConfidenceBreakdown: | |
| ``` | |
| `evidence/engine.py:576-637`. It accepts either form so *"a caller that already has the pieces | |
| does not rebuild a `SpecialistResult` to use it."* | |
| **The degradation rule β the specialist's verdict WINS.** | |
| > The specialist's own degradation verdict WINS unless explicitly overridden: it knows things | |
| > the engine does not (a zero-shot fallback, a failed input-quality gate), and overwriting | |
| > `degraded=False` here would **launder a degraded result into a confident-looking one.** | |
| > (`evidence/engine.py:589-592`) | |
| | Situation | Result | | |
| |---|---| | |
| | `degraded=None`, specialist said `degraded=True` | stays `True` | | |
| | `degraded=None`, specialist said `degraded=False` | stays `False` | | |
| | `degraded=True` explicitly | `True`, and if the specialist had not flagged it, the reason is sourced from `result.warnings[0]` β *"so the reason is a fact rather than a restatement of the boolean"* (`:618-622`) | | |
| | `degraded=True` and no reason and no warnings | falls back to the literal string `"aggregate degraded"` (`:628-629`) | | |
| | `degradation_reason` explicitly passed | that string is used verbatim | | |
| **`calibrate_result()`** (`evidence/engine.py:639-643`) is the narrower entry point: | |
| ```python | |
| def calibrate_result(self, result: SpecialistResult) -> ConfidenceBreakdown: | |
| """Calibrate a result's existing breakdown, preserving its provenance.""" | |
| return calibrate_result(result.confidence, self.calibration) | |
| ``` | |
| **`from_config()`** (`evidence/engine.py:647-665`): | |
| ```python | |
| @classmethod | |
| def from_config(cls, config, *, calibration=None, base_dir=None) -> "EvidenceEngine": | |
| return cls( | |
| max_items=int(config.get("evidence.max_items", DEFAULT_MAX_ITEMS)), | |
| calibration=calibration, | |
| ) | |
| ``` | |
| > `calibration` is explicit rather than auto-loaded: whether a fitted artifact exists is an | |
| > operational fact the caller may know better than the config file does, and silently loading | |
| > one would make the engine's behaviour depend on filesystem state. | |
| The `base_dir` parameter is accepted for signature symmetry with | |
| `evidence.confidence.load_calibration` but **is not used** by this classmethod β the engine | |
| does not resolve the artifact itself. | |
| **Module-level shorthand** (`evidence/engine.py:668-677`): | |
| ```python | |
| def aggregate_evidence(results, *, max_items=DEFAULT_MAX_ITEMS, calibration=None) -> EvidenceCollection: | |
| return EvidenceEngine(max_items=max_items, calibration=calibration).aggregate(results) | |
| ``` | |
| --- | |
| # Part C β The confidence system | |
| ## 7. The honesty rule | |
| `evidence/confidence.py:12-30` is the reason the module exists at all. It is worth quoting in | |
| full because it is the single most important design statement in this subsystem: | |
| > **THE HONESTY RULE (the reason this module exists at all)** | |
| > | |
| > A calibration that claims to be fitted when it is not is a **FALSE CLAIM OF RELIABILITY** β | |
| > strictly worse than reporting nothing, because a downstream consumer will trust it. So: | |
| > | |
| > no fitted artifact -> pass the raw score through unchanged, | |
| > set `method="uncalibrated"`, | |
| > leave `calibrated=None` | |
| > | |
| > `calibrated=None` is what makes it honest: `ConfidenceBreakdown.value` then returns `raw`, | |
| > and any consumer that wants to distinguish "we calibrated this" from "we did not" can read | |
| > `method` or `calibrated is None`. **We never fill in a plausible-looking number to make a | |
| > schema field look complete.** | |
| ### 7.1 The three states | |
| `docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.1 tabulates the decision: | |
| | State | `calibrated` | `method` | `value` | | |
| |---|---|---|---| | |
| | no fitted artifact | `None` | `"uncalibrated"` | falls back to `raw` | | |
| | fitted artifact, `T β 1` | mapped value | `"temperature_scaling"` | `calibrated` | | |
| | fitted artifact, `T = 1` | `None` | `"uncalibrated"` | falls back to `raw` | | |
| **The method strings are module constants** (`evidence/confidence.py:73-78`): | |
| ```python | |
| METHOD_TEMPERATURE = "temperature_scaling" | |
| METHOD_UNCALIBRATED = "uncalibrated" | |
| ``` | |
| > Method string reported whenever no mapping was applied. Matches the literal the existing | |
| > specialists already emit, so trace consumers see one vocabulary. | |
| That last clause is a real constraint: `VLMSpecialist._confidence_for` and | |
| `GroundingSpecialist._confidence_for` already emit `method="uncalibrated", calibrated=None` | |
| independently. The engine generalises that behaviour rather than introducing a second | |
| vocabulary (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.1). | |
| ### 7.2 `ConfidenceBreakdown` β the schema | |
| ```python | |
| class ConfidenceBreakdown(BaseModel): | |
| """Never an LLM utterance. Always measurable signals (plan section 26).""" | |
| model_config = ConfigDict(extra="forbid") | |
| raw: float = Field(ge=0.0, le=1.0) | |
| calibrated: float | None = Field(default=None, ge=0.0, le=1.0) | |
| method: str = "uncalibrated" | |
| components: dict[str, float] = Field(default_factory=dict) | |
| degraded: bool = False | |
| degradation_reason: str | None = None | |
| @property | |
| def value(self) -> float: | |
| return self.calibrated if self.calibrated is not None else self.raw | |
| ``` | |
| `core/schemas.py:259-273`. | |
| | Field | Type | Note | | |
| |---|---|---| | |
| | `raw` | `float`, `[0,1]` | the specialist's hand-weighted score β *not* a probability | | |
| | `calibrated` | `float \| None`, `[0,1]` | `None` means **no mapping was applied** β the honesty signal | | |
| | `method` | `str`, default `"uncalibrated"` | one of the two constants | | |
| | `components` | `dict[str, float]` | diagnostic signals; `docs/API_CONTRACT.md:682` β *"**Diagnostic only** β do not compute a confidence from it"* | | |
| | `degraded` | `bool` | *"Whether this confidence should be trusted"* | | |
| | `degradation_reason` | `str \| None` | why, when `degraded` | | |
| **`value` is a Python `@property` and is NOT serialised.** This is the trap | |
| `docs/API_CONTRACT.md:571` warns about: | |
| > `result.confidence.value` | **NOT a JSON field.** It is a Python `@property` on | |
| > `ConfidenceBreakdown` and is **not serialised** (verified: `model_dump()` yields only | |
| > `calibrated, components, degradation_reason, degraded, method, raw`). To get the number the | |
| > user should see, read `calibrated` if it is non-null, otherwise `raw`. | |
| `docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§D verifies this against the real model: *"`confidence.value` | |
| **is not serialised** β the six-key set matches the contract's parsed claim exactly."* | |
| ### 7.3 The four rules the frontend must follow | |
| `docs/API_CONTRACT.md:686-694`: | |
| 1. Display `calibrated` when it is not `null`; otherwise display `raw`. | |
| 2. Display `method` next to the value. `temperature_scaling` means a fitted correction was | |
| applied; `uncalibrated` means it was not. | |
| 3. **Never present a confidence as a percentage without its method.** A raw 0.99 and a | |
| calibrated 0.99 do not mean the same thing. | |
| 4. When `degraded` is `true`, show `degradation_reason`. Confidence that is degraded is not a | |
| quality signal. | |
| ## 8. Temperature scaling β the derivation | |
| ### 8.1 Why the transform is needed at all | |
| `evidence/confidence.py:5-9`: | |
| > Specialists produce a *raw* reliability score from signals they can actually point at | |
| > (`max_objectness`, input-quality autocorrelation, `used_head`, β¦). That raw score is **not a | |
| > probability: it is a hand-weighted sum, and a hand-weighted sum is not calibrated by | |
| > construction.** This module is the one place that maps raw -> calibrated. | |
| The signals named in `docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.1 are `max_objectness`, | |
| `top_mean_objectness`, `score_contrast`, `used_head`, and input-quality autocorrelation. The | |
| plan (Β§26) names per-task signals too: for grounding, *"box confidence + text/image similarity | |
| + augmentation consistency"*; for change, *"mean pixel probability + component stability + | |
| registration quality"*; for optical-SAR, *"fusion classifier margin + optical confidence + SAR | |
| confidence + cross-modal agreement"*. | |
| ### 8.2 The formula | |
| ``` | |
| logit(z) = log(z / (1 - z)) | |
| calibrated = sigmoid(logit(z) / T) | |
| ``` | |
| `evidence/confidence.py:39-47`: | |
| > The standard formulation maps a logit through `sigmoid(logit / T)`. Our raw scores are | |
| > already in (0, 1), so they are read as probabilities and converted to log-odds first. | |
| ```python | |
| def apply(self, raw: float) -> float: | |
| """Map a raw score in [0, 1] through the fitted temperature.""" | |
| return _clamp01(_sigmoid(_logit(float(raw)) / self.temperature)) | |
| ``` | |
| `evidence/confidence.py:165-167`. Three operations in order: `_logit` β divide by `T` β | |
| `_sigmoid` β `_clamp01`. | |
| | `T` | Effect | Verified | | |
| |---|---|---| | |
| | `T > 1` | **softens** β pulls scores toward the middle | `docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.3 | | |
| | `T < 1` | **sharpens** β pushes scores away from the middle | `docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.3; measured: `calibrate(0.87, T=0.9772732)` β `0.8749186077809417` | | |
| | `T = 1` | **identity** | handled as uncalibrated, Β§8.3 | | |
| **The transform is monotonic**, so calibration respaces scores without reordering them β | |
| pinned by `test_temperature_scaling_is_monotonic` (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.3). | |
| ### 8.3 `_is_effective` β why a "fitted" `T` of exactly 1.0 is reported as uncalibrated | |
| ```python | |
| #: How close to 1.0 a fitted temperature must land before it is treated as | |
| #: "does nothing". Bit-for-bit 1.0 is the honest test; this tolerance absorbs | |
| #: float round-trip through JSON without accepting a real transform. | |
| _IDENTITY_TOLERANCE = 1e-6 | |
| def _is_effective(temperature: float) -> bool: | |
| """Does this temperature actually transform anything? | |
| A temperature within `_IDENTITY_TOLERANCE` of 1.0 is the identity map. It is | |
| treated as "not fitted" so the honest report wins: an artifact that carries | |
| T=1.0 has learned nothing, and reporting `method="temperature_scaling"` for | |
| it would assert a correction that was never made. | |
| """ | |
| return abs(temperature - 1.0) > _IDENTITY_TOLERANCE | |
| ``` | |
| `evidence/confidence.py:90-93`, `:118-126`. | |
| **Verified empirically** (read-only check against the shipped module): | |
| | Input | `_is_effective` | Result | | |
| |---|---|---| | |
| | `0.9772731820958189` (the fitted T) | `True` | a real transform is applied | | |
| | `1.0` | `False` | reported as uncalibrated | | |
| **Why the tolerance and not `!= 1.0`.** The comment gives the reason: *"Bit-for-bit 1.0 is the | |
| honest test; this tolerance absorbs float round-trip through JSON without accepting a real | |
| transform."* A JSON round-trip of `1.0` can land at `1.0000000000000002`; an exact `!= 1.0` | |
| test would then accept a transform that does nothing. `1e-6` is wide enough to absorb the | |
| round-trip and far too narrow to swallow any real fitted temperature. | |
| **The identity branch records why** (`evidence/confidence.py:299-303`): | |
| ```python | |
| if calibration is not None: | |
| # A real artifact was supplied but it is the identity map. Say so | |
| # rather than dropping the information on the floor. | |
| components["calibration_identity"] = 1.0 | |
| components.update(calibration.components) | |
| ``` | |
| Measured components for `T = 1.0`: | |
| `{'calibrated_applied': 0.0, 'calibration_identity': 1.0, 'temperature': 1.0}`. | |
| `docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.2 frames this as the same failure mode as Β§3.1: *"This is | |
| a small case the brief did not name explicitly; it is the same false-claim failure mode as | |
| Β§3.1 and is handled by the same principle."* | |
| ### 8.4 `_EPS` β endpoint clamping so a hard 0.0 / 1.0 cannot become `NaN` | |
| ```python | |
| #: Scores are clamped to this closed interval before the log-odds transform. | |
| #: `logit(0)` and `logit(1)` are infinite; clamping at the boundary keeps the | |
| #: mapping finite and keeps a hard 0.0 / 1.0 from becoming a NaN. | |
| _EPS = 1e-6 | |
| def _logit(p: float) -> float: | |
| """Log-odds of a probability, with the endpoints clamped to stay finite.""" | |
| clamped = min(1.0 - _EPS, max(_EPS, p)) | |
| return math.log(clamped / (1.0 - clamped)) | |
| ``` | |
| `evidence/confidence.py:80-83`, `:108-111`. | |
| | Input `z` | Unclamped `logit(z)` | With `_EPS` clamp | Result | | |
| |---|---|---|---| | |
| | `0.0` | `log(0)` = `-inf` | `logit(1e-6)` β `-13.8155` | finite | | |
| | `1.0` | `log(1/0)` = `+inf` | `logit(1 - 1e-6)` β `+13.8155` | finite | | |
| | `0.5` | `0.0` | `0.0` | unchanged | | |
| Pinned by `test_endpoint_scores_do_not_produce_nan` (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.3). | |
| ### 8.5 `_sigmoid` β numerically stable, and why the branch exists | |
| ```python | |
| def _sigmoid(x: float) -> float: | |
| """Numerically stable logistic function. | |
| The naive form overflows for large negative `x`; the branch below keeps | |
| `exp`'s argument non-positive so the result is finite for every input. | |
| """ | |
| if x >= 0.0: | |
| return 1.0 / (1.0 + math.exp(-x)) | |
| z = math.exp(x) | |
| return z / (1.0 + z) | |
| ``` | |
| `evidence/confidence.py:96-105`. For `x >= 0` the `exp` argument is `-x <= 0`; for `x < 0` the | |
| `exp` argument is `x < 0`. **In both branches `math.exp` receives a non-positive argument**, so | |
| it can never overflow. The two forms are algebraically identical | |
| (`1/(1+e^-x) β‘ e^x/(1+e^x)`) but only one of them is safe on each side of zero. | |
| ```python | |
| def _clamp01(x: float) -> float: | |
| return min(1.0, max(0.0, x)) | |
| ``` | |
| `evidence/confidence.py:114-115`. The final guard: `sigmoid` returns `(0, 1)` mathematically, | |
| but float rounding can produce exactly `0.0` or `1.0`, and `_clamp01` keeps the value inside | |
| the schema's `ge=0.0, le=1.0` bound so Pydantic never rejects an output of this module. | |
| ### 8.6 `TemperatureCalibration` β the artifact object | |
| ```python | |
| @dataclass(frozen=True) | |
| class TemperatureCalibration: | |
| temperature: float | |
| fitted_on: str | None = None | |
| artifact: str | None = None | |
| n_samples: int | None = None | |
| ``` | |
| `evidence/confidence.py:129-150`. | |
| | Field | Type | Meaning (docstring) | | |
| |---|---|---| | |
| | `temperature` | `float` | *"the fitted scalar. `T > 1` softens β¦, `T < 1` sharpens. Validated on construction."* | | |
| | `fitted_on` | `str \| None` | *"free-text provenance, e.g. `"valid"` or a split hash. Carried into the breakdown components so a result can be traced back to the artifact that shaped it."* | | |
| | `artifact` | `str \| None` | *"path/identifier of the source file, for the same reason."* | | |
| | `n_samples` | `int \| None` | *"how many validation samples the fit used, when known."* | | |
| **Range validation** (`evidence/confidence.py:85-88`, `:152-159`): | |
| ```python | |
| _MIN_TEMPERATURE = 1e-3 | |
| _MAX_TEMPERATURE = 1e3 | |
| def __post_init__(self) -> None: | |
| t = float(self.temperature) | |
| if not math.isfinite(t) or not _MIN_TEMPERATURE <= t <= _MAX_TEMPERATURE: | |
| raise ValueError( | |
| f"temperature must be finite and in " | |
| f"[{_MIN_TEMPERATURE}, {_MAX_TEMPERATURE}], got {self.temperature!r}" | |
| ) | |
| object.__setattr__(self, "temperature", t) | |
| ``` | |
| > A temperature of zero or below is not a calibration; it is a division error. | |
| > (`evidence/confidence.py:85-88`) | |
| The check is `math.isfinite(t)` **and** the range β so `NaN` and `Β±inf` are both rejected, and | |
| `object.__setattr__` is used because the dataclass is `frozen=True`. | |
| **Two methods:** | |
| | Method | Returns | Purpose | | |
| |---|---|---| | |
| | `is_effective()` | `bool` | `_is_effective(self.temperature)` | | |
| | `apply(raw)` | `float` | `_clamp01(_sigmoid(_logit(raw) / T))` | | |
| | `components` *(property)* | `dict[str, float]` | `{"temperature": T}` plus `{"calibration_samples": n}` when known | | |
| | `describe()` | `dict[str, Any]` | `{method, temperature, effective, fitted_on, artifact, n_samples}` | | |
| **`from_dict()` β accepts two shapes** (`evidence/confidence.py:187-219`): | |
| ```python | |
| body: Any = payload | |
| if isinstance(payload.get("temperature_scaling"), dict): | |
| body = payload["temperature_scaling"] | |
| elif isinstance(payload.get("calibration"), dict): | |
| body = payload["calibration"] | |
| if not isinstance(body, dict) or "temperature" not in body: | |
| raise ValueError( | |
| "calibration artifact has no 'temperature' field " | |
| f"(source: {source or '<dict>'})" | |
| ) | |
| ``` | |
| > Accepts either the flat form (`{"temperature": 1.4, ...}`) or a nested | |
| > `{"temperature_scaling": {...}}` form, because the artifact format is owned by | |
| > `training/calibration/` and this loader should not be the thing that decides it. | |
| `fitted_on` reads `body.get("fitted_on") or body.get("split")`; `artifact` prefers the | |
| `source` path over `body.get("artifact")`; `n_samples` is coerced to `int` when present. | |
| **The shipped artifact uses the nested form** β `artifacts/calibration_v001.json` has a | |
| top-level `"temperature_scaling": {"fitted_on": "Val", "n_samples": 16441, "temperature": | |
| 0.9772731820958189}` block. This is exactly why `from_dict` accepts it. | |
| **`from_json()` β `required` semantics** (`evidence/confidence.py:221-252`): | |
| | Condition | `required=False` (default) | `required=True` | | |
| |---|---|---| | |
| | file absent | returns `None` | raises `FileNotFoundError` | | |
| | file present, valid JSON | returns the object | returns the object | | |
| | file present, **malformed JSON** | **raises `ValueError`** | **raises `ValueError`** | | |
| > `ValueError`: the file exists and is malformed (**never swallowed**: a corrupt artifact | |
| > silently treated as "no artifact" would hide an operational defect). | |
| This is the same asymmetry as the evidence cap and the config validation: **absent is a | |
| legitimate state; corrupt is a defect and must be loud.** | |
| ### 8.7 `calibrate()` β the single entry point | |
| ```python | |
| def calibrate( | |
| raw: float, | |
| calibration: TemperatureCalibration | None, | |
| *, | |
| extra_components: dict[str, float] | None = None, | |
| degraded: bool = False, | |
| degradation_reason: str | None = None, | |
| ) -> ConfidenceBreakdown: | |
| ``` | |
| `evidence/confidence.py:255-323`. **This is the single entry point used by | |
| `EvidenceEngine.confidence_for`.** | |
| **Non-finite input RAISES** (`evidence/confidence.py:289-292`): | |
| ```python | |
| value = float(raw) | |
| if not math.isfinite(value): | |
| raise ValueError(f"raw confidence must be finite, got {raw!r}") | |
| clamped = _clamp01(value) | |
| ``` | |
| > Refusing is deliberate: a non-finite score is a bug upstream, and clamping it to 0.5 would | |
| > **invent a measurement.** (`evidence/confidence.py:284-287`) | |
| Note the asymmetry with `_logit`'s endpoint clamping: an **infinite** input is refused, but a | |
| **finite** `0.0` or `1.0` is accepted and clamped *inside* the transform. The distinction is | |
| that `0.0` is a real measurement (a specialist is certain it found nothing) while `NaN` is | |
| not a measurement at all. | |
| **The honest path** (`evidence/confidence.py:296-311`): | |
| ```python | |
| if calibration is None or not calibration.is_effective(): | |
| components["calibrated_applied"] = 0.0 | |
| if calibration is not None: | |
| components["calibration_identity"] = 1.0 | |
| components.update(calibration.components) | |
| return ConfidenceBreakdown( | |
| raw=clamped, | |
| calibrated=None, | |
| method=METHOD_UNCALIBRATED, | |
| components=components, | |
| degraded=degraded, | |
| degradation_reason=degradation_reason, | |
| ) | |
| ``` | |
| **The fitted path** (`evidence/confidence.py:313-323`): | |
| ```python | |
| components["calibrated_applied"] = 1.0 | |
| components.update(calibration.components) | |
| return ConfidenceBreakdown( | |
| raw=clamped, | |
| calibrated=calibration.apply(clamped), | |
| method=METHOD_TEMPERATURE, | |
| components=components, | |
| degraded=degraded, | |
| degradation_reason=degradation_reason, | |
| ) | |
| ``` | |
| **Measured component blocks** (read-only verification against the shipped module): | |
| | Case | `components` | | |
| |---|---| | |
| | `calibrate(0.87, T=0.9772731820958189)` | `{'calibrated_applied': 1.0, 'temperature': 0.9772731820958189, 'calibration_samples': 16441.0}` | | |
| | `calibrate(0.87, T=1.0)` | `{'calibrated_applied': 0.0, 'calibration_identity': 1.0, 'temperature': 1.0}` | | |
| | `calibrate(0.87, None)` | `{'calibrated_applied': 0.0}` | | |
| The `calibrated_applied` flag is a **`float`, not a `bool`**, because `components` is typed | |
| `dict[str, float]`. `0.0` / `1.0` is the encoding. `raw` is retained in every case β the | |
| breakdown always carries both the input and the output, so nothing is destroyed by | |
| calibrating. | |
| ### 8.8 `calibrate_result()` β preserve the specialist's judgement | |
| ```python | |
| def calibrate_result(breakdown, calibration) -> ConfidenceBreakdown: | |
| return calibrate( | |
| breakdown.raw, | |
| calibration, | |
| extra_components=dict(breakdown.components), | |
| degraded=breakdown.degraded, | |
| degradation_reason=breakdown.degradation_reason, | |
| ) | |
| ``` | |
| `evidence/confidence.py:326-346`. | |
| > The specialist's own components and degradation reason are carried through untouched: **this | |
| > function adds the calibration, it does not re-judge the specialist's measurement.** | |
| Pinned by `test_specialist_degradation_survives_calibration` | |
| (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.5): *"A `degraded=True` fallback result stays degraded | |
| after calibration, with its `degradation_reason` intact."* | |
| ## 9. `load_calibration()` β the ordered candidate search | |
| ### 9.1 The frozen keys it reads | |
| ```yaml | |
| confidence: | |
| temperature_scaling: true | |
| calibration_file: calibration_v001.json | |
| ``` | |
| `configs/base.yaml:231-233`. | |
| ```python | |
| if not bool(config.get("confidence.temperature_scaling", False)): | |
| return None | |
| filename = config.get("confidence.calibration_file") | |
| if not filename: | |
| return None | |
| ``` | |
| `evidence/confidence.py:396-401`. It returns `None` β **never a fabricated default** β when: | |
| | Condition | Returns | | |
| |---|---| | |
| | the master switch is off | `None` | | |
| | no filename is configured | `None` | | |
| | the artifact is absent from every candidate | `None` | | |
| | the artifact is present but malformed | **raises `ValueError`** | | |
| > Every one of those is a legitimate deployment state (an HF Space may ship without the fitted | |
| > artifact), and each degrades to the pass-through. (`evidence/confidence.py:361-364`) | |
| ### 9.2 The ordered search | |
| ```python | |
| name = str(filename) | |
| candidates: list[Path] = [] | |
| if base_dir is not None: | |
| # An explicit base_dir is honoured first, but a REPO-anchored candidate | |
| # is still tried: several call sites in this repo pass "." meaning "the | |
| # repo", which only works when CWD happens to be the repo root. | |
| candidates.append(Path(base_dir) / name) | |
| from core.config import REPO_ROOT # local import: keeps this module cheap | |
| if base_dir is not None: | |
| candidates.append(Path(REPO_ROOT) / name) | |
| candidates.append(Path(REPO_ROOT) / "artifacts" / name) | |
| candidates.append(Path(REPO_ROOT) / "configs" / name) | |
| for candidate in candidates: | |
| if candidate.exists(): | |
| return TemperatureCalibration.from_json(candidate, required=True) | |
| # Nothing found: the master switch is on but no artifact ships. Degrade. | |
| return None | |
| ``` | |
| `evidence/confidence.py:403-424`. | |
| ```mermaid | |
| flowchart TB | |
| S["load_calibration(config, base_dir=?)"] --> SW{"confidence.temperature_scaling?"} | |
| SW -->|false| N1["return None"] | |
| SW -->|true| FN{"calibration_file set?"} | |
| FN -->|no| N2["return None"] | |
| FN -->|yes| C1{"base_dir given?"} | |
| C1 -->|yes| A["1. base_dir / name"] | |
| A --> A2{"exists?"} | |
| C1 -->|no| B | |
| A2 -->|yes| LOAD["from_json(required=True)"] | |
| A2 -->|no| B["2. REPO_ROOT / name<br/><i>only when base_dir was given</i>"] | |
| B --> B2{"exists?"} | |
| B2 -->|yes| LOAD | |
| B2 -->|no| C["3. REPO_ROOT / artifacts / name<br/><b>the real home</b>"] | |
| C --> C2{"exists?"} | |
| C2 -->|yes| LOAD | |
| C2 -->|no| D["4. REPO_ROOT / configs / name"] | |
| D --> D2{"exists?"} | |
| D2 -->|yes| LOAD | |
| D2 -->|no| N3["return None β honest degradation"] | |
| style C fill:#3fb95022,stroke:#3fb950 | |
| style LOAD fill:#1f6feb22,stroke:#1f6feb | |
| ``` | |
| **The exact candidate list, by input:** | |
| | `base_dir` | candidate 1 | candidate 2 | candidate 3 | candidate 4 | | |
| |---|---|---|---|---| | |
| | `None` | β | β | `REPO_ROOT/artifacts/name` | `REPO_ROOT/configs/name` | | |
| | `"."` | `./name` | `REPO_ROOT/name` | `REPO_ROOT/artifacts/name` | `REPO_ROOT/configs/name` | | |
| | `"configs"` | `configs/name` | `REPO_ROOT/name` | `REPO_ROOT/artifacts/name` | `REPO_ROOT/configs/name` | | |
| | `"artifacts/calibration/confidence_calibration_v001"` | that dir / name | `REPO_ROOT/name` | `REPO_ROOT/artifacts/name` | `REPO_ROOT/configs/name` | | |
| > The search is ordered and deterministic: **the first existing candidate wins**, and no | |
| > candidate existing means `None` (honest degradation, not an error). | |
| > (`evidence/confidence.py:388-389`) | |
| ### 9.3 The silent-failure shape this fixes | |
| `evidence/confidence.py:366-386` documents the defect precisely: | |
| > `base_dir` is resolved in this order (added 2026-09-22): | |
| > | |
| > 1. an explicit `base_dir` argument, when the caller passes one; | |
| > 2. `configs/` under the repository root, **but only if the artifact is actually there** β it | |
| > never is by default, so this branch exists only to keep an existing caller that parks the | |
| > artifact beside the config working; | |
| > 3. `<repo root>/artifacts/` β **the artifact's real home**, and the location | |
| > `scripts/fit_calibration.py` writes to. | |
| > | |
| > Step 3 is the important one and **it is why this function changed.** The frozen config names | |
| > a *bare filename* (`calibration_v001.json`), while the docstring examples in this | |
| > repository's own docs pass `base_dir="."` AND `base_dir="configs"` β two different | |
| > directories, neither of which is `artifacts/`. A deployment following either example would | |
| > **silently resolve to a nonexistent path, degrade to `method="uncalibrated"`, and report a | |
| > fitted artifact as never having been fitted.** That is precisely the silent-failure shape | |
| > this project's provenance discipline exists to prevent, so the search now anchors to the | |
| > repository root using the same `REPO_ROOT` that `app/serving.py` uses for the change | |
| > checkpoint. | |
| **The failure shape, stated as a table:** | |
| | Step | Pre-fix behaviour | Post-fix behaviour | | |
| |---|---|---| | |
| | caller passes `base_dir="."` from a non-repo CWD | resolves to `$CWD/calibration_v001.json` β absent β `None` β `uncalibrated` | candidate 1 fails, candidate 3 succeeds β the artifact loads | | |
| | caller passes `base_dir="configs"` | resolves to `configs/calibration_v001.json` β absent β `None` β `uncalibrated` | candidate 1 fails, candidate 3 succeeds β the artifact loads | | |
| | caller passes nothing | resolved to `$CWD/...` (`base_dir` defaulted to `Path.cwd()` pre-fix) | candidate 3 succeeds β the artifact loads | | |
| | the artifact is genuinely absent | `None` | `None` (unchanged) | | |
| | the artifact is present but corrupt | depends on the path resolved | **raises** β never swallowed | | |
| The critical property is the *symptom*: the pre-fix failure was **silent and | |
| indistinguishable from "no artifact was ever fitted"**. A fitted model reported itself as | |
| unfitted. That is the same class of defect as a fabricated confidence number β it makes the | |
| system lie about its own state. | |
| **Verified against the shipped repo** (read-only execution): | |
| | Call | Resolved artifact | | |
| |---|---| | |
| | `load_calibration(config)` | `β¦/satquery-ai/artifacts/calibration_v001.json` | | |
| | `load_calibration(config, base_dir="configs")` | `β¦/satquery-ai/artifacts/calibration_v001.json` | | |
| | `load_calibration(config, base_dir=".")` | `β¦/satquery-ai/artifacts/calibration_v001.json` | | |
| All three resolve to the **same** file via candidate 3. The artifact's own | |
| `consumer_contract.resolution` field recommends `load_calibration(config, base_dir='configs')` | |
| β which now works, and works for the reason the fix was made. | |
| **The older record, and its status.** `docs/PHASE13_EVIDENCE_ENGINE.md` Β§4.2 was written | |
| before the fix and describes the pre-fix behaviour (`base_dir` defaults to `Path.cwd()`; | |
| *"the controller must pass `base_dir` explicitly"*; *"Current behaviour on the real config, for | |
| the record: `load_calibration` returns `None`, because no artifact has been fitted yet"*). | |
| Two things changed since: the search was repo-anchored (2026-09-22), and | |
| `artifacts/calibration_v001.json` now exists. The current measured behaviour is the table | |
| above. The PHASE13 Β§4.1 expectation of | |
| `artifacts/calibration/confidence_calibration_v001/calibration_v001.json` did **not** hold β | |
| the artifact lives at `artifacts/calibration_v001.json` β which is fine because the bare | |
| filename plus the `artifacts/` candidate resolves it. | |
| ## 10. The MEASURED calibration result β a negative result, reported as one | |
| ### 10.1 The artifact | |
| `artifacts/calibration_v001.json` is the fitted artifact. Its top-level keys: | |
| `consumer_contract`, `created_utc`, `fit_diagnostics`, `metrics`, `provenance`, | |
| `reliability_diagram`, `schema`, `scope`, `temperature_scaling`, `type_mask_applied`. | |
| **`fit_diagnostics`:** | |
| | Field | Value | | |
| |---|---| | |
| | `temperature` | `0.9772731820958189` | | |
| | `log_temperature` | `-0.022989052824434128` | | |
| | `iterations` | `200` | | |
| | `n_samples` | `16441` | | |
| | `nll_before` | `0.6897411093435998` | | |
| | `nll_after` | `0.689630751845387` | | |
| | `nll_improvement` | `0.0001103574982127542` | | |
| | `effective` | `true` | | |
| | `hit_bound` | `false` | | |
| Note `log_temperature = log(0.9772731820958189) = -0.022989β¦` β the fitter optimised in | |
| log-space, and `hit_bound: false` means the optimum was interior, not at a clamp. | |
| **`metrics`:** | |
| | Field | Value | | |
| |---|---| | |
| | `ece_before` | `0.013755` | | |
| | `ece_after` | `0.014929` | | |
| | **`ece_improvement`** | **`-0.001174` (WORSE)** | | |
| | `nll_before` | `0.689741` | | |
| | `nll_after` | `0.689631` | | |
| | `nll_improvement` | `0.00011` | | |
| | `n_bins` | `15` | | |
| | `n_classes` | `19` | | |
| | `n_samples` | `16441` | | |
| **`provenance`:** | |
| | Field | Value | | |
| |---|---| | |
| | `checkpoint_path` | `artifacts/change_vqa/run/head.pt` | | |
| | `checkpoint_sha256` | `cfae5e43b97ca930f568dc5b8ae4f36b24e9ff717af226159802206ffd63a82a` | | |
| | `config_hash` | `78f1e3700da15aa1` | | |
| | `dataset_id` | `cdvqa` | | |
| | `feature_spec` | `change_feat_v1` | | |
| | `fitted_on` | `Val` | | |
| | `held_out_splits_excluded` | `["Test", "Test2"]` | | |
| | `method` | `temperature_scaling` | | |
| | `objective` | `mean_negative_log_likelihood` | | |
| | `optimizer` | `golden_section_on_log_temperature` | | |
| | `space` | `multiclass_logits` | | |
| | `iterations` | `200` | | |
| | `n_classes` | `19` | | |
| | `n_samples` | `16441` | | |
| **`scope` β what this artifact calibrates, and what it does not:** | |
| > **This temperature calibrates the R-02 change-VQA head's answer confidence. Other | |
| > specialists emit their own raw scores and are unaffected.** | |
| **`type_mask_applied`:** `false`. | |
| ### 10.2 The result, stated exactly | |
| > The raw softmax was **already near-calibrated** (ECE 0.013755) and temperature scaling made | |
| > ECE very slightly **worse** (0.014929) while improving NLL marginally. This is a | |
| > measurement, not a quality judgment, and **it must not be described as scaling being "more | |
| > accurate".** β `docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§7 | |
| `docs/FINAL_DELIVERY_TODO.md:104` records it in the measured-metrics table as | |
| `0.013755 β 0.014929 (worse)` with status **"measured (not an improvement)"**, sourced to | |
| `artifacts/calibration_v001.json:26-35`. `docs/PHASE19_FINAL_HARDENING.md:377` records the | |
| same: *"**Not claimed as an improvement.** β¦ Nothing is calibrated in a deployed path"*. | |
| | Metric | Before | After | Ξ | Reading | | |
| |---|---|---|---|---| | |
| | ECE (15 bins) | 0.013755 | 0.014929 | **β0.001174** | **WORSE** | | |
| | NLL | 0.689741 | 0.689631 | +0.000110 | marginally better | | |
| **Why both numbers matter, and neither alone.** The fitter's objective was | |
| `mean_negative_log_likelihood`, and NLL *did* improve β that is why `T = 0.9773β¦` was | |
| selected at all. But the metric a consumer reads as "is this probability trustworthy" is ECE, | |
| and **ECE got worse**. Reporting only NLL would present a fitted artifact as a success; | |
| reporting only ECE would hide why the value was chosen. Both are reported, and the headline is | |
| the ECE regression. | |
| **Why the reliability diagram is not a single score.** The artifact's own note: | |
| > Equal-width bins over predicted-class confidence. **ECE is bin-count sensitive and is not an | |
| > aggregate score.** β `artifacts/calibration_v001.json`, `reliability_diagram.note` | |
| The diagram is a 15-bin list. The first two bins (`[0.0, 0.0667)` and `[0.0667, 0.1333)`) | |
| have `count: 0` and `null` accuracy/confidence/gap β the model never predicts that low. The | |
| remaining thirteen bins are populated; the largest is `[0.9333, 1.0]` with `count: 2814`, | |
| `accuracy: 0.969794`, `confidence: 0.966895`, `gap: 0.002899`. The largest |gap| is | |
| `-0.027817` at `[0.6667, 0.7333)` (`count: 1494`). | |
| | Bin `[lo, hi)` | count | accuracy | confidence | gap | | |
| |---|---|---|---|---| | |
| | 0.0000β0.0667 | 0 | β | β | β | | |
| | 0.0667β0.1333 | 0 | β | β | β | | |
| | 0.1333β0.2000 | 8 | 0.25 | 0.188409 | +0.061591 | | |
| | 0.2000β0.2667 | 437 | 0.283753 | 0.241931 | +0.041822 | | |
| | 0.2667β0.3333 | 736 | 0.290761 | 0.299846 | β0.009085 | | |
| | 0.3333β0.4000 | 626 | 0.386581 | 0.367128 | +0.019454 | | |
| | 0.4000β0.4667 | 690 | 0.450725 | 0.434665 | +0.016060 | | |
| | 0.4667β0.5333 | 1380 | 0.534783 | 0.504472 | +0.030310 | | |
| | 0.5333β0.6000 | 1521 | 0.558185 | 0.565780 | β0.007595 | | |
| | 0.6000β0.6667 | 1395 | 0.624373 | 0.633333 | β0.008961 | | |
| | 0.6667β0.7333 | 1494 | 0.672691 | 0.700507 | **β0.027817** | | |
| | 0.7333β0.8000 | 1611 | 0.742396 | 0.766718 | β0.024322 | | |
| | 0.8000β0.8667 | 1701 | 0.833039 | 0.834257 | β0.001218 | | |
| | 0.8667β0.9333 | 2028 | 0.892998 | 0.903143 | β0.010145 | | |
| | 0.9333β1.0000 | 2814 | 0.969794 | 0.966895 | +0.002899 | | |
| *(values verbatim from `artifacts/calibration_v001.json`; the `gap` column is the artifact's | |
| own `accuracy β confidence`)* | |
| ### 10.3 A structural note on the transform, recorded rather than glossed | |
| The artifact's `consumer_contract` states: | |
| ```json | |
| "applied_as": "sigmoid(logit(z) / T) for a scalar z; softmax(logits / T) for a distribution", | |
| "class": "TemperatureCalibration", | |
| "module": "evidence.confidence", | |
| "read_keys": ["temperature", "fitted_on|split", "artifact", "n_samples"], | |
| "resolution": "load_calibration(config, base_dir='configs')" | |
| ``` | |
| The artifact's provenance records `space: "multiclass_logits"` and `n_classes: 19` β the fit | |
| happened over a 19-class softmax. The consumer class | |
| (`evidence.confidence.TemperatureCalibration.apply`) implements the **scalar** form: | |
| `_clamp01(_sigmoid(_logit(raw) / T))`. The class reads exactly the four keys the artifact | |
| declares (`temperature`, `fitted_on|split`, `artifact`, `n_samples`) and reads nothing else β | |
| the multiclass branch is a documented capability of the *artifact format*, not a code path in | |
| `evidence/confidence.py`. Both facts are stated here so the boundary is visible; no claim is | |
| made about which form a deployed path exercises, because | |
| `docs/PHASE19_FINAL_HARDENING.md:377` records that *"Nothing is calibrated in a deployed | |
| path"*. | |
| ### 10.4 The deployment state | |
| | Question | Answer | Source | | |
| |---|---|---| | |
| | Does a fitted artifact exist? | **yes** β `artifacts/calibration_v001.json` | filesystem | | |
| | Does `load_calibration(config)` find it? | **yes** β via candidate 3 | verified read-only | | |
| | Is `T` effective (non-identity)? | **yes** β `_is_effective(0.9772731820958189) == True` | verified read-only | | |
| | Does it improve ECE? | **no** β 0.013755 β 0.014929 | artifact `metrics` | | |
| | Is it wired into a served path? | `OPEN` / not claimed β *"Nothing is calibrated in a deployed path"* | `docs/PHASE19_FINAL_HARDENING.md:377` | | |
| | Which specialist does it cover? | `change_vqa` only | artifact `scope` | | |
| The artifact's `config_hash` is `78f1e3700da15aa1` β the same frozen hash as | |
| `configs/base.yaml` (Β§[07](07-configuration-freeze.md)). **The artifact is keyed to the config | |
| that produced it**, which is exactly the provenance link the config hash exists to provide. | |
| ## 11. What is deliberately absent from the confidence system | |
| ### 11.1 No LLM path | |
| > **NO LLM PATH EXISTS HERE.** There is deliberately no function that turns text into a number. | |
| > Confidence is a measurement (freeze section 5: "No LLM-generated confidence"). The only | |
| > inputs accepted are a float the specialist computed and a calibration artifact fitted on | |
| > validation data. β `evidence/confidence.py:32-37` | |
| `docs/ARCHITECTURE_FREEZE.md` Β§5 lists it as a non-negotiable: *"No LLM-generated coordinates. | |
| No LLM-generated confidence."* And the plan Β§26 opens with it: *"No LLM-generated confidence."* | |
| ### 11.2 No temperature fitter | |
| > **A fitter is NOT included:** Phase 13 fits T on validation data, and inventing one here | |
| > would be exactly the kind of unfounded number this module refuses to emit. | |
| > `TemperatureCalibration` is the *consumer* of a fitted `T`; the artifact loader reads | |
| > whatever `training/calibration/` produces. β `evidence/confidence.py:57-61` | |
| The fitter lives in `training/calibration/` (`__init__.py`, `artifact.py`, `evaluate.py`, | |
| `fitter.py`) and produced the artifact in Β§10. `evidence/confidence.py` deliberately does not | |
| contain one β the separation is what makes "the number came from a fit" a checkable claim. | |
| ### 11.3 No retrieval path | |
| There is no artifact-serving endpoint in v1 (F-16, Β§2.3). This is why `artifact_ref` is | |
| permanently `null` and why the confidence module's `artifact` field carries a **path for | |
| provenance**, not a client-facing reference. `TemperatureCalibration.artifact` is used in | |
| `describe()` for operator diagnostics; it is not serialised into a client response. | |
| --- | |
| # Part D β The execution trace and the eight events | |
| ## 12. `ExecutionTrace` β observable facts only | |
| ```python | |
| class ExecutionTrace(BaseModel): | |
| model_config = ConfigDict(extra="forbid") | |
| run_id: str = Field(default_factory=lambda: _new_id("run")) | |
| schema_version: str = SCHEMA_VERSION | |
| task: Task | None = None | |
| query: str | None = None | |
| inputs: list[str] = Field(default_factory=list) | |
| modalities: list[Modality] = Field(default_factory=list) | |
| intent: Intent | None = None | |
| validation: dict[str, Any] = Field(default_factory=dict) | |
| workflow: list[str] = Field(default_factory=list) | |
| steps: list[TraceStep] = Field(default_factory=list) | |
| selected_models: list[ModelRef] = Field(default_factory=list) | |
| parameters: dict[str, Any] = Field(default_factory=dict) | |
| outputs: list[str] = Field(default_factory=list) | |
| confidence: ConfidenceBreakdown | None = None | |
| timings: dict[str, float] = Field(default_factory=dict) | |
| fallbacks: list[str] = Field(default_factory=list) | |
| errors: list[dict[str, Any]] = Field(default_factory=list) | |
| contradiction: bool = False | |
| config_hash: str | None = None | |
| started_at: str = Field(default_factory=_utcnow) | |
| finished_at: str | None = None | |
| ``` | |
| `core/schemas.py:296-319`. The section header above it reads: | |
| **`# Execution trace (observable facts only β never chain-of-thought)`** | |
| (`core/schemas.py:277`). | |
| | Field | Type | Carries | | |
| |---|---|---| | |
| | `run_id` | `str` | `run_<12 hex>`, generated per run | | |
| | `schema_version` | `str` | `SCHEMA_VERSION` = `"1.0"` | | |
| | `task` | `Task \| None` | the routed task | | |
| | `query` | `str \| None` | the user's question, verbatim | | |
| | `inputs` | `list[str]` | asset references | | |
| | `modalities` | `list[Modality]` | inferred modalities | | |
| | `intent` | `Intent \| None` | the router's advisory output | | |
| | `validation` | `dict[str, Any]` | input-validation facts (`input_count`, `format`, `modality` in the plan Β§27 shape) | | |
| | `workflow` | `list[str]` | the planned step names | | |
| | `steps` | `list[TraceStep]` | the state machine's visits | | |
| | `selected_models` | `list[ModelRef]` | which models ran | | |
| | `parameters` | `dict[str, Any]` | parameters used | | |
| | `outputs` | `list[str]` | output kinds (`bbox`, `answer`, β¦) | | |
| | `confidence` | `ConfidenceBreakdown \| None` | the confidence breakdown | | |
| | `timings` | `dict[str, float]` | measured durations | | |
| | `fallbacks` | `list[str]` | which fallbacks fired | | |
| | `errors` | `list[dict[str, Any]]` | error records | | |
| | `contradiction` | `bool` | whether a contradiction was detected | | |
| | `config_hash` | `str \| None` | the frozen config identity β `78f1e3700da15aa1` for this revision | | |
| | `started_at` / `finished_at` | `str` | ISO-8601 UTC | | |
| **`TraceStep`** (`core/schemas.py:279-285`): | |
| ```python | |
| class TraceStep(BaseModel): | |
| model_config = ConfigDict(extra="forbid") | |
| state: ControllerState | |
| started_at: str = Field(default_factory=_utcnow) | |
| duration_ms: float | None = None | |
| detail: dict[str, Any] = Field(default_factory=dict) | |
| ``` | |
| **`ModelRef`** (`core/schemas.py:288-293`): | |
| ```python | |
| class ModelRef(BaseModel): | |
| model_config = ConfigDict(extra="forbid") | |
| name: str | |
| revision: str | None = None | |
| role: str | None = None | |
| ``` | |
| **`ControllerState`** β the nine states (`core/schemas.py:79-88`): | |
| | # | State | Meaning | | |
| |---|---|---| | |
| | 1 | `RECEIVE` | the request arrived | | |
| | 2 | `PARSE` | the query was parsed | | |
| | 3 | `VALIDATE` | inputs validated | | |
| | 4 | `PLAN` | the workflow was chosen | | |
| | 5 | `PREPROCESS` | tiling / normalisation | | |
| | 6 | `EXECUTE` | specialists ran | | |
| | 7 | `AGGREGATE` | evidence was aggregated | | |
| | 8 | `VERIFY` | confidence was computed | | |
| | 9 | `RESPOND` | the envelope was assembled | | |
| `configs/base.yaml` Β§`agent.states` declares exactly this list, in this order β so the | |
| config and the enum cannot drift. | |
| ### 12.1 There is no field for reasoning | |
| `ExecutionTrace` has **no** `reasoning`, `thought`, `rationale`, or `chain_of_thought` field, | |
| and `model_config = ConfigDict(extra="forbid")` means one cannot be smuggled in. This is a | |
| structural guarantee, not a convention. `docs/ARCHITECTURE_FREEZE.md` Β§5: | |
| > Every result carries an observable execution trace. **No chain-of-thought.** | |
| The plan Β§27 says the same at the end of its trace example: *"No chain-of-thought. Only | |
| observable execution facts."* The `summary()` methods in the evidence layer obey the same | |
| rule β `EvidenceCollection.summary()`'s docstring: *"Observable trace facts. No chain-of-thought, | |
| no interpretation."* | |
| ## 13. The eight execution events | |
| The frontend's event protocol is the integration seam between the backend and the instrument | |
| UI. It is declared in `frontend/assets/js/core.js:616-620`: | |
| ```javascript | |
| SQ.EVENT_NAMES = [ | |
| 'QUERY_RECEIVED', 'QUERY_UNDERSTOOD', 'ROUTE_SELECTED', | |
| 'SPECIALIST_STARTED', 'SPECIALIST_COMPLETED', | |
| 'EVIDENCE_GENERATED', 'CONFIDENCE_COMPUTED', 'RESULT_ASSEMBLED' | |
| ]; | |
| ``` | |
| | # | Event | What it reports | Drives the stage | | |
| |---|---|---|---| | |
| | 1 | `QUERY_RECEIVED` | the query text, the scene seed, the AOI, the GSD, the date pair | `QUERY` | | |
| | 2 | `QUERY_UNDERSTOOD` | the routed task, the parsed slots, the policy id | `UNDERSTAND` | | |
| | 3 | `ROUTE_SELECTED` | the task, the specialist list, the rule trace, the policy version | `ROUTE` | | |
| | 4 | `SPECIALIST_STARTED` | which component, its index/ordinal, its stage, its model, its device | `ANALYZE` / `GROUND` | | |
| | 5 | `SPECIALIST_COMPLETED` | which component, its duration, its output kind | (advances the active marker) | | |
| | 6 | `EVIDENCE_GENERATED` | the evidence regions, the count, the threshold, the registration RMSE | `EVIDENCE` | | |
| | 7 | `CONFIDENCE_COMPUTED` | the reported/calibrated values, the method, the reliability bins | `CONFIDENCE` | | |
| | 8 | `RESULT_ASSEMBLED` | the answer text, the task, the specialists, the evidence ids, the confidence, the provenance | `ANSWER` | | |
| **The stage rail** is a separate, coarser list (`frontend/assets/js/core.js:605-614`): | |
| ```javascript | |
| SQ.STAGES = [ | |
| { id: 'QUERY', label: 'QUERY' }, | |
| { id: 'UNDERSTAND', label: 'UNDERSTAND' }, | |
| { id: 'ROUTE', label: 'ROUTE' }, | |
| { id: 'ANALYZE', label: 'ANALYZE' }, | |
| { id: 'GROUND', label: 'GROUND' }, | |
| { id: 'EVIDENCE', label: 'EVIDENCE' }, | |
| { id: 'CONFIDENCE', label: 'CONFIDENCE' }, | |
| { id: 'ANSWER', label: 'ANSWER' } | |
| ]; | |
| ``` | |
| Eight stages for eight events β but **not a 1:1 mapping**: events 4 and 5 share the | |
| `ANALYZE`/`GROUND` stages, because a specialist's start and completion are two events about | |
| one stage visit. | |
| ### 13.1 The run engine is deliberately dumb | |
| `frontend/assets/js/core.js:8-11`: | |
| > The run engine is deliberately dumb: it renders whatever events it receives. **Nothing about | |
| > the visuals depends on the events being synthetic.** Swapping the mock driver for a | |
| > websocket / SSE feed of the same event names is the entire integration surface. | |
| ```javascript | |
| /** | |
| * THE INTEGRATION SEAM. | |
| * Feed real execution events here β same names, same payload shapes β and | |
| * every visual state in the prototype updates identically. | |
| */ | |
| ingest: function (type, payload) { ... } | |
| ``` | |
| `frontend/assets/js/core.js:737-742`. `ingest()` is a `switch` over the eight names, each | |
| case updating `state` and calling `emit()`. `emit()` records | |
| `{type, t, seq, payload}` where `t = performance.now() - t0` and `seq` is a monotonic counter | |
| (`frontend/assets/js/core.js:709-718`). | |
| **Two drivers exist, and they are not interchangeable:** | |
| | Driver | When | Payload honesty | | |
| |---|---|---| | |
| | `startMock(query, scene)` | **preview only** β when no file is selected | synthetic; marked `is-mock` | | |
| | `runLive(query)` in `frontend/assets/js/mission.js` | the production path | payloads come from the real network response | | |
| `docs/FINAL_DELIVERY_TODO.md:73` records the distinction: *"`runMock` only when no file is | |
| selected; emits empty payloads, marked `is-mock`; not in production path."* The live driver | |
| emits all eight events around real network calls (`docs/FINAL_DELIVERY_TODO.md:71`): | |
| *"`runLive` emits 8 events around real network calls; `markState` uses real values."* | |
| ### 13.2 `markState` β the trace is driven by events, never by a timer | |
| `frontend/assets/js/mission.js:404-425`: | |
| ```javascript | |
| function markState(id, note, isLive) { | |
| var n = traceNodes[id]; | |
| if (!n) return; | |
| n.node.classList.toggle('is-mock', !isLive); | |
| n.tm.textContent = note !== undefined ? note : (STATE_NOTE[id] || ''); | |
| /* Advance the trace to the furthest state reached. This is driven by the | |
| same events that carry the real result, so the bar and the node states | |
| move only when the run actually reaches a stage β never on a timer. The | |
| fill spans from the left edge to the centre of the current node. */ | |
| var idx = STATES.indexOf(id); | |
| if (idx > traceProgress) traceProgress = idx; | |
| ... | |
| if (traceFill) { | |
| traceFill.style.width = (((traceProgress + 0.5) / STATES.length) * 100) + '%'; | |
| } | |
| } | |
| ``` | |
| The design property in the comment is the important one: **the progress bar is a function of | |
| reached states, not of elapsed time.** A run that stalls shows a stalled bar. | |
| ### 13.3 The measured 94.4444 % trace fill | |
| `docs/FINAL_DELIVERY_TODO.md:72`: | |
| | Area | Status | Evidence | | |
| |---|---|---| | |
| | Execution trace (progress bar) | **VERIFIED** | `.trace__fill` width is set from real event count (measured **94.4444 %** live, 2026-09-25) | | |
| **The number is derivable from the formula**, which is why it is a real measurement rather | |
| than a coincidence. The live path reaches `RESPOND`, the ninth of nine `ControllerState` | |
| values: | |
| ``` | |
| width = ((traceProgress + 0.5) / STATES.length) * 100 | |
| = ((8 + 0.5) / 9) * 100 | |
| = (8.5 / 9) * 100 | |
| = 94.4444β¦ % | |
| ``` | |
| `STATES` is the nine-state `ControllerState` list (Β§12); `traceProgress = 8` is the | |
| zero-based index of `RESPOND`. The `+ 0.5` is what makes the bar span *to the centre of the | |
| current node* rather than to its left edge β so a fully-completed run stops at 94.4444 %, not | |
| 100 %, because the last node's centre is half a node-width short of the right edge. **A bar | |
| that read 100 % on a nine-state rail would be reporting something the run never did.** | |
| ## 14. Worked example β one item's full journey | |
| The following traces a single grounding claim from a specialist's `Box` to a client-facing | |
| evidence item. Every step is a real code path. | |
| ```mermaid | |
| sequenceDiagram | |
| participant G as GroundingSpecialist | |
| participant E as EvidenceEngine | |
| participant C as confidence.calibrate | |
| participant T as ExecutionTrace | |
| participant F as Frontend | |
| G->>G: compute Box(x1,y1,x2,y2, score, label, coordinate_system) | |
| G->>E: SpecialistResult.evidence = [Evidence(type=bounding_box, ...)] | |
| Note over E: aggregate([grounding_result]) | |
| E->>E: collect β [ev_a1b2c3d4e5f6] | |
| E->>E: _identity_key β (bounding_box, grounding, normalized_0_1, (...), 0.87) | |
| E->>E: _deduplicate β no collision | |
| E->>E: sorted(_sort_key) | |
| E->>E: _record_agreement β single specialist, no annotation | |
| E->>E: _renumber β evidence_001 | |
| E-->>T: EvidenceCollection(items=[evidence_001], sources=['grounding'], ...) | |
| E->>C: confidence_for(grounding_result) | |
| C->>C: _is_effective(0.9772731820958189) == True | |
| C->>C: _logit(0.87) / 0.9772732 β _sigmoid β _clamp01 | |
| C-->>T: ConfidenceBreakdown(raw=0.87, calibrated=0.8749186β¦, method="temperature_scaling") | |
| T->>F: emit EVIDENCE_GENERATED {regions, count, threshold, registrationRMSE} | |
| T->>F: emit CONFIDENCE_COMPUTED {calibrated, method, bins, observed} | |
| T->>F: emit RESULT_ASSEMBLED {text, evidenceIds, confidence, provenance} | |
| ``` | |
| **The state of the item at each step:** | |
| | Step | `evidence_id` | `type` | `source_specialist` | `coordinate_system` | `coordinates` | `score` | `artifact_ref` | `payload` | | |
| |---|---|---|---|---|---|---|---|---| | |
| | specialist output | `ev_a1b2c3d4e5f6` | `bounding_box` | `grounding` | `normalized_0_1` | `[0.21,0.33,0.47,0.61]` | `0.87` | `null` | `{"label": "water"}` | | |
| | after dedup | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | | |
| | after sort | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | | |
| | after agreement | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | `{"label": "water"}` *(no corroboration β single specialist)* | | |
| | after renumber | **`evidence_001`** | β¦ | β¦ | β¦ | β¦ | β¦ | β¦ | β¦ | | |
| **And the source object is untouched.** After `aggregate` returns, the original | |
| `SpecialistResult.evidence[0]` still has `evidence_id == "ev_a1b2c3d4e5f6"`. This is the purity | |
| contract (Β§5.4) made concrete: `_renumber` calls `model_copy(update={"evidence_id": ...})`, | |
| which produces a **new** object and leaves the source alone. | |
| **Verified confidence values** (read-only execution against the shipped module): | |
| | Input | `raw` | `calibrated` | `method` | | |
| |---|---|---|---| | |
| | `calibrate(0.87, T=0.9772731820958189)` | `0.87` | `0.8749186077809417` | `temperature_scaling` | | |
| | `calibrate(0.87, T=1.0)` | `0.87` | `None` | `uncalibrated` | | |
| | `calibrate(0.0, None)` | `0.0` | `None` | `uncalibrated` | | |
| --- | |
| # Part E β Honest boundaries | |
| ## 15. What is NOT RUN, OPEN, BLOCKED or REJECTED for this topic | |
| | Item | Status | Note | | |
| |---|---|---| | |
| | Calibration **improves** ECE | **`REJECTED`** | measured 0.013755 β 0.014929 (worse). The artifact is retained because it is in the frozen config, not because it helps. | | |
| | Calibration applied in a **served** path | `OPEN` | `docs/PHASE19_FINAL_HARDENING.md:377` β *"Nothing is calibrated in a deployed path"* | | |
| | Calibration covers specialists other than `change_vqa` | `NOT RUN` | the artifact's `scope` covers `change_vqa` only; every other specialist's confidence is uncalibrated | | |
| | A second fitted artifact for the other specialists | `NOT RUN` | no artifact exists | | |
| | `EvidenceEngine` wired into `core/controller.py` | `IMPLEMENTED` (Phase 14) | `docs/PHASE13_EVIDENCE_ENGINE.md` Β§8 recorded it as not-yet-wired at Phase 13; the live run path emits the eight events | | |
| | Artifact **rendering** (crops, masks, change maps served to a client) | `OPEN` | *"`artifact_ref` and the `evidence_from_*` methods provide the hooks; rendering crops, masks and change maps is controller/GUI territory."* (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§8) | | |
| | Artifact **retrieval** endpoint | `REJECTED` for v1 | F-16: no artifact-serving endpoint exists, so `artifact_ref` is permanently `null` | | |
| | `contributing_specialists` field on `Evidence` | `REJECTED` | proposed and not adopted; `core/schemas.py` unchanged (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§5) | | |
| | Change-degradation clause in `SpecialistResult` | `RESOLVED` (removed, F-16c) | the narrower replacement rule is `REJECTED` (`core/schemas.py:362-395`) | | |
| | `payload["corroborated_by"]` exercised by a live multi-specialist run | `UNKNOWN β not established from the available evidence` | the annotation is pinned by unit tests; no live run was found that produced two specialists making an identical claim | | |
| | The multiclass (`softmax(logits/T)`) form of the calibration applied in code | `UNKNOWN β not established from the available evidence` | the artifact documents the capability; `evidence/confidence.py` implements the scalar form (Β§10.3) | | |
| | Whether `_sort_key`'s coordinate tie-break is exercised with non-equal coordinates in production | `UNKNOWN β not established from the available evidence` | pinned by unit test `test_equal_scores_order_deterministically_by_coordinates` | | |
| | The exact count of live runs that populated `dropped_over_limit > 0` | `UNKNOWN β not established from the available evidence` | no run record was found | | |
| | `evidence.max_items = 32` ever being the binding constraint in a live run | `UNKNOWN β not established from the available evidence` | the cap is config-declared; no live run's `total_before_limit` was found | | |
| ## 16. Where the evidence lives | |
| | Claim | Source | | |
| |---|---| | |
| | `Evidence` fields and the spatial validator | `core/schemas.py:216-253` | | |
| | `artifact_ref` is permanently `null` (F-16) | `core/schemas.py:225-235`; `docs/API_CONTRACT.md` Β§2.4; `docs/DEPLOYMENT_ARCHITECTURE.md` Β§5.5 | | |
| | F-16c and the removed change clause | `core/schemas.py:362-395`; `docs/DEPLOYMENT_ARCHITECTURE.md:855-866` | | |
| | `EvidenceType` β all 11 members | `core/schemas.py:65-76` | | |
| | `CoordinateSystem` β 3 values | `core/schemas.py:57-62` | | |
| | Purity contract | `evidence/engine.py:38-50` | | |
| | `_identity_key` and its exclusions | `evidence/engine.py:125-142`; `:52-68` | | |
| | `_claim_key` | `evidence/engine.py:145-162` | | |
| | `_sort_key` and the four-level rationale | `evidence/engine.py:165-190` | | |
| | `_deduplicate` and the payload merge | `evidence/engine.py:352-389` | | |
| | `_record_agreement` and `AGREEMENT_KEY` | `evidence/engine.py:391-425`; `:234-238` | | |
| | `_renumber` and `ID_PREFIX` | `evidence/engine.py:427-430`; `:104-107` | | |
| | The cap and the five accounting fields | `evidence/engine.py:193-217`; `:337-350` | | |
| | `evidence_digest` | `evidence/engine.py:680-704` | | |
| | `evidence_type_for` | `evidence/engine.py:434-458` | | |
| | `evidence_from_box/region/geospatial` | `evidence/engine.py:462-572` | | |
| | `confidence_for` and the degradation-wins rule | `evidence/engine.py:576-637` | | |
| | The honesty rule | `evidence/confidence.py:12-30` | | |
| | `logit`/`sigmoid`/`_EPS`/`_clamp01` | `evidence/confidence.py:80-115` | | |
| | `_is_effective` and `_IDENTITY_TOLERANCE` | `evidence/confidence.py:90-93`, `:118-126` | | |
| | `TemperatureCalibration` and its range guard | `evidence/confidence.py:129-252` | | |
| | `calibrate` / `calibrate_result` | `evidence/confidence.py:255-346` | | |
| | `load_calibration` and the ordered search | `evidence/confidence.py:349-424` | | |
| | The 77-test phase record | `docs/PHASE13_EVIDENCE_ENGINE.md` | | |
| | The measured calibration result | `artifacts/calibration_v001.json`; `docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§7; `docs/FINAL_DELIVERY_TODO.md:104` | | |
| | The eight events | `frontend/assets/js/core.js:616-620` | | |
| | The stage rail | `frontend/assets/js/core.js:605-614` | | |
| | `markState` and the fill formula | `frontend/assets/js/mission.js:404-425` | | |
| | The 94.4444 % measurement | `docs/FINAL_DELIVERY_TODO.md:72` | | |
| | `ExecutionTrace` / `TraceStep` / `ModelRef` | `core/schemas.py:279-319` | | |
| | `ControllerState` β nine states | `core/schemas.py:79-88`; `configs/base.yaml` Β§`agent.states` | | |
| | `evidence.max_items: 32` | `configs/base.yaml` Β§`evidence` | | |
| | `confidence.*` frozen keys | `configs/base.yaml` Β§`confidence` | | |
| | Layer verbs | `docs/ARCHITECTURE_FREEZE.md` Β§5 | | |
| **Next:** [07 Configuration freeze](07-configuration-freeze.md) β the registry, the enforced | |
| invariants, and the hash `78f1e3700da15aa1` that every artifact in this chapter is keyed to. | |