Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
File size: 108,035 Bytes
8c152e4 928a5aa 8c152e4 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 1332 1333 1334 1335 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 1396 1397 1398 1399 1400 1401 1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 1459 1460 1461 1462 1463 1464 1465 1466 1467 1468 1469 1470 1471 1472 1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 1498 1499 1500 1501 1502 1503 1504 1505 1506 1507 1508 1509 1510 1511 1512 1513 1514 1515 1516 1517 1518 1519 1520 1521 1522 1523 1524 1525 1526 1527 1528 1529 1530 1531 1532 1533 1534 1535 1536 1537 1538 1539 1540 1541 1542 1543 1544 1545 1546 1547 1548 1549 1550 1551 1552 1553 1554 1555 1556 1557 1558 1559 1560 1561 1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 1572 1573 1574 1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 1590 1591 1592 1593 1594 1595 1596 1597 1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 1625 1626 1627 1628 1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 1657 1658 1659 1660 1661 1662 1663 1664 1665 1666 1667 1668 1669 1670 1671 1672 1673 1674 1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 1727 1728 1729 1730 1731 1732 1733 1734 1735 1736 1737 1738 1739 1740 1741 1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 1760 1761 1762 1763 1764 1765 1766 1767 1768 1769 1770 1771 1772 1773 1774 1775 1776 1777 1778 1779 1780 1781 1782 1783 1784 1785 1786 1787 1788 1789 1790 1791 1792 1793 1794 1795 1796 1797 1798 1799 1800 1801 1802 1803 1804 1805 1806 1807 1808 1809 1810 1811 1812 1813 1814 1815 1816 1817 1818 1819 1820 1821 1822 1823 1824 1825 1826 1827 1828 1829 1830 1831 1832 1833 1834 1835 1836 1837 1838 1839 1840 1841 1842 1843 1844 1845 1846 1847 1848 1849 1850 1851 1852 1853 1854 1855 1856 1857 1858 1859 1860 1861 1862 1863 1864 1865 1866 1867 1868 1869 1870 1871 1872 1873 1874 1875 1876 1877 1878 1879 1880 1881 1882 1883 1884 1885 1886 1887 1888 1889 1890 1891 1892 1893 1894 1895 1896 1897 1898 1899 1900 1901 1902 1903 1904 1905 1906 1907 1908 1909 1910 1911 1912 1913 1914 1915 1916 1917 1918 1919 1920 1921 1922 1923 1924 1925 1926 1927 1928 1929 1930 1931 1932 1933 1934 1935 1936 1937 1938 1939 1940 1941 1942 1943 1944 1945 1946 1947 1948 1949 1950 1951 1952 1953 1954 1955 1956 1957 1958 1959 1960 1961 1962 1963 1964 1965 1966 1967 1968 1969 1970 1971 1972 1973 1974 1975 1976 1977 1978 1979 1980 1981 1982 1983 1984 1985 1986 1987 1988 1989 1990 1991 1992 1993 1994 1995 1996 1997 1998 1999 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 2029 2030 2031 2032 2033 2034 2035 2036 2037 2038 2039 2040 2041 2042 2043 2044 2045 2046 2047 2048 2049 2050 2051 2052 2053 2054 2055 2056 2057 2058 2059 2060 2061 2062 2063 2064 2065 2066 2067 2068 2069 2070 2071 2072 2073 2074 2075 2076 2077 2078 2079 2080 2081 2082 2083 2084 2085 2086 2087 2088 2089 2090 2091 2092 2093 2094 2095 2096 2097 2098 2099 2100 2101 2102 2103 2104 2105 2106 2107 2108 2109 2110 2111 2112 2113 2114 2115 2116 2117 2118 2119 2120 2121 2122 2123 2124 2125 2126 2127 2128 2129 2130 2131 2132 2133 2134 2135 2136 2137 2138 2139 2140 2141 2142 2143 2144 2145 2146 2147 2148 2149 2150 2151 2152 2153 2154 2155 2156 2157 2158 2159 2160 2161 2162 2163 2164 2165 2166 2167 2168 2169 2170 2171 2172 2173 2174 2175 2176 2177 2178 2179 2180 2181 2182 2183 2184 2185 2186 2187 2188 2189 2190 2191 2192 2193 2194 2195 2196 2197 2198 2199 2200 2201 2202 2203 2204 2205 2206 2207 2208 2209 2210 2211 2212 2213 2214 2215 2216 2217 2218 2219 2220 2221 2222 2223 2224 2225 2226 2227 2228 2229 2230 2231 2232 2233 2234 2235 2236 2237 2238 2239 2240 2241 2242 2243 2244 2245 2246 2247 2248 2249 2250 2251 2252 2253 2254 2255 2256 2257 2258 2259 2260 2261 2262 2263 2264 2265 2266 2267 2268 2269 2270 2271 2272 2273 2274 2275 2276 2277 2278 2279 2280 2281 2282 2283 2284 2285 2286 2287 2288 2289 2290 2291 2292 2293 2294 2295 2296 2297 2298 2299 2300 2301 2302 2303 2304 2305 2306 2307 2308 2309 2310 2311 | # 06 β Evidence and Confidence
**Parent:** [Architecture hub](README.md) Β· **Sibling chapters:**
[01 System overview](01-system-overview.md) Β· [05 Specialists](05-specialists.md) Β·
[07 Configuration freeze](07-configuration-freeze.md)
**Primary sources read for this chapter (all under `C:/Users/anish/satquery-ai/`):**
| Source | What it establishes here |
|---|---|
| `evidence/engine.py` (714 lines) | the aggregation pipeline, purity contract, identity/sort/claim keys, `_deduplicate`, `_record_agreement`, `_renumber`, `evidence_type_for`, `evidence_from_box/region/geospatial`, `confidence_for`, `evidence_digest` |
| `evidence/confidence.py` (434 lines) | the honesty rule, `TemperatureCalibration`, `_is_effective`, `_logit`/`_sigmoid`, `_EPS`, `calibrate`, `calibrate_result`, `load_calibration` and its ordered candidate search |
| `core/schemas.py` (462 lines) | `Evidence`, `EvidenceType`, `ConfidenceBreakdown`, `ExecutionTrace`, `TraceStep`, `ModelRef`, `SpecialistResult`, `CoordinateSystem`, and every validator |
| `configs/base.yaml` (Β§`evidence`, Β§`confidence`) | `evidence.max_items: 32`, `confidence.temperature_scaling: true`, `confidence.calibration_file: calibration_v001.json` |
| `artifacts/calibration_v001.json` | the measured fitted temperature and the reliability diagram |
| `frontend/assets/js/core.js` | `SQ.EVENT_NAMES` β the eight execution events |
| `frontend/assets/js/mission.js` | `markState()` and the `.trace__fill` width formula |
| `docs/ARCHITECTURE_FREEZE.md` Β§1, Β§3, Β§5 | the layer verbs and the frozen non-negotiables |
| `docs/PHASE13_EVIDENCE_ENGINE.md` | the phase record for this package (77 tests) |
| `docs/API_CONTRACT.md` Β§2.4, Β§3.3, Β§4 | the client-facing shape of evidence and confidence |
| `docs/DEPLOYMENT_ARCHITECTURE.md` Β§5.5 | the F-16 owner ruling on artifact refs |
| `docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§7, Β§8 | the measured calibration and evidence integration results |
---
## 1. Where this subsystem sits
`docs/ARCHITECTURE_FREEZE.md` Β§5 gives every layer exactly one verb:
> Router *understands*; policy engine *decides*; specialists *compute*; VLM *explains*;
> **evidence engine *proves*.**
Two consecutive stages implement that last verb:
```mermaid
flowchart LR
SPEC["Specialists<br/>compute"] -->|"SpecialistResult.evidence"| EV["Evidence engine<br/>aggregate()<br/>collect β dedup β sort β<br/>annotate β renumber β cap"]
EV -->|"EvidenceCollection"| CONF["Confidence<br/>confidence_for() / calibrate()"]
CONF -->|"ConfidenceBreakdown"| RES["ResultEnvelope<br/>+ ExecutionTrace"]
style EV fill:#1f6feb22,stroke:#1f6feb
style CONF fill:#1f6feb22,stroke:#1f6feb
```
The freeze Β§1 pipeline diagram places them adjacently:
```
+---------+---------+---------+
| | | |
VQA/CAP GROUNDING CHANGE OPTICAL-SAR
| | | |
+---------+---------+---------+
|
v
Evidence engine <- this chapter, part A
|
v
Confidence calibration <- this chapter, part B
|
v
Result normaliser -> JSON + trace + PDF
```
**Division of labour, stated precisely.** `evidence/engine.py` does **not** re-derive any
specialist's claim. The module docstring is explicit (`evidence/engine.py:10-14`):
> This module is the "proves" step. Specialists each emit `Evidence` for what *they*
> computed (see `GroundingSpecialist._build_evidence`). This engine does not re-derive any
> of that. Its job is aggregation:
>
> collect across specialists -> order -> deduplicate -> renumber -> bound
**Status of this subsystem.**
| Component | Status | Basis |
|---|---|---|
| Evidence aggregation | `IMPLEMENTED` + `VERIFIED` | `evidence/engine.py`; 77 tests in `tests/unit/test_evidence_engine.py` (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§7) |
| Dedup / ordering / renumber / cap | `IMPLEMENTED` + `VERIFIED` | pinned by named tests (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§2.2βΒ§2.5) |
| Confidence honesty rule (uncalibrated pass-through) | `IMPLEMENTED` + `VERIFIED` | `evidence/confidence.py:296-311` |
| Temperature-scaling transform | `IMPLEMENTED` + `VERIFIED` | monotonicity + endpoint-safety tests (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.3) |
| Calibration artifact fitted | `MEASURED` | `artifacts/calibration_v001.json` |
| Calibration **improves** ECE | **`REJECTED` β it does not** | ECE 0.013755 β 0.014929 (worse); see Β§7 |
| Evidence engine wired into the controller | `IMPLEMENTED` (Phase 14) | `docs/PHASE13_EVIDENCE_ENGINE.md` Β§8 recorded it as "not yet wired"; the live run path emits the eight events (Β§9) |
| Artifact rendering / retrieval | `OPEN` β deliberately `null` in v1 | F-16 ruling, Β§3.2 |
---
# Part A β The evidence schema
## 2. `Evidence` β the canonical record
`Evidence` is defined once, in `core/schemas.py:216-253`. Every specialist emits it; nothing
invents a second shape (`docs/ARCHITECTURE_FREEZE.md` Β§3: *"Single schema for every
specialist. No specialist invents its own shape."*).
```python
class Evidence(BaseModel):
model_config = ConfigDict(extra="forbid")
evidence_id: str = Field(default_factory=lambda: _new_id("ev"))
type: EvidenceType
source_specialist: str
coordinate_system: CoordinateSystem | None = None
coordinates: list[float] | None = None
score: float | None = Field(default=None, ge=0.0, le=1.0)
artifact_ref: str | None = Field(
default=None,
description=(
"Reference to an externally retrievable artifact. NEVER a "
"filesystem path (F-16, owner ruling 2026-09-23): v1 exposes no "
"artifact-serving endpoint, so this is null unless a deployment "
"supplies a client-fetchable reference. An artifact may still be "
"written server-side where configured; being written is not the "
"same as being retrievable."
),
)
payload: dict[str, Any] = Field(default_factory=dict)
```
### 2.1 Every field, exhaustively
| Field | Type | Required | Default | Meaning | Note |
|---|---|---|---|---|---|
| `evidence_id` | `str` | no | `_new_id("ev")` β `ev_<12 hex>` | the item's identity | **run-local**, engine-assigned after aggregation (Β§5.8); a uuid before that |
| `type` | `EvidenceType` | **yes** | β | the *kind* of proof | closed 11-member enum, Β§3 |
| `source_specialist` | `str` | **yes** | β | which specialist made this claim | a **single string**, not a set β this is load-bearing, Β§5.7 |
| `coordinate_system` | `CoordinateSystem \| None` | no | `None` | the frame the coordinates live in | **mandatory in practice** for spatial types β see the validator, Β§4 |
| `coordinates` | `list[float] \| None` | no | `None` | the geometry | a flat list; a box is `[x1, y1, x2, y2]` |
| `score` | `float \| None` | no | `None` | reliability, bounded | `ge=0.0, le=1.0` β Pydantic rejects out-of-range |
| `artifact_ref` | `str \| None` | no | `None` | a *retrievable* artifact reference | **always `null` in v1** β Β§2.3 |
| `payload` | `dict[str, Any]` | no | `{}` | observations that *support* the claim without being it | where `corroborated_by` is written, Β§5.7 |
### 2.2 What `Evidence` deliberately does **not** have
These absences are load-bearing; each one is why a downstream design decision exists.
| Absent field | Why it is absent | Consequence |
|---|---|---|
| any ordering field (`index`, `rank`, `order`) | the engine owns order; a specialist must not pre-empt it | order is defined by `_sort_key`, Β§5.6 |
| `value` | the field is called **`score`** | `docs/PHASE19_FINAL_HARDENING.md` Β§3.6 records that the docs once said `value` and that was one of nine validated defects |
| `source` | the field is called **`source_specialist`** | same defect class; a frontend using `source` renders nothing |
| `summary` | not in the schema | same defect class |
| `contributing_specialists: list[str]` | **proposed and NOT adopted** | see Β§5.7 and `docs/PHASE13_EVIDENCE_ENGINE.md` Β§5 |
| `label` at top level | labels live in `payload` | `evidence_from_box` writes `payload["label"]` |
| a `retrievable` boolean | retrieval is not a v1 capability | `artifact_ref` being `null` *is* the signal, Β§2.3 |
`model_config = ConfigDict(extra="forbid")` is what makes all of the above enforceable: a
response carrying `value` instead of `score` is a **422**, not a silently-ignored key. The
conformance test asserts this directly β `docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§D records
"`extra="forbid"` on all eight client-facing models; `GeoMetadata` is the one documented
`extra="allow"` exception."
### 2.3 `artifact_ref` is permanently `null` in v1 β owner ruling F-16
This is the single most mis-documented field in the system, and the ruling is recorded in
three places: the field's own `description` (`core/schemas.py:225-235`), the client contract
(`docs/API_CONTRACT.md` Β§2.4 "Artifact refs β `null` in v1, and why"), and the audit
(`docs/DEPLOYMENT_ARCHITECTURE.md` Β§5.5).
**The ruling (F-16, owner, 2026-09-23): never expose filesystem paths.**
`docs/API_CONTRACT.md:586-602` states the contract in full:
> **Every `artifact_ref` and `change_map` in a v1 response is `null`.** This is a deliberate
> contract, not a missing value.
>
> F-16 (owner ruling 2026-09-23): **never expose filesystem paths.** The specialists *do*
> render their artifacts β the change map and the optical/SAR views are written
> server-side β but their location is an operator fact, not a client-facing one. A response
> that carried the server's path would disclose the deployment's directory layout to an
> unauthenticated caller, and nothing the frontend can do requires it.
>
> **No `artifact://` URI is fabricated in its place.** v1 has **no artifact-serving
> endpoint**, so a URI would be a promise the service cannot keep β strictly worse than
> `null`, because the frontend would build a link that 404s.
The docstring's own phrasing is the crispest statement of the principle:
> **being written is not the same as being retrievable.**
```mermaid
flowchart TB
subgraph Server["Server side (operator facts)"]
R["specialist renders change map / views"]
W["file written where configured<br/>change.artifact_dir"]
R --> W
end
subgraph Client["Client-facing contract (v1)"]
N["artifact_ref = null"]
P["payload statistics<br/>total_change_pixels<br/>n_components_kept<br/>threshold"]
WARN["warnings[] entry:<br/>artifact NOT retrievable"]
end
W -.->|"deliberately NOT exposed"| X["filesystem path<br/>(never sent)"]
R --> N
R --> P
R --> WARN
style X fill:#f8514922,stroke:#f85149
style N fill:#3fb95022,stroke:#3fb950
```
**What replaces the ref** (`docs/API_CONTRACT.md:604-610`):
| Removed | Replaced by |
|---|---|
| `change_map` path | `null`, plus the change statistics in the CHANGE_MAP evidence's `payload` (`total_change_pixels`, `n_components_kept`, `threshold`) |
| view `artifact_ref` path | `null`, plus `payload.rendered` / `payload.retrievable` / `payload.retrieval` |
| β | an explicit `warnings[]` entry saying the artifact is **NOT retrievable** |
**Two live carriers, and the half-implementation that followed.** The audit records that the
ruling was initially applied to only one carrier
(`docs/DEPLOYMENT_ARCHITECTURE.md:361`):
> **Two live carriers** β `change` and `croma` β and **both** had to be brought into line:
> fixing only the `croma` carrier left the ruling half implemented.
`docs/STATUS.md:21` names the lesson: *"The finding worth carrying forward: the F-16 ruling
was HALF IMPLEMENTED, and the suite was green."*
**F-16c β the consequence the fix introduced, measured and left open.** The ruling removed a
proxy and thereby changed a `degraded` signal. The `SpecialistResult` validator used to read
(`core/schemas.py:362-372`, retained as a comment):
```python
if self.task is Task.CHANGE and self.change_map is None and not self.regions:
... "change analysis produced no spatial output" ... degraded = True
```
`change_map` was a *proxy* for "a map was produced". With the ref permanently `null`, the
clause collapsed to `not regions`, and a **successful no-change analysis began reporting
`degraded: true`**. The clause was therefore **removed** under ruling F-16c
(`core/schemas.py:362-395`). The comment left in its place is worth quoting because it states
the reasoning better than a summary could:
> That conflated two different things, and the conflation WAS the defect. `degraded` means
> "the analysis could not be fully performed": no trained detector, or co-registration too
> poor to support a spatial claim. Both are set by the specialist itself
> (`specialists/change/specialist.py:345` and `:365`) and neither is this validator's to
> invent. **"No change was detected" is a NORMAL, successful outcome.**
And, on why no replacement clause was added:
> No replacement clause is added, deliberately. A narrower "empty answer => degraded" rule
> was tried and backed out: it is not what the ruling asked for, it invented a semantic the
> specialist already owns, and it made a pre-existing, unrelated fixture
> (`test_change_with_regions_is_not_degraded`, a result with regions and no answer) fail. **A
> fix that forces edits to tests it has nothing to do with is signalling over-reach, not
> diligence.** CHANGE is the one task whose `degraded` flag is now set entirely by its
> specialist.
**Status:** F-16 `CLOSED` (both carriers conform). F-16c `RESOLVED` by clause removal β the
narrower replacement rule is `REJECTED`. `mask_ref` is *not* an exemption: the audit
corrected an earlier working note that had listed it alongside `artifact_ref`
(`docs/DEPLOYMENT_ARCHITECTURE.md:868-871`).
## 3. `EvidenceType` β the closed 11-member vocabulary
```python
class EvidenceType(str, Enum):
IMAGE_CROP = "image_crop"
TILE = "tile"
BOUNDING_BOX = "bounding_box"
MASK = "mask"
CHANGE_MAP = "change_map"
OPTICAL_VIEW = "optical_view"
SAR_VIEW = "sar_view"
JOINT_FEATURE_REGION = "joint_feature_region"
STATISTIC = "statistic"
GEOLOCATION = "geolocation"
AVAILABILITY_MASK = "availability_mask" # C-1: modality trust evidence
```
`core/schemas.py:65-76`. **All eleven members, with what each one asserts:**
| # | Member | Wire value | Asserts | Typically emitted by |
|---|---|---|---|---|
| 1 | `IMAGE_CROP` | `image_crop` | a rectangular raster crop exists | preprocessing / tiling |
| 2 | `TILE` | `tile` | one tile of the tiling policy was examined | `preprocessing/tiling.py` |
| 3 | `BOUNDING_BOX` | `bounding_box` | a rectangle locates a referent | grounding, change |
| 4 | `MASK` | `mask` | a per-pixel region exists | grounding, change |
| 5 | `CHANGE_MAP` | `change_map` | a bitemporal difference map exists | change specialist |
| 6 | `OPTICAL_VIEW` | `optical_view` | the optical rendering of an optical/SAR pair | optical-SAR specialist |
| 7 | `SAR_VIEW` | `sar_view` | the SAR rendering of an optical/SAR pair | optical-SAR specialist |
| 8 | `JOINT_FEATURE_REGION` | `joint_feature_region` | a region defined in CROMA's **joint** embedding space | optical-SAR specialist |
| 9 | `STATISTIC` | `statistic` | a measured scalar or count, with no geometry | VQA, caption, unsupported, and any geometry-less region |
| 10 | `GEOLOCATION` | `geolocation` | the raster is georeferenced and *where* it sits | geospatial layer |
| 11 | `AVAILABILITY_MASK` | `availability_mask` | which modality channels were actually available | sensor adapter (C-1) |
### 3.1 `availability_mask` β the member the freeze prose omits
`docs/ARCHITECTURE_FREEZE.md` Β§3's prose list of evidence types stops at ten members and does
**not** name `availability_mask`. The enum member is nevertheless real, and the engine says
so explicitly (`evidence/engine.py:441-447`):
> The members are exactly those of `core.schemas.EvidenceType`. The freeze's section 3 prose
> list omits `availability_mask`; that member is real (C-1: modality trust evidence) and is
> included here because **a missing member would otherwise force a wrong fallback.**
Its existence follows from finding C-1: the channel-availability mask is a first-class fusion
input and is **never** a CROMA input (`core/schemas.py:7`; `docs/PHASE0_CONTRACT_VALIDATION.md`
Β§1.2). Because the mask is a first-class *input*, it is also a first-class *observation* β
"these twelve optical channels were present and these two SAR channels were not" is a fact a
consumer can audit. `docs/API_CONTRACT.md:579` lists all eleven values for the frontend.
### 3.2 Which types are spatial
The engine's own validator names the spatial set (`core/schemas.py:240-247`) β this is the
authoritative list, not a prose paraphrase:
```python
spatial = {
EvidenceType.BOUNDING_BOX,
EvidenceType.MASK,
EvidenceType.CHANGE_MAP,
EvidenceType.TILE,
EvidenceType.IMAGE_CROP,
EvidenceType.JOINT_FEATURE_REGION,
}
```
| Type | Spatial? | Why |
|---|---|---|
| `BOUNDING_BOX` | **yes** | a box is meaningless without a frame |
| `MASK` | **yes** | pixels need a frame |
| `CHANGE_MAP` | **yes** | a map is a raster |
| `TILE` | **yes** | a tile is a rectangle of a raster |
| `IMAGE_CROP` | **yes** | a crop is a rectangle of a raster |
| `JOINT_FEATURE_REGION` | **yes** | a region in a feature grid still needs a frame |
| `STATISTIC` | no | a scalar has no geometry |
| `GEOLOCATION` | no *(by this validator)* | geo bounds are already self-describing via `payload["crs"]`; `evidence_from_geospatial` sets `CoordinateSystem.GEO` when bounds are given |
| `OPTICAL_VIEW` | no *(by this validator)* | a whole-frame view is not a sub-region claim |
| `SAR_VIEW` | no *(by this validator)* | as above |
| `AVAILABILITY_MASK` | no *(by this validator)* | a channel-presence vector is not spatial |
## 4. The validator that refuses coordinates without a frame
```python
@model_validator(mode="after")
def _spatial_needs_crs(self) -> "Evidence":
spatial = {
EvidenceType.BOUNDING_BOX,
EvidenceType.MASK,
EvidenceType.CHANGE_MAP,
EvidenceType.TILE,
EvidenceType.IMAGE_CROP,
EvidenceType.JOINT_FEATURE_REGION,
}
if self.type in spatial and self.coordinates and self.coordinate_system is None:
raise ValueError(
f"evidence type '{self.type.value}' carries coordinates "
"but no coordinate_system"
)
return self
```
`core/schemas.py:238-253`. Three things about this are deliberate:
1. **It is a `model_validator(mode="after")`, not a `field_validator`.** The rule is
*cross-field* β it depends on `type`, `coordinates` and `coordinate_system` together, so
it cannot be expressed on any single field.
2. **The trigger is `self.coordinates`, not `is not None`.** An empty list is falsy and does
not trip the rule; a non-empty list does. A spatial item with **no** coordinates at all is
legal (a `MASK` whose geometry lives only in a `mask_ref`).
3. **It raises rather than defaulting.** Defaulting to `normalized_0_1` would silently
mislabel pixel or geo coordinates as normalized, which is the failure mode the finding
exists to prevent. `docs/ARCHITECTURE_FREEZE.md` Β§3 puts the requirement bluntly:
> Every spatial object carries `coordinate_system` β `{normalized_0_1, pixel, geo}`.
`docs/ARCHITECTURE_FREEZE.md` Β§2.3 states the origin: internal box coordinates are
**normalised 0β1**, **always with an explicit `coordinate_system` field** (finding C-5).
### 4.1 `CoordinateSystem` β three values, exact spellings
```python
class CoordinateSystem(str, Enum):
"""Never omit this. A bare box is meaningless without it. (C-5)"""
NORMALIZED_0_1 = "normalized_0_1"
PIXEL = "pixel"
GEO = "geo"
```
`core/schemas.py:57-62`.
| Member | Wire value | Frame |
|---|---|---|
| `NORMALIZED_0_1` | `normalized_0_1` | fractions of the frame, `[0, 1]` |
| `PIXEL` | `pixel` | integer-ish pixel indices |
| `GEO` | `geo` | a projected/geographic CRS named in `payload["crs"]` |
> **Documentation hazard, measured.** `docs/PHASE19_FINAL_HARDENING.md` Β§3.6 records that the
> docs once spelled these `normalized` / `geographic`. The real values are
> `normalized_0_1` / `geo`. *"A frontend using the documented names would send values the
> server rejects with 422."* `docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§D confirms the live enum is
> exactly `{normalized_0_1, pixel, geo}`.
`Evidence`'s `coordinate_system` is `CoordinateSystem | None` β the field is optional *in
type* but effectively mandatory *in practice*, because the validator above enforces it for
every spatial type that carries coordinates.
---
# Part B β The aggregation pipeline
## 5. `EvidenceEngine.aggregate()` β collect β dedup β sort β annotate β renumber β cap
### 5.1 Why the module exists: three problems a specialist cannot solve
The module docstring enumerates them (`evidence/engine.py:18-36`); each is real and none is
fixable inside a single specialist:
| # | Problem | Why a specialist cannot fix it |
|---|---|---|
| 1 | **Stable identity** | `Evidence.evidence_id` defaults to a random uuid β fine within one result, useless once four specialists' evidence is merged into one trace. *"Nothing can be cited."* |
| 2 | **Order** | `Evidence` has no ordering field, so the collection's order follows **specialist completion order** β the same inputs produce different JSON on different runs. *"Pure functions must not do that, and this whole system's reproducibility argument rests on it."* |
| 3 | **Duplication** | the VQA specialist emits a `STATISTIC` carrying its answer, *and* that answer travels in `SpecialistResult.answer`; two specialists that georeference the same asset emit the same `GEOLOCATION`. *"Neither is wrong, and neither knows about the other."* |
### 5.2 The pipeline, in one place
```mermaid
flowchart TB
A["results: SpecialistResult<br/>or Sequence[SpecialistResult]<br/>or evidence: Iterable[Evidence]"] --> B["collect<br/>extend(result.evidence)"]
B --> C["sources = sorted({source_specialist})<br/><i>recorded BEFORE the cap</i>"]
C --> D["_deduplicate<br/>identity = _identity_key<br/>payloads MERGED on collision"]
D --> E["sorted(key=_sort_key)<br/>type β Β· specialist β Β·<br/>score β Β· coordinates β"]
E --> F["_record_agreement<br/>_claim_key groups β<br/>payload['corroborated_by']"]
F --> G["annotated[:max_items]<br/>dropped_over_limit = total β len(capped)"]
G --> H["_renumber β evidence_001β¦"]
H --> I["EvidenceCollection"]
style D fill:#1f6feb22,stroke:#1f6feb
style F fill:#1f6feb22,stroke:#1f6feb
```
The implementation, verbatim (`evidence/engine.py:324-350`):
```python
def _canonicalise(self, raw: Iterable[Evidence]) -> EvidenceCollection:
"""Dedup -> sort -> renumber -> cap. The whole pipeline, in one place."""
items = list(raw)
sources = sorted({item.source_specialist for item in items})
deduped, dropped_duplicates = self._deduplicate(items)
ordered = sorted(deduped, key=_sort_key)
# Two specialists can make the same claim, and `deduplicate` keeps both
# (different `source_specialist` -> different key). Record the agreement
# now that ordering is fixed, so the annotation is part of the pure
# pipeline rather than a post-hoc edit a caller might forget.
annotated = self._record_agreement(ordered)
total = len(annotated)
capped = annotated[: self.max_items]
dropped_over_limit = total - len(capped)
return EvidenceCollection(
items=[
self._renumber(item, index) for index, item in enumerate(capped, start=1)
],
sources=sources,
dropped_duplicates=dropped_duplicates,
dropped_over_limit=dropped_over_limit,
total_before_limit=total,
truncated=dropped_over_limit > 0,
)
```
Note the ordering choice: **annotate comes after sort but before cap**. Annotating before
sorting would be equivalent (the annotation is a payload key, which is not in `_sort_key`),
but the docstring's stated reason is that the annotation must be *inside* the pure pipeline
rather than "a post-hoc edit a caller might forget".
### 5.3 The entry point and its argument rules
```python
def aggregate(
self,
results: SpecialistResult | Sequence[SpecialistResult] | None = None,
*,
evidence: Iterable[Evidence] | None = None,
) -> EvidenceCollection:
```
`evidence/engine.py:280-320`. Two argument invariants, both raising:
| Condition | Behaviour | Reason given in the docstring |
|---|---|---|
| both `results` and `evidence` are `None` | `ValueError("aggregate() requires results= or evidence=")` | there is nothing to aggregate |
| both are supplied | `ValueError("aggregate() accepts results= or evidence=, not both")` | ambiguity about precedence |
| empty list | returns an empty `EvidenceCollection` | *"Never raises on empty input."* |
**Only `result.evidence` is read.** The docstring explains why reading both forms would
double-count (`evidence/engine.py:292-297`):
> Specialists state their evidence twice β in `SpecialistResult.evidence` and via
> `Specialist.produce_evidence`, which currently delegates to the same list β so only
> `result.evidence` is read. Reading both would double-count every item and inflate the
> duplicate count.
### 5.4 The purity contract
```python
"""THE PURITY CONTRACT
-------------------
`aggregate` is pure and deterministic:
* source results are never mutated -- `model_copy` is used to rebuild the
collection rather than editing `Evidence.evidence_id` in place;
* ids are assigned from the sorted position, not from input order;
* dedup is order-insensitive by construction (the key is built from the
sorted specialist list);
* no clock, no RNG, no I/O.
Same inputs -> byte-identical output. `tests/unit/test_evidence_engine.py`
pins this.
"""
```
`evidence/engine.py:38-50`.
| Property | Mechanism | Pinned by |
|---|---|---|
| source non-mutation | `model_copy(update={...})` at every write β `_renumber` (`:430`), payload merge (`:385`), agreement (`:424`) | *"the original keeps its uuid and the aggregated copy is a different object"* (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§2.1) |
| ids from sorted position | `enumerate(capped, start=1)` after `sorted(...)` | Β§5.8 |
| order-insensitive dedup | key content-only; `_record_agreement` sorts `names` | Β§5.5, Β§5.7 |
| no clock / RNG / I/O | the module imports only `hashlib`, `collections.abc`, `dataclasses`, `typing` | `evidence/engine.py:80-102` |
The dedup merge rebuilds rather than mutates, and the comment says why
(`evidence/engine.py:379-382`):
> Survivor's own keys win; the loser only fills gaps. Mutating the survivor in place would be
> fine here because `merged` holds the same object, but rebuilding keeps the "never mutate a
> source result" contract true even if the caller kept a reference.
### 5.5 Deduplication β identity is *defined*, not assumed
The module docstring (`evidence/engine.py:52-68`) states the principle:
> `Evidence.evidence_id` is a uuid default, so it cannot be part of an identity key β two
> structurally identical items from two runs would never dedup. The key is therefore the
> *content* of the observation.
```python
def _identity_key(item: Evidence) -> tuple[Any, ...]:
coords = (
tuple(_round(c) for c in item.coordinates)
if item.coordinates is not None
else None
)
return (
item.type.value,
item.source_specialist,
item.coordinate_system.value if item.coordinate_system else None,
coords,
_round(item.score),
)
```
`evidence/engine.py:125-142`.
| Key component | Included? | Rationale |
|---|---|---|
| `type` | **yes** | a box and a mask are different claims even at the same coordinates |
| `source_specialist` | **yes** | keeps two specialists' identical claims as two items (Β§5.7) |
| `coordinate_system` | **yes** | the same four numbers in `pixel` and in `geo` are different claims |
| rounded `coordinates` | **yes** | disagreeing geometry β two claims; suppressing it would be a silent contradiction |
| rounded `score` | **yes** | different reliability is a different observation |
| `evidence_id` | **NO** | *"including it would mean dedup never fires in production"* |
| `payload` | **NO** | *"a detail of the claim, not a different claim"* |
| `artifact_ref` | **NO** | *"two renderings of one claim are still one claim"* |
The docstring on the exclusions is the clearest statement of intent
(`evidence/engine.py:60-68`):
> `payload` and `artifact_ref` are deliberately EXCLUDED. Two items that agree on the same
> geolocation, one carrying a `crs` in its payload and one not, have made the same claim
> about the world; deduplicating them is correct, and the surviving item's payload is merged
> with the discarded one's so the `crs` is not lost. Conversely two items that share a type
> and score but disagree on coordinates are two different claims and are both kept β which
> matters, because **suppressing a spatial disagreement would be exactly the silent
> contradiction the freeze forbids.**
**The rounding constant.**
```python
#: Rounding applied before an evidence item is turned into a dedup key. Six
#: decimals on normalized coordinates is ~1e-6 of the frame, far below any
#: meaningful spatial difference, but enough to absorb float noise from two
#: specialists rounding the same value two different ways.
_KEY_PRECISION = 6
def _round(value: float | None) -> float | None:
return None if value is None else round(float(value), _KEY_PRECISION)
```
`evidence/engine.py:114-122`. `_round` is `None`-safe: a `None` coordinate stays `None` and
does not become `0.0` in the *identity* key (unlike the sort key, Β§5.6, where a missing
coordinate becomes `0.0` because a total order needs a value).
**Payloads are MERGED on collision.**
```python
dropped += 1
# Survivor's own keys win; the loser only fills gaps.
payload = {**item.payload, **existing.payload}
if payload != existing.payload:
updated = existing.model_copy(update={"payload": payload})
seen[key] = updated
merged[merged.index(existing)] = updated
```
`evidence/engine.py:378-387`. The spread order is `{**loser, **survivor}` β because Python
dict unpacking is last-wins, `existing.payload` (the survivor's) overrides the loser's on
key collision. This is the "survivor's own keys win; the loser only fills gaps" rule
expressed directly in the merge order.
| Behaviour | Result |
|---|---|
| survivor had `crs`, loser did not | survivor's `crs` survives |
| loser had `crs`, survivor did not | `crs` is **filled in** from the loser β the information is not lost |
| both had `crs` with different values | survivor's wins; the loser's value is dropped |
| no payload difference | no `model_copy` is performed (the `if payload != existing.payload` guard) |
`docs/PHASE13_EVIDENCE_ENGINE.md` Β§2.3 names the test that pins this:
`test_deduplication_ignores_payload_differences_and_merges_them`.
**What dedup does *not* merge.** Because the key includes `source_specialist`, *"only
same-specialist repeats are merged here; two specialists agreeing is handled by
`_record_agreement`, which preserves the second specialist's identity rather than discarding
it"* (`evidence/engine.py:356-361`).
### 5.6 The total sort order β and why each level exists
```python
def _sort_key(item: Evidence) -> tuple[Any, ...]:
if item.score is None:
score_key: tuple[int, float] = (1, 0.0)
else:
score_key = (0, -float(item.score))
coords = (
tuple(_round(c) or 0.0 for c in item.coordinates)
if item.coordinates is not None
else ()
)
return (item.type.value, item.source_specialist, score_key, coords)
```
`evidence/engine.py:165-190`.
| # | Key | Direction | Why it exists (verbatim from the docstring) |
|---|---|---|---|
| 1 | `type` | ascending | *"groups like with like, so a reader sees all the boxes together"* |
| 2 | `source_specialist` | ascending | *"stable and meaningful, unlike the uuid"* |
| 3 | `score` | **descending** | *"within a type, the strongest claim leads"* β implemented as `-score` so the tuple stays uniformly ascending |
| 4 | `coordinates` | ascending | *"the tie-breaker that makes the order total. Without it two items of the same type, specialist and score would sort by Python's stable-sort insertion order, which reintroduces input-order dependence."* |
`docs/PHASE13_EVIDENCE_ENGINE.md` Β§2.2 restates key 4 and names the pinning test:
> Key 4 is not decoration. Without it, two items of the same type, specialist and score would
> sort by Python's stable-sort insertion order β which reintroduces exactly the input-order
> dependence the sort exists to remove. Pinned by
> `test_equal_scores_order_deterministically_by_coordinates`.
**Unscored items sort last.** *"Items with no score sort AFTER items with a score: a missing
score is not evidence of strength."* (`evidence/engine.py:178-179`). The mechanism is a
`(0, β¦)` / `(1, β¦)` discriminator in `score_key`: a scored item leads with `0`, an unscored
item leads with `1`, so every scored item precedes every unscored item **within the same
`(type, specialist)` group**.
```mermaid
flowchart LR
subgraph G1["type=bounding_box, specialist=grounding"]
direction TB
A1["score 0.91"] --> A2["score 0.62"] --> A3["score None<br/><i>unscored last</i>"]
end
subgraph G2["type=statistic, specialist=vqa"]
direction TB
B1["score 0.74"]
end
G1 --> G2
style A3 fill:#d2992222,stroke:#d29922
```
### 5.7 `_record_agreement` β corroboration, and why nothing is ever removed
**The problem.** `source_specialist` is a single `str`, not a set (the schema is frozen), so
two specialists making the same claim are two *different* items under `_identity_key` and
**both survive**. The docstring calls this *"the correct conservative behaviour"*
(`evidence/engine.py:70-77`):
> That is the correct conservative behaviour: collapsing them would require the engine to pick
> a winner, and the engine has no basis for that. What it does instead is keep both and record
> the disagreement-free agreement in the collection's `sources` and in each item's `payload`
> via `_record_agreement`.
**The claim key β deliberately distinct from the identity key.**
```python
def _claim_key(item: Evidence) -> tuple[Any, ...]:
"""Identity of the *claim*, ignoring which specialist made it.
Used only to detect corroboration. Deliberately distinct from
`_identity_key`: that one answers "is this the same item", this one answers
"are these two specialists saying the same thing about the world".
"""
coords = (
tuple(_round(c) for c in item.coordinates)
if item.coordinates is not None
else None
)
return (
item.type.value,
item.coordinate_system.value if item.coordinate_system else None,
coords,
_round(item.score),
)
```
`evidence/engine.py:145-162`. The **only** difference from `_identity_key` is the absent
`source_specialist` component.
```mermaid
flowchart TB
I["_identity_key =<br/>(type, specialist, crs, coords, score)"] --> IQ{"same item?"}
C["_claim_key =<br/>(type, crs, coords, score)"] --> CQ{"same claim about the world?"}
IQ -->|"yes β merge payloads"| DEDUP["_deduplicate"]
CQ -->|"yes, β₯2 distinct specialists β annotate"| AGR["_record_agreement"]
style I fill:#1f6feb22,stroke:#1f6feb
style C fill:#3fb95022,stroke:#3fb950
```
**The annotation.**
```python
AGREEMENT_KEY = "corroborated_by"
...
out = list(items)
for indexes in groups.values():
if len(indexes) < 2:
continue
names = sorted({items[i].source_specialist for i in indexes})
if len(names) < 2:
# Same specialist restating itself across the collection. The
# identity key would normally have merged these, so reaching
# here means the payloads differed; it is not corroboration.
continue
for i in indexes:
item = out[i]
others = [n for n in names if n != item.source_specialist]
payload = {**item.payload, EvidenceCollection.AGREEMENT_KEY: others}
out[i] = item.model_copy(update={"payload": payload})
return out
```
`evidence/engine.py:391-425`.
**Why the payload, and not a schema field.** The constant's own comment states it
(`evidence/engine.py:234-238`):
> Payload key under which `_record_agreement` stashes the specialists that made the identical
> claim. **A single string field cannot hold a set, so the agreement is recorded in the
> payload** β which is exactly what the payload is for: observations that support the evidence
> but are not the claim.
The `contributing_specialists: list[str]` field that would have made this first-class was
**proposed and rejected** β `docs/PHASE13_EVIDENCE_ENGINE.md` Β§5 records the full
ARCHITECTURE CHANGE entry. The three reasons:
| # | Reason |
|---|---|
| 1 | Adding a field is a frozen-contract change, and the freeze rule requires evidence that the current design is insufficient. No such evidence exists: identity comes from `evidence_id`; provenance is already served by `EvidenceCollection.sources` + `Evidence.source_specialist`; corroboration is recorded in `payload["corroborated_by"]`. |
| 2 | Because `source_specialist` is a single string, two specialists making the identical claim are two items and **both survive**. Collapsing them would require the engine to pick a winner, and the engine has no basis for that. |
| 3 | Reversing this later is cheap: the engine already computes `_claim_key()`, which groups items by claim ignoring the specialist. Adopting the field would mean emitting one item per group instead of N. |
**Decision: proposal `REJECTED`; `core/schemas.py` unchanged.** Adopted interim =
`source_specialist` (single) + `EvidenceCollection.sources` (run-level provenance) +
`payload["corroborated_by"]` (per-claim corroboration). *"Revisit only if a downstream
consumer needs corroboration as a first-class, queryable field."*
**Nothing is ever removed.** The docstring is unambiguous (`evidence/engine.py:401-403`):
> **Nothing is removed:** dropping a member would throw away a specialist's attribution, and
> the freeze does not permit silently discarding a specialist's evidence.
**The two guards inside the loop** are worth naming because they are the two ways a
non-corroboration could masquerade as one:
| Guard | Condition | Why |
|---|---|---|
| group size | `len(indexes) < 2` β skip | a lone item is not agreement |
| distinct specialists | `len(names) < 2` β skip | the same specialist restating itself is *not* corroboration; if this is reached, the payloads differed (otherwise `_identity_key` would have merged them) |
`corroborated_by` is always the **other** specialists (`others = [n for n in names if n !=
item.source_specialist]`), so an item never lists itself.
**Excluded from the digest.** *"The corroboration annotation in the payload IS excluded,
because it is a derived observation about the collection, not part of the claim."*
(`evidence/engine.py:692-693`) β consistent with `payload` being excluded from
`_identity_key`.
### 5.8 Renumbering to `evidence_001β¦`
```python
ID_PREFIX = "evidence"
@staticmethod
def _renumber(item: Evidence, index: int) -> Evidence:
"""Give one item its canonical id, leaving the source untouched."""
return item.model_copy(update={"evidence_id": f"{ID_PREFIX}_{index:03d}"})
```
`evidence/engine.py:104-107`, `:427-430`.
| Property | Value | Reason |
|---|---|---|
| format | `evidence_001` β¦ | zero-padded to three digits |
| why three digits | *"so lexical sort matches numeric sort up to 999 items -- well past the `evidence.max_items` bound of 32"* (`:104-107`) |
| assignment basis | **sorted position**, not input order | `enumerate(capped, start=1)` after `sorted` |
| scope | **run-local** | *"Ids restart at `001` for each `aggregate` call β they are run-local citations, not global identities."* (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§2.4) |
| source untouched | `model_copy` | the source keeps its `ev_<uuid>` |
**Why this matters at all** (`evidence/engine.py:21-25`):
> `Evidence.evidence_id` defaults to a random uuid, which is fine within one result but
> useless when the controller merges four specialists' evidence into the trace of a single
> run. Nothing can be cited. The engine renumbers to `evidence_001`, `evidence_002`, β¦ so a
> downstream artefact (report, UI, audit) can reference one item deterministically.
### 5.9 The cap, and loss accounting that is never silent
```python
DEFAULT_MAX_ITEMS = 32
```
`evidence/engine.py:109-112`, mirroring `configs/base.yaml`:
```yaml
evidence:
max_items: 32
coordinate_system_default: normalized_0_1
```
`EvidenceEngine.__init__` **refuses** a non-positive cap (`evidence/engine.py:267-276`):
```python
if max_items < 1:
raise ValueError(f"max_items must be >= 1, got {max_items}")
```
> a cap of zero would silently discard every piece of evidence, which is a policy decision
> this engine has no business making on its own. (`evidence/engine.py:258-261`)
**The accounting fields, all five:**
| Field | Type | Meaning | Non-zero means |
|---|---|---|---|
| `sources` | `list[str]` | sorted specialist names that contributed anything, **BEFORE** the cap | provenance β *"Provenance is a property of the run, not of the surviving items."* |
| `dropped_duplicates` | `int` | items merged into an existing claim | **normal and healthy** β two specialists agreed |
| `dropped_over_limit` | `int` | items the cap discarded | **a warning** β a specialist's evidence did not survive |
| `total_before_limit` | `int` | deduplicated count prior to capping | the denominator for the two drop counts |
| `truncated` | `bool` | `dropped_over_limit > 0` | a boolean summary so a trace can branch without recomputing |
`evidence/engine.py:196-217`; restated in `docs/PHASE13_EVIDENCE_ENGINE.md` Β§2.5.
> Losses are **counted, never silent**.
And on why `sources` is recorded pre-cap:
> `sources` is recorded pre-cap deliberately: provenance is a property of the *run*, not of
> the surviving items. A specialist whose evidence did not fit under the cap still ran, and
> the trace must still say so. (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§2.5)
This is the same principle as `core/config.py`'s refusal to silently truncate and
`evidence/confidence.py`'s refusal to fabricate a number: **a bounded output must declare
what the bound cost.**
### 5.10 `EvidenceCollection` β the returned container
```python
@dataclass(frozen=True)
class EvidenceCollection:
items: list[Evidence] = field(default_factory=list)
sources: list[str] = field(default_factory=list)
dropped_duplicates: int = 0
dropped_over_limit: int = 0
total_before_limit: int = 0
truncated: bool = False
```
`evidence/engine.py:193-250`. It is `frozen=True` β a dataclass decorator that blocks
attribute *rebinding* on the collection, complementing the `model_copy` discipline that keeps
the individual `Evidence` items unmutated.
| Method | Returns | Purpose |
|---|---|---|
| `__len__` | `int` | `len(items)` |
| `__iter__` | iterator | iterate the items directly |
| `ids()` | `list[str]` | `[item.evidence_id for item in self.items]` |
| `by_type(t)` | `list[Evidence]` | filter by `EvidenceType` (identity comparison: `item.type is evidence_type`) |
| `by_specialist(name)` | `list[Evidence]` | filter by `source_specialist` |
| `summary()` | `dict[str, Any]` | the observable trace facts |
| `AGREEMENT_KEY` | `"corroborated_by"` | the payload key for corroboration |
**`summary()` β the trace-facing view.**
```python
def summary(self) -> dict[str, Any]:
"""Observable trace facts. No chain-of-thought, no interpretation."""
return {
"returned": len(self.items),
"total_before_limit": self.total_before_limit,
"dropped_duplicates": self.dropped_duplicates,
"dropped_over_limit": self.dropped_over_limit,
"truncated": self.truncated,
"sources": list(self.sources),
"types": sorted({item.type.value for item in self.items}),
}
```
`evidence/engine.py:240-250`. Note `types` is a **sorted set** of the type *values* actually
present in the surviving items β so a consumer can see which evidence kinds survived the cap
without walking the list. The docstring's first line is a constraint, not a description: *"No
chain-of-thought, no interpretation."* This mirrors `docs/ARCHITECTURE_FREEZE.md` Β§5's *"Every
result carries an observable execution trace. No chain-of-thought."*
### 5.11 `evidence_digest()` β reproducibility as an assertion
```python
def evidence_digest(collection: EvidenceCollection | Sequence[Evidence]) -> str:
items = (
collection.items
if isinstance(collection, EvidenceCollection)
else list(collection)
)
hasher = hashlib.sha256()
for item in items:
hasher.update(repr(_identity_key(item)).encode("utf-8"))
hasher.update(b"\n")
return hasher.hexdigest()
```
`evidence/engine.py:680-704`.
| Design choice | Consequence |
|---|---|
| built from `_identity_key`, **in order** | changes when the *claims* change |
| **not** from `evidence_id` | does **not** change when a specialist restates the same claim under a new uuid |
| **not** from `payload` | payload is a detail of the claim, not the claim |
| `repr(...)` + `b"\n"` separator | unambiguous framing; no two distinct sequences can hash the same |
| corroboration annotation excluded | *"it is a derived observation about the collection, not part of the claim"* |
The docstring states the contract it enables (`evidence/engine.py:687-690`):
> That is what makes it useful as a reproducibility assertion:
>
> aggregate(inputs_a) is reproducible iff digest(a) == digest(b)
>
> against the same input, **even across processes where the uuid defaults differ.**
### 5.12 The canonical vocabulary: `evidence_type_for`
```python
@staticmethod
def evidence_type_for(result: SpecialistResult) -> EvidenceType:
from core.schemas import Task
mapping = {
Task.GROUNDING: EvidenceType.BOUNDING_BOX,
Task.CHANGE: EvidenceType.CHANGE_MAP,
Task.OPTICAL_SAR: EvidenceType.JOINT_FEATURE_REGION,
Task.VQA: EvidenceType.STATISTIC,
Task.CAPTION: EvidenceType.STATISTIC,
Task.UNSUPPORTED: EvidenceType.STATISTIC,
}
return mapping.get(result.task, EvidenceType.STATISTIC)
```
`evidence/engine.py:434-458`.
> This is the ONE place that knows the mapping from specialist task to evidence vocabulary, so
> no specialist needs to branch on it. Note it returns the *primary* type; a specialist's own
> `produce_evidence` emits its real per-artefact types, and those are preserved by `aggregate`.
| Task | Primary `EvidenceType` | Why |
|---|---|---|
| `grounding` | `BOUNDING_BOX` | the task *is* localisation |
| `change` | `CHANGE_MAP` | the task *is* a change map |
| `optical_sar` | `JOINT_FEATURE_REGION` | the claim lives in CROMA's joint embedding space |
| `vqa` | `STATISTIC` | an answer is a measured scalar, not geometry |
| `caption` | `STATISTIC` | as above |
| `unsupported` | `STATISTIC` | there is no claim to localise |
| *(fallback)* | `STATISTIC` | `.get(..., STATISTIC)` β an unknown task degrades to a non-spatial claim rather than a wrong spatial one |
**`change_vqa` is absent from the mapping.** `Task.CHANGE_VQA` exists in the enum
(`core/schemas.py:46`) but has no row, so it falls through to `STATISTIC` β which is correct:
change-VQA returns a short *answer*, not a map. The docstring's note that this returns the
*primary* type matters here: a change-VQA result may still carry `CHANGE_MAP` evidence
emitted by the detector it shares, and `aggregate` preserves that.
### 5.13 The three `evidence_from_*` constructors
These exist so the controller does not hand-build `Evidence` objects β the module docstring
calls them *"Convenience for the controller when it needs evidence for a box that a specialist
emitted but did not itself describe."*
#### `evidence_from_box`
```python
def evidence_from_box(self, box, *, source_specialist, payload=None, asset_ref=None) -> Evidence:
return Evidence(
type=EvidenceType.BOUNDING_BOX,
source_specialist=source_specialist,
coordinate_system=box.coordinate_system,
coordinates=[box.x1, box.y1, box.x2, box.y2],
score=box.score,
artifact_ref=asset_ref,
payload={"label": box.label, **(payload or {})},
)
```
`evidence/engine.py:462-485`.
> The `coordinate_system` is copied from the box, **never assumed** β `core.schemas` requires
> it and `Evidence` rejects spatial coordinates without one.
The coordinates are flattened to `[x1, y1, x2, y2]` β matching the flat-`Box` shape that
`docs/PHASE19_FINAL_HARDENING.md` Β§3.6 records as a corrected documentation defect.
#### `evidence_from_region` β a three-way branch
`evidence/engine.py:487-540`. A region may carry a mask, a box, both, or neither, and each
case maps to a different evidence type:
```mermaid
flowchart TB
R["Region"] --> M{"region.mask_ref?"}
M -->|yes| MASK["EvidenceType.MASK<br/>artifact_ref = asset_ref or mask_ref<br/>coords = box if box else None"]
M -->|no| B{"region.box?"}
B -->|yes| BOX["EvidenceType.BOUNDING_BOX<br/>coords from box<br/>score = region.score or box.score"]
B -->|no| STAT["EvidenceType.STATISTIC<br/>no coordinates<br/>score = region.score"]
style MASK fill:#1f6feb22,stroke:#1f6feb
style BOX fill:#1f6feb22,stroke:#1f6feb
style STAT fill:#d2992222,stroke:#d29922
```
> A region carrying a mask reference becomes `MASK` evidence; one with only a box becomes
> `BOUNDING_BOX` evidence. A region with neither is reported as a `STATISTIC` **rather than a
> spatial claim, because there is no geometry to prove.** (`evidence/engine.py:496-500`)
Two details worth naming:
- The `MASK` branch's `coordinates` are `None` when `region.box is None` β a mask without a
bounding box is legal, and the validator (Β§4) only fires when coordinates are *present*.
- The `BOUNDING_BOX` branch reads `coordinate_system` from **`region.box.coordinate_system`**
(the box's own frame), not from `region.coordinate_system`. The frame of the geometry is
the frame of the geometry.
#### `evidence_from_geospatial` β may return `None`
```python
def evidence_from_geospatial(self, geo, *, source_specialist, bounds=None, score=None) -> Evidence | None:
if not geo.has_crs and not geo.crs:
return None
return Evidence(
type=EvidenceType.GEOLOCATION,
source_specialist=source_specialist,
coordinate_system=CoordinateSystem.GEO if bounds else None,
coordinates=list(bounds) if bounds else None,
score=score,
payload={
"crs": geo.crs,
"bounds": list(geo.bounds) if geo.bounds else None,
"width": geo.width,
"height": geo.height,
"is_georeferenced": geo.is_georeferenced,
},
)
```
`evidence/engine.py:542-572`.
> Returns `None` when there is no CRS to report. Emitting a `GEOLOCATION` item without a CRS
> would assert a placement the data does not support, which is the same failure mode the
> grounding degenerate-box guard exists to prevent.
This is the honesty rule applied to geometry: **a claim that cannot be supported is not
emitted, rather than emitted with a placeholder.** Note the guard is `not geo.has_crs and not
geo.crs` β either signal is sufficient, so a raster that reports a CRS string without the
`has_crs` flag still produces evidence.
## 6. `confidence_for()` β calibrating through the engine's artifact
```python
def confidence_for(
self,
result: SpecialistResult | ConfidenceBreakdown,
*,
extra_components: dict[str, float] | None = None,
degraded: bool | None = None,
degradation_reason: str | None = None,
) -> ConfidenceBreakdown:
```
`evidence/engine.py:576-637`. It accepts either form so *"a caller that already has the pieces
does not rebuild a `SpecialistResult` to use it."*
**The degradation rule β the specialist's verdict WINS.**
> The specialist's own degradation verdict WINS unless explicitly overridden: it knows things
> the engine does not (a zero-shot fallback, a failed input-quality gate), and overwriting
> `degraded=False` here would **launder a degraded result into a confident-looking one.**
> (`evidence/engine.py:589-592`)
| Situation | Result |
|---|---|
| `degraded=None`, specialist said `degraded=True` | stays `True` |
| `degraded=None`, specialist said `degraded=False` | stays `False` |
| `degraded=True` explicitly | `True`, and if the specialist had not flagged it, the reason is sourced from `result.warnings[0]` β *"so the reason is a fact rather than a restatement of the boolean"* (`:618-622`) |
| `degraded=True` and no reason and no warnings | falls back to the literal string `"aggregate degraded"` (`:628-629`) |
| `degradation_reason` explicitly passed | that string is used verbatim |
**`calibrate_result()`** (`evidence/engine.py:639-643`) is the narrower entry point:
```python
def calibrate_result(self, result: SpecialistResult) -> ConfidenceBreakdown:
"""Calibrate a result's existing breakdown, preserving its provenance."""
return calibrate_result(result.confidence, self.calibration)
```
**`from_config()`** (`evidence/engine.py:647-665`):
```python
@classmethod
def from_config(cls, config, *, calibration=None, base_dir=None) -> "EvidenceEngine":
return cls(
max_items=int(config.get("evidence.max_items", DEFAULT_MAX_ITEMS)),
calibration=calibration,
)
```
> `calibration` is explicit rather than auto-loaded: whether a fitted artifact exists is an
> operational fact the caller may know better than the config file does, and silently loading
> one would make the engine's behaviour depend on filesystem state.
The `base_dir` parameter is accepted for signature symmetry with
`evidence.confidence.load_calibration` but **is not used** by this classmethod β the engine
does not resolve the artifact itself.
**Module-level shorthand** (`evidence/engine.py:668-677`):
```python
def aggregate_evidence(results, *, max_items=DEFAULT_MAX_ITEMS, calibration=None) -> EvidenceCollection:
return EvidenceEngine(max_items=max_items, calibration=calibration).aggregate(results)
```
---
# Part C β The confidence system
## 7. The honesty rule
`evidence/confidence.py:12-30` is the reason the module exists at all. It is worth quoting in
full because it is the single most important design statement in this subsystem:
> **THE HONESTY RULE (the reason this module exists at all)**
>
> A calibration that claims to be fitted when it is not is a **FALSE CLAIM OF RELIABILITY** β
> strictly worse than reporting nothing, because a downstream consumer will trust it. So:
>
> no fitted artifact -> pass the raw score through unchanged,
> set `method="uncalibrated"`,
> leave `calibrated=None`
>
> `calibrated=None` is what makes it honest: `ConfidenceBreakdown.value` then returns `raw`,
> and any consumer that wants to distinguish "we calibrated this" from "we did not" can read
> `method` or `calibrated is None`. **We never fill in a plausible-looking number to make a
> schema field look complete.**
### 7.1 The three states
`docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.1 tabulates the decision:
| State | `calibrated` | `method` | `value` |
|---|---|---|---|
| no fitted artifact | `None` | `"uncalibrated"` | falls back to `raw` |
| fitted artifact, `T β 1` | mapped value | `"temperature_scaling"` | `calibrated` |
| fitted artifact, `T = 1` | `None` | `"uncalibrated"` | falls back to `raw` |
**The method strings are module constants** (`evidence/confidence.py:73-78`):
```python
METHOD_TEMPERATURE = "temperature_scaling"
METHOD_UNCALIBRATED = "uncalibrated"
```
> Method string reported whenever no mapping was applied. Matches the literal the existing
> specialists already emit, so trace consumers see one vocabulary.
That last clause is a real constraint: `VLMSpecialist._confidence_for` and
`GroundingSpecialist._confidence_for` already emit `method="uncalibrated", calibrated=None`
independently. The engine generalises that behaviour rather than introducing a second
vocabulary (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.1).
### 7.2 `ConfidenceBreakdown` β the schema
```python
class ConfidenceBreakdown(BaseModel):
"""Never an LLM utterance. Always measurable signals (plan section 26)."""
model_config = ConfigDict(extra="forbid")
raw: float = Field(ge=0.0, le=1.0)
calibrated: float | None = Field(default=None, ge=0.0, le=1.0)
method: str = "uncalibrated"
components: dict[str, float] = Field(default_factory=dict)
degraded: bool = False
degradation_reason: str | None = None
@property
def value(self) -> float:
return self.calibrated if self.calibrated is not None else self.raw
```
`core/schemas.py:259-273`.
| Field | Type | Note |
|---|---|---|
| `raw` | `float`, `[0,1]` | the specialist's hand-weighted score β *not* a probability |
| `calibrated` | `float \| None`, `[0,1]` | `None` means **no mapping was applied** β the honesty signal |
| `method` | `str`, default `"uncalibrated"` | one of the two constants |
| `components` | `dict[str, float]` | diagnostic signals; `docs/API_CONTRACT.md:682` β *"**Diagnostic only** β do not compute a confidence from it"* |
| `degraded` | `bool` | *"Whether this confidence should be trusted"* |
| `degradation_reason` | `str \| None` | why, when `degraded` |
**`value` is a Python `@property` and is NOT serialised.** This is the trap
`docs/API_CONTRACT.md:571` warns about:
> `result.confidence.value` | **NOT a JSON field.** It is a Python `@property` on
> `ConfidenceBreakdown` and is **not serialised** (verified: `model_dump()` yields only
> `calibrated, components, degradation_reason, degraded, method, raw`). To get the number the
> user should see, read `calibrated` if it is non-null, otherwise `raw`.
`docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§D verifies this against the real model: *"`confidence.value`
**is not serialised** β the six-key set matches the contract's parsed claim exactly."*
### 7.3 The four rules the frontend must follow
`docs/API_CONTRACT.md:686-694`:
1. Display `calibrated` when it is not `null`; otherwise display `raw`.
2. Display `method` next to the value. `temperature_scaling` means a fitted correction was
applied; `uncalibrated` means it was not.
3. **Never present a confidence as a percentage without its method.** A raw 0.99 and a
calibrated 0.99 do not mean the same thing.
4. When `degraded` is `true`, show `degradation_reason`. Confidence that is degraded is not a
quality signal.
## 8. Temperature scaling β the derivation
### 8.1 Why the transform is needed at all
`evidence/confidence.py:5-9`:
> Specialists produce a *raw* reliability score from signals they can actually point at
> (`max_objectness`, input-quality autocorrelation, `used_head`, β¦). That raw score is **not a
> probability: it is a hand-weighted sum, and a hand-weighted sum is not calibrated by
> construction.** This module is the one place that maps raw -> calibrated.
The signals named in `docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.1 are `max_objectness`,
`top_mean_objectness`, `score_contrast`, `used_head`, and input-quality autocorrelation. The
plan (Β§26) names per-task signals too: for grounding, *"box confidence + text/image similarity
+ augmentation consistency"*; for change, *"mean pixel probability + component stability +
registration quality"*; for optical-SAR, *"fusion classifier margin + optical confidence + SAR
confidence + cross-modal agreement"*.
### 8.2 The formula
```
logit(z) = log(z / (1 - z))
calibrated = sigmoid(logit(z) / T)
```
`evidence/confidence.py:39-47`:
> The standard formulation maps a logit through `sigmoid(logit / T)`. Our raw scores are
> already in (0, 1), so they are read as probabilities and converted to log-odds first.
```python
def apply(self, raw: float) -> float:
"""Map a raw score in [0, 1] through the fitted temperature."""
return _clamp01(_sigmoid(_logit(float(raw)) / self.temperature))
```
`evidence/confidence.py:165-167`. Three operations in order: `_logit` β divide by `T` β
`_sigmoid` β `_clamp01`.
| `T` | Effect | Verified |
|---|---|---|
| `T > 1` | **softens** β pulls scores toward the middle | `docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.3 |
| `T < 1` | **sharpens** β pushes scores away from the middle | `docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.3; measured: `calibrate(0.87, T=0.9772732)` β `0.8749186077809417` |
| `T = 1` | **identity** | handled as uncalibrated, Β§8.3 |
**The transform is monotonic**, so calibration respaces scores without reordering them β
pinned by `test_temperature_scaling_is_monotonic` (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.3).
### 8.3 `_is_effective` β why a "fitted" `T` of exactly 1.0 is reported as uncalibrated
```python
#: How close to 1.0 a fitted temperature must land before it is treated as
#: "does nothing". Bit-for-bit 1.0 is the honest test; this tolerance absorbs
#: float round-trip through JSON without accepting a real transform.
_IDENTITY_TOLERANCE = 1e-6
def _is_effective(temperature: float) -> bool:
"""Does this temperature actually transform anything?
A temperature within `_IDENTITY_TOLERANCE` of 1.0 is the identity map. It is
treated as "not fitted" so the honest report wins: an artifact that carries
T=1.0 has learned nothing, and reporting `method="temperature_scaling"` for
it would assert a correction that was never made.
"""
return abs(temperature - 1.0) > _IDENTITY_TOLERANCE
```
`evidence/confidence.py:90-93`, `:118-126`.
**Verified empirically** (read-only check against the shipped module):
| Input | `_is_effective` | Result |
|---|---|---|
| `0.9772731820958189` (the fitted T) | `True` | a real transform is applied |
| `1.0` | `False` | reported as uncalibrated |
**Why the tolerance and not `!= 1.0`.** The comment gives the reason: *"Bit-for-bit 1.0 is the
honest test; this tolerance absorbs float round-trip through JSON without accepting a real
transform."* A JSON round-trip of `1.0` can land at `1.0000000000000002`; an exact `!= 1.0`
test would then accept a transform that does nothing. `1e-6` is wide enough to absorb the
round-trip and far too narrow to swallow any real fitted temperature.
**The identity branch records why** (`evidence/confidence.py:299-303`):
```python
if calibration is not None:
# A real artifact was supplied but it is the identity map. Say so
# rather than dropping the information on the floor.
components["calibration_identity"] = 1.0
components.update(calibration.components)
```
Measured components for `T = 1.0`:
`{'calibrated_applied': 0.0, 'calibration_identity': 1.0, 'temperature': 1.0}`.
`docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.2 frames this as the same failure mode as Β§3.1: *"This is
a small case the brief did not name explicitly; it is the same false-claim failure mode as
Β§3.1 and is handled by the same principle."*
### 8.4 `_EPS` β endpoint clamping so a hard 0.0 / 1.0 cannot become `NaN`
```python
#: Scores are clamped to this closed interval before the log-odds transform.
#: `logit(0)` and `logit(1)` are infinite; clamping at the boundary keeps the
#: mapping finite and keeps a hard 0.0 / 1.0 from becoming a NaN.
_EPS = 1e-6
def _logit(p: float) -> float:
"""Log-odds of a probability, with the endpoints clamped to stay finite."""
clamped = min(1.0 - _EPS, max(_EPS, p))
return math.log(clamped / (1.0 - clamped))
```
`evidence/confidence.py:80-83`, `:108-111`.
| Input `z` | Unclamped `logit(z)` | With `_EPS` clamp | Result |
|---|---|---|---|
| `0.0` | `log(0)` = `-inf` | `logit(1e-6)` β `-13.8155` | finite |
| `1.0` | `log(1/0)` = `+inf` | `logit(1 - 1e-6)` β `+13.8155` | finite |
| `0.5` | `0.0` | `0.0` | unchanged |
Pinned by `test_endpoint_scores_do_not_produce_nan` (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.3).
### 8.5 `_sigmoid` β numerically stable, and why the branch exists
```python
def _sigmoid(x: float) -> float:
"""Numerically stable logistic function.
The naive form overflows for large negative `x`; the branch below keeps
`exp`'s argument non-positive so the result is finite for every input.
"""
if x >= 0.0:
return 1.0 / (1.0 + math.exp(-x))
z = math.exp(x)
return z / (1.0 + z)
```
`evidence/confidence.py:96-105`. For `x >= 0` the `exp` argument is `-x <= 0`; for `x < 0` the
`exp` argument is `x < 0`. **In both branches `math.exp` receives a non-positive argument**, so
it can never overflow. The two forms are algebraically identical
(`1/(1+e^-x) β‘ e^x/(1+e^x)`) but only one of them is safe on each side of zero.
```python
def _clamp01(x: float) -> float:
return min(1.0, max(0.0, x))
```
`evidence/confidence.py:114-115`. The final guard: `sigmoid` returns `(0, 1)` mathematically,
but float rounding can produce exactly `0.0` or `1.0`, and `_clamp01` keeps the value inside
the schema's `ge=0.0, le=1.0` bound so Pydantic never rejects an output of this module.
### 8.6 `TemperatureCalibration` β the artifact object
```python
@dataclass(frozen=True)
class TemperatureCalibration:
temperature: float
fitted_on: str | None = None
artifact: str | None = None
n_samples: int | None = None
```
`evidence/confidence.py:129-150`.
| Field | Type | Meaning (docstring) |
|---|---|---|
| `temperature` | `float` | *"the fitted scalar. `T > 1` softens β¦, `T < 1` sharpens. Validated on construction."* |
| `fitted_on` | `str \| None` | *"free-text provenance, e.g. `"valid"` or a split hash. Carried into the breakdown components so a result can be traced back to the artifact that shaped it."* |
| `artifact` | `str \| None` | *"path/identifier of the source file, for the same reason."* |
| `n_samples` | `int \| None` | *"how many validation samples the fit used, when known."* |
**Range validation** (`evidence/confidence.py:85-88`, `:152-159`):
```python
_MIN_TEMPERATURE = 1e-3
_MAX_TEMPERATURE = 1e3
def __post_init__(self) -> None:
t = float(self.temperature)
if not math.isfinite(t) or not _MIN_TEMPERATURE <= t <= _MAX_TEMPERATURE:
raise ValueError(
f"temperature must be finite and in "
f"[{_MIN_TEMPERATURE}, {_MAX_TEMPERATURE}], got {self.temperature!r}"
)
object.__setattr__(self, "temperature", t)
```
> A temperature of zero or below is not a calibration; it is a division error.
> (`evidence/confidence.py:85-88`)
The check is `math.isfinite(t)` **and** the range β so `NaN` and `Β±inf` are both rejected, and
`object.__setattr__` is used because the dataclass is `frozen=True`.
**Two methods:**
| Method | Returns | Purpose |
|---|---|---|
| `is_effective()` | `bool` | `_is_effective(self.temperature)` |
| `apply(raw)` | `float` | `_clamp01(_sigmoid(_logit(raw) / T))` |
| `components` *(property)* | `dict[str, float]` | `{"temperature": T}` plus `{"calibration_samples": n}` when known |
| `describe()` | `dict[str, Any]` | `{method, temperature, effective, fitted_on, artifact, n_samples}` |
**`from_dict()` β accepts two shapes** (`evidence/confidence.py:187-219`):
```python
body: Any = payload
if isinstance(payload.get("temperature_scaling"), dict):
body = payload["temperature_scaling"]
elif isinstance(payload.get("calibration"), dict):
body = payload["calibration"]
if not isinstance(body, dict) or "temperature" not in body:
raise ValueError(
"calibration artifact has no 'temperature' field "
f"(source: {source or '<dict>'})"
)
```
> Accepts either the flat form (`{"temperature": 1.4, ...}`) or a nested
> `{"temperature_scaling": {...}}` form, because the artifact format is owned by
> `training/calibration/` and this loader should not be the thing that decides it.
`fitted_on` reads `body.get("fitted_on") or body.get("split")`; `artifact` prefers the
`source` path over `body.get("artifact")`; `n_samples` is coerced to `int` when present.
**The shipped artifact uses the nested form** β `artifacts/calibration_v001.json` has a
top-level `"temperature_scaling": {"fitted_on": "Val", "n_samples": 16441, "temperature":
0.9772731820958189}` block. This is exactly why `from_dict` accepts it.
**`from_json()` β `required` semantics** (`evidence/confidence.py:221-252`):
| Condition | `required=False` (default) | `required=True` |
|---|---|---|
| file absent | returns `None` | raises `FileNotFoundError` |
| file present, valid JSON | returns the object | returns the object |
| file present, **malformed JSON** | **raises `ValueError`** | **raises `ValueError`** |
> `ValueError`: the file exists and is malformed (**never swallowed**: a corrupt artifact
> silently treated as "no artifact" would hide an operational defect).
This is the same asymmetry as the evidence cap and the config validation: **absent is a
legitimate state; corrupt is a defect and must be loud.**
### 8.7 `calibrate()` β the single entry point
```python
def calibrate(
raw: float,
calibration: TemperatureCalibration | None,
*,
extra_components: dict[str, float] | None = None,
degraded: bool = False,
degradation_reason: str | None = None,
) -> ConfidenceBreakdown:
```
`evidence/confidence.py:255-323`. **This is the single entry point used by
`EvidenceEngine.confidence_for`.**
**Non-finite input RAISES** (`evidence/confidence.py:289-292`):
```python
value = float(raw)
if not math.isfinite(value):
raise ValueError(f"raw confidence must be finite, got {raw!r}")
clamped = _clamp01(value)
```
> Refusing is deliberate: a non-finite score is a bug upstream, and clamping it to 0.5 would
> **invent a measurement.** (`evidence/confidence.py:284-287`)
Note the asymmetry with `_logit`'s endpoint clamping: an **infinite** input is refused, but a
**finite** `0.0` or `1.0` is accepted and clamped *inside* the transform. The distinction is
that `0.0` is a real measurement (a specialist is certain it found nothing) while `NaN` is
not a measurement at all.
**The honest path** (`evidence/confidence.py:296-311`):
```python
if calibration is None or not calibration.is_effective():
components["calibrated_applied"] = 0.0
if calibration is not None:
components["calibration_identity"] = 1.0
components.update(calibration.components)
return ConfidenceBreakdown(
raw=clamped,
calibrated=None,
method=METHOD_UNCALIBRATED,
components=components,
degraded=degraded,
degradation_reason=degradation_reason,
)
```
**The fitted path** (`evidence/confidence.py:313-323`):
```python
components["calibrated_applied"] = 1.0
components.update(calibration.components)
return ConfidenceBreakdown(
raw=clamped,
calibrated=calibration.apply(clamped),
method=METHOD_TEMPERATURE,
components=components,
degraded=degraded,
degradation_reason=degradation_reason,
)
```
**Measured component blocks** (read-only verification against the shipped module):
| Case | `components` |
|---|---|
| `calibrate(0.87, T=0.9772731820958189)` | `{'calibrated_applied': 1.0, 'temperature': 0.9772731820958189, 'calibration_samples': 16441.0}` |
| `calibrate(0.87, T=1.0)` | `{'calibrated_applied': 0.0, 'calibration_identity': 1.0, 'temperature': 1.0}` |
| `calibrate(0.87, None)` | `{'calibrated_applied': 0.0}` |
The `calibrated_applied` flag is a **`float`, not a `bool`**, because `components` is typed
`dict[str, float]`. `0.0` / `1.0` is the encoding. `raw` is retained in every case β the
breakdown always carries both the input and the output, so nothing is destroyed by
calibrating.
### 8.8 `calibrate_result()` β preserve the specialist's judgement
```python
def calibrate_result(breakdown, calibration) -> ConfidenceBreakdown:
return calibrate(
breakdown.raw,
calibration,
extra_components=dict(breakdown.components),
degraded=breakdown.degraded,
degradation_reason=breakdown.degradation_reason,
)
```
`evidence/confidence.py:326-346`.
> The specialist's own components and degradation reason are carried through untouched: **this
> function adds the calibration, it does not re-judge the specialist's measurement.**
Pinned by `test_specialist_degradation_survives_calibration`
(`docs/PHASE13_EVIDENCE_ENGINE.md` Β§3.5): *"A `degraded=True` fallback result stays degraded
after calibration, with its `degradation_reason` intact."*
## 9. `load_calibration()` β the ordered candidate search
### 9.1 The frozen keys it reads
```yaml
confidence:
temperature_scaling: true
calibration_file: calibration_v001.json
```
`configs/base.yaml:231-233`.
```python
if not bool(config.get("confidence.temperature_scaling", False)):
return None
filename = config.get("confidence.calibration_file")
if not filename:
return None
```
`evidence/confidence.py:396-401`. It returns `None` β **never a fabricated default** β when:
| Condition | Returns |
|---|---|
| the master switch is off | `None` |
| no filename is configured | `None` |
| the artifact is absent from every candidate | `None` |
| the artifact is present but malformed | **raises `ValueError`** |
> Every one of those is a legitimate deployment state (an HF Space may ship without the fitted
> artifact), and each degrades to the pass-through. (`evidence/confidence.py:361-364`)
### 9.2 The ordered search
```python
name = str(filename)
candidates: list[Path] = []
if base_dir is not None:
# An explicit base_dir is honoured first, but a REPO-anchored candidate
# is still tried: several call sites in this repo pass "." meaning "the
# repo", which only works when CWD happens to be the repo root.
candidates.append(Path(base_dir) / name)
from core.config import REPO_ROOT # local import: keeps this module cheap
if base_dir is not None:
candidates.append(Path(REPO_ROOT) / name)
candidates.append(Path(REPO_ROOT) / "artifacts" / name)
candidates.append(Path(REPO_ROOT) / "configs" / name)
for candidate in candidates:
if candidate.exists():
return TemperatureCalibration.from_json(candidate, required=True)
# Nothing found: the master switch is on but no artifact ships. Degrade.
return None
```
`evidence/confidence.py:403-424`.
```mermaid
flowchart TB
S["load_calibration(config, base_dir=?)"] --> SW{"confidence.temperature_scaling?"}
SW -->|false| N1["return None"]
SW -->|true| FN{"calibration_file set?"}
FN -->|no| N2["return None"]
FN -->|yes| C1{"base_dir given?"}
C1 -->|yes| A["1. base_dir / name"]
A --> A2{"exists?"}
C1 -->|no| B
A2 -->|yes| LOAD["from_json(required=True)"]
A2 -->|no| B["2. REPO_ROOT / name<br/><i>only when base_dir was given</i>"]
B --> B2{"exists?"}
B2 -->|yes| LOAD
B2 -->|no| C["3. REPO_ROOT / artifacts / name<br/><b>the real home</b>"]
C --> C2{"exists?"}
C2 -->|yes| LOAD
C2 -->|no| D["4. REPO_ROOT / configs / name"]
D --> D2{"exists?"}
D2 -->|yes| LOAD
D2 -->|no| N3["return None β honest degradation"]
style C fill:#3fb95022,stroke:#3fb950
style LOAD fill:#1f6feb22,stroke:#1f6feb
```
**The exact candidate list, by input:**
| `base_dir` | candidate 1 | candidate 2 | candidate 3 | candidate 4 |
|---|---|---|---|---|
| `None` | β | β | `REPO_ROOT/artifacts/name` | `REPO_ROOT/configs/name` |
| `"."` | `./name` | `REPO_ROOT/name` | `REPO_ROOT/artifacts/name` | `REPO_ROOT/configs/name` |
| `"configs"` | `configs/name` | `REPO_ROOT/name` | `REPO_ROOT/artifacts/name` | `REPO_ROOT/configs/name` |
| `"artifacts/calibration/confidence_calibration_v001"` | that dir / name | `REPO_ROOT/name` | `REPO_ROOT/artifacts/name` | `REPO_ROOT/configs/name` |
> The search is ordered and deterministic: **the first existing candidate wins**, and no
> candidate existing means `None` (honest degradation, not an error).
> (`evidence/confidence.py:388-389`)
### 9.3 The silent-failure shape this fixes
`evidence/confidence.py:366-386` documents the defect precisely:
> `base_dir` is resolved in this order (added 2026-09-22):
>
> 1. an explicit `base_dir` argument, when the caller passes one;
> 2. `configs/` under the repository root, **but only if the artifact is actually there** β it
> never is by default, so this branch exists only to keep an existing caller that parks the
> artifact beside the config working;
> 3. `<repo root>/artifacts/` β **the artifact's real home**, and the location
> `scripts/fit_calibration.py` writes to.
>
> Step 3 is the important one and **it is why this function changed.** The frozen config names
> a *bare filename* (`calibration_v001.json`), while the docstring examples in this
> repository's own docs pass `base_dir="."` AND `base_dir="configs"` β two different
> directories, neither of which is `artifacts/`. A deployment following either example would
> **silently resolve to a nonexistent path, degrade to `method="uncalibrated"`, and report a
> fitted artifact as never having been fitted.** That is precisely the silent-failure shape
> this project's provenance discipline exists to prevent, so the search now anchors to the
> repository root using the same `REPO_ROOT` that `app/serving.py` uses for the change
> checkpoint.
**The failure shape, stated as a table:**
| Step | Pre-fix behaviour | Post-fix behaviour |
|---|---|---|
| caller passes `base_dir="."` from a non-repo CWD | resolves to `$CWD/calibration_v001.json` β absent β `None` β `uncalibrated` | candidate 1 fails, candidate 3 succeeds β the artifact loads |
| caller passes `base_dir="configs"` | resolves to `configs/calibration_v001.json` β absent β `None` β `uncalibrated` | candidate 1 fails, candidate 3 succeeds β the artifact loads |
| caller passes nothing | resolved to `$CWD/...` (`base_dir` defaulted to `Path.cwd()` pre-fix) | candidate 3 succeeds β the artifact loads |
| the artifact is genuinely absent | `None` | `None` (unchanged) |
| the artifact is present but corrupt | depends on the path resolved | **raises** β never swallowed |
The critical property is the *symptom*: the pre-fix failure was **silent and
indistinguishable from "no artifact was ever fitted"**. A fitted model reported itself as
unfitted. That is the same class of defect as a fabricated confidence number β it makes the
system lie about its own state.
**Verified against the shipped repo** (read-only execution):
| Call | Resolved artifact |
|---|---|
| `load_calibration(config)` | `β¦/satquery-ai/artifacts/calibration_v001.json` |
| `load_calibration(config, base_dir="configs")` | `β¦/satquery-ai/artifacts/calibration_v001.json` |
| `load_calibration(config, base_dir=".")` | `β¦/satquery-ai/artifacts/calibration_v001.json` |
All three resolve to the **same** file via candidate 3. The artifact's own
`consumer_contract.resolution` field recommends `load_calibration(config, base_dir='configs')`
β which now works, and works for the reason the fix was made.
**The older record, and its status.** `docs/PHASE13_EVIDENCE_ENGINE.md` Β§4.2 was written
before the fix and describes the pre-fix behaviour (`base_dir` defaults to `Path.cwd()`;
*"the controller must pass `base_dir` explicitly"*; *"Current behaviour on the real config, for
the record: `load_calibration` returns `None`, because no artifact has been fitted yet"*).
Two things changed since: the search was repo-anchored (2026-09-22), and
`artifacts/calibration_v001.json` now exists. The current measured behaviour is the table
above. The PHASE13 Β§4.1 expectation of
`artifacts/calibration/confidence_calibration_v001/calibration_v001.json` did **not** hold β
the artifact lives at `artifacts/calibration_v001.json` β which is fine because the bare
filename plus the `artifacts/` candidate resolves it.
## 10. The MEASURED calibration result β a negative result, reported as one
### 10.1 The artifact
`artifacts/calibration_v001.json` is the fitted artifact. Its top-level keys:
`consumer_contract`, `created_utc`, `fit_diagnostics`, `metrics`, `provenance`,
`reliability_diagram`, `schema`, `scope`, `temperature_scaling`, `type_mask_applied`.
**`fit_diagnostics`:**
| Field | Value |
|---|---|
| `temperature` | `0.9772731820958189` |
| `log_temperature` | `-0.022989052824434128` |
| `iterations` | `200` |
| `n_samples` | `16441` |
| `nll_before` | `0.6897411093435998` |
| `nll_after` | `0.689630751845387` |
| `nll_improvement` | `0.0001103574982127542` |
| `effective` | `true` |
| `hit_bound` | `false` |
Note `log_temperature = log(0.9772731820958189) = -0.022989β¦` β the fitter optimised in
log-space, and `hit_bound: false` means the optimum was interior, not at a clamp.
**`metrics`:**
| Field | Value |
|---|---|
| `ece_before` | `0.013755` |
| `ece_after` | `0.014929` |
| **`ece_improvement`** | **`-0.001174` (WORSE)** |
| `nll_before` | `0.689741` |
| `nll_after` | `0.689631` |
| `nll_improvement` | `0.00011` |
| `n_bins` | `15` |
| `n_classes` | `19` |
| `n_samples` | `16441` |
**`provenance`:**
| Field | Value |
|---|---|
| `checkpoint_path` | `artifacts/change_vqa/run/head.pt` |
| `checkpoint_sha256` | `cfae5e43b97ca930f568dc5b8ae4f36b24e9ff717af226159802206ffd63a82a` |
| `config_hash` | `78f1e3700da15aa1` |
| `dataset_id` | `cdvqa` |
| `feature_spec` | `change_feat_v1` |
| `fitted_on` | `Val` |
| `held_out_splits_excluded` | `["Test", "Test2"]` |
| `method` | `temperature_scaling` |
| `objective` | `mean_negative_log_likelihood` |
| `optimizer` | `golden_section_on_log_temperature` |
| `space` | `multiclass_logits` |
| `iterations` | `200` |
| `n_classes` | `19` |
| `n_samples` | `16441` |
**`scope` β what this artifact calibrates, and what it does not:**
> **This temperature calibrates the R-02 change-VQA head's answer confidence. Other
> specialists emit their own raw scores and are unaffected.**
**`type_mask_applied`:** `false`.
### 10.2 The result, stated exactly
> The raw softmax was **already near-calibrated** (ECE 0.013755) and temperature scaling made
> ECE very slightly **worse** (0.014929) while improving NLL marginally. This is a
> measurement, not a quality judgment, and **it must not be described as scaling being "more
> accurate".** β `docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§7
`docs/FINAL_DELIVERY_TODO.md:104` records it in the measured-metrics table as
`0.013755 β 0.014929 (worse)` with status **"measured (not an improvement)"**, sourced to
`artifacts/calibration_v001.json:26-35`. `docs/PHASE19_FINAL_HARDENING.md:377` records the
same: *"**Not claimed as an improvement.** β¦ Nothing is calibrated in a deployed path"*.
| Metric | Before | After | Ξ | Reading |
|---|---|---|---|---|
| ECE (15 bins) | 0.013755 | 0.014929 | **β0.001174** | **WORSE** |
| NLL | 0.689741 | 0.689631 | +0.000110 | marginally better |
**Why both numbers matter, and neither alone.** The fitter's objective was
`mean_negative_log_likelihood`, and NLL *did* improve β that is why `T = 0.9773β¦` was
selected at all. But the metric a consumer reads as "is this probability trustworthy" is ECE,
and **ECE got worse**. Reporting only NLL would present a fitted artifact as a success;
reporting only ECE would hide why the value was chosen. Both are reported, and the headline is
the ECE regression.
**Why the reliability diagram is not a single score.** The artifact's own note:
> Equal-width bins over predicted-class confidence. **ECE is bin-count sensitive and is not an
> aggregate score.** β `artifacts/calibration_v001.json`, `reliability_diagram.note`
The diagram is a 15-bin list. The first two bins (`[0.0, 0.0667)` and `[0.0667, 0.1333)`)
have `count: 0` and `null` accuracy/confidence/gap β the model never predicts that low. The
remaining thirteen bins are populated; the largest is `[0.9333, 1.0]` with `count: 2814`,
`accuracy: 0.969794`, `confidence: 0.966895`, `gap: 0.002899`. The largest |gap| is
`-0.027817` at `[0.6667, 0.7333)` (`count: 1494`).
| Bin `[lo, hi)` | count | accuracy | confidence | gap |
|---|---|---|---|---|
| 0.0000β0.0667 | 0 | β | β | β |
| 0.0667β0.1333 | 0 | β | β | β |
| 0.1333β0.2000 | 8 | 0.25 | 0.188409 | +0.061591 |
| 0.2000β0.2667 | 437 | 0.283753 | 0.241931 | +0.041822 |
| 0.2667β0.3333 | 736 | 0.290761 | 0.299846 | β0.009085 |
| 0.3333β0.4000 | 626 | 0.386581 | 0.367128 | +0.019454 |
| 0.4000β0.4667 | 690 | 0.450725 | 0.434665 | +0.016060 |
| 0.4667β0.5333 | 1380 | 0.534783 | 0.504472 | +0.030310 |
| 0.5333β0.6000 | 1521 | 0.558185 | 0.565780 | β0.007595 |
| 0.6000β0.6667 | 1395 | 0.624373 | 0.633333 | β0.008961 |
| 0.6667β0.7333 | 1494 | 0.672691 | 0.700507 | **β0.027817** |
| 0.7333β0.8000 | 1611 | 0.742396 | 0.766718 | β0.024322 |
| 0.8000β0.8667 | 1701 | 0.833039 | 0.834257 | β0.001218 |
| 0.8667β0.9333 | 2028 | 0.892998 | 0.903143 | β0.010145 |
| 0.9333β1.0000 | 2814 | 0.969794 | 0.966895 | +0.002899 |
*(values verbatim from `artifacts/calibration_v001.json`; the `gap` column is the artifact's
own `accuracy β confidence`)*
### 10.3 A structural note on the transform, recorded rather than glossed
The artifact's `consumer_contract` states:
```json
"applied_as": "sigmoid(logit(z) / T) for a scalar z; softmax(logits / T) for a distribution",
"class": "TemperatureCalibration",
"module": "evidence.confidence",
"read_keys": ["temperature", "fitted_on|split", "artifact", "n_samples"],
"resolution": "load_calibration(config, base_dir='configs')"
```
The artifact's provenance records `space: "multiclass_logits"` and `n_classes: 19` β the fit
happened over a 19-class softmax. The consumer class
(`evidence.confidence.TemperatureCalibration.apply`) implements the **scalar** form:
`_clamp01(_sigmoid(_logit(raw) / T))`. The class reads exactly the four keys the artifact
declares (`temperature`, `fitted_on|split`, `artifact`, `n_samples`) and reads nothing else β
the multiclass branch is a documented capability of the *artifact format*, not a code path in
`evidence/confidence.py`. Both facts are stated here so the boundary is visible; no claim is
made about which form a deployed path exercises, because
`docs/PHASE19_FINAL_HARDENING.md:377` records that *"Nothing is calibrated in a deployed
path"*.
### 10.4 The deployment state
| Question | Answer | Source |
|---|---|---|
| Does a fitted artifact exist? | **yes** β `artifacts/calibration_v001.json` | filesystem |
| Does `load_calibration(config)` find it? | **yes** β via candidate 3 | verified read-only |
| Is `T` effective (non-identity)? | **yes** β `_is_effective(0.9772731820958189) == True` | verified read-only |
| Does it improve ECE? | **no** β 0.013755 β 0.014929 | artifact `metrics` |
| Is it wired into a served path? | `OPEN` / not claimed β *"Nothing is calibrated in a deployed path"* | `docs/PHASE19_FINAL_HARDENING.md:377` |
| Which specialist does it cover? | `change_vqa` only | artifact `scope` |
The artifact's `config_hash` is `78f1e3700da15aa1` β the same frozen hash as
`configs/base.yaml` (Β§[07](07-configuration-freeze.md)). **The artifact is keyed to the config
that produced it**, which is exactly the provenance link the config hash exists to provide.
## 11. What is deliberately absent from the confidence system
### 11.1 No LLM path
> **NO LLM PATH EXISTS HERE.** There is deliberately no function that turns text into a number.
> Confidence is a measurement (freeze section 5: "No LLM-generated confidence"). The only
> inputs accepted are a float the specialist computed and a calibration artifact fitted on
> validation data. β `evidence/confidence.py:32-37`
`docs/ARCHITECTURE_FREEZE.md` Β§5 lists it as a non-negotiable: *"No LLM-generated coordinates.
No LLM-generated confidence."* And the plan Β§26 opens with it: *"No LLM-generated confidence."*
### 11.2 No temperature fitter
> **A fitter is NOT included:** Phase 13 fits T on validation data, and inventing one here
> would be exactly the kind of unfounded number this module refuses to emit.
> `TemperatureCalibration` is the *consumer* of a fitted `T`; the artifact loader reads
> whatever `training/calibration/` produces. β `evidence/confidence.py:57-61`
The fitter lives in `training/calibration/` (`__init__.py`, `artifact.py`, `evaluate.py`,
`fitter.py`) and produced the artifact in Β§10. `evidence/confidence.py` deliberately does not
contain one β the separation is what makes "the number came from a fit" a checkable claim.
### 11.3 No retrieval path
There is no artifact-serving endpoint in v1 (F-16, Β§2.3). This is why `artifact_ref` is
permanently `null` and why the confidence module's `artifact` field carries a **path for
provenance**, not a client-facing reference. `TemperatureCalibration.artifact` is used in
`describe()` for operator diagnostics; it is not serialised into a client response.
---
# Part D β The execution trace and the eight events
## 12. `ExecutionTrace` β observable facts only
```python
class ExecutionTrace(BaseModel):
model_config = ConfigDict(extra="forbid")
run_id: str = Field(default_factory=lambda: _new_id("run"))
schema_version: str = SCHEMA_VERSION
task: Task | None = None
query: str | None = None
inputs: list[str] = Field(default_factory=list)
modalities: list[Modality] = Field(default_factory=list)
intent: Intent | None = None
validation: dict[str, Any] = Field(default_factory=dict)
workflow: list[str] = Field(default_factory=list)
steps: list[TraceStep] = Field(default_factory=list)
selected_models: list[ModelRef] = Field(default_factory=list)
parameters: dict[str, Any] = Field(default_factory=dict)
outputs: list[str] = Field(default_factory=list)
confidence: ConfidenceBreakdown | None = None
timings: dict[str, float] = Field(default_factory=dict)
fallbacks: list[str] = Field(default_factory=list)
errors: list[dict[str, Any]] = Field(default_factory=list)
contradiction: bool = False
config_hash: str | None = None
started_at: str = Field(default_factory=_utcnow)
finished_at: str | None = None
```
`core/schemas.py:296-319`. The section header above it reads:
**`# Execution trace (observable facts only β never chain-of-thought)`**
(`core/schemas.py:277`).
| Field | Type | Carries |
|---|---|---|
| `run_id` | `str` | `run_<12 hex>`, generated per run |
| `schema_version` | `str` | `SCHEMA_VERSION` = `"1.0"` |
| `task` | `Task \| None` | the routed task |
| `query` | `str \| None` | the user's question, verbatim |
| `inputs` | `list[str]` | asset references |
| `modalities` | `list[Modality]` | inferred modalities |
| `intent` | `Intent \| None` | the router's advisory output |
| `validation` | `dict[str, Any]` | input-validation facts (`input_count`, `format`, `modality` in the plan Β§27 shape) |
| `workflow` | `list[str]` | the planned step names |
| `steps` | `list[TraceStep]` | the state machine's visits |
| `selected_models` | `list[ModelRef]` | which models ran |
| `parameters` | `dict[str, Any]` | parameters used |
| `outputs` | `list[str]` | output kinds (`bbox`, `answer`, β¦) |
| `confidence` | `ConfidenceBreakdown \| None` | the confidence breakdown |
| `timings` | `dict[str, float]` | measured durations |
| `fallbacks` | `list[str]` | which fallbacks fired |
| `errors` | `list[dict[str, Any]]` | error records |
| `contradiction` | `bool` | whether a contradiction was detected |
| `config_hash` | `str \| None` | the frozen config identity β `78f1e3700da15aa1` for this revision |
| `started_at` / `finished_at` | `str` | ISO-8601 UTC |
**`TraceStep`** (`core/schemas.py:279-285`):
```python
class TraceStep(BaseModel):
model_config = ConfigDict(extra="forbid")
state: ControllerState
started_at: str = Field(default_factory=_utcnow)
duration_ms: float | None = None
detail: dict[str, Any] = Field(default_factory=dict)
```
**`ModelRef`** (`core/schemas.py:288-293`):
```python
class ModelRef(BaseModel):
model_config = ConfigDict(extra="forbid")
name: str
revision: str | None = None
role: str | None = None
```
**`ControllerState`** β the nine states (`core/schemas.py:79-88`):
| # | State | Meaning |
|---|---|---|
| 1 | `RECEIVE` | the request arrived |
| 2 | `PARSE` | the query was parsed |
| 3 | `VALIDATE` | inputs validated |
| 4 | `PLAN` | the workflow was chosen |
| 5 | `PREPROCESS` | tiling / normalisation |
| 6 | `EXECUTE` | specialists ran |
| 7 | `AGGREGATE` | evidence was aggregated |
| 8 | `VERIFY` | confidence was computed |
| 9 | `RESPOND` | the envelope was assembled |
`configs/base.yaml` Β§`agent.states` declares exactly this list, in this order β so the
config and the enum cannot drift.
### 12.1 There is no field for reasoning
`ExecutionTrace` has **no** `reasoning`, `thought`, `rationale`, or `chain_of_thought` field,
and `model_config = ConfigDict(extra="forbid")` means one cannot be smuggled in. This is a
structural guarantee, not a convention. `docs/ARCHITECTURE_FREEZE.md` Β§5:
> Every result carries an observable execution trace. **No chain-of-thought.**
The plan Β§27 says the same at the end of its trace example: *"No chain-of-thought. Only
observable execution facts."* The `summary()` methods in the evidence layer obey the same
rule β `EvidenceCollection.summary()`'s docstring: *"Observable trace facts. No chain-of-thought,
no interpretation."*
## 13. The eight execution events
The frontend's event protocol is the integration seam between the backend and the instrument
UI. It is declared in `frontend/assets/js/core.js:616-620`:
```javascript
SQ.EVENT_NAMES = [
'QUERY_RECEIVED', 'QUERY_UNDERSTOOD', 'ROUTE_SELECTED',
'SPECIALIST_STARTED', 'SPECIALIST_COMPLETED',
'EVIDENCE_GENERATED', 'CONFIDENCE_COMPUTED', 'RESULT_ASSEMBLED'
];
```
| # | Event | What it reports | Drives the stage |
|---|---|---|---|
| 1 | `QUERY_RECEIVED` | the query text, the scene seed, the AOI, the GSD, the date pair | `QUERY` |
| 2 | `QUERY_UNDERSTOOD` | the routed task, the parsed slots, the policy id | `UNDERSTAND` |
| 3 | `ROUTE_SELECTED` | the task, the specialist list, the rule trace, the policy version | `ROUTE` |
| 4 | `SPECIALIST_STARTED` | which component, its index/ordinal, its stage, its model, its device | `ANALYZE` / `GROUND` |
| 5 | `SPECIALIST_COMPLETED` | which component, its duration, its output kind | (advances the active marker) |
| 6 | `EVIDENCE_GENERATED` | the evidence regions, the count, the threshold, the registration RMSE | `EVIDENCE` |
| 7 | `CONFIDENCE_COMPUTED` | the reported/calibrated values, the method, the reliability bins | `CONFIDENCE` |
| 8 | `RESULT_ASSEMBLED` | the answer text, the task, the specialists, the evidence ids, the confidence, the provenance | `ANSWER` |
**The stage rail** is a separate, coarser list (`frontend/assets/js/core.js:605-614`):
```javascript
SQ.STAGES = [
{ id: 'QUERY', label: 'QUERY' },
{ id: 'UNDERSTAND', label: 'UNDERSTAND' },
{ id: 'ROUTE', label: 'ROUTE' },
{ id: 'ANALYZE', label: 'ANALYZE' },
{ id: 'GROUND', label: 'GROUND' },
{ id: 'EVIDENCE', label: 'EVIDENCE' },
{ id: 'CONFIDENCE', label: 'CONFIDENCE' },
{ id: 'ANSWER', label: 'ANSWER' }
];
```
Eight stages for eight events β but **not a 1:1 mapping**: events 4 and 5 share the
`ANALYZE`/`GROUND` stages, because a specialist's start and completion are two events about
one stage visit.
### 13.1 The run engine is deliberately dumb
`frontend/assets/js/core.js:8-11`:
> The run engine is deliberately dumb: it renders whatever events it receives. **Nothing about
> the visuals depends on the events being synthetic.** Swapping the mock driver for a
> websocket / SSE feed of the same event names is the entire integration surface.
```javascript
/**
* THE INTEGRATION SEAM.
* Feed real execution events here β same names, same payload shapes β and
* every visual state in the prototype updates identically.
*/
ingest: function (type, payload) { ... }
```
`frontend/assets/js/core.js:737-742`. `ingest()` is a `switch` over the eight names, each
case updating `state` and calling `emit()`. `emit()` records
`{type, t, seq, payload}` where `t = performance.now() - t0` and `seq` is a monotonic counter
(`frontend/assets/js/core.js:709-718`).
**Two drivers exist, and they are not interchangeable:**
| Driver | When | Payload honesty |
|---|---|---|
| `startMock(query, scene)` | **preview only** β when no file is selected | synthetic; marked `is-mock` |
| `runLive(query)` in `frontend/assets/js/mission.js` | the production path | payloads come from the real network response |
`docs/FINAL_DELIVERY_TODO.md:73` records the distinction: *"`runMock` only when no file is
selected; emits empty payloads, marked `is-mock`; not in production path."* The live driver
emits all eight events around real network calls (`docs/FINAL_DELIVERY_TODO.md:71`):
*"`runLive` emits 8 events around real network calls; `markState` uses real values."*
### 13.2 `markState` β the trace is driven by events, never by a timer
`frontend/assets/js/mission.js:404-425`:
```javascript
function markState(id, note, isLive) {
var n = traceNodes[id];
if (!n) return;
n.node.classList.toggle('is-mock', !isLive);
n.tm.textContent = note !== undefined ? note : (STATE_NOTE[id] || '');
/* Advance the trace to the furthest state reached. This is driven by the
same events that carry the real result, so the bar and the node states
move only when the run actually reaches a stage β never on a timer. The
fill spans from the left edge to the centre of the current node. */
var idx = STATES.indexOf(id);
if (idx > traceProgress) traceProgress = idx;
...
if (traceFill) {
traceFill.style.width = (((traceProgress + 0.5) / STATES.length) * 100) + '%';
}
}
```
The design property in the comment is the important one: **the progress bar is a function of
reached states, not of elapsed time.** A run that stalls shows a stalled bar.
### 13.3 The measured 94.4444 % trace fill
`docs/FINAL_DELIVERY_TODO.md:72`:
| Area | Status | Evidence |
|---|---|---|
| Execution trace (progress bar) | **VERIFIED** | `.trace__fill` width is set from real event count (measured **94.4444 %** live, 2026-09-25) |
**The number is derivable from the formula**, which is why it is a real measurement rather
than a coincidence. The live path reaches `RESPOND`, the ninth of nine `ControllerState`
values:
```
width = ((traceProgress + 0.5) / STATES.length) * 100
= ((8 + 0.5) / 9) * 100
= (8.5 / 9) * 100
= 94.4444β¦ %
```
`STATES` is the nine-state `ControllerState` list (Β§12); `traceProgress = 8` is the
zero-based index of `RESPOND`. The `+ 0.5` is what makes the bar span *to the centre of the
current node* rather than to its left edge β so a fully-completed run stops at 94.4444 %, not
100 %, because the last node's centre is half a node-width short of the right edge. **A bar
that read 100 % on a nine-state rail would be reporting something the run never did.**
## 14. Worked example β one item's full journey
The following traces a single grounding claim from a specialist's `Box` to a client-facing
evidence item. Every step is a real code path.
```mermaid
sequenceDiagram
participant G as GroundingSpecialist
participant E as EvidenceEngine
participant C as confidence.calibrate
participant T as ExecutionTrace
participant F as Frontend
G->>G: compute Box(x1,y1,x2,y2, score, label, coordinate_system)
G->>E: SpecialistResult.evidence = [Evidence(type=bounding_box, ...)]
Note over E: aggregate([grounding_result])
E->>E: collect β [ev_a1b2c3d4e5f6]
E->>E: _identity_key β (bounding_box, grounding, normalized_0_1, (...), 0.87)
E->>E: _deduplicate β no collision
E->>E: sorted(_sort_key)
E->>E: _record_agreement β single specialist, no annotation
E->>E: _renumber β evidence_001
E-->>T: EvidenceCollection(items=[evidence_001], sources=['grounding'], ...)
E->>C: confidence_for(grounding_result)
C->>C: _is_effective(0.9772731820958189) == True
C->>C: _logit(0.87) / 0.9772732 β _sigmoid β _clamp01
C-->>T: ConfidenceBreakdown(raw=0.87, calibrated=0.8749186β¦, method="temperature_scaling")
T->>F: emit EVIDENCE_GENERATED {regions, count, threshold, registrationRMSE}
T->>F: emit CONFIDENCE_COMPUTED {calibrated, method, bins, observed}
T->>F: emit RESULT_ASSEMBLED {text, evidenceIds, confidence, provenance}
```
**The state of the item at each step:**
| Step | `evidence_id` | `type` | `source_specialist` | `coordinate_system` | `coordinates` | `score` | `artifact_ref` | `payload` |
|---|---|---|---|---|---|---|---|---|
| specialist output | `ev_a1b2c3d4e5f6` | `bounding_box` | `grounding` | `normalized_0_1` | `[0.21,0.33,0.47,0.61]` | `0.87` | `null` | `{"label": "water"}` |
| after dedup | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged |
| after sort | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged |
| after agreement | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | `{"label": "water"}` *(no corroboration β single specialist)* |
| after renumber | **`evidence_001`** | β¦ | β¦ | β¦ | β¦ | β¦ | β¦ | β¦ |
**And the source object is untouched.** After `aggregate` returns, the original
`SpecialistResult.evidence[0]` still has `evidence_id == "ev_a1b2c3d4e5f6"`. This is the purity
contract (Β§5.4) made concrete: `_renumber` calls `model_copy(update={"evidence_id": ...})`,
which produces a **new** object and leaves the source alone.
**Verified confidence values** (read-only execution against the shipped module):
| Input | `raw` | `calibrated` | `method` |
|---|---|---|---|
| `calibrate(0.87, T=0.9772731820958189)` | `0.87` | `0.8749186077809417` | `temperature_scaling` |
| `calibrate(0.87, T=1.0)` | `0.87` | `None` | `uncalibrated` |
| `calibrate(0.0, None)` | `0.0` | `None` | `uncalibrated` |
---
# Part E β Honest boundaries
## 15. What is NOT RUN, OPEN, BLOCKED or REJECTED for this topic
| Item | Status | Note |
|---|---|---|
| Calibration **improves** ECE | **`REJECTED`** | measured 0.013755 β 0.014929 (worse). The artifact is retained because it is in the frozen config, not because it helps. |
| Calibration applied in a **served** path | `OPEN` | `docs/PHASE19_FINAL_HARDENING.md:377` β *"Nothing is calibrated in a deployed path"* |
| Calibration covers specialists other than `change_vqa` | `NOT RUN` | the artifact's `scope` covers `change_vqa` only; every other specialist's confidence is uncalibrated |
| A second fitted artifact for the other specialists | `NOT RUN` | no artifact exists |
| `EvidenceEngine` wired into `core/controller.py` | `IMPLEMENTED` (Phase 14) | `docs/PHASE13_EVIDENCE_ENGINE.md` Β§8 recorded it as not-yet-wired at Phase 13; the live run path emits the eight events |
| Artifact **rendering** (crops, masks, change maps served to a client) | `OPEN` | *"`artifact_ref` and the `evidence_from_*` methods provide the hooks; rendering crops, masks and change maps is controller/GUI territory."* (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§8) |
| Artifact **retrieval** endpoint | `REJECTED` for v1 | F-16: no artifact-serving endpoint exists, so `artifact_ref` is permanently `null` |
| `contributing_specialists` field on `Evidence` | `REJECTED` | proposed and not adopted; `core/schemas.py` unchanged (`docs/PHASE13_EVIDENCE_ENGINE.md` Β§5) |
| Change-degradation clause in `SpecialistResult` | `RESOLVED` (removed, F-16c) | the narrower replacement rule is `REJECTED` (`core/schemas.py:362-395`) |
| `payload["corroborated_by"]` exercised by a live multi-specialist run | `UNKNOWN β not established from the available evidence` | the annotation is pinned by unit tests; no live run was found that produced two specialists making an identical claim |
| The multiclass (`softmax(logits/T)`) form of the calibration applied in code | `UNKNOWN β not established from the available evidence` | the artifact documents the capability; `evidence/confidence.py` implements the scalar form (Β§10.3) |
| Whether `_sort_key`'s coordinate tie-break is exercised with non-equal coordinates in production | `UNKNOWN β not established from the available evidence` | pinned by unit test `test_equal_scores_order_deterministically_by_coordinates` |
| The exact count of live runs that populated `dropped_over_limit > 0` | `UNKNOWN β not established from the available evidence` | no run record was found |
| `evidence.max_items = 32` ever being the binding constraint in a live run | `UNKNOWN β not established from the available evidence` | the cap is config-declared; no live run's `total_before_limit` was found |
## 16. Where the evidence lives
| Claim | Source |
|---|---|
| `Evidence` fields and the spatial validator | `core/schemas.py:216-253` |
| `artifact_ref` is permanently `null` (F-16) | `core/schemas.py:225-235`; `docs/API_CONTRACT.md` Β§2.4; `docs/DEPLOYMENT_ARCHITECTURE.md` Β§5.5 |
| F-16c and the removed change clause | `core/schemas.py:362-395`; `docs/DEPLOYMENT_ARCHITECTURE.md:855-866` |
| `EvidenceType` β all 11 members | `core/schemas.py:65-76` |
| `CoordinateSystem` β 3 values | `core/schemas.py:57-62` |
| Purity contract | `evidence/engine.py:38-50` |
| `_identity_key` and its exclusions | `evidence/engine.py:125-142`; `:52-68` |
| `_claim_key` | `evidence/engine.py:145-162` |
| `_sort_key` and the four-level rationale | `evidence/engine.py:165-190` |
| `_deduplicate` and the payload merge | `evidence/engine.py:352-389` |
| `_record_agreement` and `AGREEMENT_KEY` | `evidence/engine.py:391-425`; `:234-238` |
| `_renumber` and `ID_PREFIX` | `evidence/engine.py:427-430`; `:104-107` |
| The cap and the five accounting fields | `evidence/engine.py:193-217`; `:337-350` |
| `evidence_digest` | `evidence/engine.py:680-704` |
| `evidence_type_for` | `evidence/engine.py:434-458` |
| `evidence_from_box/region/geospatial` | `evidence/engine.py:462-572` |
| `confidence_for` and the degradation-wins rule | `evidence/engine.py:576-637` |
| The honesty rule | `evidence/confidence.py:12-30` |
| `logit`/`sigmoid`/`_EPS`/`_clamp01` | `evidence/confidence.py:80-115` |
| `_is_effective` and `_IDENTITY_TOLERANCE` | `evidence/confidence.py:90-93`, `:118-126` |
| `TemperatureCalibration` and its range guard | `evidence/confidence.py:129-252` |
| `calibrate` / `calibrate_result` | `evidence/confidence.py:255-346` |
| `load_calibration` and the ordered search | `evidence/confidence.py:349-424` |
| The 77-test phase record | `docs/PHASE13_EVIDENCE_ENGINE.md` |
| The measured calibration result | `artifacts/calibration_v001.json`; `docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§7; `docs/FINAL_DELIVERY_TODO.md:104` |
| The eight events | `frontend/assets/js/core.js:616-620` |
| The stage rail | `frontend/assets/js/core.js:605-614` |
| `markState` and the fill formula | `frontend/assets/js/mission.js:404-425` |
| The 94.4444 % measurement | `docs/FINAL_DELIVERY_TODO.md:72` |
| `ExecutionTrace` / `TraceStep` / `ModelRef` | `core/schemas.py:279-319` |
| `ControllerState` β nine states | `core/schemas.py:79-88`; `configs/base.yaml` Β§`agent.states` |
| `evidence.max_items: 32` | `configs/base.yaml` Β§`evidence` |
| `confidence.*` frozen keys | `configs/base.yaml` Β§`confidence` |
| Layer verbs | `docs/ARCHITECTURE_FREEZE.md` Β§5 |
**Next:** [07 Configuration freeze](07-configuration-freeze.md) β the registry, the enforced
invariants, and the hash `78f1e3700da15aa1` that every artifact in this chapter is keyed to.
|