Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
File size: 262,643 Bytes
a83f144 b762b81 a83f144 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 1332 1333 1334 1335 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 1396 1397 1398 1399 1400 1401 1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 1459 1460 1461 1462 1463 1464 1465 1466 1467 1468 1469 1470 1471 1472 1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 1498 1499 1500 1501 1502 1503 1504 1505 1506 1507 1508 1509 1510 1511 1512 1513 1514 1515 1516 1517 1518 1519 1520 1521 1522 1523 1524 1525 1526 1527 1528 1529 1530 1531 1532 1533 1534 1535 1536 1537 1538 1539 1540 1541 1542 1543 1544 1545 1546 1547 1548 1549 1550 1551 1552 1553 1554 1555 1556 1557 1558 1559 1560 1561 1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 1572 1573 1574 1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 1590 1591 1592 1593 1594 1595 1596 1597 1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 1625 1626 1627 1628 1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 1657 1658 1659 1660 1661 1662 1663 1664 1665 1666 1667 1668 1669 1670 1671 1672 1673 1674 1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 1727 1728 1729 1730 1731 1732 1733 1734 1735 1736 1737 1738 1739 1740 1741 1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 1760 1761 1762 1763 1764 1765 1766 1767 1768 1769 1770 1771 1772 1773 1774 1775 1776 1777 1778 1779 1780 1781 1782 1783 1784 1785 1786 1787 1788 1789 1790 1791 1792 1793 1794 1795 1796 1797 1798 1799 1800 1801 1802 1803 1804 1805 1806 1807 1808 1809 1810 1811 1812 1813 1814 1815 1816 1817 1818 1819 1820 1821 1822 1823 1824 1825 1826 1827 1828 1829 1830 1831 1832 1833 1834 1835 1836 1837 1838 1839 1840 1841 1842 1843 1844 1845 1846 1847 1848 1849 1850 1851 1852 1853 1854 1855 1856 1857 1858 1859 1860 1861 1862 1863 1864 1865 1866 1867 1868 1869 1870 1871 1872 1873 1874 1875 1876 1877 1878 1879 1880 1881 1882 1883 1884 1885 1886 1887 1888 1889 1890 1891 1892 1893 1894 1895 1896 1897 1898 1899 1900 1901 1902 1903 1904 1905 1906 1907 1908 1909 1910 1911 1912 1913 1914 1915 1916 1917 1918 1919 1920 1921 1922 1923 1924 1925 1926 1927 1928 1929 1930 1931 1932 1933 1934 1935 1936 1937 1938 1939 1940 1941 1942 1943 1944 1945 1946 1947 1948 1949 1950 1951 1952 1953 1954 1955 1956 1957 1958 1959 1960 1961 1962 1963 1964 1965 1966 1967 1968 1969 1970 1971 1972 1973 1974 1975 1976 1977 1978 1979 1980 1981 1982 1983 1984 1985 1986 1987 1988 1989 1990 1991 1992 1993 1994 1995 1996 1997 1998 1999 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 2029 2030 2031 2032 2033 2034 2035 2036 2037 2038 2039 2040 2041 2042 2043 2044 2045 2046 2047 2048 2049 2050 2051 2052 2053 2054 2055 2056 2057 2058 2059 2060 2061 2062 2063 2064 2065 2066 2067 2068 2069 2070 2071 2072 2073 2074 2075 2076 2077 2078 2079 2080 2081 2082 2083 2084 2085 2086 2087 2088 2089 2090 2091 2092 2093 2094 2095 2096 2097 2098 2099 2100 2101 2102 2103 2104 2105 2106 2107 2108 2109 2110 2111 2112 2113 2114 2115 2116 2117 2118 2119 2120 2121 2122 2123 2124 2125 2126 2127 2128 2129 2130 2131 2132 2133 2134 2135 2136 2137 2138 2139 2140 2141 2142 2143 2144 2145 2146 2147 2148 2149 2150 2151 2152 2153 2154 2155 2156 2157 2158 2159 2160 2161 2162 2163 2164 2165 2166 2167 2168 2169 2170 2171 2172 2173 2174 2175 2176 2177 2178 2179 2180 2181 2182 2183 2184 2185 2186 2187 2188 2189 2190 2191 2192 2193 2194 2195 2196 2197 2198 2199 2200 2201 2202 2203 2204 2205 2206 2207 2208 2209 2210 2211 2212 2213 2214 2215 2216 2217 2218 2219 2220 2221 2222 2223 2224 2225 2226 2227 2228 2229 2230 2231 2232 2233 2234 2235 2236 2237 2238 2239 2240 2241 2242 2243 2244 2245 2246 2247 2248 2249 2250 2251 2252 2253 2254 2255 2256 2257 2258 2259 2260 2261 2262 2263 2264 2265 2266 2267 2268 2269 2270 2271 2272 2273 2274 2275 2276 2277 2278 2279 2280 2281 2282 2283 2284 2285 2286 2287 2288 2289 2290 2291 2292 2293 2294 2295 2296 2297 2298 2299 2300 2301 2302 2303 2304 2305 2306 2307 2308 2309 2310 2311 2312 2313 2314 2315 2316 2317 2318 2319 2320 2321 2322 2323 2324 2325 2326 2327 2328 2329 2330 2331 2332 2333 2334 2335 2336 2337 2338 2339 2340 2341 2342 2343 2344 2345 2346 2347 2348 2349 2350 2351 2352 2353 2354 2355 2356 2357 2358 2359 2360 2361 2362 2363 2364 2365 2366 2367 2368 2369 2370 2371 2372 2373 2374 2375 2376 2377 2378 2379 2380 2381 2382 2383 2384 2385 2386 2387 2388 2389 2390 2391 2392 2393 2394 2395 2396 2397 2398 2399 2400 2401 2402 2403 2404 2405 2406 2407 2408 2409 2410 2411 2412 2413 2414 2415 2416 2417 2418 2419 2420 2421 2422 2423 2424 2425 2426 2427 2428 2429 2430 2431 2432 2433 2434 2435 2436 2437 2438 2439 2440 2441 2442 2443 2444 2445 2446 2447 2448 2449 2450 2451 2452 2453 2454 2455 2456 2457 2458 2459 2460 2461 2462 2463 2464 2465 2466 2467 2468 2469 2470 2471 2472 2473 2474 2475 2476 2477 2478 2479 2480 2481 2482 2483 2484 2485 2486 2487 2488 2489 2490 2491 2492 2493 2494 2495 2496 2497 2498 2499 2500 2501 2502 2503 2504 2505 2506 2507 2508 2509 2510 2511 2512 2513 2514 2515 2516 2517 2518 2519 2520 2521 2522 2523 2524 2525 2526 2527 2528 2529 2530 2531 2532 2533 2534 2535 2536 2537 2538 2539 2540 2541 2542 2543 2544 2545 2546 2547 2548 2549 2550 2551 2552 2553 2554 2555 2556 2557 2558 2559 2560 2561 2562 2563 2564 2565 2566 2567 2568 2569 2570 2571 2572 2573 2574 2575 2576 2577 2578 2579 2580 2581 2582 2583 2584 2585 2586 2587 2588 2589 2590 2591 2592 2593 2594 2595 2596 2597 2598 2599 2600 2601 2602 2603 2604 2605 2606 2607 2608 2609 2610 2611 2612 2613 2614 2615 2616 2617 2618 2619 2620 2621 2622 2623 2624 2625 2626 2627 2628 2629 2630 2631 2632 2633 2634 2635 2636 2637 2638 2639 2640 2641 2642 2643 2644 2645 2646 2647 2648 2649 2650 2651 2652 2653 2654 2655 2656 2657 2658 2659 2660 2661 2662 2663 2664 2665 2666 2667 2668 2669 2670 2671 2672 2673 2674 2675 2676 2677 2678 2679 2680 2681 2682 2683 2684 2685 2686 2687 2688 2689 2690 2691 2692 2693 2694 2695 2696 2697 2698 2699 2700 2701 2702 2703 2704 2705 2706 2707 2708 2709 2710 2711 2712 2713 2714 2715 2716 2717 2718 2719 2720 2721 2722 2723 2724 2725 2726 2727 2728 2729 2730 2731 2732 2733 2734 2735 2736 2737 2738 2739 2740 2741 2742 2743 2744 2745 2746 2747 2748 2749 2750 2751 2752 2753 2754 2755 2756 2757 2758 2759 2760 2761 2762 2763 2764 2765 2766 2767 2768 2769 2770 2771 2772 2773 2774 2775 2776 2777 2778 2779 2780 2781 2782 2783 2784 2785 2786 2787 2788 2789 2790 2791 2792 2793 2794 2795 2796 2797 2798 2799 2800 2801 2802 2803 2804 2805 2806 2807 2808 2809 2810 2811 2812 2813 2814 2815 2816 2817 2818 2819 2820 2821 2822 2823 2824 2825 2826 2827 2828 2829 2830 2831 2832 2833 2834 2835 2836 2837 2838 2839 2840 2841 2842 2843 2844 2845 2846 2847 2848 2849 2850 2851 2852 2853 2854 2855 2856 2857 2858 2859 2860 2861 2862 2863 2864 2865 2866 2867 2868 2869 2870 2871 2872 2873 2874 2875 2876 2877 2878 2879 2880 2881 2882 2883 2884 2885 2886 2887 2888 2889 2890 2891 2892 2893 2894 2895 2896 2897 2898 2899 2900 2901 2902 2903 2904 2905 2906 2907 2908 2909 2910 2911 2912 2913 2914 2915 2916 2917 2918 2919 2920 2921 2922 2923 2924 2925 2926 2927 2928 2929 2930 2931 2932 2933 2934 2935 2936 2937 2938 2939 2940 2941 2942 2943 2944 2945 2946 2947 2948 2949 2950 2951 2952 2953 2954 2955 2956 2957 2958 2959 2960 2961 2962 2963 2964 2965 2966 2967 2968 2969 2970 2971 2972 2973 2974 2975 2976 2977 2978 2979 2980 2981 2982 2983 2984 2985 2986 2987 2988 2989 2990 2991 2992 2993 2994 2995 2996 2997 2998 2999 3000 3001 3002 3003 3004 3005 3006 3007 3008 3009 3010 3011 3012 3013 3014 3015 3016 3017 3018 3019 3020 3021 3022 3023 3024 3025 3026 3027 3028 3029 3030 3031 3032 3033 3034 3035 3036 3037 3038 3039 3040 3041 3042 3043 3044 3045 3046 3047 3048 3049 3050 3051 3052 3053 3054 3055 3056 3057 3058 3059 3060 3061 3062 3063 3064 3065 3066 3067 3068 3069 3070 3071 3072 3073 3074 3075 3076 3077 3078 3079 3080 3081 3082 3083 3084 3085 3086 3087 3088 3089 3090 3091 3092 3093 3094 3095 3096 3097 3098 3099 3100 3101 3102 3103 3104 3105 3106 3107 3108 3109 3110 3111 3112 3113 3114 3115 3116 3117 3118 3119 3120 3121 3122 3123 3124 3125 3126 3127 3128 3129 3130 3131 3132 3133 3134 3135 3136 3137 3138 3139 3140 3141 3142 3143 3144 3145 3146 3147 3148 3149 3150 3151 3152 3153 3154 3155 3156 3157 3158 3159 3160 3161 3162 3163 3164 3165 3166 3167 3168 3169 3170 3171 3172 3173 3174 3175 3176 3177 3178 3179 3180 3181 3182 3183 3184 3185 3186 3187 3188 3189 3190 3191 3192 3193 3194 3195 3196 3197 3198 3199 3200 3201 3202 3203 3204 3205 3206 3207 3208 3209 3210 3211 3212 3213 3214 3215 3216 3217 3218 3219 3220 3221 3222 3223 3224 3225 3226 3227 3228 3229 3230 3231 3232 3233 3234 3235 3236 3237 3238 3239 3240 3241 3242 3243 3244 3245 3246 3247 3248 3249 3250 3251 3252 3253 3254 3255 3256 3257 3258 3259 3260 3261 3262 3263 3264 3265 3266 3267 3268 3269 3270 3271 3272 3273 3274 3275 3276 3277 3278 3279 3280 3281 3282 3283 3284 3285 3286 3287 3288 3289 3290 3291 3292 3293 3294 3295 3296 3297 3298 3299 3300 3301 3302 3303 3304 3305 3306 3307 3308 3309 3310 3311 3312 3313 3314 3315 3316 3317 3318 3319 3320 3321 3322 3323 3324 3325 3326 3327 3328 3329 3330 3331 3332 3333 3334 3335 3336 3337 3338 3339 3340 3341 3342 3343 3344 3345 3346 3347 3348 3349 3350 3351 3352 3353 3354 3355 3356 3357 3358 3359 3360 3361 3362 3363 3364 3365 3366 3367 3368 3369 3370 3371 3372 3373 3374 3375 3376 3377 3378 3379 3380 3381 3382 3383 3384 3385 3386 3387 3388 3389 3390 3391 3392 3393 3394 3395 3396 3397 3398 3399 3400 3401 3402 3403 3404 3405 3406 3407 3408 3409 3410 3411 3412 3413 3414 3415 3416 3417 3418 3419 3420 3421 3422 3423 3424 3425 3426 3427 3428 3429 3430 3431 3432 3433 3434 3435 3436 3437 3438 3439 3440 3441 3442 3443 3444 3445 3446 3447 3448 3449 3450 3451 3452 3453 3454 3455 3456 3457 3458 3459 3460 3461 3462 3463 3464 3465 3466 3467 3468 3469 3470 3471 3472 3473 3474 3475 3476 3477 3478 3479 3480 3481 3482 3483 3484 3485 3486 3487 3488 3489 3490 3491 3492 3493 3494 3495 3496 3497 3498 3499 3500 3501 3502 3503 3504 3505 3506 3507 3508 3509 3510 3511 3512 3513 3514 3515 3516 3517 3518 3519 3520 3521 3522 3523 3524 3525 3526 3527 3528 3529 3530 3531 3532 3533 3534 3535 3536 3537 3538 3539 3540 3541 3542 3543 3544 3545 3546 3547 3548 3549 3550 3551 3552 3553 3554 3555 3556 3557 3558 3559 3560 3561 3562 3563 3564 3565 3566 3567 3568 3569 3570 3571 3572 3573 3574 3575 3576 3577 3578 3579 3580 3581 3582 3583 3584 3585 3586 3587 3588 3589 3590 3591 3592 3593 3594 3595 3596 3597 3598 3599 3600 3601 3602 3603 3604 3605 3606 3607 3608 3609 3610 3611 3612 3613 3614 3615 3616 3617 3618 3619 3620 3621 3622 3623 3624 3625 3626 3627 3628 3629 3630 3631 3632 3633 3634 3635 3636 3637 3638 3639 3640 3641 3642 3643 3644 3645 3646 3647 3648 3649 3650 3651 3652 3653 3654 3655 3656 3657 3658 3659 3660 3661 3662 3663 3664 3665 3666 3667 3668 3669 3670 3671 3672 3673 3674 3675 3676 3677 3678 3679 3680 3681 3682 3683 3684 3685 3686 3687 3688 3689 3690 3691 3692 3693 3694 3695 3696 3697 3698 3699 3700 3701 3702 3703 3704 3705 3706 3707 3708 3709 3710 3711 3712 3713 3714 3715 3716 3717 3718 3719 3720 3721 3722 3723 3724 3725 3726 3727 3728 3729 3730 3731 3732 3733 3734 3735 3736 3737 3738 3739 3740 3741 3742 3743 3744 3745 3746 3747 3748 3749 3750 3751 3752 3753 3754 3755 3756 3757 3758 3759 3760 3761 3762 3763 3764 3765 3766 3767 3768 3769 3770 3771 3772 3773 3774 3775 3776 3777 3778 3779 3780 3781 3782 3783 3784 3785 3786 3787 3788 3789 3790 3791 3792 3793 3794 3795 3796 3797 3798 3799 3800 3801 3802 3803 3804 3805 3806 3807 3808 3809 3810 3811 3812 3813 3814 3815 3816 3817 3818 3819 3820 3821 3822 3823 3824 3825 3826 3827 3828 3829 3830 3831 3832 3833 3834 3835 3836 3837 3838 3839 3840 3841 3842 3843 3844 3845 3846 3847 3848 3849 3850 3851 3852 3853 3854 3855 3856 3857 3858 3859 3860 3861 3862 3863 3864 3865 3866 3867 3868 3869 3870 3871 3872 3873 3874 3875 3876 3877 3878 3879 3880 3881 3882 3883 3884 3885 3886 3887 3888 3889 3890 3891 3892 3893 3894 3895 3896 3897 3898 3899 3900 3901 3902 3903 3904 3905 3906 3907 3908 3909 3910 3911 3912 3913 3914 3915 3916 3917 3918 3919 3920 3921 3922 3923 3924 3925 3926 3927 3928 3929 3930 3931 3932 3933 3934 3935 3936 3937 3938 3939 3940 3941 3942 3943 3944 3945 3946 3947 3948 3949 3950 3951 3952 3953 3954 3955 3956 3957 3958 3959 3960 3961 3962 3963 3964 3965 3966 3967 3968 3969 3970 3971 3972 3973 3974 3975 3976 3977 3978 3979 3980 3981 3982 3983 3984 3985 3986 3987 3988 3989 3990 3991 3992 3993 3994 3995 3996 3997 3998 3999 4000 4001 4002 4003 4004 4005 4006 4007 4008 4009 4010 4011 4012 4013 4014 4015 4016 4017 4018 4019 4020 4021 4022 4023 4024 4025 4026 4027 4028 4029 4030 4031 4032 4033 4034 4035 4036 4037 4038 4039 4040 4041 4042 4043 4044 4045 4046 4047 4048 4049 4050 4051 4052 4053 4054 4055 4056 4057 4058 4059 4060 4061 4062 4063 4064 4065 4066 4067 4068 4069 4070 4071 4072 4073 4074 4075 4076 4077 4078 4079 4080 4081 4082 4083 4084 4085 4086 4087 4088 4089 4090 4091 4092 4093 4094 4095 4096 4097 4098 4099 4100 4101 4102 4103 4104 4105 4106 4107 4108 4109 4110 4111 4112 4113 4114 4115 4116 4117 4118 4119 4120 4121 4122 4123 4124 4125 4126 4127 4128 4129 4130 4131 4132 4133 4134 4135 4136 4137 4138 4139 4140 4141 4142 4143 4144 4145 4146 4147 4148 4149 4150 4151 4152 4153 4154 4155 4156 4157 4158 4159 4160 4161 4162 4163 4164 4165 4166 4167 4168 4169 4170 4171 4172 4173 4174 4175 4176 4177 4178 4179 4180 4181 4182 4183 4184 4185 4186 4187 4188 4189 4190 4191 4192 4193 4194 4195 4196 4197 4198 4199 4200 4201 4202 4203 4204 4205 4206 4207 4208 4209 4210 4211 4212 4213 4214 4215 4216 4217 4218 4219 4220 4221 4222 4223 4224 4225 4226 4227 4228 4229 4230 4231 4232 4233 4234 4235 4236 4237 4238 4239 4240 4241 4242 4243 4244 4245 4246 4247 4248 4249 4250 4251 4252 4253 4254 4255 4256 4257 4258 4259 4260 4261 4262 4263 4264 4265 4266 4267 4268 4269 4270 4271 4272 4273 4274 4275 4276 4277 4278 4279 4280 4281 4282 4283 4284 4285 4286 4287 4288 4289 4290 4291 4292 4293 4294 4295 4296 4297 4298 4299 4300 4301 4302 4303 4304 4305 4306 4307 4308 4309 4310 4311 4312 4313 4314 4315 4316 4317 4318 4319 4320 4321 4322 4323 4324 4325 4326 4327 4328 4329 4330 4331 4332 4333 4334 4335 4336 4337 4338 4339 4340 4341 4342 4343 4344 4345 4346 4347 4348 4349 4350 4351 4352 4353 4354 4355 4356 4357 4358 4359 4360 4361 4362 4363 4364 4365 4366 4367 4368 4369 4370 4371 4372 4373 4374 4375 4376 4377 4378 4379 4380 4381 4382 4383 4384 4385 4386 4387 4388 4389 4390 4391 4392 4393 4394 4395 4396 4397 4398 4399 4400 4401 4402 4403 4404 4405 4406 4407 4408 4409 4410 4411 4412 4413 4414 4415 4416 4417 4418 4419 4420 4421 4422 4423 4424 4425 4426 4427 4428 4429 4430 4431 4432 4433 4434 4435 4436 4437 4438 4439 4440 4441 4442 4443 4444 4445 4446 4447 4448 4449 4450 4451 4452 4453 4454 4455 4456 4457 4458 4459 4460 4461 4462 4463 4464 4465 4466 4467 4468 4469 4470 4471 4472 4473 4474 4475 4476 4477 4478 4479 4480 4481 4482 4483 4484 4485 4486 4487 4488 4489 4490 4491 4492 4493 4494 4495 4496 4497 4498 4499 4500 4501 4502 4503 4504 4505 4506 4507 4508 4509 4510 4511 4512 4513 4514 4515 4516 4517 4518 4519 4520 4521 4522 4523 4524 4525 4526 4527 4528 4529 4530 4531 4532 4533 4534 4535 4536 4537 4538 4539 4540 4541 4542 4543 4544 4545 4546 4547 4548 4549 4550 4551 4552 4553 4554 4555 4556 4557 4558 4559 4560 4561 4562 4563 4564 4565 4566 4567 4568 4569 4570 4571 4572 4573 4574 4575 4576 4577 4578 4579 4580 4581 4582 4583 4584 4585 4586 4587 4588 4589 4590 4591 4592 4593 4594 4595 4596 4597 4598 4599 4600 4601 4602 4603 4604 4605 4606 4607 4608 4609 4610 4611 4612 4613 4614 4615 4616 4617 4618 4619 4620 4621 4622 4623 4624 4625 4626 4627 4628 4629 4630 4631 4632 4633 4634 4635 4636 4637 4638 4639 4640 4641 4642 4643 4644 4645 4646 4647 4648 4649 4650 4651 4652 4653 4654 4655 4656 4657 4658 4659 4660 4661 4662 4663 4664 4665 4666 4667 4668 4669 4670 4671 4672 4673 4674 4675 4676 4677 4678 4679 4680 4681 4682 4683 4684 4685 4686 4687 4688 4689 4690 4691 4692 4693 4694 4695 4696 4697 4698 4699 4700 4701 4702 4703 4704 4705 4706 4707 4708 4709 4710 4711 4712 4713 4714 4715 4716 4717 4718 4719 4720 4721 4722 4723 4724 4725 4726 4727 4728 4729 4730 4731 4732 4733 4734 4735 4736 4737 4738 4739 4740 4741 4742 4743 4744 4745 4746 4747 4748 4749 4750 4751 4752 4753 4754 4755 4756 4757 4758 4759 4760 4761 4762 4763 4764 4765 4766 4767 4768 4769 4770 4771 4772 4773 4774 4775 4776 4777 4778 4779 4780 4781 4782 4783 4784 4785 4786 4787 4788 4789 4790 4791 4792 4793 4794 4795 4796 4797 4798 4799 4800 4801 4802 4803 4804 4805 4806 4807 4808 4809 4810 4811 4812 4813 4814 4815 4816 4817 4818 4819 4820 4821 4822 4823 4824 4825 4826 4827 4828 4829 4830 4831 4832 4833 4834 4835 4836 4837 4838 4839 4840 4841 4842 4843 4844 4845 4846 4847 4848 4849 4850 4851 4852 4853 4854 4855 4856 4857 4858 4859 4860 4861 4862 4863 4864 4865 4866 4867 4868 4869 4870 4871 4872 4873 4874 4875 4876 4877 4878 4879 4880 4881 4882 4883 4884 4885 4886 4887 4888 4889 4890 4891 4892 4893 4894 4895 4896 4897 4898 4899 4900 4901 4902 4903 4904 4905 4906 4907 4908 4909 4910 4911 4912 4913 4914 4915 4916 4917 4918 4919 4920 4921 4922 4923 4924 4925 4926 4927 4928 4929 4930 4931 4932 4933 4934 4935 4936 4937 4938 4939 4940 4941 4942 4943 4944 4945 4946 4947 4948 4949 4950 4951 4952 4953 4954 4955 4956 4957 4958 4959 4960 4961 4962 4963 4964 4965 4966 4967 4968 4969 4970 4971 4972 4973 4974 4975 4976 4977 4978 4979 4980 4981 4982 4983 4984 4985 4986 4987 4988 4989 4990 4991 4992 4993 4994 4995 4996 4997 4998 4999 5000 5001 5002 5003 5004 5005 5006 5007 5008 5009 5010 5011 5012 5013 5014 5015 5016 5017 5018 5019 5020 5021 5022 5023 5024 5025 5026 5027 5028 5029 5030 5031 5032 5033 5034 5035 5036 5037 5038 5039 5040 5041 5042 5043 5044 5045 5046 5047 5048 5049 5050 5051 5052 5053 5054 5055 5056 5057 5058 5059 5060 5061 5062 5063 5064 5065 5066 5067 5068 5069 5070 5071 5072 5073 5074 5075 5076 5077 5078 5079 5080 5081 5082 5083 5084 5085 5086 5087 5088 5089 5090 5091 5092 5093 5094 5095 5096 5097 5098 5099 5100 5101 5102 5103 5104 5105 5106 5107 5108 5109 5110 5111 5112 5113 5114 5115 5116 5117 5118 5119 5120 5121 5122 5123 5124 5125 5126 5127 5128 5129 5130 5131 5132 5133 5134 5135 5136 5137 5138 5139 5140 5141 5142 5143 5144 5145 5146 5147 5148 5149 5150 5151 5152 5153 5154 5155 5156 5157 5158 5159 5160 5161 5162 5163 5164 5165 5166 5167 5168 5169 5170 5171 5172 5173 5174 5175 5176 5177 5178 5179 5180 5181 5182 5183 5184 5185 5186 5187 5188 5189 5190 5191 5192 5193 5194 5195 5196 5197 5198 5199 5200 5201 5202 5203 5204 5205 5206 5207 5208 5209 5210 5211 5212 5213 5214 5215 5216 5217 5218 5219 5220 5221 5222 5223 5224 5225 5226 5227 5228 5229 5230 5231 5232 5233 5234 5235 5236 5237 5238 5239 5240 5241 5242 5243 5244 5245 5246 5247 5248 5249 5250 5251 5252 5253 5254 5255 5256 5257 5258 5259 5260 5261 5262 5263 5264 5265 5266 5267 5268 5269 5270 5271 5272 5273 5274 5275 5276 5277 5278 5279 5280 5281 5282 5283 5284 5285 5286 5287 5288 5289 5290 5291 5292 5293 5294 5295 5296 5297 5298 5299 5300 5301 5302 5303 5304 5305 5306 5307 5308 5309 5310 5311 5312 5313 5314 5315 5316 5317 5318 5319 5320 5321 5322 5323 5324 5325 5326 5327 5328 5329 5330 5331 5332 5333 5334 5335 5336 5337 5338 5339 5340 5341 5342 5343 5344 5345 5346 5347 5348 5349 5350 5351 5352 5353 5354 5355 5356 5357 5358 5359 5360 5361 5362 5363 5364 5365 5366 5367 5368 5369 5370 5371 5372 5373 5374 5375 5376 5377 5378 5379 5380 5381 5382 5383 5384 5385 5386 5387 5388 5389 5390 5391 5392 5393 5394 5395 5396 5397 5398 5399 5400 5401 5402 5403 5404 5405 5406 5407 5408 5409 5410 5411 5412 5413 5414 5415 5416 5417 5418 5419 5420 5421 5422 5423 5424 5425 5426 5427 5428 5429 5430 5431 5432 5433 5434 5435 5436 5437 5438 5439 5440 5441 5442 5443 5444 5445 5446 5447 5448 5449 5450 5451 5452 5453 5454 5455 5456 5457 5458 5459 5460 5461 5462 5463 5464 5465 5466 5467 5468 5469 5470 5471 5472 5473 5474 5475 5476 5477 5478 5479 5480 5481 5482 5483 5484 5485 5486 5487 5488 5489 5490 5491 5492 5493 5494 5495 5496 5497 5498 5499 5500 5501 5502 5503 5504 5505 5506 5507 5508 5509 5510 5511 5512 5513 5514 5515 5516 5517 5518 5519 5520 5521 5522 5523 5524 5525 5526 5527 5528 5529 5530 5531 5532 5533 5534 5535 5536 5537 5538 5539 5540 5541 5542 5543 5544 5545 5546 5547 5548 5549 5550 5551 5552 5553 5554 5555 5556 5557 5558 5559 5560 5561 5562 5563 5564 5565 5566 5567 5568 5569 5570 5571 5572 5573 5574 5575 5576 5577 5578 5579 5580 5581 5582 5583 5584 5585 5586 5587 5588 5589 5590 5591 5592 5593 5594 5595 5596 5597 5598 5599 5600 5601 5602 5603 5604 5605 5606 5607 5608 5609 5610 5611 5612 5613 5614 5615 5616 5617 5618 5619 5620 5621 5622 5623 5624 5625 5626 5627 5628 5629 5630 5631 5632 5633 5634 5635 5636 5637 5638 5639 5640 5641 5642 5643 5644 5645 5646 5647 5648 5649 5650 5651 5652 5653 5654 5655 5656 5657 5658 5659 5660 5661 5662 5663 5664 5665 5666 5667 5668 5669 5670 5671 5672 5673 5674 5675 5676 5677 5678 5679 5680 5681 5682 5683 5684 5685 5686 5687 5688 5689 5690 5691 5692 5693 5694 5695 5696 5697 5698 5699 5700 5701 5702 5703 5704 5705 5706 5707 5708 5709 5710 5711 5712 5713 5714 5715 5716 5717 5718 5719 5720 5721 5722 5723 5724 5725 5726 5727 5728 5729 5730 5731 5732 5733 5734 5735 5736 5737 5738 5739 5740 5741 5742 5743 5744 5745 5746 5747 5748 5749 5750 5751 5752 5753 5754 5755 5756 5757 5758 5759 5760 5761 5762 5763 5764 5765 5766 5767 5768 5769 5770 5771 5772 5773 5774 5775 5776 5777 5778 5779 5780 5781 5782 5783 5784 5785 5786 5787 5788 | # 04 β The Router
**Parent:** [Architecture hub](README.md) Β· **Sibling chapters:**
[01 System overview](01-system-overview.md) Β· [03 Request lifecycle](03-request-lifecycle.md) Β·
[05 Specialists](05-specialists.md) Β· [06 Evidence and confidence](06-evidence-and-confidence.md) Β·
[07 Configuration freeze](07-configuration-freeze.md)
**Primary sources read for this chapter (all under `C:/Users/anish/satquery-ai/` unless stated):**
| Source | Lines | What it establishes here |
|---|---|---|
| `router/encoder.py` | 190 | `FrozenEncoder`, `VERIFIED_TOKENIZER_MAX_LENGTH = 256`, `VERIFIED_EMBEDDING_DIM = 384`, the `max_length β€ 256` refusal, `build_encoder`, the frozen-revision contract probe |
| `router/adapter.py` | 174 | `IntentAdapter` β all seven sub-modules, `AdapterOutput`, `_init_weights`, the forward shape guards, `config_dict` / `from_config_dict` / `num_parameters` |
| `router/label_space.py` | 129 | `TASK_CLASSES`, `MODALITY_CLASSES`, `BINARY_HEADS`, `TEMPORAL_TASKS`, `DUAL_MODALITY_TASKS`, `SPATIAL_TASKS`, `default_attributes`, `default_modality` |
| `router/fallback.py` | 333 | the seven term tables, the eight ordered rules, every rule confidence, `self_check()`'s ten curated cases, `is_available()` |
| `router/classifier.py` | 475 | `IntentRouter`, `confidence_threshold = 0.70`, `_predict_learned` and the coherence repair, `route()`'s three-step order, `RouterPrediction.to_trace()`, `save_adapter` / `load_adapter` |
| `router/dataset.py` | 707 | `RouterExample`, `CURATED` (70 entries), `TEMPLATES` (45 entries), `SUBJECTS`, `HARD_NEGATIVE_PREFIX`, `RouterCorpus.validate/dedupe`, `split_by_group`, `split_leakage_report` |
| `router/train.py` | 753 | `train_router`, `SplitMetrics`, `TrainingResult.gate2_passed`, `embed_corpus_cached`, class weights, loss composition, `_stub_embeddings` |
| `artifacts/router/router_adapter_v001/metadata.json` | 626 | the shipped adapter artifact: 51,725 params, corpus 576 / 54 groups, all three split metric blocks, the full 60-epoch history |
| `artifacts/router/threshold_sweep_val.json` | 50 rows | the release-level threshold sweep: val only, `n_test_examples_scored: 0`, `test_split_touched: false`, shipped threshold 0.70, selected 0.76 |
| `core/planner.py` | 662 | `PolicyPlanner` β `LEXICAL_FALLBACK_DISCOUNT = 0.75`, `SETTLED_CONFIDENCE = 0.60`, `TASK_CAPABILITY`, `CAPABILITY_ASSETS`, `PlanStep` / `PlanRefusal` / `ExecutionPlan`, the closed refusal list, the Β§3.5 widening rules |
| `core/schemas.py` | 462 | `Task` (7 members), `Modality` (4), `Intent` + its `_consistency` validator |
| `configs/base.yaml` Β§`router` | β | every configured value: model, revision, `max_length: 128`, `embedding_dim: 384`, `hidden_dim: 128`, `dropout: 0.10`, `confidence_threshold: 0.70`, `num_tasks: 6`, and the six training keys |
| `frontend/assets/js/mission.js` | 1β300 | `interpret()` (line 70) and `chooseTask()` (line 184) β the two-stage frontend router, with the four defect-fix comments verbatim |
| `docs/PHASE4_ROUTER_REPORT.md` | 179 | the phase record: Gate 2 PASS, F4-1/F4-2/F4-3, four defects found by running, the standing caveats |
| `release/repo/docs/RESEARCH_NOTES.md` Β§3, Β§6 | β | the router-defect case study, the live-pass table, the two run ids, and the `interpret()`/`chooseTask()` asymmetry ruling |
| `release/repo/README.md` Β§"Routing and the execution trace", Β§"Known limitations" rows 2 and 8 | β | the two-stage description, the five-head table, the confidence-gate diagram, the router residuals |
---
## 1. Where the router sits, and the verb it owns
`docs/ARCHITECTURE_FREEZE.md` Β§5 assigns each layer exactly one verb:
> Router *understands*; policy engine *decides*; specialists *compute*; VLM *explains*;
> **evidence engine *proves*.**
The router is the first verb. It takes a **string** and produces a **typed reading of what the
string is asking for** β nothing more. It does not choose what runs, it does not touch pixels, and
it does not answer anything. `core/schemas.py:95` states the contract in the model's own docstring:
```python
class Intent(BaseModel):
"""Output of the learned router. Advisory only β the controller decides."""
```
That single line is load-bearing for the whole chapter. Everything downstream of the router is
allowed β required β to overrule it. `core/planner.py:13-15` restates it as the planner's reason to
exist:
> **The planner is the only component permitted to choose what runs. Its input is a
> `RouterPrediction`; its output is an `ExecutionPlan`.** [β¦] The router's `Intent` is an *input to
> a decision*, never the decision.
### 1.1 The router is one of two stages, and the split is deliberate
Routing happens twice in this system, in two languages, with two different jobs. This is the single
most misread part of the architecture, so it is stated first.
| Stage | Where | Blind to asset count? | Job |
|---|---|---|---|
| **Reading** β `interpret()` | `frontend/assets/js/mission.js:70` | **yes** β text only | what does the question *mean*? |
| **Dispatch** β `chooseTask()` | `frontend/assets/js/mission.js:184` | **no** β reads `assetCount` | what can actually be *computed* with the assets present? |
| **Reading** β `IntentRouter.route()` | `router/classifier.py:311` | **yes** | the server-side analogue: query β `Intent` |
| **Dispatch** β `PolicyPlanner.plan()` | `core/planner.py:276` | **no** β reads `len(request.assets)` | the server-side analogue: `Intent` β `ExecutionPlan` |
The frontend pair and the server pair are structurally the same shape and were written
independently. Β§35βΒ§38 covers the frontend pair in full; Β§39βΒ§43 covers the server pair.
### 1.2 The pipeline
```mermaid
flowchart TB
Q["query : str"] --> ENC["FrozenEncoder.encode_one<br/>MiniLM-L6-v2 @ 1110a243fdf4<br/>384-d, frozen"]
Q --> FB["lexical_route(query)<br/>deterministic, ordered rules"]
ENC --> AD["IntentAdapter<br/>5 heads"]
AD --> GATE{"task softmax max<br/>β₯ confidence_threshold<br/>(0.70)"}
GATE -->|yes| LEARNED["Intent(source='learned')"]
GATE -->|no| CMP{"fallback confidence<br/>β₯ learned confidence?"}
FB --> CMP
CMP -->|yes| FALL["Intent(source='lexical_fallback')"]
CMP -->|no| LEARNED
LEARNED --> PLAN["PolicyPlanner.plan()<br/>core/planner.py:276"]
FALL --> PLAN
PLAN --> EP["ExecutionPlan<br/>steps | refusal"]
style ENC fill:#1f6feb22,stroke:#1f6feb
style AD fill:#1f6feb22,stroke:#1f6feb
style PLAN fill:#8957e522,stroke:#8957e5
```
The encoder and adapter are the **learned** path. `lexical_route` is the **fallback** path. Both
always run when both are available; Β§23 explains exactly how `route()` arbitrates between them.
### 1.3 Status of every component in this chapter
Statuses follow `DOCS_STYLE_GUIDE.md` Β§2. Nothing below is upgraded.
| Component | Status | Basis |
|---|---|---|
| `FrozenEncoder` β load, freeze, guard `max_length β€ 256` | `IMPLEMENTED` + `VERIFIED` | `router/encoder.py`; the contract probe in the module docstring (2026-09-16, sentence-transformers 6.0.1) |
| MiniLM revision pin `1110a243fdf4` | `VERIFIED` reachable | `router/encoder.py:7`; recorded again in the shipped `metadata.json` |
| `IntentAdapter` β 5 heads, 51,725 params | `IMPLEMENTED` + `MEASURED` | `router/adapter.py`; `metadata.json` `adapter.num_parameters: 51725` |
| Adapter trained on the shipped corpus | `MEASURED` | `metadata.json` β 60 epochs, 4.92 s, `artifacts/router/router_adapter_v001` |
| Gate 2 (task accuracy β₯ 0.95, all 6 classes measured) on the **training-time test split** | `MEASURED` β reported PASS at 0.975, n=80 | `metadata.json.metrics.test`; `docs/PHASE4_ROUTER_REPORT.md` |
| Release-level **threshold sweep** | `MEASURED` β **val only**, `n_test_examples_scored: 0` | `artifacts/router/threshold_sweep_val.json` |
| Router **test split touched by the release evaluation** | **`NOT RUN`** | `threshold_sweep_val.json` β `test_split_touched: false`; `README.md` limitation 2 |
| Lexical fallback β all eight rules | `IMPLEMENTED` + `VERIFIED` | `router/fallback.py`; `self_check()` ten cases |
| Fallback agrees with the learned router on the curated hard negatives | `VERIFIED` β 0.800 hard-negative accuracy on test | `metadata.json.metrics.test.hard_negative_accuracy: 0.8` |
| `IntentRouter.route()` two-path arbitration | `IMPLEMENTED` + `VERIFIED` | `router/classifier.py:311-339` |
| `PolicyPlanner` β pure, no torch, no filesystem | `IMPLEMENTED` | `core/planner.py`; the `TYPE_CHECKING`-only import of `RouterPrediction` |
| Frontend `interpret()` / `chooseTask()` two-stage split | `IMPLEMENTED` + `VERIFIED` live | `frontend/assets/js/mission.js`; 3 passes Γ 8 cases, 24 runs |
| Router-defect fix (`built`, `\barea\b`, `new`, `chang` stem) | `RESOLVED` β deployed and verified live | `docs/RESEARCH_NOTES.md` Β§3.3; run ids `run_467ffa406f22`, `run_46980ba55c62` |
| Router calibration | **`NOT RUN`** β the 0.70 threshold is uncalibrated by the phase's own admission | `docs/PHASE4_ROUTER_REPORT.md` Β§"Standing caveats" |
| Router **accuracy** as a benchmark claim | **`NOT RUN`** β the corpus is synthetic (576 examples / 54 groups) | `docs/PHASE4_ROUTER_REPORT.md`; `threshold_sweep_val.json` β `corpus_limited: true` |
| Router **lexical residuals** (`"What is the new runway?"`, `"How much built-up area was added?"`) | **`OPEN`** | `README.md` limitation 8 |
### 1.4 The three vocabularies, and why there are three
A reader who compares `router/label_space.py` to `core/schemas.py` to `GET /api/capabilities` finds
three different sets of task names. That is not drift. `README.md:136-151` documents it as
deliberate, and it is worth restating because every later chapter depends on it.
| Vocabulary | Where | Members | Count |
|---|---|---|---|
| `core.schemas.Task` | `core/schemas.py:35-47` | `vqa`, `caption`, `grounding`, `change`, `optical_sar`, `change_vqa`, `unsupported` | **7** |
| `router.label_space.TASK_CLASSES` | `router/label_space.py:26-33` | `vqa`, `caption`, `grounding`, `change`, `optical_sar`, `unsupported` | **6** |
| `GET /api/capabilities` | live capability contract | `vqa`, `caption`, `grounding`, `change`, `change_vqa`, `optical_sar` | **6** |
The differences are each a design decision:
- **`unsupported` is in `Task` and in `TASK_CLASSES` but not in the capability list.** `unsupported`
is a **routing outcome** β "this is not a satellite-imagery question" β not a servable capability.
`core/planner.py:120` says so in a comment on `TASK_CAPABILITY`: *"`UNSUPPORTED` is deliberately
absent β it is a refusal, not a capability."*
- **`change_vqa` is in `Task` and in the capability list but not in `TASK_CLASSES`.** It is reached
through the change *family* rather than being a separate router class: the router has no
`change_vqa` head output, and `core/planner.py:470-491` widens a `change` + `language_output`
prediction into a `change_vqa` step. `core/schemas.py:41-46` explains why the three are not
interchangeable:
```python
#: R-02. Two temporally corresponding assets plus a change-oriented question
#: in, a short answer out. Distinct from CHANGE, which is the change
#: *detector* and returns a spatial change map with no language output, and
#: distinct from VQA, which answers about ONE asset. The three are not
#: interchangeable and the planner must not substitute one for another.
CHANGE_VQA = "change_vqa"
```
- **The reconciliation happens in one place.** Because `AnalysisRequest` is `extra="forbid"` and
`force_task` is typed `Task | None` (`core/schemas.py:412-418`), any string outside the `Task`
enum is a 422. The frontend therefore maps the router's reading to the server enum through a
single table, `ROUTER_TASK_TO_SERVER` (`frontend/assets/js/mission.js:148-155`).
`router/classifier.py` guards the router-side alignment at import time, so a label added to
`label_space.py` without a schema mapping cannot ship silently:
```python
def _assert_schema_alignment() -> None:
missing_tasks = [t for t in TASK_CLASSES if t not in _TASK_TO_SCHEMA]
if missing_tasks:
raise RoutingError(
f"label_space TASK_CLASSES contains labels with no schema mapping: "
f"{missing_tasks}"
)
missing_mods = [m for m in MODALITY_CLASSES if m not in _MODALITY_TO_SCHEMA]
...
_assert_schema_alignment()
```
(`router/classifier.py:62-77`.) The guard runs at **module import**, not at first call, so the
failure is a startup failure rather than a per-request one.
---
# Part A β The frozen encoder
## 2. `FrozenEncoder` β the contract, probed not assumed
`router/encoder.py` opens with the verified contract, and the phrase *verified* is literal: these
were probed, not read from documentation.
```
Verified contract (Phase 4 probe, 2026-09-16, sentence-transformers 6.0.1):
SentenceTransformer(
model_name_or_path='sentence-transformers/all-MiniLM-L6-v2',
revision='1110a243fdf4', # pinned, verified reachable
device='cpu',
)
.get_sentence_embedding_dimension() -> 384
.tokenizer.model_max_length -> 256 (finding F4-1)
.max_seq_length -> 256
params -> 22,713,216
encode(64 queries, CPU) -> 0.118 s
```
(`router/encoder.py:3-14`.)
| Property | Value | Source |
|---|---|---|
| Model | `sentence-transformers/all-MiniLM-L6-v2` | `configs/base.yaml` Β§`router.model`; `metadata.json.encoder.model` |
| Revision | `1110a243fdf4` | `configs/base.yaml` Β§`router.revision`; `metadata.json.encoder.revision` |
| Embedding dimension | **384** | probe; `VERIFIED_EMBEDDING_DIM = 384` |
| Tokenizer ceiling | **256** | probe; `VERIFIED_TOKENIZER_MAX_LENGTH = 256` |
| Configured `max_length` | **128** | `configs/base.yaml` Β§`router.max_length`; `metadata.json.encoder.max_length` |
| Parameters | **22,713,216** | probe; `metadata.json.encoder.parameters` |
| Approximate on-disk size | 90.9 MB | `configs/base.yaml` comment on the revision pin |
| Throughput | 0.118 s for 64 queries, CPU | probe |
| Device | `cpu` (config default `auto` β `config.device_preference`) | `router/encoder.py:168-182` |
### 2.1 The two module constants
```python
#: The tokenizer's own ceiling, verified by probe. Truncating above this is a
#: silent no-op, so the encoder refuses rather than pretending.
VERIFIED_TOKENIZER_MAX_LENGTH = 256
#: Verified embedding dimension for all-MiniLM-L6-v2.
VERIFIED_EMBEDDING_DIM = 384
```
(`router/encoder.py:30-35`.)
## 3. Finding F4-1 β `max_length` 128 is not the model's limit
This is the most-quoted finding in the router, and it is worth stating precisely, because the
distinction between *the model's ceiling* and *our truncation* is exactly the kind of thing that
silently rots.
**What the config said.** `max_length: 128`.
**What the probe found.** The MiniLM tokenizer's own ceiling is **256**
(`tokenizer.model_max_length β 256`, `max_seq_length β 256`).
**Why 128 and not 256.** The truncation is deliberate and sits *well inside* the ceiling. The
configured rationale, verbatim from `configs/base.yaml`:
```yaml
# Finding F4-1: the MiniLM tokenizer's own ceiling is 256 (verified by probe).
# 128 is a deliberate truncation well inside that ceiling, not the model
# limit. Satellite queries are short; halving the sequence halves attention
# cost for no measurable accuracy loss. The encoder asserts this value is
# <= 256, because truncating above the ceiling is a silent no-op.
max_length: 128
```
**Why the encoder must *refuse* rather than merely document.** Truncating above the tokenizer's
ceiling is a **silent no-op** β the tokenizer will not truncate past what it can encode, so a
`max_length` of 512 would appear to be honoured while doing nothing at all. That is a control that
looks like it works and does not. `FrozenEncoder.__init__` therefore raises:
```python
if max_length < 1:
raise ModelLoadError(f"max_length must be >= 1, got {max_length}")
if max_length > VERIFIED_TOKENIZER_MAX_LENGTH:
raise ModelLoadError(
f"max_length={max_length} exceeds the MiniLM tokenizer ceiling of "
f"{VERIFIED_TOKENIZER_MAX_LENGTH}; truncation would be a silent no-op"
)
```
(`router/encoder.py:54-60`.)
**The subtlety most readers miss.** The configured value is not merely validated β it is *applied*.
Setting `max_seq_length` rewrites the tokenizer's truncation limit, so the config value is enforced
rather than documented:
```python
# Apply the configured truncation for real (finding F4-1). Setting
# max_seq_length rewrites the tokenizer's truncation limit, so the
# config value is enforced rather than merely documented.
self._model.max_seq_length = max_length
```
(`router/encoder.py:96-99`.)
There is then a second, quieter guard: some tokenizers report a **sentinel** for "no limit"
(a very large integer). The class normalises that sentinel back to the verified ceiling so
`tokenizer_max_length` never reports a nonsense number:
```python
tokenizer_limit = int(getattr(self._model.tokenizer, "model_max_length", 0))
if tokenizer_limit and tokenizer_limit > 10**6:
# Some tokenizers report a sentinel for "no limit".
tokenizer_limit = VERIFIED_TOKENIZER_MAX_LENGTH
self._tokenizer_max_length = tokenizer_limit
```
(`router/encoder.py:101-105`.)
**Status.** `MEASURED` β the probe is recorded with its date (2026-09-16) and its
sentence-transformers version (6.0.1) in the module docstring. The **assertion** is `IMPLEMENTED`
and its refusal path is covered by the routing test suite
(`tests/routing/test_router.py`, 82 tests, per `docs/PHASE4_ROUTER_REPORT.md` Β§Files).
## 4. The freeze β the whole architectural premise
The encoder is not fine-tuned. It is loaded, put in eval mode, and every parameter's gradient is
switched off:
```python
# Freeze. This is the whole architectural premise.
self._model.eval()
for param in self._model.parameters():
param.requires_grad_(False)
```
(`router/encoder.py:91-94`.)
The class docstring states what this buys:
> Deliberately narrow: it encodes text and reports its dimensions. It does not train, does not
> expose the underlying model, and does not permit gradient flow. Everything downstream treats it
> as a pure function.
(`router/encoder.py:40-43`.)
Because the encoder is frozen, **embeddings are a pure function of the query text**. That single
fact is what makes finding F4-2 possible (Β§8).
### 4.1 `encode()` β the inference surface
```python
def encode(
self,
texts: Sequence[str],
batch_size: int = 64,
normalize: bool = True,
) -> np.ndarray:
"""Encode texts to a (n, embedding_dim) float32 array.
The returned array is detached: it carries no autograd history. That is
intentional β the adapter consumes it as a fixed input feature.
"""
if not texts:
return np.zeros((0, self.embedding_dim), dtype=np.float32)
cleaned = [t if isinstance(t, str) else str(t) for t in texts]
try:
embeddings = self._model.encode(
cleaned,
batch_size=batch_size,
show_progress_bar=False,
convert_to_numpy=True,
normalize_embeddings=normalize,
)
except Exception as exc: # noqa: BLE001
raise ModelLoadError(
f"encoder inference failed: {exc}", specialist="router"
) from exc
return np.asarray(embeddings, dtype=np.float32)
```
(`router/encoder.py:127-155`.) Four details are worth naming:
1. **Empty input returns an empty array of the right shape**, not an error β
`np.zeros((0, 384))`. A caller batching over a filtered list does not need a special case.
2. **Non-string inputs are coerced** rather than rejected (`str(t)`). The corpus is
`RouterExample` objects, and `RouterCorpus.texts()` already yields strings; the coercion is a
boundary guard.
3. **`normalize_embeddings=normalize`** β the default is `True`, so the vectors are unit-length.
Nothing downstream re-normalises.
4. **`show_progress_bar=False`** is pinned, so training output is not polluted by a progress bar
on every cache miss.
`encode_one` is the single-query convenience wrapper:
```python
def encode_one(self, text: str, normalize: bool = True) -> np.ndarray:
return self.encode([text], batch_size=1, normalize=normalize)[0]
```
(`router/encoder.py:157-158`.) This is what `IntentRouter._predict_learned` calls.
### 4.2 The four read-only properties
| Property | Returns | Backing |
|---|---|---|
| `embedding_dim` | `int` | `self._model.get_sentence_embedding_dimension()` β **asked of the model**, not read from config |
| `tokenizer_max_length` | `int` | the normalised sentinel-aware value (Β§3) |
| `load_seconds` | `float \| None` | wall time of the `SentenceTransformer(...)` construction |
| `num_parameters` | `int` | `sum(p.numel() for p in self._model.parameters())` |
The `embedding_dim` property deliberately queries the model rather than trusting
`router.embedding_dim` from the config. `IntentRouter.from_config` then **compares the two**:
```python
encoder = build_encoder(config, device=device)
if adapter is not None and adapter.input_dim != encoder.embedding_dim:
raise ModelLoadError(
f"adapter input_dim={adapter.input_dim} but the loaded encoder "
f"produces {encoder.embedding_dim}-d embeddings",
specialist="router",
)
```
(`router/classifier.py:179-185`.) A config that says 384 while the loaded encoder emits something
else fails loudly at construction rather than producing garbage at inference.
### 4.3 `build_encoder()` β config-driven construction with a device fallback
```python
def build_encoder(config, device: str | None = None) -> FrozenEncoder:
"""Construct the encoder from the central configuration registry."""
if device is None:
configured = config.get("router.device", "auto")
if configured == "auto":
device = config.device_preference
else:
device = configured
return FrozenEncoder(
model_name=config.require("router.model"),
revision=config.require("router.revision"),
max_length=int(config.get("router.max_length", 128)),
device=device,
)
```
(`router/encoder.py:168-182`.)
Three points:
- **`model` and `revision` use `config.require`, not `config.get`.** A missing model name or
revision is a hard failure β the pin is not optional.
- **`max_length` defaults to 128** if the key is absent, so the encoder degrades to the configured
behaviour rather than to the ceiling.
- **`device: auto` resolves through `config.device_preference`.** This is how a CPU-only Space and a
GPU host share one config file. `configs/base.yaml` sets `router.device: auto`.
## 5. Finding F4-2 β the router needs no GPU
The finding, stated by the training module itself:
> The encoder is FROZEN. Embeddings are therefore a pure function of the query text. So we embed
> the whole corpus ONCE (measured: 0.118 s / 64 queries on CPU), cache the vectors to disk, and
> train the 50,822-parameter adapter on the cached matrix (measured: 20 epochs / 4,096 vectors in
> 0.28 s).
>
> Router training needs NO GPU. The plan's Kaggle budget ("Router | CPU/T4 | <1 h") is roughly
> three orders of magnitude pessimistic. Phase 4 runs and completes locally.
(`router/train.py:3-12`.)
`docs/PHASE4_ROUTER_REPORT.md` Β§F4-2 repeats it with the same numbers and adds the outcome:
**"The plan's budget β 'Router | CPU/T4 | <1 h' β was roughly three orders of magnitude
pessimistic. Phase 4 ran to completion locally at zero quota cost."**
**Measured evidence for the claim, and its limits.**
| Quantity | Value | Source | Status |
|---|---|---|---|
| Encoder throughput | 0.118 s / 64 queries, CPU | probe, `router/encoder.py:14` | `MEASURED` |
| Adapter training | 20 epochs / 4,096 vectors in 0.28 s | probe, `router/encoder.py:22-23` | `MEASURED` |
| **Full 60-epoch run, 576-example corpus** | **4.92 s total** | `metadata.json.duration_seconds: 4.92` | `MEASURED` |
| GPU used | **none** | `threshold_sweep_val.json.environment.cuda_available: false` | `MEASURED` |
The 0.28 s / 20 epochs figure is the **probe**, on a synthetic 4,096 Γ 384 matrix. The **real** run
is the 4.92 s in the shipped artifact β a different and larger quantity (60 epochs, cache
fingerprint computation, three split evaluations, artifact write). Both are true; they are not the
same measurement, and this document does not merge them.
**Caveat, stated because the docstring itself states it.** The `router/train.py` docstring says
"the 50,822-parameter adapter". The shipped artifact records **51,725**. Β§8 resolves this by
arithmetic; it is a documentation discrepancy in a comment, not a second model.
---
# Part B β The five-head adapter
## 6. `IntentAdapter` β the only trainable part of the router
`router/adapter.py` opens with the architecture as an ASCII diagram, which is the clearest
statement of the shape anywhere in the repository:
```
embedding (384)
|
LayerNorm
|
Linear(384 -> hidden_dim) default hidden_dim = 128
|
GELU
|
Dropout
|
+--> task_head Linear(hidden, 6)
+--> modality_head Linear(hidden, 4)
+--> temporal_head Linear(hidden, 1) logit
+--> spatial_head Linear(hidden, 1) logit
+--> language_head Linear(hidden, 1) logit
```
(`router/adapter.py:5-20`.)
Read that diagram carefully: **there is exactly one shared trunk, and the five heads all read from
it.** The heads are independent of one another β no head consumes another head's output. That
independence is what makes the coherence-repair logic in Β§21 necessary.
### 6.1 The constructor, in full
```python
def __init__(
self,
input_dim: int = 384,
hidden_dim: int = 128,
dropout: float = 0.1,
num_tasks: int = NUM_TASKS,
num_modalities: int = NUM_MODALITIES,
) -> None:
super().__init__()
if input_dim < 1 or hidden_dim < 1:
raise ValueError(
f"input_dim and hidden_dim must be positive, got {input_dim}, {hidden_dim}"
)
if not 0.0 <= dropout < 1.0:
raise ValueError(f"dropout must be in [0, 1), got {dropout}")
self.input_dim = input_dim
self.hidden_dim = hidden_dim
self.dropout_p = dropout
self.num_tasks = num_tasks
self.num_modalities = num_modalities
self.input_norm = nn.LayerNorm(input_dim)
self.trunk = nn.Sequential(
nn.Linear(input_dim, hidden_dim),
nn.GELU(),
nn.Dropout(dropout),
)
self.task_head = nn.Linear(hidden_dim, num_tasks)
self.modality_head = nn.Linear(hidden_dim, num_modalities)
# One logit per binary head. P(yes) = sigmoid(logit).
self.temporal_head = nn.Linear(hidden_dim, 1)
self.spatial_head = nn.Linear(hidden_dim, 1)
self.language_head = nn.Linear(hidden_dim, 1)
self._init_weights()
```
(`router/adapter.py:64-102`.)
**Every sub-module, exhaustively.**
| Attribute | Type | Shape / config | Role |
|---|---|---|---|
| `input_norm` | `nn.LayerNorm` | normalised shape `(input_dim,)` = `(384,)` | normalises the frozen embedding before the trunk |
| `trunk` | `nn.Sequential` | `Linear(384β128) β GELU β Dropout(0.1)` | the single shared representation |
| `task_head` | `nn.Linear` | `(128 β 6)` | softmax over `TASK_CLASSES` |
| `modality_head` | `nn.Linear` | `(128 β 4)` | softmax over `MODALITY_CLASSES` |
| `temporal_head` | `nn.Linear` | `(128 β 1)` | one logit; `P(yes) = sigmoid(logit)` |
| `spatial_head` | `nn.Linear` | `(128 β 1)` | one logit |
| `language_head` | `nn.Linear` | `(128 β 1)` | one logit |
`dropout_p` is stored under that name rather than `dropout` so it does not shadow the
`nn.Dropout` module convention; it is what `config_dict()` serialises.
**Two validation guards, both at construction.** `input_dim` and `hidden_dim` must be positive;
`dropout` must be in `[0, 1)` β note the half-open interval, so `dropout=1.0` (which would zero
the trunk entirely) is rejected.
### 6.2 `AdapterOutput` β raw logits, no activation
```python
@dataclass
class AdapterOutput:
"""Raw logits from all five heads. No softmax applied here."""
task_logits: torch.Tensor # (B, NUM_TASKS)
modality_logits: torch.Tensor # (B, NUM_MODALITIES)
temporal_logit: torch.Tensor # (B,)
spatial_logit: torch.Tensor # (B,)
language_logit: torch.Tensor # (B,)
def binary_logit(self, head: str) -> torch.Tensor:
if head == "temporal":
return self.temporal_logit
if head == "spatial_output":
return self.spatial_logit
if head == "language_output":
return self.language_logit
raise KeyError(f"unknown binary head: {head!r}")
```
(`router/adapter.py:41-58`.)
Three design facts:
1. **The two multi-class heads return `(B, C)` logits; the three binary heads return `(B,)`.** The
`squeeze(-1)` in `forward` is what removes the trailing unit dimension. This matters for loss
wiring: `BCEWithLogitsLoss` wants `(B,)` targets, and `nn.CrossEntropyLoss` wants `(B, C)`
logits with `(B,)` integer targets.
2. **No activation is applied.** `softmax` and `sigmoid` are applied by the consumer β the loss
during training, and `_predict_learned` at inference. This keeps the numerically-stable
`*_with_logits` losses usable.
3. **`binary_logit(head)` keys on the label-space names** β `"temporal"`, `"spatial_output"`,
`"language_output"` β which are exactly the members of `BINARY_HEADS`
(`router/label_space.py:63`). `evaluate_split` and the training loop both iterate
`for head in BINARY_HEADS` and call this method, so the attribute names `spatial_logit` and the
key `"spatial_output"` are reconciled in one place instead of at every call site.
### 6.3 Weight initialisation β and why it is not the default
```python
def _init_weights(self) -> None:
"""Small-std init on the heads keeps the initial sigmoid near 0.5.
Without this the binary heads can start saturated, and BCE gradients
vanish before the task head has learned anything useful.
"""
for module in (self.task_head, self.modality_head,
self.temporal_head, self.spatial_head, self.language_head):
nn.init.normal_(module.weight, std=0.02)
nn.init.zeros_(module.bias)
```
(`router/adapter.py:104-113`.)
This is a deliberately non-default initialisation, and the reason is a real failure mode:
a single-logit binary head whose weights start large can produce a saturated sigmoid on the first
forward pass, and a saturated sigmoid has a vanishing BCE gradient. The task head would then be
learning against a trunk that is receiving almost no signal from half its objectives.
Note precisely **what** is re-initialised: the **five heads only**. `input_norm` and `trunk` keep
PyTorch's defaults. The init is applied to `weight` (normal, Ο = 0.02) and `bias` (zeros).
**Cross-reference.** `router/adapter.py:22-23` records the probe cost as "50,822 parameters".
The arithmetic in Β§8 gives 51,725. The `std=0.02` figure in this chapter is read from
`router/adapter.py:112`; it is **not** the same as the `20.0` objectness weight in the grounding
head (`specialists/grounding/head.py`), which is a different subsystem and a different quantity.
### 6.4 `forward()` β and its two shape guards
```python
def forward(self, embeddings: torch.Tensor) -> AdapterOutput:
"""Args:
embeddings: (B, input_dim) float tensor. Must already be detached
from the frozen encoder β the adapter does not back-propagate
into MiniLM.
"""
if embeddings.dim() != 2:
raise ValueError(
f"expected a 2-D (batch, dim) tensor, got shape {tuple(embeddings.shape)}"
)
if embeddings.shape[1] != self.input_dim:
raise ValueError(
f"expected input_dim={self.input_dim}, got {embeddings.shape[1]}"
)
hidden = self.trunk(self.input_norm(embeddings))
return AdapterOutput(
task_logits=self.task_head(hidden),
modality_logits=self.modality_head(hidden),
temporal_logit=self.temporal_head(hidden).squeeze(-1),
spatial_logit=self.spatial_head(hidden).squeeze(-1),
language_logit=self.language_head(hidden).squeeze(-1),
)
```
(`router/adapter.py:117-140`.)
**The two guards are not decoration.** A 1-D tensor would silently broadcast through the
`Linear` layers and produce a plausible-looking but wrong output; a mismatched feature width
would produce a `RuntimeError` deep inside `nn.Linear` with a shape trace that does not name the
cause. Both are caught here with a message that names the actual and expected dims.
**The order of operations is fixed: normalise, then trunk.** `self.trunk(self.input_norm(...))` β
LayerNorm is applied to the frozen embedding, *then* the linear projection. Reversing them would
normalise the hidden layer instead, which is a different model.
### 6.5 Serialisation β the adapter carries its own architecture
```python
def config_dict(self) -> dict[str, Any]:
return {
"input_dim": self.input_dim,
"hidden_dim": self.hidden_dim,
"dropout": self.dropout_p,
"num_tasks": self.num_tasks,
"num_modalities": self.num_modalities,
"binary_heads": list(BINARY_HEADS),
}
@classmethod
def from_config_dict(cls, payload: dict[str, Any]) -> "IntentAdapter":
return cls(
input_dim=int(payload["input_dim"]),
hidden_dim=int(payload["hidden_dim"]),
dropout=float(payload["dropout"]),
num_tasks=int(payload["num_tasks"]),
num_modalities=int(payload["num_modalities"]),
)
def num_parameters(self) -> int:
return sum(p.numel() for p in self.parameters())
```
(`router/adapter.py:144-165`.)
`config_dict()` is written **into the checkpoint** by `save_adapter`, and `from_config_dict()` is
what makes `load_adapter` able to reconstruct the architecture without being told it. The
checkpoint format is:
```python
torch.save(
{
"state_dict": adapter.state_dict(),
"config": adapter.config_dict(),
},
path / ADAPTER_WEIGHTS,
)
```
(`router/classifier.py:362-368`.) `load_adapter` then **refuses** a checkpoint with no embedded
config, rather than guessing:
```python
config = payload.get("config")
if not config:
raise ModelLoadError(
"adapter checkpoint has no embedded config; it cannot be reconstructed "
"without guessing the architecture",
specialist="router",
)
```
(`router/classifier.py:438-444`.) And it loads with `strict=True`, so a state dict that does not
match the reconstructed architecture is an error, not a partial load:
```python
adapter = IntentAdapter.from_config_dict(config)
try:
adapter.load_state_dict(payload["state_dict"], strict=True)
except Exception as exc: # noqa: BLE001
raise ModelLoadError(
f"adapter state_dict does not match the reconstructed architecture: {exc}",
specialist="router",
) from exc
```
(`router/classifier.py:446-453`.)
**The `binary_heads` key asymmetry.** `config_dict()` *emits* `"binary_heads"`, but
`from_config_dict()` does **not** read it. That is intentional and correct: the binary heads are
three fixed `Linear(128, 1)` layers, so their names are not a reconstructable parameter β they are
a label-space invariant. The key is written for **auditability** (a reader of `metadata.json` can
see which binary heads the artifact was built for) rather than for reconstruction. The shipped
`metadata.json` records it twice, under both `adapter` and `adapter_config`:
```json
"binary_heads": ["temporal", "spatial_output", "language_output"],
"dropout": 0.1,
"hidden_dim": 128,
"input_dim": 384,
"num_modalities": 4,
"num_parameters": 51725,
"num_tasks": 6
```
## 7. The label space, exhaustively
`router/label_space.py` is the single source of truth for the router's output ontology. Its
docstring states why it is a module rather than a set of constants duplicated per file:
> Single source of truth for the router's output ontology. Both the dataset generator and the
> training script import from here, so a class added in one place cannot silently desynchronise
> the other.
(`router/label_space.py:1-6`.)
### 7.1 The task head β six classes, in index order
```python
TASK_CLASSES: tuple[str, ...] = (
"vqa",
"caption",
"grounding",
"change",
"optical_sar",
"unsupported",
)
TASK_TO_INDEX: dict[str, int] = {name: i for i, name in enumerate(TASK_CLASSES)}
NUM_TASKS: int = len(TASK_CLASSES)
```
(`router/label_space.py:26-35`.)
**The tuple order is the index order.** `TASK_TO_INDEX` is derived from the tuple by `enumerate`,
and `train.py` uses `TASK_TO_INDEX[e.task]` to build targets while `evaluate_split` uses
`TASK_CLASSES[i]` to name predictions. Reordering the tuple would silently re-map every label β
which is why the tuple is defined once and everything else is derived.
| Index | Label | Means | Configured in `base.yaml`? |
|---|---|---|---|
| 0 | `vqa` | a free-form question about a single scene | yes, `router.tasks[0]` |
| 1 | `caption` | a description of a single scene | yes, `router.tasks[1]` |
| 2 | `grounding` | *where* β a spatial output | yes, `router.tasks[2]` |
| 3 | `change` | *what changed* between two acquisitions | yes, `router.tasks[3]` |
| 4 | `optical_sar` | joint classification from an optical + SAR pair | yes, `router.tasks[4]` |
| 5 | `unsupported` | not a satellite-imagery question | yes, `router.tasks[5]` |
The config repeats the same six in the same order under `router.tasks`
(`configs/base.yaml` Β§`router.tasks`), so a reader of the config sees the ontology without opening
Python.
### 7.2 The three task families
```python
#: Tasks that inherently require two acquisitions.
TEMPORAL_TASKS: frozenset[str] = frozenset({"change"})
#: Tasks that inherently require two modalities.
DUAL_MODALITY_TASKS: frozenset[str] = frozenset({"optical_sar"})
#: Tasks whose primary output is spatial rather than textual.
SPATIAL_TASKS: frozenset[str] = frozenset({"grounding"})
```
(`router/label_space.py:37-44`.)
Each is a `frozenset` with exactly one member today. They are sets rather than scalars so that
adding a second member does not require touching the consumers, and they are read by
`default_attributes` below.
### 7.3 The modality head β four classes
```python
MODALITY_CLASSES: tuple[str, ...] = (
"optical",
"sar",
"optical_sar",
"unknown",
)
MODALITY_TO_INDEX: dict[str, int] = {name: i for i, name in enumerate(MODALITY_CLASSES)}
NUM_MODALITIES: int = len(MODALITY_CLASSES)
```
(`router/label_space.py:50-57`.) Four classes, and `unknown` is a first-class member rather than a
sentinel β a query with no modality marker is genuinely `unknown`, not "optical by default".
### 7.4 The binary heads β three names
```python
BINARY_HEADS: tuple[str, ...] = ("temporal", "spatial_output", "language_output")
NUM_BINARY_HEADS: int = len(BINARY_HEADS)
```
(`router/label_space.py:63-64`.) The docstring explains the single-logit choice:
> `task` and `modality` use softmax cross-entropy. The three binary heads use a single logit with
> `BCEWithLogitsLoss` β a 2-way softmax would waste a parameter and make the loss harder to
> weight.
(`router/label_space.py:15-17`.) `README.md:478-479` repeats the reasoning.
**The names are not the same as the head attributes.** `BINARY_HEADS` uses `spatial_output` and
`language_output`; the module attributes are `spatial_head` and `language_head`. `AdapterOutput.
binary_logit()` (Β§6.2) is the one place that reconciles them.
### 7.5 The two helper functions
```python
def default_attributes(task: str) -> dict[str, bool]:
"""Per-task attribute defaults, used to sanity-check generated examples.
These are *defaults*, not invariants: `change` with `spatial_output=False`
is a perfectly legal request ("what changed?"). They exist so the dataset
generator and the fallback agree on what a task usually implies.
"""
return {
"temporal": task in TEMPORAL_TASKS,
"spatial_output": task in SPATIAL_TASKS,
"language_output": task != "unsupported",
}
def default_modality(task: str) -> str:
if task in DUAL_MODALITY_TASKS:
return "optical_sar"
if task == "unsupported":
return "unknown"
return "unknown"
```
(`router/label_space.py:89-108`.)
Note `default_modality`'s **third branch is redundant**: `task == "unsupported"` already returns
`"unknown"`, and the final `return "unknown"` covers it. The function is equivalent to
`"optical_sar" if task in DUAL_MODALITY_TASKS else "unknown"`. This is a cosmetic redundancy, not a
defect β the two branches produce the same value β and it is recorded here only because a reader
diffing the branches will notice it.
Also note the docstring's own warning: these are **defaults, not invariants**. `default_attributes`
is *not* what `RouterCorpus.validate()` enforces; the validator enforces a stricter and different
set (Β§11.2).
### 7.6 The module's remaining surface
`task_index(name)`, `modality_index(name)`, `is_valid_task(name)` and `is_valid_modality(name)` are
thin lookups. `task_index` and `modality_index` **raise `KeyError`** on an unknown label, whereas
`is_valid_*` return a bool. `router/fallback.py` uses the `is_valid_*` pair inside `self_check()` to
assert that no rule ever emits a label outside the space.
## 8. Parameter accounting β 51,725, and why a comment says 50,822
The shipped artifact records `adapter.num_parameters: 51725`
(`artifacts/router/router_adapter_v001/metadata.json`), and `docs/PHASE4_ROUTER_REPORT.md:37`
agrees: "51,725 adapter params on a frozen 22,713,216-param encoder". Two other places say
**50,822**: the `router/adapter.py` module docstring (`router/adapter.py:22`) and the
`router/train.py` docstring (`router/train.py:8`), plus `configs/base.yaml`'s comment on F4-2.
The arithmetic is unambiguous, and it settles the question:
| Sub-module | Parameter count | Working |
|---|---|---|
| `input_norm` (`LayerNorm(384)`) | **768** | `384 Γ 2` (weight + bias) |
| `trunk[0]` (`Linear(384 β 128)`) | **49,280** | `384 Γ 128` weights + `128` bias |
| `task_head` (`Linear(128 β 6)`) | **774** | `128 Γ 6` weights + `6` bias |
| `modality_head` (`Linear(128 β 4)`) | **516** | `128 Γ 4` weights + `4` bias |
| `temporal_head` (`Linear(128 β 1)`) | **129** | `128` weights + `1` bias |
| `spatial_head` (`Linear(128 β 1)`) | **129** | `128` weights + `1` bias |
| `language_head` (`Linear(128 β 1)`) | **129** | `128` weights + `1` bias |
| **Total** | **51,725** | 768 + 49,280 + 774 + 516 + 129 + 129 + 129 |
The `nn.GELU()` and `nn.Dropout()` layers in the trunk contribute **zero** parameters, which is why
the trunk's total (49,280) equals its `Linear` layer alone.
**Status.** The 51,725 figure is `MEASURED` β it is what `num_parameters()` returns for the shipped
weights, and it is what the artifact records. The 50,822 figure is a **stale comment**; it is
reported here rather than quietly corrected, because the docs are public and a reader who greps
the repository will find both. The README's F4-2 paragraph still carries 50,822
(`README.md:508`). Treating this as an **`OPEN` documentation discrepancy** is the honest
description: the *code* is right, three *comments* are stale, and no behaviour depends on either
number.
## 9. The loss β three objectives, weighted
The training loss composes three terms. The docstring names them:
```
task cross-entropy, class-weighted (classes are imbalanced)
modality cross-entropy
binary x3 BCE-with-logits, one logit per head
Weights live in config under `router.training`.
```
(`router/train.py:22-28`.)
The implementation, verbatim from the inner loop:
```python
loss_task = ce(out.task_logits, train_tensors["task"][idx])
loss_mod = ce_mod(out.modality_logits, train_tensors["modality"][idx])
binary_targets = train_tensors["binary"][idx]
loss_bin = (
bce(out.temporal_logit, binary_targets[:, 0])
+ bce(out.spatial_logit, binary_targets[:, 1])
+ bce(out.language_logit, binary_targets[:, 2])
) / len(BINARY_HEADS)
loss = (
task_loss_weight * loss_task
+ modality_loss_weight * loss_mod
+ binary_loss_weight * loss_bin
)
```
(`router/train.py:606-620`.)
**Three details that matter.**
1. **The binary term is the *mean* over the three heads**, not the sum. The division by
`len(BINARY_HEADS)` keeps the term on the same scale as the two cross-entropies, so
`binary_loss_weight` means the same thing regardless of how many binary heads exist.
2. **The binary targets are read by column index, in `BINARY_HEADS` order.** `binary_targets[:, 0]`
is `temporal`, `[:, 1]` is `spatial_output`, `[:, 2]` is `language_output` β the order comes from
`_to_tensors`, which builds the matrix by iterating `BINARY_HEADS`
(`router/train.py:301-305`). The two orders are the same tuple, so they cannot drift.
3. **Only the task cross-entropy is class-weighted.** `ce` carries `weight=class_weights`;
`ce_mod` is an unweighted `nn.CrossEntropyLoss()` and `bce` is a plain `BCEWithLogitsLoss()`
(`router/train.py:582-584`). The class imbalance that motivated weighting is a *task-class*
imbalance; the modality and binary targets are not weighted.
**Gradient clipping.** Every step clips the global norm to 1.0:
```python
optimizer.zero_grad(set_to_none=True)
loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
optimizer.step()
```
(`router/train.py:622-625`.) `set_to_none=True` frees the gradient buffers rather than zeroing
them, which is both faster and stricter β a parameter that receives no gradient ends with `None`,
not a zero tensor.
**The configured weights** (`configs/base.yaml` Β§`router.training`; `metadata.json.hyperparameters`):
| Key | Value | Effect |
|---|---|---|
| `task_loss_weight` | `1.0` | the primary objective, full weight |
| `modality_loss_weight` | `0.3` | auxiliary; the modality label is often `unknown` |
| `binary_loss_weight` | `0.5` | auxiliary; three heads, mean-reduced |
The ordering β task 1.0 > binary 0.5 > modality 0.3 β encodes the priority: the task label is what
the controller keys off; the binary heads describe aspects of the request; the modality label is
the least informative head for this corpus.
---
# Part C β The corpus
## 10. `RouterExample` β one labelled query, with a leakage boundary
```python
@dataclass
class RouterExample:
"""One labelled query.
`group` is the leakage boundary: all examples sharing a group must land in
the same split. `source` distinguishes hand-written from generated material
so the two can be reported separately.
"""
text: str
task: str
modality: str = "unknown"
temporal: bool = False
spatial_output: bool = False
language_output: bool = True
group: str = ""
source: str = "template"
def __post_init__(self) -> None:
if self.task not in TASK_TO_INDEX:
raise ValueError(f"unknown task label: {self.task!r}")
if self.modality not in MODALITY_TO_INDEX:
raise ValueError(f"unknown modality label: {self.modality!r}")
if not self.text.strip():
raise ValueError("example text must not be empty")
if not self.group:
raise ValueError(f"example {self.text!r} has no group tag")
```
(`router/dataset.py:87-113`.)
| Field | Type | Default | Meaning | Constraint |
|---|---|---|---|---|
| `text` | `str` | β | the query as a user would type it | must be non-empty after `strip()` |
| `task` | `str` | β | one of `TASK_CLASSES` | validated against `TASK_TO_INDEX` |
| `modality` | `str` | `"unknown"` | one of `MODALITY_CLASSES` | validated against `MODALITY_TO_INDEX` |
| `temporal` | `bool` | `False` | needs two acquisitions | β |
| `spatial_output` | `bool` | `False` | the answer must contain coordinates | β |
| `language_output` | `bool` | `True` | the answer must contain prose | β |
| `group` | `str` | `""` | **the leakage boundary** | **must be non-empty** β `__post_init__` refuses a blank group |
| `source` | `str` | `"template"` | `"curated"` or `"template"` | not validated (a free string) |
**The `group` requirement is the load-bearing one.** A blank group would make `split_by_group`
unable to guarantee that a near-duplicate pair travels together, so it is a construction-time
error rather than a split-time surprise.
`label_tuple()` is the training-time projection of an example into tensors:
```python
def label_tuple(self) -> tuple[int, int, tuple[float, ...]]:
"""(task_index, modality_index, (temporal, spatial, language)) as floats."""
binaries = tuple(
1.0 if getattr(self, head) else 0.0 for head in BINARY_HEADS
)
return TASK_TO_INDEX[self.task], MODALITY_TO_INDEX[self.modality], binaries
```
(`router/dataset.py:115-120`.) Note `getattr(self, head)` β the binary head names in
`BINARY_HEADS` (`temporal`, `spatial_output`, `language_output`) are **exactly the dataclass field
names**, so no mapping table is needed. `to_dict()` serialises all eight fields
(`router/dataset.py:122-132`).
## 11. `RouterCorpus` β the container and its four checks
```python
@dataclass
class RouterCorpus:
"""A labelled, group-tagged corpus ready for splitting and caching."""
examples: list[RouterExample] = field(default_factory=list)
```
(`router/dataset.py:389-393`.)
| Method | Returns | What it does |
|---|---|---|
| `__len__` | `int` | number of examples |
| `__iter__` | iterator | iterates `examples` |
| `texts()` | `list[str]` | `[e.text for e in self.examples]` β the encoder's input |
| `groups()` | `set[str]` | the distinct group ids |
| `task_counts()` | `dict[str, int]` | **initialised over all six `TASK_CLASSES`**, so an absent class reports `0` rather than being missing |
| `source_counts()` | `dict[str, int]` | counts by `source` (only keys that occur) |
| `positives(head)` | `int` | `sum(1 for e in self.examples if getattr(e, head))` |
| `dedupe()` | `(RouterCorpus, list[str])` | removes exact duplicate texts; returns conflicts |
| `validate()` | `list[str]` | structural checks; empty list means clean |
**`task_counts()` initialising every class to zero is deliberate.** A missing key and a zero count
are different things, and a validator that iterates `counts[name]` would raise `KeyError` on an
absent class instead of reporting it. Because the dict is pre-seeded, `validate()` can report
`"task class 'vqa' has no examples"` instead of crashing.
### 11.1 `dedupe()` β and why a duplicate is a conflict, not a silent merge
```python
def dedupe(self) -> tuple["RouterCorpus", list[str]]:
"""Remove exact duplicate texts, keeping the first occurrence.
A query appearing under two different labels is a labelling conflict and
is reported rather than silently resolved.
"""
seen: dict[str, str] = {}
kept: list[RouterExample] = []
conflicts: list[str] = []
for example in self.examples:
key = example.text.strip().lower()
if key in seen:
if seen[key] != example.task:
conflicts.append(
f"{example.text!r}: {seen[key]} vs {example.task}"
)
continue
seen[key] = example.task
kept.append(example)
return RouterCorpus(kept), conflicts
```
(`router/dataset.py:422-443`.)
Three behaviours to name:
1. **The comparison key is `text.strip().lower()`** β case- and whitespace-insensitive. So
`"What changed?"` and `"what changed? "` are the same example.
2. **The first occurrence wins.** The corpus is assembled with `CURATED` first
(`router/dataset.py:525-526`), so a hand-checked curated example always beats a generated
template that happens to produce the same text.
3. **Same text, different task β a reported conflict.** `train_router` turns any conflict into a
hard failure:
```python
corpus, conflicts = corpus.dedupe()
if conflicts:
raise RoutingError(
"router corpus contains labelling conflicts (same text, two tasks):\n - "
+ "\n - ".join(conflicts)
)
```
(`router/train.py:488-493`.) This is why the hard-negative families are legal: they contain the
*same text* under *different groups* only when the texts differ (`"Describe the water body."`
vs `"Show me the water body."`). A genuine same-text-two-labels pair would abort training.
### 11.2 `validate()` β five structural checks
```python
def validate(self) -> list[str]:
"""Structural sanity checks. Returns a list of problems (empty = good)."""
```
(`router/dataset.py:445-446`.) The checks, in order:
**(a) The corpus is non-empty.** An empty corpus returns immediately with a single problem.
**(b) Groups may span tasks only if they are hard-negative families.**
```python
spanning = {
group: sorted({e.task for e in examples})
for group, examples in _group_index(self).items()
if len({e.task for e in examples}) > 1
}
for group, tasks in spanning.items():
if not group.startswith(HARD_NEGATIVE_PREFIX):
problems.append(
f"group {group!r} spans multiple tasks {tasks} but is not a "
f"hard-negative family (expected prefix "
f"{HARD_NEGATIVE_PREFIX!r}); either split it or rename it"
)
```
(`router/dataset.py:465-476`.) The comment above it is the clearest statement of the design:
> A group is a LEAKAGE boundary, not a label constraint: every example in a group goes to the same
> split, whatever its label. That is exactly what the hard-negative families need β
> `hn_desc_vs_show` contains both "Describe the water body." (caption) and "Show me the water body."
> (grounding), and splitting those two apart would destroy the whole point of the pair.
>
> So a group MAY span tasks. What must never happen is two examples with the same TEXT carrying
> different labels; that is a labelling conflict and `dedupe()` reports it.
(`router/dataset.py:453-464`.)
**(c) Every task class is represented.**
```python
counts = self.task_counts()
for name in TASK_CLASSES:
if counts[name] == 0:
problems.append(f"task class {name!r} has no examples")
```
(`router/dataset.py:478-482`.)
**(d) An `unsupported` example must never claim a capability.**
```python
for e in self.examples:
if e.task == "unsupported" and (e.temporal or e.spatial_output or e.language_output):
problems.append(
f"unsupported example {e.text!r} claims a capability"
)
```
(`router/dataset.py:484-489`.) This is the *strict* rule β and it is stricter than
`default_attributes`, which sets `language_output: task != "unsupported"` (also false for
unsupported). The two agree; `validate()` is the enforcement.
**(e) Temporal / spatial / modality coherence.**
```python
for e in self.examples:
if e.task == "change" and not e.temporal:
problems.append(f"change example {e.text!r} is not marked temporal")
if e.task == "grounding" and not e.spatial_output:
problems.append(
f"grounding example {e.text!r} is not marked spatial_output"
)
if e.task == "optical_sar" and e.modality != "optical_sar":
problems.append(
f"optical_sar example {e.text!r} has modality {e.modality!r}"
)
```
(`router/dataset.py:491-502`.)
Note the **asymmetry** between (e) and `default_attributes`: a `change` example must be `temporal`,
and a `grounding` example must be `spatial_output` β but a `change` example may be
`spatial_output=True` or `False` (both appear in `CURATED`, Β§12). The validator enforces the
*necessary* attribute, not the full default.
## 12. `CURATED` β the 70 hand-written examples
The corpus docstring states the design of the two sources:
> 1. A hand-written CURATED set. Short, natural, exactly the phrasings a user types. This is the
> part that has to be right; everything else is volume.
> 2. Template-generated examples from a shared template table. Each (task, template) pair carries a
> GROUP id.
(`router/dataset.py:3-9`.)
`CURATED` is a `tuple[RouterExample, ...]` of **70 declared entries** across six task families plus
three hard-negative families. Every entry is a `RouterExample` with all eight fields positional:
`(text, task, modality, temporal, spatial_output, language_output, group, source)`.
### 12.1 The curated examples by family
**Caption β 8 examples, group `cur_cap`, all `source="curated"`:**
| Text | Modality | temporal | spatial | language |
|---|---|---|---|---|
| `Describe this image.` | `unknown` | F | F | T |
| `Describe this satellite image.` | `optical` | F | F | T |
| `Give me a description of the scene.` | `unknown` | F | F | T |
| `What do you see in this image?` | `unknown` | F | F | T |
| `Write a caption for this image.` | `unknown` | F | F | T |
| `Summarize what is visible here.` | `unknown` | F | F | T |
| `Describe the land cover in this scene.` | `optical` | F | F | T |
| `Tell me about this image in detail.` | `unknown` | F | F | T |
**VQA β 12 examples, group `cur_vqa`, all `modality="optical"`:**
`What land cover is visible?` Β· `How many buildings are in this image?` Β·
`Is there water in this scene?` Β· `What is the dominant land cover?` Β·
`Are there any roads visible?` Β· `What type of vegetation is present?` Β·
`Is this an urban or rural area?` Β· `Does this image contain a river?` Β·
`How many vehicles can you count?` Β· `What is the weather like in this image?` Β·
`Is this image from an agricultural area?` Β· `What season does this image show?`
All twelve: `temporal=F, spatial_output=F, language_output=T`.
**Grounding β 10 examples, group `cur_grd`, all `modality="optical"`, all `spatial_output=True`:**
| Text |
|---|
| `Show me the water body.` |
| `Locate the buildings.` |
| `Where is the road?` |
| `Highlight the vegetation.` |
| `Find the bridge in this image.` |
| `Point out the harbor.` |
| `Draw a box around the river.` |
| `Can you show me where the water body is?` |
| `Mark the location of the airport.` |
| `Where are the trees located?` |
**Change β 12 examples, group `cur_chg`, all `modality="optical"`, all `temporal=True`.** Four of
the twelve carry `spatial_output=True`:
| Text | `spatial_output` |
|---|---|
| `What changed between these images?` | F |
| `What changed?` | F |
| `Has the urban area increased?` | F |
| `Compare the before and after images.` | F |
| `Describe the changes.` | F |
| `Did any buildings appear?` | F |
| `Where did the change happen?` | **T** |
| `Show me the changed regions.` | **T** |
| `What changed and show me where?` | **T** |
| `Highlight the areas that changed.` | **T** |
| `Locate where new construction appeared.` | **T** |
| `How much forest was lost between the two dates?` | F |
This table is the single most useful object in the corpus, because it is the concrete evidence that
**`change` does not imply `spatial_output`**. The docstring for `default_attributes` says exactly
this: *"`change` with `spatial_output=False` is a perfectly legal request ('what changed?')"*
(`router/label_space.py:92-94`).
**Optical-SAR β 8 examples, group `cur_osr`, all `modality="optical_sar"`, all
`temporal=F, spatial_output=F, language_output=T`:**
`Compare the optical and radar images.` Β· `Compare optical and SAR to identify built-up regions.` Β·
`Use both the radar and optical images.` Β· `Jointly analyse the optical and SAR pair.` Β·
`What can the radar tell us that the optical cannot?` Β· `Fuse the optical and SAR data.` Β·
`Analyse the co-registered optical and radar pair.` Β·
`Show built-up areas using radar and optical together.`
Note that `Compare optical and SAR to identify built-up regions.` and
`Show built-up areas using radar and optical together.` both contain **built-up** β and both are
labelled `optical_sar`, not `change`. These two curated examples are the corpus-side record of the
very bug the router later exhibited (Β§44βΒ§47): the *word* "built-up" is land-cover vocabulary, and the
modality pair is what makes the request a fusion request.
**Unsupported β 6 examples, group `cur_uns`, all `modality="unknown"`, all three booleans
`False`:**
`Book me a flight to Delhi.` Β· `What is the weather tomorrow?` Β·
`Write me a poem about satellites.` Β· `Who won the cricket match?` Β·
`Send an email to my supervisor.` Β· `What is the capital of France?`
### 12.2 The three hard-negative families
The corpus docstring names the pairs that matter, and the reason they are first-class:
> Hard negatives are first-class here, not an afterthought. The pairs that matter:
>
> "describe the water body" -> caption (describe = language)
> "show me the water body" -> grounding (show = spatial)
>
> "what changed" -> change, spatial_output=False
> "where did the change happen"-> change, spatial_output=True
>
> "compare these two images" -> change (two images, temporal)
> "compare optical and radar" -> optical_sar (two modalities)
>
> Every one of those is a single-token difference in the input and a different label on the output.
> A router that has not seen such pairs will fail them.
(`router/dataset.py:18-31`.)
**`hn_desc_vs_show` β 6 examples.** The describe/show contrast, three subjects:
| Text | Task | spatial_output |
|---|---|---|
| `Describe the water body.` | `caption` | F |
| `Show me the water body.` | `grounding` | T |
| `Describe the road.` | `caption` | F |
| `Show me the road.` | `grounding` | T |
| `Describe the buildings.` | `caption` | F |
| `Show me the buildings.` | `grounding` | T |
**`hn_what_vs_where` β 4 examples.** The what/where contrast on the same task:
| Text | Task | spatial_output |
|---|---|---|
| `What changed?` | `change` | F |
| `Where did the change happen?` | `change` | **T** |
| `Describe the changes.` | `change` | F |
| `Show me where the changes are.` | `change` | **T** |
**`hn_temporal_vs_modality` β 4 examples.** The temporal/modality contrast β the hardest family,
because the surface forms are near-identical and the *semantics* differ:
| Text | Task | temporal | modality |
|---|---|---|---|
| `Compare these two images.` | `change` | **T** | `optical` |
| `Compare optical and radar.` | `optical_sar` | F | `optical_sar` |
| `What is different between these two dates?` | `change` | **T** | `optical` |
| `What is different between the optical and SAR views?` | `optical_sar` | F | `optical_sar` |
The `HARD_NEGATIVE_PREFIX` constant is what makes these three families special:
```python
#: Groups whose name starts with this are hard-negative FAMILIES: sets of
#: deliberately confusable queries that must travel together through any split.
#: `hn_desc_vs_show` holds both "Describe the water body." (caption) and
#: "Show me the water body." (grounding) β splitting those apart would remove
#: the only signal that teaches the router the difference.
HARD_NEGATIVE_PREFIX = "hn_"
```
(`router/dataset.py:53-58`.)
**The prefix does three jobs at once:** `validate()` uses it to permit a group to span tasks;
`split_by_group` uses it to force the family into the **test** split; and `evaluate_split` uses it
to compute the `hard_negative_accuracy` metric. One string, three consumers, one definition.
### 12.3 Declared versus recorded counts
`CURATED` declares **70** entries. `metadata.json.corpus.by_source` records **66** curated and
**510** template, totalling **576** across **54** groups. The difference is `dedupe()`, which runs
before the split (Β§11.1) and removes exact duplicate texts while keeping the first occurrence.
Four of the duplicates are *within* `CURATED`, and they are visible in the source: the same text
appears once in a task family and once in a hard-negative family, deliberately, so the family is
self-contained.
| Duplicated text | First occurrence (family) | Second occurrence (family) |
|---|---|---|
| `Show me the water body.` | `cur_grd` (grounding) | `hn_desc_vs_show` (grounding) |
| `What changed?` | `cur_chg` (change) | `hn_what_vs_where` (change) |
| `Where did the change happen?` | `cur_chg` (change) | `hn_what_vs_where` (change) |
| `Describe the changes.` | `cur_chg` (change) | `hn_what_vs_where` (change) |
A further set is *cross-source*: a curated example whose text a template reproduces exactly, such as
`Show me the water body.` (curated) versus template `t_grd_a` = `"Show me the {}."` with the subject
`water body`. Because `CURATED` is extended into the corpus **first**
(`router/dataset.py:525-526`), the curated entry survives and the generated duplicate is dropped.
**Status.** The **recorded** counts (66 / 510 / 576 / 54) are `MEASURED` β they are in the shipped
artifact. The **mechanism** of the reduction (exact-duplicate removal by `dedupe()`) is
`IMPLEMENTED` and verified by the source. The **exact per-text enumeration** of every dropped
duplicate is `UNKNOWN β not established from the available evidence`: reconstructing it requires
running `build_corpus()` + `dedupe()` and diffing, which this document did not do. The four
within-`CURATED` duplicates above are visible by inspection of the tuple and are therefore stated;
the cross-source ones are stated as a mechanism with one example, not as a complete list.
## 13. `TEMPLATES` and `SUBJECTS` β volume with a group tag
The template table is a tuple of 7-tuples:
```python
#: (group_id, task, template, modality, temporal, spatial, language)
#: `{}` is replaced by a subject from the matching subject list.
TEMPLATES: tuple[tuple[str, str, str, str, bool, bool, bool], ...] = (
```
(`router/dataset.py:233-235`.)
`TEMPLATES` holds **45 entries**. Expanded against `SUBJECTS`, each template yields one example per
subject, and **every example from a given template shares that template's group id** β which is the
whole point (Β§14).
### 13.1 Every template, by family
**Caption β 8 templates** (`router/dataset.py:237-241`, `302-304`):
| Group | Template | modality | temporal | spatial | language |
|---|---|---|---|---|---|
| `t_cap_a` | `Describe the {}.` | optical | F | F | T |
| `t_cap_b` | `Give a detailed description of the {}.` | optical | F | F | T |
| `t_cap_c` | `What does the {} look like?` | optical | F | F | T |
| `t_cap_d` | `Write a caption describing the {}.` | optical | F | F | T |
| `t_cap_e` | `Summarize the {} visible in this image.` | optical | F | F | T |
| `t_cap_f` | `Provide an overview of the {}.` | optical | F | F | T |
| `t_cap_g` | `Explain what is happening in the {}.` | optical | F | F | T |
| `t_cap_h` | `Give me a short report on the {}.` | optical | F | F | T |
**VQA β 5 templates** (`router/dataset.py:244-248`):
| Group | Template |
|---|---|
| `t_vqa_a` | `Is the {} present in this image?` |
| `t_vqa_b` | `How many {} are visible?` |
| `t_vqa_c` | `What kind of {} is shown here?` |
| `t_vqa_d` | `Can you tell if this contains {}?` |
| `t_vqa_e` | `Are there any {} in the scene?` |
All five: `optical`, `F, F, T`.
**Grounding β 8 templates** (`router/dataset.py:251-256`, `307`, `317`):
| Group | Template |
|---|---|
| `t_grd_a` | `Show me the {}.` |
| `t_grd_b` | `Locate the {}.` |
| `t_grd_c` | `Where is the {}?` |
| `t_grd_d` | `Highlight the {}.` |
| `t_grd_e` | `Draw a box around the {}.` |
| `t_grd_f` | `Point out the {} on the map.` |
| `t_grd_g` | `Give me the coordinates of the {}.` |
| `t_grd_h` | `Trace the outline of the {}.` |
All eight: `optical`, `F, T, T`.
**Change β 10 templates** (`router/dataset.py:259-264`, `295-298`):
| Group | Template | spatial |
|---|---|---|
| `t_chg_a` | `What happened to the {} between the two images?` | F |
| `t_chg_b` | `Has the {} changed?` | F |
| `t_chg_c` | `Describe how the {} changed over time.` | F |
| `t_chg_d` | `Where did the {} change?` | **T** |
| `t_chg_e` | `Show the regions where the {} changed.` | **T** |
| `t_chg_f` | `How much did the {} increase or decrease?` | F |
| `t_chg_g` | `Compare the two acquisitions of the {}.` | F |
| `t_chg_h` | `What is different about the {} between the two dates?` | F |
| `t_chg_i` | `Did the {} expand or shrink?` | F |
| `t_chg_j` | `Point out where the {} was modified.` | **T** |
All ten: `optical`, `temporal=T`, `language=T`.
**Optical-SAR β 4 templates** (`router/dataset.py:267-270`):
| Group | Template |
|---|---|
| `t_osr_a` | `Compare the optical and radar views of the {}.` |
| `t_osr_b` | `Use both sensors to identify the {}.` |
| `t_osr_c` | `How does the {} appear differently in SAR and optical?` |
| `t_osr_d` | `Analyse the {} using the co-registered optical and SAR pair.` |
All four: `optical_sar`, `F, F, T`.
**Unsupported β 10 templates** (`router/dataset.py:282-291`):
| Group | Template |
|---|---|
| `t_uns_a` | `Tell me a joke about {}.` |
| `t_uns_b` | `What is the population of {}?` |
| `t_uns_c` | `Recommend a restaurant near {}.` |
| `t_uns_d` | `Book me a flight to {}.` |
| `t_uns_e` | `Who won the match in {}?` |
| `t_uns_f` | `Translate this sentence into {}.` |
| `t_uns_g` | `Write a poem about {}.` |
| `t_uns_h` | `What is the weather in {} tomorrow?` |
| `t_uns_i` | `Send an email about {}.` |
| `t_uns_j` | `Summarize this news article about {}.` |
All ten: `unknown`, `F, F, F`.
### 13.2 The `unsupported` template-count comment is a defect record
The comment above the `unsupported` block is one of the most valuable passages in the router
source, because it explains a **measured** failure and the fix that followed:
```python
# Volume matters here more than anywhere else. Group-level splitting keeps
# every template on one side of the boundary, so a class with only four
# templates ends up with two of them available for training β and a template
# the model never saw cannot be learned, only guessed at. With four templates
# the "Recommend a restaurant near X" phrasing fell entirely into test and
# the router, having never seen it, read "near <scene noun>" as a grounding
# request. That is correct behaviour on insufficient data, so the fix is more
# distinct phrasings, not a looser gate.
```
(`router/dataset.py:272-281`.)
The generalisable rule is stated in the same block:
> **Template COUNT is what buys group-level trainability; subject count only buys volume.**
(`router/dataset.py:354-355`.) This is the second half of the `unsupported` defect in
`docs/PHASE4_ROUTER_REPORT.md` Β§2 β the same edit pass that grew `unsupported` from 4 to 10
templates also grew its subject list to 20, taking the class to **30.4 %** of the corpus and
forcing its inverse-frequency class weight down to **0.548**, which destabilised the surrounding
classes. The fix was to hold at 10 templates Γ 10 subjects β 14 %.
### 13.3 `SUBJECTS` β ten to fifteen nouns per family
```python
SUBJECTS: dict[str, tuple[str, ...]] = {
"caption": (
"scene", "image", "landscape", "area", "region",
"terrain", "land cover", "cityscape", "coastline", "farmland",
),
"vqa": (
"water body", "road", "building", "forest", "river",
"bridge", "harbor", "vehicle", "tree", "field",
"airport", "railway", "dam", "lake", "industrial area",
),
"grounding": (
"water body", "building", "road", "river", "forest",
"bridge", "harbor", "lake", "airport", "dam",
"railway line", "parking lot", "swimming pool", "stadium", "quarry",
),
"change": (
"urban area", "forest", "water body", "built-up area", "vegetation",
"farmland", "road network", "construction site", "coastline", "wetland",
),
"optical_sar": (
"built-up area", "water body", "forest", "agricultural field", "urban region",
"flooded area", "road network", "industrial zone", "coastline", "wetland",
),
"unsupported": (
"the moon", "the ocean", "galaxies", "the ISS", "satellites",
"Delhi", "Mumbai", "cricket", "French", "the Himalayas",
),
}
```
(`router/dataset.py:320-360`.)
| Family | Subject count | Template count | Instances (`count Γ count`) |
|---|---|---|---|
| `caption` | 10 | 8 | 80 |
| `vqa` | 15 | 5 | 75 |
| `grounding` | 15 | 8 | 120 |
| `change` | 10 | 10 | 100 |
| `optical_sar` | 10 | 4 | 40 |
| `unsupported` | 10 | 10 | 100 |
| **Total** | β | **45** | **515** |
**The `unsupported` subject list is the one with a stated design, and it mixes two kinds of noun:**
> * satellite-imagery nouns used in a NON-imagery frame (the moon, the ocean, the ISS) β the
> valuable half. "Recommend a restaurant near the ocean" contains a scene noun but is not an
> imagery request, so it teaches that the TEMPLATE dominates the subject.
> * unambiguous out-of-domain entities (Delhi, cricket, French)
(`router/dataset.py:344-351`.)
That first kind is a **deliberate hard negative inside the unsupported class**: it isolates the
template as the label-bearing feature, so the router cannot learn "contains a scene noun β
imagery request".
**The arithmetic gap.** 70 declared curated + 515 template instances = **585**, but the recorded
corpus is **576** (66 + 510). The 9-item reduction is `dedupe()` (Β§12.3). The recorded
per-family `by_task` totals are:
| Task | Recorded count | Declared total before dedupe |
|---|---|---|
| `caption` | 91 | 8 + 80 = 88 |
| `change` | 115 | 12 + 100 = 112 |
| `grounding` | 128 | 10 + 120 = 130 |
| `optical_sar` | 50 | 8 + 40 = 48 |
| `unsupported` | 105 | 6 + 100 = 106 |
| `vqa` | 87 | 12 + 75 = 87 |
| **Total** | **576** | **571** |
Note that the recorded totals **exceed** the declared totals for four families β which is the
opposite of what pure deduplication would produce. This means the declared `TEMPLATES`/`SUBJECTS`
tables as read here **do not fully account for the shipped corpus**: the shipped
`metadata.json` was produced on 2026-09-16, and the corpus source has clearly been edited since
(the `unsupported` comment records an expansion from 4 to 10 templates, and `caption` gained
`t_cap_f/g/h`). **The per-family provenance of the shipped 576-example corpus is therefore
`UNKNOWN β not established from the available evidence`.** What *is* established: the shipped
artifact's own counts (66 curated, 510 template, 576 total, 54 groups, and the six `by_task`
counts), the current declared table (70 curated, 45 templates), and the fact that `train_router`
runs `dedupe()` before splitting. A reader who needs the exact shipped corpus should regenerate it
with `build_corpus(seed=42)` against the 2026-09-16 revision of `router/dataset.py`.
**This is stated rather than smoothed over deliberately.** The style guide's rule 6 is that an
honest gap beats a confident invention, and the tempting move here β asserting "dedupe removed 9
duplicates" β is contradicted by the arithmetic.
### 13.4 `_generate_templates()` β one group per template
```python
def _generate_templates() -> list[RouterExample]:
"""Expand the template table into examples, one group per template."""
out: list[RouterExample] = []
for group, task, template, modality, temporal, spatial, language in TEMPLATES:
subjects = SUBJECTS.get(task, SUBJECTS["caption"])
for subject in subjects:
out.append(
RouterExample(
text=template.format(subject),
task=task,
modality=modality,
temporal=temporal,
spatial_output=spatial,
language_output=language,
group=group,
source="template",
)
)
return out
```
(`router/dataset.py:363-381`.)
Two details: the **group is the template id**, shared by every subject expansion (this is what
makes the group a leakage boundary at the *template* level, not the *example* level); and the
subject list falls back to `SUBJECTS["caption"]` for an unknown task, so a new task added to
`TEMPLATES` without a subject list still generates rather than crashing.
### 13.5 `build_corpus()` β assembly and the ablation switches
```python
def build_corpus(
n_template_repeats: int = 1,
seed: int = 42,
include_curated: bool = True,
include_templates: bool = True,
) -> RouterCorpus:
"""Assemble the full router corpus.
Args:
n_template_repeats: duplicate the template block N times with shuffled
subject ordering. Volume without new phrasings; keep at 1 unless the
adapter is clearly underfitting.
seed: controls subject shuffling.
include_curated / include_templates: ablation switches.
"""
rng = random.Random(seed)
examples: list[RouterExample] = []
if include_curated:
examples.extend(CURATED)
if include_templates:
for _ in range(max(1, n_template_repeats)):
block = _generate_templates()
rng.shuffle(block)
examples.extend(block)
return RouterCorpus(examples)
```
(`router/dataset.py:507-534`.)
**Three things to note.** `CURATED` is extended **first**, which is what makes the curated entry
win a dedupe collision. The shuffle is applied to the *template block* only, so curated order is
deterministic. And `n_template_repeats` repeats the same **phrasings** with shuffled subject order
β the docstring is explicit that this buys volume, not new phrasings: *"Volume without new
phrasings; keep at 1 unless the adapter is clearly underfitting."* The shipped run used the default
`n_template_repeats=1`.
## 14. Finding F4-3 β splits must be by group, never by example
The dataset docstring states the finding with its worked example:
> The group tag is the important part (finding F4-3). `"show me the water body"` and
> `"show me the road"` come from the same template and differ by one token. Splitting them across
> train/val makes validation trivially easy and gives a fake accuracy number. Splitting by group
> keeps every template on exactly one side β the router-level analogue of scene-level splitting,
> and the same class of bug Gate 1 exists to catch.
(`router/dataset.py:11-16`.)
`router/train.py` restates it and adds the enforcement:
> Groups are template ids and curated families. Splitting by example would put "show me the water
> body" in train and "show me the road" in val β same template, one token apart β and report a
> fake accuracy. `split_by_group` mirrors `evaluation.leakage.assign_splits_by_scene` deliberately,
> and the train function REFUSES to proceed if the split report is not clean.
(`router/train.py:14-20`.)
The refusal is real code:
```python
report = split_leakage_report(split)
if not report["clean"]:
raise RoutingError(
"router split failed the group-leakage audit; refusing to train: "
f"{report['groups_across_splits']}"
)
```
(`router/train.py:502-507`.)
### 14.1 `split_by_group()` β the signature and the two guarantees
```python
def split_by_group(
corpus: RouterCorpus,
train_ratio: float = 0.75,
val_ratio: float = 0.15,
seed: int = 42,
hard_negatives_to_test: bool = True,
) -> dict[str, list[RouterExample]]:
"""Split by GROUP, never by example (finding F4-3).
Mirrors `evaluation.leakage.assign_splits_by_scene` deliberately: the failure
mode is identical, so the guard should look identical too. Groups are sorted
before shuffling so the result does not depend on corpus order.
Two guarantees beyond a naive ratio split:
1. **Every task class reaches train.** A random group split can strand a
class entirely in val/test, and then its head can never learn it. Groups
are walked in shuffled order and any group that covers a not-yet-covered
task is pulled into train first, before the ratio is topped up.
2. **Hard-negative families are held out.** They are the cases the router
most needs measured, so by default they go to test rather than training
on the exact pairs being scored.
"""
```
(`router/dataset.py:537-566`.)
**Input validation, before any work:**
```python
if not 0.0 < train_ratio < 1.0:
raise ValueError(f"train_ratio must be in (0,1), got {train_ratio}")
if not 0.0 <= val_ratio < 1.0:
raise ValueError(f"val_ratio must be in [0,1), got {val_ratio}")
if train_ratio + val_ratio >= 1.0:
raise ValueError(
f"train_ratio + val_ratio must be < 1, got {train_ratio} + {val_ratio}"
)
```
(`router/dataset.py:567-574`.) Note the asymmetry: `train_ratio` is strictly interior
`(0, 1)` β a 100 %-train split is meaningless β while `val_ratio` may be `0`.
**Forced test groups:**
```python
groups = sorted(by_group)
forced_test = {
g for g in groups
if hard_negatives_to_test and g.startswith(HARD_NEGATIVE_PREFIX)
}
pool = [g for g in groups if g not in forced_test]
if not pool:
raise ValueError(
"every group is a hard-negative family; there is nothing to train on"
)
```
(`router/dataset.py:580-591`.)
**Stratification by dominant task β and the comment that records the defect:**
```python
# Stratify by dominant task ----------------------------------------
# Group-level splitting is what prevents leakage. It is NOT sufficient on
# its own: a naive group split can strand an entire task class in one
# partition. That is not hypothetical β with 38 groups and a flat 75/15
# allocation, `vqa` (the most important mandatory task) and `unsupported`
# both ended up with ZERO test examples, so the acceptance gate could not
# measure them and their recall printed as a misleading 0.000.
#
# Stratifying by task fixes that WITHOUT weakening leakage safety: a group
# still travels to exactly one split, whatever its labels. Only the choice
# of which split changes.
test_ratio = 1.0 - train_ratio - val_ratio
strata: dict[str, list[str]] = {}
for group in pool:
strata.setdefault(_dominant_task(by_group[group]), []).append(group)
```
(`router/dataset.py:593-608`.)
`_dominant_task` is the stratum chooser, and its tie-break is deterministic:
```python
def _dominant_task(examples: list["RouterExample"]) -> str:
"""Most common task in a group, ties broken deterministically.
Used only to choose which stratum a group belongs to when balancing splits.
Hard-negative families deliberately span tasks, so they need a stratum too;
the alphabetically-first task is as good a tie-break as any.
"""
counts: dict[str, int] = {}
for e in examples:
counts[e.task] = counts.get(e.task, 0) + 1
return sorted(counts.items(), key=lambda kv: (-kv[1], kv[0]))[0][0]
```
(`router/dataset.py:74-84`.) The sort key `(-count, name)` means: highest count first, and on a tie,
alphabetically-first name. Deterministic and reproducible.
**The per-stratum allocation, with its three degenerate cases:**
```python
for task in sorted(strata):
group_list = list(strata[task])
rng.shuffle(group_list)
n = len(group_list)
if n == 1:
# Cannot be spread without losing the class from training. Training
# must win: an unlearnable class is worse than an unmeasured one.
assignment[group_list[0]] = "train"
continue
if n == 2:
# One to train (learnable), one to test (measurable). Validation
# coverage for this task comes from the corpus at large.
assignment[group_list[0]] = "train"
assignment[group_list[1]] = "test"
continue
n_test = max(1, round(n * test_ratio))
n_val = max(1, round(n * val_ratio))
# Keep at least one group for training, whatever the ratios ask for.
while n_test + n_val >= n:
if n_test > 1:
n_test -= 1
elif n_val > 1:
n_val -= 1
else:
break
n_train = n - n_test - n_val
for i, group in enumerate(group_list):
if i < n_train:
assignment[group] = "train"
elif i < n_train + n_val:
assignment[group] = "val"
else:
assignment[group] = "test"
```
(`router/dataset.py:613-649`.)
**The `n == 1` rule is the explicit priority statement:** *an unlearnable class is worse than an
unmeasured one*. A single-group stratum goes to **train**, so `Gate 2`'s "every class measured"
condition can still fail β which is the correct failure, because an unmeasured class is at least
visible in `unmeasured_test_tasks` (Β§18.1).
**The shuffle is per stratum**, not global, so adding a group to one task family does not perturb
the assignment of another.
**The output is emitted in sorted group order, not shuffle order:**
```python
# Emit in sorted group order. The seed chooses *which* group goes where; it
# must not control the order of the returned lists, or downstream code
# taking records[:N] would silently depend on the seed.
out: dict[str, list[RouterExample]] = {"train": [], "val": [], "test": []}
for group in groups:
out[assignment[group]].extend(by_group[group])
return out
```
(`router/dataset.py:660-666`.) This is a subtle but important guarantee: **the seed chooses the
partition, but not the ordering within it.** A caller who slices `split["train"][:100]` gets the
same 100 examples regardless of seed.
### 14.2 `split_leakage_report()` β the audit
```python
def split_leakage_report(split: dict[str, list[RouterExample]]) -> dict[str, object]:
"""Verify no group crosses a split boundary. Returns a report dict."""
group_splits: dict[str, set[str]] = {}
for name, examples in split.items():
for e in examples:
group_splits.setdefault(e.group, set()).add(name)
offenders = {g: sorted(s) for g, s in group_splits.items() if len(s) > 1}
return {
"clean": not offenders,
"groups_across_splits": offenders,
"sizes": {name: len(examples) for name, examples in split.items()},
"task_counts": {
name: _count_tasks(examples) for name, examples in split.items()
},
"unique_groups": {name: len({e.group for e in examples})
for name, examples in split.items()},
}
```
(`router/dataset.py:669-687`.)
The report has five keys, and the shipped artifact records all five
(`metadata.json.split`):
```json
"clean": true,
"groups_across_splits": {},
"sizes": { "test": 80, "train": 410, "val": 86 },
"task_counts": {
"test": { "caption": 13, "change": 13, "grounding": 17, "optical_sar": 12, "unsupported": 10, "vqa": 15 },
"train": { "caption": 70, "change": 82, "grounding": 97, "optical_sar": 28, "unsupported": 76, "vqa": 57 },
"val": { "caption": 8, "change": 20, "grounding": 14, "optical_sar": 10, "unsupported": 19, "vqa": 15 }
},
"unique_groups": { "test": 9, "train": 37, "val": 8 }
```
**Every one of the six task classes has non-zero test support.** That is the stratification
working, and it is the precondition for `gate2_passed` (Β§18.2) being able to report anything at all.
| Split | Examples | Unique groups | Notes |
|---|---|---|---|
| train | 410 | 37 | 71.2 % of 576 |
| val | 86 | 8 | 14.9 % |
| test | 80 | 9 | 13.9 % β **includes all three `hn_*` families** |
| **total** | **576** | **54** | β |
The group counts (37 + 8 + 9 = 54) sum exactly to the corpus's 54 groups, which is the audit's
`clean: true` in arithmetic form: every group appears in exactly one split.
**The test split's 9 groups include the 3 hard-negative families**, so the 80 test examples
contain all 14 curated `hn_*` examples. That is why `hard_negative_accuracy` can be computed at all
(Β§16.4) and why it is the most honest generalisation number in the artifact.
---
# Part D β Training
## 15. `train_router()` β the full signature
```python
def train_router(
corpus: RouterCorpus | None = None,
encoder=None,
epochs: int = 60,
batch_size: int = 64,
learning_rate: float = 1e-3,
weight_decay: float = 0.01,
task_loss_weight: float = 1.0,
modality_loss_weight: float = 0.3,
binary_loss_weight: float = 0.5,
hidden_dim: int = 128,
dropout: float = 0.1,
seed: int = 42,
device: str = "cpu",
artifact_dir: str | Path | None = None,
cache_dir: str | Path | None = None,
config_hash: str | None = None,
n_template_repeats: int = 1,
val_ratio: float = 0.15,
hard_negatives_to_test: bool = True,
verbose: bool = True,
) -> TrainingResult:
```
(`router/train.py:444-465`.)
| Parameter | Default | Shipped value | Source of the shipped value |
|---|---|---|---|
| `corpus` | `None` β `build_corpus(...)` | built from source | `metadata.json.corpus` |
| `encoder` | `None` β stub | the frozen MiniLM encoder | `metadata.json.encoder_type: "frozen_sentence_transformer"` |
| `epochs` | 60 | **60** | `metadata.json.hyperparameters.epochs` |
| `batch_size` | 64 | **64** | `metadata.json.hyperparameters.batch_size` |
| `learning_rate` | `1e-3` | **0.001** | `metadata.json.hyperparameters.learning_rate` |
| `weight_decay` | 0.01 | **0.01** | `metadata.json.hyperparameters.weight_decay` |
| `task_loss_weight` | 1.0 | **1.0** | `metadata.json.hyperparameters.task_loss_weight` |
| `modality_loss_weight` | 0.3 | **0.3** | `metadata.json.hyperparameters.modality_loss_weight` |
| `binary_loss_weight` | 0.5 | **0.5** | `metadata.json.hyperparameters.binary_loss_weight` |
| `hidden_dim` | 128 | **128** | `metadata.json.adapter.hidden_dim` |
| `dropout` | 0.1 | **0.1** | `metadata.json.adapter.dropout` |
| `seed` | 42 | **42** | `metadata.json.seed` |
| `device` | `"cpu"` | `cpu` | `threshold_sweep_val.json.environment.requested_device` |
| `artifact_dir` | `None` β `artifacts/router/router_adapter_v001` | that path | `metadata.json` location |
| `cache_dir` | `None` β `artifacts/router/cache` | that path | `artifacts/router/cache/` exists |
| `config_hash` | `None` | `615478910dc266bf` | `metadata.json.config_hash` |
| `n_template_repeats` | 1 | 1 (default) | `metadata.json.corpus.total: 576` is consistent with 1 |
| `val_ratio` | 0.15 | **0.15** | `configs/base.yaml` Β§`router.training.val_ratio` |
| `hard_negatives_to_test` | `True` | **true** | `configs/base.yaml` Β§`router.training.hard_negatives_to_test` |
| `verbose` | `True` | β | prints every 10 epochs |
**Two `config_hash` values exist and they are different numbers for different objects.** The
adapter artifact records `config_hash: "615478910dc266bf"`, while the threshold sweep and the
release-wide frozen config hash is `78f1e3700da15aa1`. `threshold_sweep_val.json` records **both**,
under different keys β `adapter_config_hash: "615478910dc266bf"` and
`config_hash: "78f1e3700da15aa1"`. So the two are not in conflict: one hashes the adapter's
architecture config, the other the frozen system configuration. The style guide's
"Frozen config hash `78f1e3700da15aa1`" refers to the latter.
### 15.1 The five-stage body, in order
```python
started = time.time()
random.seed(seed)
np.random.seed(seed)
torch.manual_seed(seed)
```
(`router/train.py:473-476`.) **All three RNGs are seeded**, not just torch's β because the shuffle
uses `random.Random(seed)` (Β§15.3), the numpy path is used by the stub embeddings, and torch seeds
the weight init.
**Stage 1 β corpus validation and dedupe.** `validate()` problems abort with a `RoutingError`
naming every problem; `dedupe()` conflicts abort with the conflict list (Β§11.1).
**Stage 2 β split, with the leakage guard proven to have fired.** `split_by_group(...)` then
`split_leakage_report(...)`, and a non-clean report aborts. Then a second guard:
```python
train_ex, val_ex, test_ex = split["train"], split["val"], split["test"]
if not train_ex or not val_ex or not test_ex:
raise RoutingError(
f"split produced an empty partition: train={len(train_ex)} "
f"val={len(val_ex)} test={len(test_ex)}"
)
```
(`router/train.py:509-514`.) The shipped split is 410 / 86 / 80 β none empty.
**Stage 3 β embeddings.** Either the cached frozen-encoder path or the stub path (Β§17).
**Stage 4 β model, class weights, optimizer, losses, the epoch loop.**
**Stage 5 β final evaluation, metadata assembly, artifact write, return.**
### 15.2 The embedding cache and its fingerprint
```python
def _corpus_fingerprint(corpus: RouterCorpus, encoder_id: str, max_length: int) -> str:
payload = {
"version": CORPUS_CACHE_VERSION,
"encoder": encoder_id,
"max_length": max_length,
"texts": sorted(e.text for e in corpus),
}
blob = json.dumps(payload, sort_keys=True).encode()
return hashlib.sha256(blob).hexdigest()
```
(`router/train.py:221-229`.) `CORPUS_CACHE_VERSION = "v1"` (`router/train.py:71`).
The cache key includes **the sorted text set**, not just a count β so changing one query
invalidates the cache. `embed_corpus_cached` explains why this matters:
> The cache key includes the encoder id AND the corpus text set, so changing either invalidates it.
> A stale cache silently training on the wrong vectors would be worse than no cache.
(`router/train.py:239-243`.)
```python
encoder_id = f"{encoder.model_name}@{encoder.revision}"
fingerprint = _corpus_fingerprint(corpus, encoder_id, encoder.max_length)
vectors_file = cache_path / f"embeddings_{fingerprint[:16]}.npy"
meta_file = cache_path / f"embeddings_{fingerprint[:16]}.json"
if vectors_file.exists() and meta_file.exists() and not force:
try:
cached = np.load(vectors_file)
meta = json.loads(meta_file.read_text(encoding="utf-8"))
if cached.shape[0] == len(corpus) and meta.get("fingerprint") == fingerprint:
return cached.astype(np.float32)
except Exception: # noqa: BLE001 - a bad cache is a miss, not an error
pass
```
(`router/train.py:248-260`.)
**Three robustness details, each deliberate:**
1. **A corrupt cache is a cache miss, not an error.** The `except Exception: pass` is commented as
such β a truncated `.npy` or malformed JSON causes a re-embed, not a crash.
2. **The row count is checked *and* the fingerprint is checked.** Either alone would be weaker: the
fingerprint catches a changed corpus, the shape check catches a truncated file whose metadata
still parses.
3. **The `.npy` and `.json` are keyed by the same 16-hex prefix**, so they cannot be mismatched.
The sidecar metadata records the provenance of the cache:
```python
meta_file.write_text(
json.dumps(
{
"fingerprint": fingerprint,
"encoder": encoder_id,
"max_length": encoder.max_length,
"count": int(vectors.shape[0]),
"dim": int(vectors.shape[1]),
"seconds": round(time.time() - started, 3),
},
indent=2,
sort_keys=True,
),
encoding="utf-8",
)
```
(`router/train.py:267-281`.)
**The row-index bookkeeping is by text, not by position:**
```python
# index bookkeeping: embeddings follow corpus order -----------------
text_to_row = {e.text: i for i, e in enumerate(corpus)}
all_ex = corpus.examples
def rows_for(examples: list[RouterExample]) -> np.ndarray:
idx = [text_to_row[e.text] for e in examples]
return embeddings[np.asarray(idx, dtype=np.int64)]
```
(`router/train.py:547-557`.) Embeddings are computed over `corpus.texts()` in corpus order, and the
split lists are subsets of the same objects; the `text_to_row` map is what re-associates a split
example with its vector. A subtle consequence: because the map is keyed on `text`, a corpus that
survived `dedupe()` is required β two examples with the same text would map to one row. `dedupe()`
has already guaranteed that.
### 15.3 Class weights, computed on the train split only
```python
# Class weights: inverse frequency on the TRAIN split only. Computing them
# on the full corpus would leak val/test label distribution into training.
train_task_counts = np.zeros(NUM_TASKS, dtype=np.float64)
for e in train_ex:
train_task_counts[TASK_TO_INDEX[e.task]] += 1
counts = np.maximum(train_task_counts, 1.0)
class_weights = torch.tensor(
(counts.sum() / (NUM_TASKS * counts)), dtype=torch.float32, device=device
)
```
(`router/train.py:568-576`.)
The formula is the standard "balanced" weighting: `w_c = N / (C Β· n_c)`, so a class with the mean
count gets weight 1.0, a rarer class gets more, and the weights average to 1.0 across classes.
`np.maximum(train_task_counts, 1.0)` floors an absent class at 1 to avoid a division by zero β
though `validate()` has already guaranteed every class has examples.
**Computing on the train split only is a leakage control**, and the comment says so. Using the full
corpus would let the val/test label distribution influence the loss weights.
**The shipped weights are not recorded in the artifact.** `metadata.json.hyperparameters` carries
the loss weights but not the derived `class_weights`. The values are therefore **`UNKNOWN β not
established from the available evidence`**; they are computable from
`metadata.json.split.task_counts.train`, and this document does not compute them.
### 15.4 Optimizer and scheduler
```python
optimizer = torch.optim.AdamW(
model.parameters(), lr=learning_rate, weight_decay=weight_decay
)
scheduler = torch.optim.lr_scheduler.CosineAnnealingLR(optimizer, T_max=max(1, epochs))
ce = nn.CrossEntropyLoss(weight=class_weights)
ce_mod = nn.CrossEntropyLoss()
bce = nn.BCEWithLogitsLoss()
```
(`router/train.py:578-584`.)
**AdamW** with `lr=0.001`, `weight_decay=0.01` β decoupled weight decay, which is the "W" in the
name and the reason it is preferred over `Adam(weight_decay=...)`.
**`CosineAnnealingLR` with `T_max = epochs` (60).** The shipped history confirms the cosine schedule
exactly: the learning rate starts at 0.0009993147673772868 (epoch 1), passes exactly `0.00075` at
epoch 20, exactly `0.0005` at epoch 30, exactly `0.00025` at epoch 40, and reaches exactly `0.0` at
epoch 60 (`metadata.json.history`). Those four round numbers at 20/30/40/60 are the signature of
`T_max=60` with no warm restarts.
**Best-state selection.** Every epoch, the model is evaluated on val and the best val task accuracy
is snapshotted:
```python
if val_metrics.task_accuracy > best_val:
best_val = val_metrics.task_accuracy
best_state = {k: v.detach().clone() for k, v in model.state_dict().items()}
...
if best_state is not None:
model.load_state_dict(best_state)
```
(`router/train.py:643-656`.) The shipped history shows the best val task accuracy first reached at
**epoch 8** (0.9651162790697675) and never exceeded β every epoch from 8 to 60 reports exactly that
value. So the shipped adapter's weights are the **epoch-8** snapshot, and the remaining 52 epochs
did not improve the selected metric. This is `MEASURED`, directly from the history array.
Note the strict `>`: a tie does not replace the snapshot, so the *earliest* epoch achieving the
best value wins. With a flat plateau from epoch 8 onward, that is epoch 8.
### 15.5 The epoch loop, and the shuffle
```python
rng = random.Random(seed)
for epoch in range(epochs):
model.train()
order = list(range(n_train))
rng.shuffle(order)
epoch_loss = 0.0
batches = 0
for start in range(0, n_train, batch_size):
idx = torch.tensor(order[start:start + batch_size], dtype=torch.long, device=device)
out = model(train_tensors["x"][idx])
...
scheduler.step()
val_metrics = evaluate_split(model, val_ex, val_x, "val", device)
history.append(
{
"epoch": float(epoch + 1),
"loss": epoch_loss / max(1, batches),
"val_task_accuracy": val_metrics.task_accuracy,
"val_combined_accuracy": val_metrics.combined_accuracy,
"lr": float(optimizer.param_groups[0]["lr"]),
}
)
```
(`router/train.py:592-641`.)
**Four details:**
1. **The shuffle is a `random.Random(seed)` instance, not the global RNG.** Because `train_router`
also seeds the global `random` at the top, both are deterministic β but the local instance means
the shuffle stream is independent of anything else that touches `random` during the run.
2. **`scheduler.step()` is called once per epoch**, after the batch loop, before the val evaluation.
The history's `lr` field is therefore the rate used *for the next* epoch's first batch.
3. **`history` is appended every epoch**, and the shipped artifact contains **60 rows** β one per
epoch, matching `epochs: 60`.
4. **`val_metrics` is recomputed inside the loop** and again after the best-state restore. The
post-restore recomputation is what produces the final `metrics.val` block, which is why the
recorded `val.task_accuracy` (0.9651) matches the epoch-8 history value exactly.
## 16. Metrics β `SplitMetrics` and `evaluate_split`
```python
@dataclass
class SplitMetrics:
"""Per-split evaluation results."""
name: str
n: int
task_accuracy: float
modality_accuracy: float
binary_accuracy: dict[str, float]
combined_accuracy: float
macro_f1_task: float
per_task_recall: dict[str, float] = field(default_factory=dict)
#: How many test examples each class actually had. A recall without
#: support is a number that means nothing; they are reported together.
task_support: dict[str, int] = field(default_factory=dict)
hard_negative_accuracy: float | None = None
```
(`router/train.py:79-94`.)
| Field | Meaning | Rounding on serialisation |
|---|---|---|
| `name` | `"train"` / `"val"` / `"test"` | β |
| `n` | examples in the split | β |
| `task_accuracy` | argmax over `task_logits` == target | 4 dp |
| `modality_accuracy` | argmax over `modality_logits` == target | 4 dp |
| `binary_accuracy` | per head: `sigmoid(logit) >= 0.5` == target | 4 dp each |
| `combined_accuracy` | **every head correct** for the example to count | 4 dp |
| `macro_f1_task` | macro-averaged F1 over the six task classes | 4 dp |
| `per_task_recall` | recall per class **that has support > 0** | 4 dp each |
| `task_support` | example count per class, **all six keys** | β |
| `hard_negative_accuracy` | accuracy over `hn_*` examples, `None` if none | 4 dp or `null` |
### 16.1 The support-alongside-recall decision
```python
def _recall_per_class(
y_true: list[int], y_pred: list[int], n_classes: int
) -> tuple[list[float], list[int]]:
"""Per-class recall AND per-class support.
Support is returned alongside recall because a class with zero test
examples has a recall of 0.0 that means nothing. Returning them together
makes that impossible to misread: `vqa 0.000` without support is exactly
the kind of number that gets quoted as a failure when it is an absence.
"""
```
(`router/train.py:332-341`.) This is the code-level fix for defect 1 of
`docs/PHASE4_ROUTER_REPORT.md` Β§"Four real defects" β the phase where `vqa` landed with 72 training
examples and **0 test examples**, so its recall printed as `0.000` and the gate reported "BELOW
TARGET (0.909)" over a model never asked about `vqa` at all.
Note the deliberate asymmetry in the two fields: `per_task_recall` **omits** a zero-support class
(`if support[i] > 0`), while `task_support` **includes all six** keys. So a class with no test
examples appears in `task_support` as `0` and is absent from `per_task_recall` β which is exactly
what the shipped artifact shows for no class, because stratification gave all six test support.
### 16.2 The local macro-F1
```python
def _macro_f1(y_true: list[int], y_pred: list[int], n_classes: int) -> float:
"""Macro-averaged F1. Implemented locally to avoid a sklearn dependency in
the training path (sklearn is installed for dev, not required at runtime)."""
if not y_true:
return 0.0
f1s: list[float] = []
for c in range(n_classes):
tp = sum(1 for t, p in zip(y_true, y_pred) if t == c and p == c)
fp = sum(1 for t, p in zip(y_true, y_pred) if t != c and p == c)
fn = sum(1 for t, p in zip(y_true, y_pred) if t == c and p != c)
precision = tp / (tp + fp) if (tp + fp) else 0.0
recall = tp / (tp + fn) if (tp + fn) else 0.0
f1 = (2 * precision * recall / (precision + recall)) if (precision + recall) else 0.0
f1s.append(f1)
return sum(f1s) / len(f1s)
```
(`router/train.py:315-329`.)
**The denominator is `n_classes` = 6, always** β including a class with zero predictions and zero
support, which contributes `0.0`. This is the same denominator convention that makes the
optical-SAR macro-F1 (0.434161) low by construction with 5 absent classes
(`README.md:827-829`) β a shared convention across the project, worth noting because it means a
macro score here is not comparable to a macro score computed over present classes only.
### 16.3 `evaluate_split()` β and the `combined_accuracy` definition
```python
@torch.no_grad()
def evaluate_split(
model: IntentAdapter,
examples: list[RouterExample],
embeddings: np.ndarray,
name: str,
device: str,
batch_size: int = 128,
) -> SplitMetrics:
```
(`router/train.py:352-360`.)
The empty-split short-circuit is explicit:
```python
if not examples:
return SplitMetrics(
name=name, n=0, task_accuracy=0.0, modality_accuracy=0.0,
binary_accuracy={h: 0.0 for h in BINARY_HEADS},
combined_accuracy=0.0, macro_f1_task=0.0,
task_support={task: 0 for task in TASK_CLASSES},
)
```
(`router/train.py:363-369`.) Note `task_support` is populated with six zeros, so an empty split
still produces a complete support dict rather than a missing one.
**Batched inference at 128**, larger than the training batch of 64 β evaluation has no gradient and
no optimizer state, so it can afford larger batches. `@torch.no_grad()` decorates the whole
function.
**`combined_accuracy` is an AND over all five heads:**
```python
# Combined: every head must be right for an example to count.
combined_hits = 0
for i in range(n):
ok = task_pred[i] == task_true[i] and modality_pred[i] == modality_true[i]
if ok:
for j, head in enumerate(BINARY_HEADS):
if int(tensors["binary"][i, j].item()) != binary_pred[head][i]:
ok = False
break
combined_hits += int(ok)
```
(`router/train.py:401-410`.) This is the strictest of the reported metrics, and it is why the
shipped test `combined_accuracy` (0.9625) is lower than `task_accuracy` (0.975): five examples got
the task right but at least one binary head wrong. On val the gap is much wider β 0.7907 combined
against 0.9651 task β because the val `modality_accuracy` is only 0.8256, so most of the
combined failures are modality errors, not task errors.
### 16.4 The hard-negative metric
```python
# Hard negatives: examples whose group starts with "hn_" (curated) β these
# are the one-token-difference pairs the router exists to get right.
hn_examples = [e for e in examples if e.group.startswith(HARD_NEGATIVE_PREFIX)]
hn_acc: float | None = None
if hn_examples:
hn_idx = [i for i, e in enumerate(examples) if e.group.startswith(HARD_NEGATIVE_PREFIX)]
hits = sum(1 for i in hn_idx if task_pred[i] == task_true[i])
hn_acc = hits / len(hn_idx)
```
(`router/train.py:414-421`.)
**Three properties:** it is `None` when the split has no `hn_*` examples (which is why `train` and
`val` both record `null` in the shipped metadata β the families are forced to test); it measures
**task accuracy only**, not the binary heads; and it is computed over the `hn_*` examples in
whatever split is being evaluated.
| Split | `hard_negative_accuracy` | Why |
|---|---|---|
| train | `null` | no `hn_*` groups β they are forced to test |
| val | `null` | same |
| **test** | **0.8** | 14 `hn_*` examples, 11 correct |
The test value of 0.8 means **2 of the 14 curated hard-negative examples were mis-classified at the
task level.** The artifact does not record *which* two. That is `UNKNOWN β not established from the
available evidence`, and it is the single most useful thing a future run could add, because those
two examples are precisely the router's remaining frontier.
## 17. The stub path β tests only, never deployed
```python
def _stub_embeddings(corpus: RouterCorpus, dim: int = 384) -> np.ndarray:
"""Deterministic bag-of-hashes pseudo-embedding.
Exists so `train_router` can be exercised in unit tests without downloading
a 90 MB model. It is NOT a semantic encoder: two paraphrases get unrelated
vectors. Any artifact trained with it is marked `encoder_type: "stub"` and
must never be deployed.
"""
out = np.zeros((len(corpus), dim), dtype=np.float32)
for row, example in enumerate(corpus):
for token in example.text.lower().split():
h = int(hashlib.sha256(token.encode()).hexdigest()[:8], 16)
out[row, h % dim] += 1.0
norms = np.linalg.norm(out, axis=1, keepdims=True)
norms[norms == 0] = 1.0
return out / norms
```
(`router/train.py:728-743`.)
**The mechanism:** each token is hashed to a bucket in `[0, 384)` and incremented, then the row is
L2-normalised β a hashing-trick bag-of-words with no semantic structure. Two paraphrases get
**unrelated** vectors, which is exactly why it is unusable as a router and adequate as a test
fixture.
**Three guards around it:** the `encoder_type` is recorded as `"stub"`; `encoder_meta` records
`model: "stub"`, `revision: "none"`, `max_length: 0`, `parameters: 0`; and the docstring's
"must never be deployed" is repeated in `docs/PHASE4_ROUTER_REPORT.md` Β§"Standing caveats". The
shipped artifact records `encoder_type: "frozen_sentence_transformer"` β so the shipped adapter is
**not** a stub artifact.
**A subtlety in the stub path's corpus hash:** when `encoder is None`, `corpus_hash` is computed
with `max_length=0` and the encoder id `"stub@none"`:
```python
corpus_hash = _corpus_fingerprint(
corpus, encoder_meta["model"] + "@" + str(encoder_meta["revision"]), 0
)
```
(`router/train.py:543-545`.) The shipped artifact's `corpus.hash` is
`8054810736ef97c3db873b2d7073948773a8982f1d16411833d39e15e1871e83`, computed with `max_length=0`
regardless of the encoder used β a small inconsistency worth naming, since it means the recorded
corpus hash is **not** the same value the embedding cache keys on (which uses the real
`max_length=128`). Both are deterministic; they are simply different hashes over different payloads.
## 18. `TrainingResult` β and the gate
```python
@dataclass
class TrainingResult:
"""Everything the caller needs to decide the next step."""
artifact_dir: Path
metadata: dict[str, Any]
train_metrics: SplitMetrics
val_metrics: SplitMetrics
test_metrics: SplitMetrics
split_report: dict[str, Any]
corpus_hash: str
duration_seconds: float
history: list[dict[str, float]] = field(default_factory=list)
```
(`router/train.py:114-126`.)
### 18.1 `unmeasured_test_tasks` β the visibility property
```python
@property
def unmeasured_test_tasks(self) -> list[str]:
"""Task classes with no test examples. The gate cannot speak for these."""
return [
task for task in TASK_CLASSES
if self.test_metrics.task_support.get(task, 0) == 0
]
```
(`router/train.py:128-134`.) A class absent from the test set is **named**, not silently scored as
zero. The shipped result has an empty list β all six classes have test support (13 / 13 / 17 / 12 /
10 / 15).
### 18.2 `gate2_passed` β the two-part condition
```python
@property
def gate2_passed(self) -> bool:
"""Gate 2 passes only if EVERY task class was measured AND cleared 0.95.
Without the first condition a class can be absent from the test set and
the gate reports a pass over a model that was never asked about it.
"""
return not self.unmeasured_test_tasks and self.test_metrics.task_accuracy >= 0.95
```
(`router/train.py:136-143`.)
**Two conditions, ANDed, and the first is the novel one.** A naive gate would check only
`task_accuracy >= 0.95`; that gate passed at 0.909 over five of six classes in the defective run.
The two-part gate cannot pass while any class is unmeasured, which is what forces the stratification
fix rather than allowing it to be skipped.
The gate's own threshold is **0.95**, and it is a hardcoded literal in the property β not a config
key. The shipped test accuracy is **0.975**, so `gate2_passed` is `True`, and
`docs/PHASE4_ROUTER_REPORT.md` records the verdict block verbatim:
```
GATE 2 ROUTER ACCEPTANCE: task accuracy >= 0.95, all classes measured
classes with test support : 6/6
-> PASS (0.975)
```
### 18.3 `summary()` β the printed report
`TrainingResult.summary()` (`router/train.py:145-213`) builds a fixed-width text report with five
blocks: a header with corpus / groups / hash / encoder / adapter params / epochs / duration; the
split sizes and the leakage-clean flag; a metrics table with one row per split; the hard-negative
accuracy when present; the per-task recall table with an `n/a` row and the explicit
`(no test examples)` marker for a zero-support class; and the gate verdict.
Two details in the gate block are worth quoting because they are the honesty mechanism in
human-readable form:
```python
measured = len(TASK_CLASSES) - len(self.unmeasured_test_tasks)
lines.append(f" classes with test support : {measured}/{len(TASK_CLASSES)}")
if self.unmeasured_test_tasks:
lines.append(
f" classes NOT measurable : {', '.join(self.unmeasured_test_tasks)}"
)
acc = self.test_metrics.task_accuracy
if self.unmeasured_test_tasks:
lines.append(
f" -> NOT MEASURABLE ({acc:.3f} over the {measured} classes that "
f"have test data)"
)
else:
verdict = "PASS" if acc >= 0.95 else "BELOW TARGET"
lines.append(f" -> {verdict} ({acc:.3f})")
```
(`router/train.py:195-209`.) An unmeasured class produces the phrase **"NOT MEASURABLE"**, never
"BELOW TARGET" β the two are different facts and the report distinguishes them.
---
# Part E β Inference
## 19. `RouterPrediction` β the decision plus its explanation
```python
@dataclass
class RouterPrediction:
"""A router decision plus everything needed to explain it in a trace."""
intent: Intent
task_probs: dict[str, float] = field(default_factory=dict)
modality_probs: dict[str, float] = field(default_factory=dict)
binary_probs: dict[str, float] = field(default_factory=dict)
used_fallback: bool = False
fallback_rule: str | None = None
matched_terms: tuple[str, ...] = ()
above_threshold: bool = True
def to_trace(self) -> dict[str, Any]:
"""Observable facts only. No chain-of-thought."""
return {
"task": self.intent.task.value,
"modality": self.intent.modality.value,
"temporal": self.intent.temporal,
"spatial_output": self.intent.spatial_output,
"language_output": self.intent.language_output,
"confidence": round(self.intent.confidence, 4),
"source": self.intent.source,
"above_threshold": self.above_threshold,
"used_fallback": self.used_fallback,
"fallback_rule": self.fallback_rule,
}
```
(`router/classifier.py:80-106`.)
| Field | Populated by the learned path? | Populated by the fallback? |
|---|---|---|
| `intent` | yes β `Intent(source="learned")` | yes β `Intent(source="lexical_fallback")` |
| `task_probs` | yes β all six softmax probabilities | **no** β empty dict |
| `modality_probs` | yes β all four | **no** β empty dict |
| `binary_probs` | yes β the three sigmoid values | **no** β empty dict |
| `used_fallback` | `False` | `True` |
| `fallback_rule` | `None` | the rule id, e.g. `"spatial"` |
| `matched_terms` | `()` | the terms that fired |
| `above_threshold` | `confidence >= 0.70` | `match.confidence >= 0.70` |
**`to_trace()` deliberately omits the probability dicts.** It emits the task, modality, three
booleans, the rounded confidence, the source, `above_threshold`, `used_fallback` and
`fallback_rule` β and nothing else. `README.md:601-603` states the rule:
> The router's own trace projection is the model of this: it emits the task, modality, the three
> booleans, the rounded confidence, the source, `above_threshold`, `used_fallback` and
> `fallback_rule` β and nothing else.
The docstring's *"Observable facts only. No chain-of-thought."* is the enforcement note. Note that
`matched_terms` is **not** in the trace projection even though it is on the dataclass β the
fallback's matched substrings are diagnostic metadata for a caller, not part of the trace contract.
## 20. `IntentRouter` β construction and the `from_config` warning
```python
class IntentRouter:
"""Loads the adapter (and optionally the encoder) and routes queries."""
def __init__(
self,
adapter: IntentAdapter | None = None,
encoder: FrozenEncoder | None = None,
confidence_threshold: float = 0.70,
device: str = "cpu",
) -> None:
if not 0.0 <= confidence_threshold <= 1.0:
raise RoutingError(
f"confidence_threshold must be in [0,1], got {confidence_threshold}"
)
self.adapter = adapter
self.encoder = encoder
self.confidence_threshold = confidence_threshold
self.device = device
if self.adapter is not None:
self.adapter.eval()
self.adapter.to(device)
```
(`router/classifier.py:109-130`.)
**`confidence_threshold = 0.70`** β the default matches `configs/base.yaml`
Β§`router.confidence_threshold: 0.70`, and `from_config` reads the config value rather than relying
on the default.
**An adapter, once passed, is immediately put in eval mode and moved to the device.** So the caller
cannot accidentally route with a training-mode adapter.
### 20.1 `from_config()` and the warning that is the whole point
```python
@classmethod
def from_config(
cls,
config,
adapter_path: str | Path | None = None,
load_encoder: bool = True,
device: str | None = None,
) -> "IntentRouter":
"""Build from the central config, optionally loading a trained adapter.
IMPORTANT: `adapter_path` defaults to None, so the default router runs
the LEXICAL FALLBACK, not the trained adapter. That default is
deliberate -- it keeps tests fast and offline -- but it means a caller
who never passes `adapter_path` gets fallback answers while believing
the trained model is serving. The two are distinguishable but not
obviously so: on the spec section 29 examples the fallback returns
confidence 0.850-0.920 against the trained model's 0.780-1.000.
Check `router.has_adapter` (and `router.adapter_source`) before
trusting a routing decision as model-backed.
"""
```
(`router/classifier.py:134-154`.)
**This warning is the most consequential paragraph in the router.** A default-constructed router is
a **lexical** router. The two paths produce answers in **overlapping confidence bands**, and the
fallback can be **more** confident than the trained model β so confidence alone cannot distinguish
them. `core/planner.py:48-62` records this as the reason the planner applies a provenance discount:
> A lexical fallback at 0.9 is not the same evidence as a learned model at 0.9: one is a regex that
> matched, the other is a learned distribution. Treating them identically would let a matched
> keyword outrank the model it fell back from.
**Two dimension checks, both loud:**
```python
if adapter_path is not None:
adapter, metadata = load_adapter(adapter_path, device=device)
# An adapter trained against a different input dimension cannot be
# used with this encoder; fail loudly rather than producing garbage.
expected_dim = int(config.get("router.embedding_dim", 384))
if adapter.input_dim != expected_dim:
raise ModelLoadError(
f"adapter input_dim={adapter.input_dim} does not match the "
f"configured embedding dim {expected_dim}",
specialist="router",
)
if load_encoder:
from router.encoder import build_encoder
encoder = build_encoder(config, device=device)
if adapter is not None and adapter.input_dim != encoder.embedding_dim:
raise ModelLoadError(
f"adapter input_dim={adapter.input_dim} but the loaded encoder "
f"produces {encoder.embedding_dim}-d embeddings",
specialist="router",
)
```
(`router/classifier.py:164-185`.)
**Check one compares the adapter to the *config*; check two compares it to the *loaded encoder*.**
Both are needed: check one catches a config that disagrees with the adapter before the encoder is
loaded at all; check two catches a config that is *right* while the actual encoder is different.
The `from router.encoder import build_encoder` is a **deferred import inside the method** so that
`load_encoder=False` does not pay the import cost.
### 20.2 The three properties
```python
@property
def has_adapter(self) -> bool:
return self.adapter is not None
@property
def adapter_source(self) -> str:
"""Which path `route()` will take: 'trained' or 'lexical_fallback'.
...
"""
return "trained" if self.has_adapter else "lexical_fallback"
@property
def has_encoder(self) -> bool:
return self.encoder is not None
```
(`router/classifier.py:196-213`.)
`adapter_source` returns `"trained"` when an adapter is present β **regardless of whether an
encoder is loaded**. So a router with an adapter but no encoder reports `"trained"` while `route()`
actually takes the fallback path (because `has_encoder` is `False`, so `_predict_learned` is never
called). This is a genuine edge case worth naming: the property answers *"is a trained adapter
attached?"*, not *"will the learned path run?"*. The two differ exactly when
`has_adapter and not has_encoder`. `core/planner.py:50` relies on `adapter_source` for provenance, so
a caller in that state would mislabel the source. It is `IMPLEMENTED` as written; whether it is a
defect depends on the intended contract, and the property's docstring β *"Which path `route()` will
take"* β reads as the stronger claim. Recorded as an **`OPEN`** observation, not asserted as a bug.
## 21. `_predict_learned()` β the forward pass and the coherence repair
```python
def _predict_learned(self, query: str) -> RouterPrediction | None:
"""Run the encoder + adapter. Returns None when either is unavailable."""
if self.adapter is None or self.encoder is None:
return None
embedding = self.encoder.encode_one(query, normalize=True)
tensor = torch.from_numpy(embedding).unsqueeze(0).to(self.device)
with torch.no_grad():
out = self.adapter(tensor)
task_probs = torch.softmax(out.task_logits, dim=-1)[0]
modality_probs = torch.softmax(out.modality_logits, dim=-1)[0]
temporal_p = torch.sigmoid(out.temporal_logit)[0]
spatial_p = torch.sigmoid(out.spatial_logit)[0]
language_p = torch.sigmoid(out.language_logit)[0]
task_idx = int(torch.argmax(task_probs).item())
modality_idx = int(torch.argmax(modality_probs).item())
confidence = float(task_probs[task_idx].item())
```
(`router/classifier.py:217-235`.)
**The confidence is the max softmax probability of the task head**, not a modality or binary
probability. That is a deliberate choice: the task is what the controller keys off, so the gate
gates on the task.
**`encode_one` β `unsqueeze(0)` β `torch.no_grad()`.** The `unsqueeze(0)` adds the batch dimension
the adapter's forward guard requires; the `no_grad` is the inference-mode guard.
**The coherence repair, in full:**
```python
# Coherence repair. The heads are independent by construction, so a
# confident task label with an incoherent binary head is possible. We
# resolve in favour of the task label, because the task is what the
# controller keys off β and we record that we did so.
temporal = binary_probs["temporal"] >= 0.5
spatial = binary_probs["spatial_output"] >= 0.5
language = binary_probs["language_output"] >= 0.5
if task_label in ("change",):
temporal = True
if task_label == "grounding":
spatial = True
if task_label == "unsupported":
language = False
spatial = False
temporal = False
if task_label == "optical_sar" and modality_label != "optical_sar":
modality_label = "optical_sar"
```
(`router/classifier.py:246-263`.)
**Four repairs, and each maps to a label-space invariant:**
| Repair | Rule | Label-space source |
|---|---|---|
| `change` β `temporal = True` | a change task always needs two acquisitions | `TEMPORAL_TASKS = {"change"}` |
| `grounding` β `spatial = True` | grounding's output is spatial | `SPATIAL_TASKS = {"grounding"}` |
| `unsupported` β all three `False` | unsupported claims no capability | `RouterCorpus.validate()` check (d) |
| `optical_sar` β `modality = "optical_sar"` | a dual-modality task has a dual-modality label | `DUAL_MODALITY_TASKS = {"optical_sar"}` |
**Note what is *not* repaired:** a `caption` task with `spatial_output=True` is left alone; a `vqa`
task with `temporal=True` is left alone. The repairs are the *necessary* implications only, matching
the asymmetry in `RouterCorpus.validate()` (Β§11.2(e)). The comment's *"we resolve in favour of the
task label β¦ and we record that we did so"* is slightly stronger than the code: the repair happens,
but **nothing records that a repair fired** β `RouterPrediction.to_trace()` has no
`coherence_repaired` field. The repair is therefore silent. This is a small gap between the comment
and the implementation, named here rather than smoothed over.
**The `Intent` construction is wrapped in a try/except that converts a Pydantic error into a
`RoutingError`:**
```python
try:
intent = Intent(
task=_TASK_TO_SCHEMA[task_label],
modality=_MODALITY_TO_SCHEMA[modality_label],
temporal=temporal,
spatial_output=spatial,
language_output=language,
confidence=confidence,
source="learned",
)
except Exception as exc: # noqa: BLE001 - pydantic validation
raise RoutingError(
f"learned router produced an invalid intent: {exc}",
context={"query_length": len(query), "task": task_label},
) from exc
```
(`router/classifier.py:265-279`.) **The error context deliberately carries `query_length`, not the
query.** That is a privacy-conscious choice: a routing failure does not put the user's text into an
error record.
**`Intent`'s own validator will also fire** for two cases the repair already handled β the
`_consistency` model validator (`core/schemas.py:107-114`) forces `temporal=True` for `CHANGE` and
`CHANGE_VQA`, and forces `modality=OPTICAL_SAR` for an `OPTICAL_SAR` task with `UNKNOWN` modality.
So the repair and the schema validator **agree**, and the repair is what keeps the schema validator
from having to do the work.
## 22. `_predict_fallback()` β the lexical path
```python
def _predict_fallback(self, query: str) -> RouterPrediction:
match: LexicalMatch = lexical_route(query)
intent = Intent(
task=_TASK_TO_SCHEMA[match.task],
modality=_MODALITY_TO_SCHEMA[match.modality],
temporal=match.temporal,
spatial_output=match.spatial_output,
language_output=match.language_output,
confidence=match.confidence,
source="lexical_fallback",
)
return RouterPrediction(
intent=intent,
used_fallback=True,
fallback_rule=match.rule,
matched_terms=match.matched_terms,
above_threshold=match.confidence >= self.confidence_threshold,
)
```
(`router/classifier.py:292-309`.)
**No try/except here** β because `lexical_route` is total: every path returns a valid
`LexicalMatch`, and `self_check()` asserts the label validity of the curated cases. The
`matched_terms` and `fallback_rule` are carried through, which is what makes the fallback
**inspectable** where the learned path is not.
## 23. `route()` β the three-step arbitration
```python
def route(self, query: str) -> RouterPrediction:
"""Route a query to an Intent.
Order:
1. learned router, if confident enough -> use it
2. otherwise lexical fallback
3. if the fallback also lands below threshold -> return it anyway,
flagged, and let the controller decide. The router never refuses
to answer; `above_threshold` carries the uncertainty.
"""
if not isinstance(query, str) or not query.strip():
raise RoutingError("query must be a non-empty string")
learned: RouterPrediction | None = None
if self.has_adapter and self.has_encoder:
learned = self._predict_learned(query)
if learned is not None and learned.above_threshold:
return learned
fallback = self._predict_fallback(query)
if learned is None:
return fallback
# Both ran. Prefer whichever is more confident; the fallback wins ties
# on interpretability (its rule and matched terms are inspectable).
if fallback.intent.confidence >= learned.intent.confidence:
return fallback
return learned
```
(`router/classifier.py:311-339`.)
**Trace every branch.** This is the algorithm of the entire router in 28 lines.
| Situation | What runs | What is returned |
|---|---|---|
| no adapter **and** no encoder | fallback only | fallback |
| adapter + encoder, learned β₯ 0.70 | learned only (early return) | learned |
| adapter + encoder, learned < 0.70, fallback β₯ learned | both, then compare | fallback |
| adapter + encoder, learned < 0.70, fallback < learned | both, then compare | learned (flagged `above_threshold=False`) |
| adapter + encoder, **learned < 0.70 and fallback also < 0.70** | both, then compare | whichever is larger, flagged below threshold |
**The early return is the important control-flow fact.** When the learned router clears the gate,
the fallback **never runs** β so on a confident learned route, `used_fallback` is `False`,
`fallback_rule` is `None`, and `task_probs` / `modality_probs` / `binary_probs` are all populated.
**When both run, the tie-break favours the fallback**, on the stated ground of *interpretability*:
the fallback's rule and matched terms are inspectable, the learned path's are not. Note that
`>=` means a **tie goes to the fallback**.
**The router never raises on a low-confidence route.** It returns the prediction with
`above_threshold=False`, and `core/planner.py:43-46` explains why the planner consumes that bool
rather than re-reading the threshold:
> So the planner consumes `above_threshold` (a bool), never the numeric threshold. Re-reading
> `router.confidence_threshold` here would be a second, independent gate over the same quantity β
> two places bound to one knob, which is how a threshold ends up meaning two different things.
**One input guard only:** a non-string or blank query raises `RoutingError`. Nothing else about the
query is validated here; length, language and content are the specialist's concern.
`route_batch(queries)` is a trivial list comprehension over `route` (`router/classifier.py:341-342`)
β it is **not** batched inference, and the name should not be read as such. Each query gets its own
`encode_one` call.
## 24. Persistence β `save_adapter` and `load_adapter`
```python
ADAPTER_WEIGHTS = "adapter.pt"
ADAPTER_METADATA = "metadata.json"
```
(`router/classifier.py:349-350`.)
```python
def save_adapter(
directory: str | Path,
adapter: IntentAdapter,
metadata: dict[str, Any],
) -> Path:
"""Write the adapter plus its metadata. Metadata is mandatory."""
path = Path(directory)
path.mkdir(parents=True, exist_ok=True)
torch.save(
{
"state_dict": adapter.state_dict(),
"config": adapter.config_dict(),
},
path / ADAPTER_WEIGHTS,
)
payload = dict(metadata)
payload.setdefault("adapter_config", adapter.config_dict())
payload.setdefault("num_parameters", adapter.num_parameters())
(path / ADAPTER_METADATA).write_text(
json.dumps(payload, indent=2, sort_keys=True, default=str), encoding="utf-8"
)
return path
```
(`router/classifier.py:353-376`.)
**Three details:**
1. **Metadata is mandatory** β the parameter has no default. A checkpoint without metadata cannot
be written by this function.
2. **`setdefault`, not assignment.** If `train_router` already put `adapter_config` or
`num_parameters` in the metadata, they are preserved. The shipped artifact carries
`adapter_config` and a top-level `num_parameters` alongside the `adapter` block β three
representations of the same facts, which is why the file is verbose but self-describing.
3. **`sort_keys=True`** β the JSON is deterministic, so two identical runs produce byte-identical
metadata (given identical inputs). `default=str` handles any non-JSON-serialisable value by
stringifying rather than raising.
### 24.1 `load_adapter()` and the doubled-path trap
```python
def load_adapter(
directory: str | Path,
device: str = "cpu",
) -> tuple[IntentAdapter, dict[str, Any]]:
"""Load a saved adapter and its metadata.
Args:
directory: the artifact DIRECTORY written by `save_adapter`, OR the
weights file itself. Both are accepted: `save_adapter` writes
`<dir>/adapter.pt`, so a caller who was handed the weight path --
the obvious thing to try, since the name reads like a weight file --
would otherwise have `.pt` appended a second time.
...
"""
path = Path(directory)
# Accept the weights file directly. Without this, passing `.../adapter.pt`
# produces `<...>/adapter.pt/adapter.pt`, whose doubled filename is the
# tell but is easy to misread as a genuinely missing file.
if path.is_file():
if path.name != ADAPTER_WEIGHTS and path.suffix != ".pt":
raise ModelLoadError(
f"expected an adapter directory or a '{ADAPTER_WEIGHTS}' file, "
f"got the file {path}",
specialist="router",
)
weights = path
meta_path = path.parent / ADAPTER_METADATA
else:
weights = path / ADAPTER_WEIGHTS
meta_path = path / ADAPTER_METADATA
```
(`router/classifier.py:379-413`.)
**This is a fix for a real usability failure, documented as such.** `save_adapter` writes
`<dir>/adapter.pt`; a caller handed `.../adapter.pt` would naively get
`<dir>/adapter.pt/adapter.pt`. The loader accepts either form.
**A second, more specific error for the same trap:**
```python
if not weights.exists():
# Name the specific confusion when a `.pt` path was given that exists
# as a directory or not at all, rather than reporting a doubled path.
if path.suffix == ".pt":
raise ModelLoadError(
f"no adapter weights at {weights}: {path} looks like a weights "
f"file, not an adapter directory. Pass the file itself (it must "
f"exist and end in .pt) or the directory containing "
f"'{ADAPTER_WEIGHTS}'.",
specialist="router",
context={"given": str(path), "resolved": str(weights)},
)
raise ModelLoadError(
f"adapter weights not found at {weights}", specialist="router"
)
```
(`router/classifier.py:415-429`.) The error's `context` dict carries both the given and the resolved
path, so a caller can see the doubling rather than infer it.
**Metadata is optional at load time, and its absence is not fatal:**
```python
metadata: dict[str, Any] = {}
if meta_path.exists():
try:
metadata = json.loads(meta_path.read_text(encoding="utf-8"))
except json.JSONDecodeError:
metadata = {"_warning": "metadata.json was not valid JSON"}
```
(`router/classifier.py:458-463`.) A malformed `metadata.json` yields a `_warning` key rather than an
exception β because the metadata is descriptive, while the **weights and the embedded config** are
what the model needs to run.
**The asymmetry between the two files is the design:** `adapter.pt` is **required** and its absence
or mismatch is fatal; `metadata.json` is **advisory** and its absence is silent. `router.pt`'s
embedded `config` is what makes the weights self-sufficient.
---
# Part F β The lexical fallback
## 25. What the fallback is, and the four rules it obeys
```python
"""SatQuery AI β deterministic lexical intent fallback.
The learned router handles natural phrasing. This module exists for the case
the router itself flags as low-confidence, and for environments where the
encoder cannot be loaded at all (a CPU-only Space with no cached weights).
Design rules:
* Purely lexical. No model, no embeddings, no randomness.
* Ordered rules, highest specificity first. The first match wins.
* Never invents capability. If nothing matches, it returns `unsupported`
with low confidence rather than guessing a task.
* Must agree with the learned router on the curated hard negatives. If the
two disagree on `"describe the water body"` vs `"show me the water body"`,
the fallback is wrong, not the router.
"""
```
(`router/fallback.py:1-16`.)
**The fourth rule is unusual and important.** It assigns the burden of disagreement: if the fallback
and the router disagree on a curated hard negative, **the fallback is wrong**. The learned router was
trained on those pairs; the fallback is hand-written. `self_check()` (Β§29) is the mechanism that
holds the fallback to it.
### 25.1 `LexicalMatch` β the returned evidence
```python
@dataclass(frozen=True)
class LexicalMatch:
"""One rule hit, with the evidence that produced it."""
task: str
modality: str
temporal: bool
spatial_output: bool
language_output: bool
confidence: float
rule: str
matched_terms: tuple[str, ...]
```
(`router/fallback.py:29-40`.) **`frozen=True`** β a match is immutable, so a caller cannot mutate a
rule's output. The eight fields are exactly the five `Intent` slots plus `confidence`, plus the two
diagnostic fields (`rule`, `matched_terms`).
| Field | Example | Meaning |
|---|---|---|
| `task` | `"grounding"` | one of `TASK_CLASSES` |
| `modality` | `"optical"` | one of `MODALITY_CLASSES` |
| `temporal` | `False` | β |
| `spatial_output` | `True` | β |
| `language_output` | `True` | β |
| `confidence` | `0.88` | the rule's fixed confidence |
| `rule` | `"spatial"` | the rule id β one of 10 values (Β§27.10) |
| `matched_terms` | `("where", "locate")` | the substrings that fired |
## 26. The seven term tables, in full
Every table is a `tuple[str, ...]`, matched by **substring containment** after lower-casing and
stripping the query (`router/fallback.py:121-126`):
```python
def _hits(text: str, terms: tuple[str, ...]) -> tuple[str, ...]:
return tuple(term for term in terms if term in text)
def _contains_any(text: str, terms: tuple[str, ...]) -> bool:
return any(term in text for term in terms)
```
**`in`, not a regex, not a word-boundary match.** So `"where"` matches inside `"somewhere"`, and
`"what"` matches inside `"whatever"`. This is a deliberate simplicity trade-off β the fallback is
the *fallback*, and the learned router handles nuance β but it is the mechanism behind the
residuals in Β§48.
### 26.1 `_SPATIAL_TERMS` β 14 terms, including bare `"where"`
```python
#: "show me where" phrasings. Spatial intent.
_SPATIAL_TERMS: tuple[str, ...] = (
# Bare "where" is included deliberately. "Where did the change happen?"
# must resolve to change + spatial_output=True, and the temporal branch
# consults this table only AFTER the temporal terms have already matched.
# Without it, "where did" carries no spatial signal at all.
"where", "where is", "where are", "where's",
"locate", "location of",
"show me where", "show where",
"highlight", "mark the", "point out", "point to",
"draw a box", "draw box", "bounding box",
"find the", "find all",
"show me the", "show the region",
)
```
(`router/fallback.py:47-60`.)
**Count: 19 entries**, not 14. The comment explains the two load-bearing ones:
- **Bare `"where"`** is the term that makes `"Where did the change happen?"` produce
`change + spatial_output=True`. It is safe to include bare because the spatial table is consulted
**after** the temporal table (Β§27.3), so a temporal query never reaches it as a *task* selector β
only as a *spatial flag*.
- **`"show me the"`** is the term that separates `"show me the water body"` (grounding) from
`"describe the water body"` (caption) β the first curated hard-negative family.
Note `"draw box"` as well as `"draw a box"` β the informal phrasing without the article.
### 26.2 `_TEMPORAL_TERMS` β 20 terms
```python
#: Temporal phrasings.
_TEMPORAL_TERMS: tuple[str, ...] = (
"changed", "change", "changes", "changing",
"before and after", "before-and-after",
"between these two", "between the two",
# "Compare these two images." is a temporal request. The superficially
# similar "Compare the optical and radar images." is claimed by the
# dual-modality table, which is consulted first.
"these two images", "two images", "compare these", "compare the two",
"over time", "temporal", "time series", "time-series",
"differences between the two dates", "two dates",
"has the", "did the", "have the",
"compared to before", "since then",
)
```
(`router/fallback.py:62-75`.)
**Count: 24 entries.** The comment records the near-miss that the precedence rule resolves:
`"Compare these two images."` (temporal) versus `"Compare the optical and radar images."`
(dual-modality). The surface form is near-identical; the dual-modality table is consulted first, so
the second is claimed correctly.
**Note that `"change"` is a substring of `"changed"`, `"changes"` and `"changing"`** β so the four
entries `changed`/`change`/`changes`/`changing` are redundant under substring matching: `"change"`
alone would match all four. They are listed explicitly for readability and for the matched-terms
output, which will report whichever entries the loop found. This is harmless redundancy, not a bug.
### 26.3 `_DUAL_MODALITY_TERMS` β 18 entries
```python
#: Dual-modality phrasings.
_DUAL_MODALITY_TERMS: tuple[str, ...] = (
"optical and sar", "optical and radar", "sar and optical",
"radar and optical", "radar and optical", "both images",
"both sensors", "both modalities", "two modalities",
"sar image", "radar image", "radar data", "sar data",
"co-registered", "coregistered", "fuse", "fusion",
"jointly", "multi-modal", "multimodal",
)
```
(`router/fallback.py:77-85`.)
**Count: 20 entries, with `"radar and optical"` listed twice** β a duplicated entry, harmless under
`_hits` (which would report it twice in `matched_terms`) but a genuine source redundancy. Recorded
here as a cosmetic observation.
The `explicit_pair` subset used by rule 1 is:
```python
explicit_pair = any(
term in dual_hits
for term in ("optical and sar", "optical and radar", "sar and optical",
"radar and optical", "both sensors", "both modalities",
"two modalities", "fuse", "fusion", "jointly",
"multi-modal", "multimodal", "co-registered", "coregistered")
)
```
(`router/fallback.py:175-181`.) Note that `"both images"`, `"sar image"`, `"radar image"`,
`"radar data"` and `"sar data"` are in the table but **not** in `explicit_pair` β so a query
containing only `"both images"` does **not** trigger the dual-modality rule. The comment above the
rule explains why: *"'both images' alone is ambiguous; require a modality word too."*
(`router/fallback.py:174`.) This is the `"two images"` (temporal) versus `"both images"` (ambiguous)
distinction, and it is what keeps `"Compare these two images."` on the temporal branch.
### 26.4 `_CAPTION_TERMS` β 11 entries
```python
#: Caption phrasings.
_CAPTION_TERMS: tuple[str, ...] = (
"describe", "description", "caption", "write a caption",
"summarize", "summarise", "summary of",
"tell me about this image", "what do you see",
"what can you see", "explain this image",
)
```
(`router/fallback.py:87-93`.) **Count: 11.** Note `"summarize"` and `"summarise"` are both present
for the en-US / en-GB split β a genuine, necessary pair, unlike the redundant `change` family.
### 26.5 `_VQA_TERMS` β 15 entries, including bare `"what"`
```python
#: VQA phrasings.
_VQA_TERMS: tuple[str, ...] = (
# Bare "what" is needed because "What land cover is visible?" carries no
# other interrogative marker. Precedence protects it: the temporal,
# spatial and caption tables are all consulted before this one, so
# "what changed" and "what do you see" never reach here.
"what",
"how many", "how much", "is there", "are there",
"what is", "what are", "what kind", "what type",
"which", "does this", "do you see", "can you tell",
"is this", "are these", "identify the type",
)
```
(`router/fallback.py:95-106`.)
**Count: 16.** Bare `"what"` is the load-bearing entry, and the comment explains both why it is
needed and why it is safe: the temporal, spatial and caption tables are all consulted first, so
`"what changed"` and `"what do you see"` never reach the VQA rule.
**This is the same bare-interrogative pattern as bare `"where"`**, and it is the pattern most
vulnerable to the substring-matching limitation. `"what"` inside `"whatever"`, `"what's"` and
`"somewhat"` all fire.
### 26.6 `_SAR_TERMS` and `_OPTICAL_TERMS`
```python
#: SAR-only phrasings.
_SAR_TERMS: tuple[str, ...] = (
"sar image", "sar data", "radar image", "radar data",
"synthetic aperture radar", "backscatter",
)
#: Optical-only phrasings.
_OPTICAL_TERMS: tuple[str, ...] = (
"optical image", "optical data", "multispectral", "true colour",
"true color", "rgb image", "visible image",
)
```
(`router/fallback.py:108-118`.) **Six and seven entries.** These two tables do **not** select a task
β they select a **modality**, and they are consulted in rules 4 and 5, *after* the temporal and
spatial rules. Note `"true colour"` and `"true color"` are both present, the second necessary
spelling pair.
**Both tables overlap `_DUAL_MODALITY_TERMS`.** `"sar image"` and `"radar image"` appear in both
`_SAR_TERMS` and `_DUAL_MODALITY_TERMS`, and `"sar data"`/`"radar data"` likewise. The precedence
rule resolves it: rule 1 (dual-modality) is checked first, so `"sar image"` alone triggers
`explicit_pair` only if a *pair* term is also present. If it is not, the query falls through to
rule 4 and becomes a **SAR-only VQA** request. That is the intended reading of
`"What does this sar image show?"` β a single-SAR-modality question, not a fusion request.
## 27. The eight ordered rules
`lexical_route(query)` is one function with eight sequential returns. The docstring states the
precedence and its two motivating examples:
```python
def lexical_route(query: str) -> LexicalMatch:
"""Classify a query with ordered lexical rules.
Rule order encodes precedence, which matters because the phrasings overlap:
"show me where the change happened"
has both a spatial term AND a temporal term -> change + spatial
"compare optical and radar to locate built-up areas"
has dual-modality AND spatial -> optical_sar (spatial doesn't apply
to the joint workflow, whose output is a classification)
Precedence: dual_modality > temporal > spatial > caption > vqa > unsupported
"""
```
(`router/fallback.py:134-147`.)
**Note the docstring's precedence line versus the code's actual rule order.** The docstring lists six
levels; the code has **eight** rules, and the two extra ones β `sar_single` and `optical_single` β
sit **between** `spatial` and `caption`:
```
implemented order: dual_modality > temporal > spatial > sar_single > optical_single > caption > vqa > unsupported
docstring order: dual_modality > temporal > spatial > caption > vqa > unsupported
README order: dual_modality > temporal > spatial > caption > vqa > unsupported
```
The docstring and `README.md:537` agree with each other and both **omit the two modality rules**.
The code is the authority, and the full eight-level precedence is the one stated above. This is a
documentation gap, not a behavioural one β the modality rules always *precede* `caption`, so a query
that would match both a modality term and a caption term is classified by modality. Recorded here so
a reader working from the README is not surprised by the code.
### 27.1 Rule 0 β the empty query
```python
if not isinstance(query, str) or not query.strip():
return LexicalMatch(
task="unsupported",
modality="unknown",
temporal=False,
spatial_output=False,
language_output=False,
confidence=0.0,
rule="empty_query",
matched_terms=(),
)
text = query.lower().strip()
```
(`router/fallback.py:148-160`.)
**`confidence = 0.0`** β the lowest confidence in the module, and `task="unsupported"`. An empty
query is not a low-confidence guess; it is a definite non-request. Note this branch also catches a
**non-string** input, so `lexical_route(None)` returns a match rather than raising. The
`text` normalisation (lower-case + strip) is computed once, before any table is consulted.
**All seven tables are evaluated before any rule fires:**
```python
dual_hits = _hits(text, _DUAL_MODALITY_TERMS)
temporal_hits = _hits(text, _TEMPORAL_TERMS)
spatial_hits = _hits(text, _SPATIAL_TERMS)
caption_hits = _hits(text, _CAPTION_TERMS)
vqa_hits = _hits(text, _VQA_TERMS)
sar_hits = _hits(text, _SAR_TERMS)
optical_hits = _hits(text, _OPTICAL_TERMS)
```
(`router/fallback.py:162-168`.) This is why `spatial_hits` is available to the temporal rule (rule 2
uses it to set `spatial_output`) even though the spatial *rule* (rule 3) has not been reached.
### 27.2 Rule 1 β `dual_modality`, confidence 0.92
```python
if dual_hits:
explicit_pair = any(...)
if explicit_pair:
return LexicalMatch(
task="optical_sar",
modality="optical_sar",
temporal=False,
spatial_output=False,
language_output=True,
confidence=0.92,
rule="dual_modality",
matched_terms=dual_hits,
)
```
(`router/fallback.py:170-192`.) **The highest confidence in the module (0.92)** and the only rule
that sets `modality="optical_sar"`. `spatial_output=False` is forced β the comment explains: the
joint workflow's output is a classification, not a box, so a spatial term in the query does not make
the output spatial. **The `if dual_hits:` outer guard is redundant** given the `if explicit_pair:`
inner guard (an empty `dual_hits` makes `explicit_pair` false), but it short-circuits the generator
expression. Harmless.
### 27.3 Rule 2 β `temporal` / `temporal_spatial`, confidence 0.85 / 0.90
```python
if temporal_hits:
spatial = _contains_any(text, _SPATIAL_TERMS)
return LexicalMatch(
task="change",
modality="optical",
temporal=True,
spatial_output=spatial,
language_output=True,
confidence=0.90 if spatial else 0.85,
rule="temporal_spatial" if spatial else "temporal",
matched_terms=temporal_hits + spatial_hits,
)
```
(`router/fallback.py:194-206`.) **This is the only rule with two confidences and two rule ids.** The
`spatial` flag is computed from the spatial table regardless of rule order, which is the mechanism
that makes `"Where did the change happen?"` produce `change + spatial_output=True` at 0.90.
`modality="optical"` is hardcoded β a temporal request is assumed optical unless rule 1 claimed it.
**`matched_terms=temporal_hits + spatial_hits`** β both hit sets are concatenated, so the diagnostic
shows the temporal terms first.
### 27.4 Rule 3 β `spatial`, confidence 0.88
```python
if spatial_hits:
return LexicalMatch(
task="grounding",
modality="optical",
temporal=False,
spatial_output=True,
language_output=True,
confidence=0.88,
rule="spatial",
matched_terms=spatial_hits,
)
```
(`router/fallback.py:208-219`.) The plain grounding rule. Because rule 2 has already claimed every
temporal query, a query reaching here has a spatial term and no temporal term.
### 27.5 Rule 4 β `sar_single`, confidence 0.75
```python
if sar_hits and not optical_hits:
return LexicalMatch(
task="vqa",
modality="sar",
temporal=False,
spatial_output=False,
language_output=True,
confidence=0.75,
rule="sar_single",
matched_terms=sar_hits,
)
```
(`router/fallback.py:221-232`.) **The task is `vqa`** β a single-modality question about a SAR
image. The guard `and not optical_hits` is what distinguishes a SAR-only query from a dual-modality
one; a query mentioning both falls through to rule 5 (which also fails its `and not sar_hits`) and
then to caption/vqa.
### 27.6 Rule 5 β `optical_single`, confidence 0.72
```python
if optical_hits and not sar_hits:
return LexicalMatch(
task="vqa",
modality="optical",
temporal=False,
spatial_output=False,
language_output=True,
confidence=0.72,
rule="optical_single",
matched_terms=optical_hits,
)
```
(`router/fallback.py:234-245`.) **The lowest task-producing confidence (0.72)**, and the mirror of
rule 4. The two guards are mutually exclusive, so at most one of rules 4 and 5 can fire.
### 27.7 Rule 6 β `caption`, confidence 0.85
```python
if caption_hits:
return LexicalMatch(
task="caption",
modality="unknown",
temporal=False,
spatial_output=False,
language_output=True,
confidence=0.85,
rule="caption",
matched_terms=caption_hits,
)
```
(`router/fallback.py:247-258`.) `modality="unknown"` β a caption request does not imply a sensor.
### 27.8 Rule 7 β `vqa`, confidence 0.78
```python
if vqa_hits:
return LexicalMatch(
task="vqa",
modality="unknown",
temporal=False,
spatial_output=False,
language_output=True,
confidence=0.78,
rule="vqa",
matched_terms=vqa_hits,
)
```
(`router/fallback.py:260-271`.)
### 27.9 Rule 8 β `no_match`, confidence 0.30
```python
# -- 8. nothing matched: refuse rather than guess -------------------
return LexicalMatch(
task="unsupported",
modality="unknown",
temporal=False,
spatial_output=False,
language_output=False,
confidence=0.30,
rule="no_match",
matched_terms=(),
)
```
(`router/fallback.py:273-283`.)
**This is the rule that implements "never invents capability".** A query matching nothing is
`unsupported` at **0.30** β below the 0.70 gate, so `above_threshold` is `False` and the planner
sees an uncertain refusal rather than a confident one. `language_output=False` is required: an
`unsupported` result must claim no capability, exactly as `RouterCorpus.validate()` check (d)
requires of the corpus.
**The 0.30 confidence is the design's honesty mechanism.** A "no match" is not a 0.0 (which would
claim certainty that the query is nonsense) and not a 0.5 (which would claim uncertainty about a
definite non-match). 0.30 is "probably not answerable, and I have no positive evidence either way".
### 27.10 The complete rule table
| # | Rule id | Fires when | Task | Modality | temp | spat | lang | Confidence |
|---|---|---|---|---|---|---|---|---|
| 0 | `empty_query` | query is not a non-blank string | `unsupported` | `unknown` | F | F | F | **0.00** |
| 1 | `dual_modality` | an `explicit_pair` dual-modality term is present | `optical_sar` | `optical_sar` | F | F | T | **0.92** |
| 2 | `temporal_spatial` | a temporal term **and** a spatial term | `change` | `optical` | **T** | **T** | T | **0.90** |
| 2 | `temporal` | a temporal term, no spatial term | `change` | `optical` | **T** | F | T | **0.85** |
| 3 | `spatial` | a spatial term | `grounding` | `optical` | F | **T** | T | **0.88** |
| 4 | `sar_single` | a SAR term and no optical term | `vqa` | `sar` | F | F | T | **0.75** |
| 5 | `optical_single` | an optical term and no SAR term | `vqa` | `optical` | F | F | T | **0.72** |
| 6 | `caption` | a caption term | `caption` | `unknown` | F | F | T | **0.85** |
| 7 | `vqa` | a VQA term | `vqa` | `unknown` | F | F | T | **0.78** |
| 8 | `no_match` | nothing matched | `unsupported` | `unknown` | F | F | F | **0.30** |
**Ten rule ids across eight rule positions** (rule 2 has two ids). The confidence band is
**0.72β0.92** for task-producing rules, plus the two boundary values 0.00 and 0.30.
**The band overlaps and can exceed the trained model's band.** `router/classifier.py:148-150`
records the measured comparison: *"on the spec section 29 examples the fallback returns
confidence 0.850-0.920 against the trained model's 0.780-1.000."* That overlap is why
`core/planner.py` applies `LEXICAL_FALLBACK_DISCOUNT = 0.75` to its own reading (Β§43) rather than
trusting the raw number.
## 28. `is_available()` β the fallback's one dependency claim
```python
def is_available() -> bool:
"""The fallback has no model dependency and is always available."""
return True
```
(`router/fallback.py:286-288`.) A constant `True` with a docstring that is the reason. The fallback
is the **degradation path** β it is what runs when the encoder cannot load, so it must never itself
have a load dependency. This is the function a registry would call to decide whether a lexical
capability exists; here it is unconditional.
## 29. `self_check()` β the ten curated cases
```python
def self_check() -> list[str]:
"""Assert the fallback agrees with the curated hard negatives.
Returns a list of failures. An empty list means the fallback is internally
consistent with the label space.
"""
cases: tuple[tuple[str, str, bool | None], ...] = (
("Describe this image.", "caption", None),
("What land cover is visible?", "vqa", None),
("Show me the water body.", "grounding", None),
("Describe the water body.", "caption", None),
("What changed between these images?", "change", None),
("Where did the change happen?", "change", True),
("What changed?", "change", False),
("Compare the optical and radar images.", "optical_sar", None),
("Locate the buildings.", "grounding", None),
("Book me a flight to Delhi.", "unsupported", None),
)
```
(`router/fallback.py:291-308`.)
Each case is `(query, expected_task, expected_spatial)` where `expected_spatial` is `None` when only
the task is asserted. **All ten cases, with the rule that satisfies each:**
| # | Query | Expected task | Expected `spatial_output` | Satisfying rule | Resulting confidence |
|---|---|---|---|---|---|
| 1 | `Describe this image.` | `caption` | β | `caption` | 0.85 |
| 2 | `What land cover is visible?` | `vqa` | β | `vqa` (bare `"what"`) | 0.78 |
| 3 | `Show me the water body.` | `grounding` | β | `spatial` (`"show me the"`) | 0.88 |
| 4 | `Describe the water body.` | `caption` | β | `caption` | 0.85 |
| 5 | `What changed between these images?` | `change` | β | `temporal` (spatial false) | 0.85 |
| 6 | `Where did the change happen?` | `change` | **True** | `temporal_spatial` | 0.90 |
| 7 | `What changed?` | `change` | **False** | `temporal` | 0.85 |
| 8 | `Compare the optical and radar images.` | `optical_sar` | β | `dual_modality` | 0.92 |
| 9 | `Locate the buildings.` | `grounding` | β | `spatial` (`"locate"`) | 0.88 |
| 10 | `Book me a flight to Delhi.` | `unsupported` | β | `no_match` | 0.30 |
**Cases 3 and 4 are the hard-negative pair** β `"Show me the water body."` (grounding) versus
`"Describe the water body."` (caption) β and they are the reason the design rule says *"the fallback
is wrong, not the router"* if the two disagree. **Cases 6 and 7 are the second hard-negative pair** β
the same task with `spatial_output` flipped.
**Case 5 needs a note.** `"What changed between these images?"` contains `"these two images"`? No β
it contains `"between these two"`? No. It contains `"change"` and `"these images"`. The temporal
table's `"these two images"` and `"two images"` entries do **not** match `"these images"`. So the
rule fires on `"change"` (a substring of `"changed"`) alone, giving `temporal` (not
`temporal_spatial`, since no spatial term is present) β confidence 0.85, `spatial_output=False`.
`self_check()` does not assert `spatial_output` for case 5 (`None`), so the case passes regardless.
**Case 10 is the refusal case** β an out-of-domain query that must reach `no_match`. It is in the
same table because "never invents capability" is as testable a property as any classification.
**The five assertions per case:**
```python
failures: list[str] = []
for query, expected_task, expected_spatial in cases:
match = lexical_route(query)
if match.task != expected_task:
failures.append(
f"{query!r}: expected task {expected_task!r}, got {match.task!r} "
f"(rule={match.rule})"
)
if expected_spatial is not None and match.spatial_output != expected_spatial:
failures.append(
f"{query!r}: expected spatial_output={expected_spatial}, "
f"got {match.spatial_output}"
)
if not is_valid_task(match.task):
failures.append(f"{query!r}: produced invalid task {match.task!r}")
if not is_valid_modality(match.modality):
failures.append(f"{query!r}: produced invalid modality {match.modality!r}")
if match.task not in TASK_CLASSES:
failures.append(f"{query!r}: task outside the label space")
return failures
```
(`router/fallback.py:310-330`.)
**The last three assertions are a label-space invariant check on every case**, and the third is
redundant with the first (`is_valid_task` and `task not in TASK_CLASSES` test the same membership β
`is_valid_task` is literally `name in TASK_TO_INDEX`, and `TASK_CLASSES` is what `TASK_TO_INDEX` is
derived from). The redundancy is defensive, not harmful.
**The return convention is worth naming: `self_check()` returns a list of failures, and an empty
list means pass.** This is the same convention as `RouterCorpus.validate()`. A caller checks
truthiness of the returned list.
**Status.** `IMPLEMENTED` and exercised by `tests/routing/test_router.py`. The **measured** result of
`self_check()` on the shipped code is **`UNKNOWN β not established from the available evidence`**:
this document did not execute it, and no artifact records its output. What *is* established is that
the learned router and the fallback agree on the canonical queries
(`docs/PHASE4_ROUTER_REPORT.md` Β§"Standing caveats": *"both paths currently agree on every canonical
query"*), and that the learned router's test hard-negative accuracy is 0.800.
---
# Part G β The measured result
## 30. The shipped adapter artifact
The router ships exactly one trained artifact, at `artifacts/router/router_adapter_v001/`:
| File | Role |
|---|---|
| `adapter.pt` | the weights + the embedded architecture config |
| `metadata.json` | 626 lines of provenance, hyperparameters, split report, metrics and the full 60-epoch history |
`docs/PHASE4_ROUTER_REPORT.md:36-37` records the size: **"223.8 KB (51,725 adapter params on a
frozen 22,713,216-param encoder)."**
### 30.1 The artifact's identity
| Key | Value |
|---|---|
| `artifact` | `router_adapter` |
| `created_at` | `2026-09-16T08:24:38.236767+00:00` |
| `duration_seconds` | **4.92** |
| `seed` | `42` |
| `config_hash` | `615478910dc266bf` |
| `num_parameters` | `51725` |
| `encoder_type` | `frozen_sentence_transformer` |
| `encoder.model` | `sentence-transformers/all-MiniLM-L6-v2` |
| `encoder.revision` | `1110a243fdf4` |
| `encoder.max_length` | `128` |
| `encoder.parameters` | `22713216` |
| `encoder.type` | `frozen_sentence_transformer` |
**The `created_at` date is worth noting against the code.** The adapter was trained on
**2026-09-16**, and the threshold sweep was created on **2026-09-21**
(`threshold_sweep_val.json.created_at: "2026-09-21T05:56:32.394830+00:00"`). The corpus source in
the repository has clearly been edited after 2026-09-16 (Β§13.3), so **the shipped adapter's training
data is not byte-identical to the current `router/dataset.py`**. The adapter is a real, measured
artifact; it is simply not reproducible from the current source without reverting
`router/dataset.py`. This is stated because reproducibility claims depend on it (Β§50).
### 30.2 The corpus the adapter was trained on
```json
"corpus": {
"by_source": { "curated": 66, "template": 510 },
"by_task": {
"caption": 91, "change": 115, "grounding": 128,
"optical_sar": 50, "unsupported": 105, "vqa": 87
},
"groups": 54,
"hash": "8054810736ef97c3db873b2d7073948773a8982f1d16411833d39e15e1871e83",
"positives": { "language_output": 471, "spatial_output": 164, "temporal": 115 },
"total": 576
}
```
**The three `positives` counts are the binary-head label balance**, and they explain the
`combined_accuracy` gap:
| Head | Positives | Positives / 576 | Interpretation |
|---|---|---|---|
| `language_output` | **471** | 81.8 % | most requests want prose |
| `spatial_output` | **164** | 28.5 % | a minority want coordinates |
| `temporal` | **115** | 20.0 % | one fifth need two acquisitions |
`temporal = 115` exactly equals the `change` class count (115), which is the corpus-level proof that
**every `change` example is marked temporal and no non-`change` example is** β the invariant
`RouterCorpus.validate()` check (e) enforces. `spatial_output = 164` exceeds the `grounding` count
(128) by 36, which is the number of `change` examples marked spatial: 36 of the 115 change examples
carry `spatial_output=True`, i.e. 31.3 % β consistent with the 4-of-12 (33.3 %) ratio in the curated
change family (Β§12.1).
**The corpus is synthetic and the phase says so.** `docs/PHASE4_ROUTER_REPORT.md` Β§"Standing
caveats": *"The corpus is synthetic. 576 examples, 54 groups, hand-written and templated.
`hard-negative accuracy 0.800` is the most honest generalisation number here; the 0.975 headline is
partly earned on templates the split kept in training. Treat the router as working, not as
benchmarked."*
## 31. Gate 2, and the three split metric blocks
The gate's verdict, from `docs/PHASE4_ROUTER_REPORT.md` Β§Verdict:
```
GATE 2 ROUTER ACCEPTANCE: task accuracy >= 0.95, all classes measured
classes with test support : 6/6
-> PASS (0.975)
```
### 31.1 The per-split table, from the artifact
| split | `n` | task | modality | spatial | temporal | language | combined | macro F1 |
|---|---|---|---|---|---|---|---|---|
| train | 410 | 1.000 | 1.000 | 0.9707 | 1.000 | 1.000 | 0.9707 | 1.000 |
| val | 86 | **0.9651** | 0.8256 | 0.9302 | 1.000 | 0.9884 | 0.7907 | 0.9545 |
| **test** | 80 | **0.975** | 1.000 | 1.000 | 0.9875 | 1.000 | **0.9625** | **0.976** |
`docs/PHASE4_ROUTER_REPORT.md:17-21` presents the same table rounded to three decimals (val task
0.965, val modality 0.826; test task 0.975, test modality 1.000). The artifact's four-decimal values
are the primary source; the report's three-decimal values are the same numbers rounded.
**Two observations the table forces.**
1. **`val` is the weakest split on every head except `temporal`.** Val task accuracy 0.9651 against
test 0.975; val modality 0.8256 against test 1.000. Val is not a subset of test and the two are
not comparable β val has 20 `change` and 19 `unsupported` examples against test's 13 and 10
(`metadata.json.split.task_counts`), a different class mix.
2. **The val `modality_accuracy` of 0.8256 is what drags `combined_accuracy` to 0.7907.** The val
task accuracy is 0.9651 and the three binary heads are 0.9302β1.000, so the combined figure's
gap is dominated by modality errors. The val `per_task_recall` shows `optical_sar` at **0.7**
(7 of 10) β the only class below 1.0 on val β which is consistent with modality being the
weakest val head.
### 31.2 Per-class recall and support, test
| class | recall | `n` | Misses |
|---|---|---|---|
| `vqa` | 1.000 | 15 | 0 |
| `caption` | **0.8462** | 13 | **2** |
| `grounding` | 1.000 | 17 | 0 |
| `change` | 1.000 | 13 | 0 |
| `optical_sar` | 1.000 | 12 | 0 |
| `unsupported` | 1.000 | 10 | 0 |
| **total** | β | **80** | **2** |
**`caption` is the only class with a miss**, and it accounts for exactly the 2 errors in the
0.975 headline (78/80 = 0.975). The phase's caveat names it:
> **`caption` recall is 0.846** β 2 of 13 missed. Not investigated; within noise at n=13 and below
> this corpus's significance floor.
(`docs/PHASE4_ROUTER_REPORT.md` Β§"Standing caveats".) **2 of 13 at n=13 is not a finding.** With 13
examples, one miss moves recall by 0.077; the class is below the corpus's significance floor.
### 31.3 `macro_f1_task` β 0.976 on test
Computed by the local `_macro_f1` with `n_classes = 6`. Because all six classes have support, no
class contributes a structural `0.0`, and the macro F1 (0.976) sits just above the accuracy (0.975)
β a sign that the two `caption` errors did not disproportionately hurt a rare class.
## 32. The 60-epoch history β the training trajectory, verbatim
`metadata.json.history` contains **60 rows**, one per epoch. The shape of the curve is the whole
story of this model, so the key epochs are tabulated rather than summarised.
| epoch | loss | val task acc | val combined acc | lr |
|---|---|---|---|---|
| 1 | 2.276186 | 0.5814 | 0.1744 | 0.00099931 |
| 2 | 1.524010 | 0.7558 | 0.1977 | 0.00099726 |
| 3 | 0.852199 | 0.8721 | 0.1977 | 0.00099384 |
| 4 | 0.419326 | 0.9186 | 0.3372 | 0.00098907 |
| 5 | 0.219211 | 0.9302 | 0.6395 | 0.00098296 |
| 6 | 0.122340 | 0.9535 | 0.7558 | 0.00097553 |
| 7 | 0.074093 | 0.9535 | 0.7558 | 0.00096679 |
| **8** | **0.050349** | **0.9651** | **0.7907** | 0.00095677 |
| 9 | 0.037641 | 0.9651 | 0.7907 | 0.00094550 |
| 10 | 0.030423 | 0.9651 | 0.7558 | 0.00093301 |
| 20 | 0.008119 | 0.9651 | 0.8023 | **0.00075** |
| 30 | 0.004611 | 0.9651 | 0.8140 | **0.0005** |
| 40 | 0.003455 | 0.9651 | 0.8140 | **0.00025** |
| 50 | 0.002699 | 0.9651 | 0.8140 | 0.00006699 |
| 59 | 0.003170 | 0.9651 | 0.8140 | 0.00000069 |
| 60 | 0.003065 | 0.9651 | 0.8140 | **0.0** |
**Five facts this table establishes.**
1. **The task head converges by epoch 8.** Val task accuracy reaches 0.9651 at epoch 8 and never
moves again β 53 further epochs at the same value. The strict `>` best-state rule means the
**epoch-8 snapshot is the shipped model** (Β§15.4).
2. **The loss keeps falling while the metric is flat.** Loss goes 0.0503 β 0.0031 across epochs
8β60, a 16Γ reduction, with zero val-task improvement. That is the signature of a model
**overfitting the training set's template variants** while the group-held-out val set stays
saturated β exactly what the group-split design is meant to expose.
3. **`val_combined_accuracy` does improve after epoch 8** β 0.7907 β 0.8140 by epoch 21 β because
the *binary heads* keep improving even though the task head does not. But `best_state` selects on
`task_accuracy` only, so that later improvement is **not** in the shipped weights. The shipped
adapter's val combined accuracy is therefore the epoch-8 value, 0.7907, which matches
`metadata.json.metrics.val.combined_accuracy`.
4. **The cosine schedule is exact.** `lr` is exactly `0.00075` at epoch 20, `0.0005` at epoch 30,
`0.00025` at epoch 40, and exactly `0.0` at epoch 60 β the analytic cosine values for
`T_max = 60`, confirming no warm restarts and no `eta_min`.
5. **The first epoch's loss (2.276) is consistent with the small-std init.** With five heads
initialised at Ο = 0.02 (Β§6.3), the initial task cross-entropy is near `ln(6) = 1.79` plus the
modality and binary terms, and the weighted total starts above 2.0 β not saturated, not
degenerate. The init is doing what its docstring claims.
## 33. The threshold sweep β and what "the test split was not run" means
`artifacts/router/threshold_sweep_val.json` is the **release-level** router artifact, and it is the
one `README.md` cites as the source of the router number. It is a sweep over 50 thresholds, and its
own metadata is unusually explicit about its limits. The `note` field, verbatim:
> corpus-limited: val n=86 vs plan >=500. This is NOT a calibration -- the corpus is synthetic and
> too small (min per-class support 8, caption) and val carries 0 hard negatives (hn_* families are
> held out to TEST by design). Selecting a threshold here yields a justified default, not a
> calibrated value. The corpus was NOT padded with generated queries. Backlog P1-9's 'n=80' is the
> TEST split; the sweep target is val n=86. **The test split was NOT touched.**
### 33.1 The sweep's configuration
| Key | Value |
|---|---|
| `split` | `val` |
| `n_val` | 86 |
| `n_val_examples_scored` | **86** |
| `n_test_examples_scored` | **0** |
| `test_split_touched` | **`false`** |
| `hard_negatives_in_val` | **0** |
| `thresholds` | 0.50 β¦ 0.99, step 0.01 β **50 rows** |
| `select_by` | `covered_accuracy` |
| `shipped_threshold` | **0.70** |
| `shipped_row` | `{coverage: 0.848837, covered_task_accuracy: 0.972603, fallback_rate: 0.151163, n_covered: 73, threshold: 0.7}` |
| `selected` | `{coverage: 0.790698, covered_task_accuracy: 1.0, fallback_rate: 0.209302, n_covered: 68, threshold: 0.76}` |
| `overall_ungated_accuracy` | **0.965116** |
| `delta_vs_shipped` | `{coverage: -0.0581, covered_task_accuracy: 0.0274}` |
| `corpus_limited` | `true` |
| `corpus_total` / `corpus_groups` | 576 / 54 |
| `plan_min_val_queries` | **500** |
| `plan_min_hard_negatives` | **100** |
| `val_min_support` | **8** |
| `seconds` | 0.206 |
| `adapter_config_hash` | `615478910dc266bf` |
| `config_hash` | `78f1e3700da15aa1` |
| `environment` | `{cuda_available: false, platform: "Windows-10-10.0.26200-SP0", python: "3.11.16", requested_device: "cpu", torch: "2.14.0+cpu"}` |
### 33.2 The sweep's extremes, and what they show
| Threshold | `n_covered` | `coverage` | `covered_task_accuracy` | `fallback_rate` |
|---|---|---|---|---|
| 0.50 (first row) | 84 | 0.976744 | 0.964286 | 0.023256 |
| 0.53 | 83 | 0.965116 | 0.963855 | 0.034884 |
| 0.55 | 81 | 0.94186 | **0.975309** | 0.05814 |
| **0.70 (shipped)** | **73** | **0.848837** | **0.972603** | **0.151163** |
| **0.76 (selected)** | **68** | **0.790698** | **1.000000** | **0.209302** |
| 0.99 (last row) | 8 | 0.093023 | 1.000000 | 0.906977 |
**The monotone structure is the point.** As the threshold rises, fewer queries clear the gate
(`n_covered` 84 β 8), coverage falls (0.977 β 0.093), the fallback rate rises (0.023 β 0.907), and
the *covered* task accuracy rises (0.964 β 1.000) β because the gate keeps only the queries the
model is sure about.
**The shipped threshold (0.70) is not the sweep's optimum.** The sweep's `select_by:
covered_accuracy` would pick **0.76**, at which all 68 covered queries are correct
(`covered_task_accuracy: 1.000000`). The `delta_vs_shipped` block quantifies the difference: 0.76
trades **β0.0581 coverage** for **+0.0274 covered accuracy**. The shipped 0.70 was kept anyway.
**The reason is in the note:** the val split has **min per-class support 8** (caption) and
**0 hard negatives**, against a plan minimum of 500 val queries and 100 hard negatives. The sweep's
own conclusion is that *"selecting a threshold here yields a justified default, not a calibrated
value."* Choosing 0.76 on an 8-example-per-class validation set would be fitting to noise.
### 33.3 Reconciling "Gate 2 PASS (0.975)" with "the test split was NOT RUN"
These two statements appear to conflict and do not. The resolution is entirely in
`threshold_sweep_val.json`, which records both `n_test_examples_scored: 0` and
`test_split_touched: false` while also recording `split_sizes: {test: 80, train: 410, val: 86}` β so
the sweep **knew about** the test split and deliberately did not score it.
| Statement | Source | What it is about |
|---|---|---|
| test task accuracy **0.975**, `n=80`, Gate 2 **PASS** | `metadata.json.metrics.test`; `docs/PHASE4_ROUTER_REPORT.md` | the **training-time** evaluation, produced by `train_router`'s own `evaluate_split(..., "test", ...)` call |
| **`test_split_touched: false`**, `n_test_examples_scored: 0` | `threshold_sweep_val.json` | the **release-level threshold sweep**, which scored val only |
| router "overall ungated accuracy" **0.965116**, val, n = 86, **"TEST NOT RUN"** | `README.md` metric table; `DOCS_STYLE_GUIDE.md` Β§3 | the **release-level claim**, which quotes the val number and declines to promote the training-time test score |
**The honest description, stated precisely:**
- The router **was** evaluated on a test split, once, at training time, and the result is 0.975 with
n = 80, all six classes measured. That number is in the shipped artifact.
- The **release-level** router evaluation β the threshold sweep that `README.md` cites as the
router's source artifact β was run on **val only**. The test split was **not touched** by it.
- The release therefore quotes **0.965116** (the val, ungated number) as the router's headline and
labels the test as **`NOT RUN`**, because the *release-level* evaluation of the test split is the
thing that does not exist.
- `DOCS_STYLE_GUIDE.md` Β§3 states the release rule as **"0.965116 is validation, ungated, n = 86;
the test split was NOT RUN"**, and this chapter obeys it.
**What is *not* established:** whether the release team considers the training-time test score
(0.975) a valid test-set result, or whether it is discounted because it is an artifact-internal
number. `README.md` limitation 2 says the test set *"was never run"*, which is stricter than the
artifact supports. This chapter reports both facts and does not resolve the discrepancy β it is
recorded as an **`OPEN`** documentation question, because resolving it would require an owner
ruling, not an inference from the files.
**One more detail from the sweep that is easy to miss.** `n_val_examples_scored` is **86**, and the
val `task_counts` sum to 86 (8 + 20 + 14 + 10 + 19 + 15 = 86) β so the sweep scored **every** val
example, and `coverage` is the fraction that cleared the gate rather than the fraction evaluated.
## 34. The router's standing caveats
`docs/PHASE4_ROUTER_REPORT.md` Β§"Standing caveats" is the phase's own list. It is reproduced here
because a reader of this chapter should not have to open a second file to find the limits.
| Caveat | Statement |
|---|---|
| **The corpus is synthetic** | 576 examples, 54 groups, hand-written and templated. **Hard-negative accuracy 0.800** is the most honest generalisation number; the 0.975 headline is partly earned on templates the split kept in training. **"Treat the router as working, not as benchmarked."** |
| **`caption` recall is 0.846** | 2 of 13 missed. Not investigated; within noise at n=13 and below this corpus's significance floor. |
| **The stub-encoder path exists only for tests** | Any artifact it produces is marked `encoder_type: "stub"` and **must never be deployed**. |
| **Threshold 0.70 is uncalibrated** | It gates the lexical fallback, and both paths currently agree on every canonical query. *"Calibration belongs in Phase 13 with real validation data, not here."* |
**One caveat the phase states and this chapter must repeat:** the phase's F4-2 section says the
adapter has "51,725 parameters" while the `router/train.py` docstring says 50,822. Β§8 resolves the
arithmetic in favour of 51,725; the phase report itself already uses the correct figure.
---
# Part H β The frontend router: `interpret()` vs `chooseTask()`
## 35. Two functions, two information sets
The frontend console runs its own router, and it is split in two. The file header states the whole
design:
```
TWO DRIVERS, ONE EVENT SEAM
---------------------------
* LIVE (default when files are chosen): the browser uploads the user's own
imagery to `POST /api/assets`, receives asset IDs, posts them to
`POST /api/infer`, and feeds the REAL result into the same eight events.
Every value on screen then traces to the server's own response.
* PREVIEW (no files chosen): the deterministic router still runs so the
instrument is legible, but the specialist/result stages stay honestly empty
("awaiting backend") instead of pretending an analysis happened.
What is NEVER done in either mode: fabricating an answer, a confidence value,
an evidence record, a run id or a coordinate. If the backend is unreachable
the page says which step failed and shows the server's own message.
```
(`frontend/assets/js/mission.js:1-18`.)
**The router is what makes the PREVIEW path legible.** With no assets, no specialist can run, but the
page can still show *what the question was understood to be* β because `interpret()` is pure lexical
text processing with no model dependency.
The comment above `interpret()` states the contract:
```
Maps the question to an Intent-shaped reading, using only lexical rules.
This is the *interpretation* the page exposes β clearly a mock router,
never a measurement. Slots follow the real schema's Intent type.
```
(`frontend/assets/js/mission.js:66-69`.)
| Function | Line | Inputs | Blind to asset count? | Returns |
|---|---|---|---|---|
| `interpret(query)` | `mission.js:70` | the query text | **yes** | a reading: `task`, `modality`, `temporal`, `spatial_output`, `evidence`, `source`, `language_output` |
| `chooseTask(intent, query, assetCount)` | `mission.js:184` | the reading, the query, the **asset count** | **no** | `{task, substituted, wanted, reason}` |
**The asymmetry is intentional and documented.** `docs/RESEARCH_NOTES.md` Β§6 records the ruling:
> For *"What changed between the earlier and later image?"* with **one** asset attached, the console
> **reads** `change` while dispatch correctly falls back to **`change_vqa`**. This is not a bug: the
> reading describes the question's intent, the dispatch respects what can actually be computed with
> the assets present. It is documented so it is not mistaken for a defect.
`RESOLVED` β and the reason it must be documented is that a reader who sees "read `change`, dispatched
`change_vqa`" in a trace will otherwise file it as a bug.
**The server-side pair has the same shape** (Β§39): `IntentRouter.route()` reads, `PolicyPlanner.plan()`
decides. The frontend pair and the server pair were written independently and converged on the same
two-stage split β which is itself evidence that the split is the right one.
## 36. `interpret()` in full
```javascript
function interpret(query) {
var q = (query || '').toLowerCase();
var where = /where|locate|position|which part|bound|outline|coordinate/.test(q);
/* The change stem is matched WITHOUT a trailing \b: `\bchang\b` cannot
match "changed", "changes" or "changing", because there is no word
boundary between the stem and its inflection. With the boundary, the
page's own default question ("What changed here?") fell through to the
vqa branch β so the change path was unreachable from the UI that exists
to reach it. The other terms keep their boundaries; they are whole words. */
var changeStem = /chang/.test(q);
/* The temporal slot is `required` ONLY for a genuine change/pair marker.
Pair wording names two dates outright, so it can never be a location
question and is tested on its own. */
var pairWording = /\b(between|versus|vs\.?|pair|temporal|before|after|differ\w*|expand\w*|grow\w*|encroach\w*|lost|removed)\b/.test(q);
/* `new` is a place descriptor as often as a change marker: the repo ships
eo/new-airport.jpg, so "Where is the new airport?" is a real question.
It counts as a change ONLY when the query is not a `where` question.
`built` was dropped outright -- "built-up areas" is land-cover
vocabulary, not a change marker. While it sat in the temporal set, the
location question "Where are the built-up areas in this image?" was read
as a change request and the server answered it with the degenerate
one-word "River". Measured live, 2026-09-25. */
var newAsChange = /\bnew\b/.test(q) && !where;
var temporal = (changeStem || pairWording || newAsChange) ? 'required' : 'none';
var sar = /\bsar\b|radar|backscatter|insar/.test(q);
/* A CAPTION task is a request to DESCRIBE the whole scene, not to answer a
specific question about it. The lexical markers below are deliberately
narrow so a question like "what is the building" stays on the VQA branch:
caption fires only on explicit describe/caption/summarise wording or on the
"what is in / do you see" open forms. This is the path that was missing β
descriptive queries previously fell through to `vqa`, so the page never
asked the server for a caption even though `caption` is a valid task. */
var caption = /describe|caption|summari[sz]e|give me a (description|summary)|what (do|can) you see|what (is|are) (in|shown in|depicted in) (this|the) (image|scene|picture)|tell me about (this|the) (image|scene|picture)|what does this (image|scene|picture) (show|contain)/.test(q);
var task, modality, evidence, spatial;
if (sar) { task = 'optical_sar'; modality = 'optical_sar'; evidence = 'joint_feature_region'; spatial = where ? 'box' : 'scene'; }
else if (temporal === 'required') { task = 'change'; modality = 'optical'; evidence = 'change_region'; spatial = where ? 'box' : 'scene'; }
else if (where) { task = 'grounding'; modality = 'optical'; evidence = 'bounding_box'; spatial = 'box'; }
else if (caption) { task = 'caption'; modality = 'optical'; evidence = 'caption'; spatial = 'scene'; }
else { task = 'vqa'; modality = 'optical'; evidence = 'statistic'; spatial = 'scene'; }
return {
task: task, modality: modality, temporal: temporal,
spatial_output: spatial, evidence: evidence,
source: 'lexical_fallback', language_output: 'en'
};
}
```
(`frontend/assets/js/mission.js:70-114`.)
### 36.1 The four lexical predicates
| Variable | Pattern | Kind |
|---|---|---|
| `where` | `/where|locate|position|which part|bound|outline|coordinate/` | unanchored substring alternation |
| `changeStem` | `/chang/` | **stem, no trailing `\b`** |
| `pairWording` | `/\b(between|versus|vs\.?|pair|temporal|before|after|differ\w*|expand\w*|grow\w*|encroach\w*|lost|removed)\b/` | word-bounded alternation |
| `newAsChange` | `/\bnew\b/` **AND NOT** `where` | word-bounded, gated |
| `sar` | `/\bsar\b|radar|backscatter|insar/` | mixed |
| `caption` | the long alternation above | unanchored, multi-form |
**`where` is unanchored** β `"where"` matches inside `"somewhere"`, and `"bound"` matches inside
`"boundary"`. This is the same substring-matching limitation as the Python fallback (Β§26), and it is
the source of the residuals in Β§48.
### 36.2 The dispatch chain β a strict if/else ladder
The five branches are evaluated **in order**, and the first match wins:
| # | Condition | task | modality | evidence | spatial |
|---|---|---|---|---|---|
| 1 | `sar` | `optical_sar` | `optical_sar` | `joint_feature_region` | `box` if `where` else `scene` |
| 2 | `temporal === 'required'` | `change` | `optical` | `change_region` | `box` if `where` else `scene` |
| 3 | `where` | `grounding` | `optical` | `bounding_box` | `box` |
| 4 | `caption` | `caption` | `optical` | `caption` | `scene` |
| 5 | otherwise | `vqa` | `optical` | `statistic` | `scene` |
**The precedence mirrors the Python fallback's** β dual-modality (`sar`) first, then temporal, then
spatial, then caption, then VQA β with one difference: the frontend has **no SAR-only or
optical-only rule**, so `sar` alone (a single modality word) sends the request to `optical_sar`
rather than to a SAR-modality VQA. That is a genuine behavioural difference between the frontend
router and the server fallback, and it is the reason a bare `"radar"` in a query reads as a fusion
request in the console. It is `IMPLEMENTED` as written; whether it is intended is not recorded, and
this chapter does not assert either way.
**The `evidence` and `spatial` slots are the frontend's own vocabulary**, not the server's
`EvidenceType` enum. `'bounding_box'` and `'change_region'` and `'statistic'` are display strings for
the console's reading panel; the server's `EvidenceType` uses `bounding_box`, `change_map` and
`statistic` (`core/schemas.py:65-76`). So `'change_region'` is **not** a server evidence type β the
console's reading is a *display* projection, and the server's evidence is authoritative.
**`language_output: 'en'`** β a string, not a boolean, where the server's `Intent.language_output` is
a `bool`. The reading is an `Intent`-**shaped** object, not an `Intent`; the comment says *"Slots
follow the real schema's Intent type"*, meaning the slot names, not the types. `source:
'lexical_fallback'` does match the server's literal (`core/schemas.py:105`).
## 37. `chooseTask()` in full
The comment block above `chooseTask` is the clearest statement anywhere of *why* dispatch is
asset-aware:
```
THE FRONTEND MUST SPEAK THE SERVER'S VOCABULARY.
`core.schemas.Task` accepts exactly: vqa, caption, grounding, change,
optical_sar, change_vqa, unsupported. Any other string is a 422 from
`AnalysisRequest`, because the model is `extra="forbid"` and `force_task`
is a `Task | None`. The legacy prototype names (SAR_ANALYSIS,
CHANGE_ANALYSIS, VLM_QA, VLM_CAPTION) are NOT valid and would fail.
So the router's reading is mapped to the server enum here, in one place.
`change` -> `change_vqa` when the question is quantificational, because the
server's `change` returns a spatial change map with NO language output while
`change_vqa` returns a short answer -- and the page's Answer block promises
an answer. Sending `change` for "how much changed" would fill the map and
leave the answer empty, which looks like a bug and is really a wrong task.
PAIR-AWARE, BECAUSE THE DEPLOYMENT SAYS SO
------------------------------------------
`/api/capabilities` declares `requires_pair` and `max_assets` per task, and
the server enforces them: with one asset, `change` answers
`invalid_request` ("change requires exactly 2 assets (T1 and T2); got 1")
and the whole envelope comes back `degraded: true`. Measured live, 2026-09-25.
So the task cannot be chosen from the question alone -- it depends on how
many files were actually selected. The page's OWN default question ("What
changed here?") reads as a change request, and a user who selects one image
and presses Run would get a degraded non-answer for asking a reasonable
thing. That is a wiring defect, not a user error.
The rule below therefore: request a pair-requiring task ONLY when a second
asset exists; otherwise fall back to a task that accepts one asset, and
report the substitution rather than hiding it.
```
(`frontend/assets/js/mission.js:116-147`.)
### 37.1 The three constants
```javascript
var ROUTER_TASK_TO_SERVER = {
vqa: 'vqa',
caption: 'caption',
grounding: 'grounding',
change: 'change',
optical_sar: 'optical_sar',
change_vqa: 'change_vqa'
};
/*: Tasks the deployment reports as `requires_pair: true`. Mirrors
`capabilities_payload()`; the live capabilities are fetched at boot and
override this when available, so a server-side change is picked up rather
than hardcoded wrong. */
var PAIRED_TASKS = { change: true, change_vqa: true, optical_sar: true };
/*: What to ask for when a pair-requiring task was chosen but only one asset
is available. `vqa` accepts a single asset and answers a question about it,
which is the honest closest behaviour available for one image. */
var SINGLE_ASSET_FALLBACK = 'vqa';
```
(`frontend/assets/js/mission.js:148-166`.)
**`PAIRED_TASKS` is a fallback for the live capability contract, not the source of truth.** The
comment says the live capabilities are fetched at boot and override it:
```javascript
var pairRequired = {}; // task -> bool, from /api/capabilities when reachable
var singleAssetTasks = {}; // task -> bool
function requiresPair(task) {
if (Object.prototype.hasOwnProperty.call(pairRequired, task)) return pairRequired[task];
return !!PAIRED_TASKS[task];
}
```
(`frontend/assets/js/mission.js:168-174`.)
**So a server-side change to `requires_pair` is picked up at boot, and the hardcoded table is only
the offline default.** That is a genuinely careful piece of design: the frontend does not hardcode
the deployment's contract as fact, it hardcodes it as a *fallback*.
**`SINGLE_ASSET_FALLBACK = 'vqa'`** β and the comment states the reason it is `vqa` and not
`caption`: `vqa` *"accepts a single asset and answers a question about it, which is the honest
closest behaviour available for one image."* A caption would also accept one asset, but a caption is
not an answer to a question.
### 37.2 The function
```javascript
function chooseTask(intent, query, assetCount) {
var wanted = ROUTER_TASK_TO_SERVER[intent.task] || 'caption';
if (wanted === 'change') {
var q = (query || '').toLowerCase();
/* `\barea\b`, not bare `area`: without the boundary the substring matched
inside "areas", so "Where are the built-up areas ..." (already a
mis-read change question) was upgraded again to change_vqa. The
boundary keeps the quantifier reading for a real "how much area
changed" while refusing the plural land-cover noun. */
if (/\bhow (much|many|large)\b|\bhas\b|\bdid\b|percent|\barea\b|quantif/.test(q)) {
wanted = 'change_vqa';
}
}
if (assetCount < 2 && requiresPair(wanted)) {
/* No pair available: ask for a single-asset task instead, and name the
substitution so the UI can show it. Silently sending a pair-requiring
task would produce a degraded envelope whose reason ("requires exactly
2 assets") the user never sees. */
return {
task: SINGLE_ASSET_FALLBACK,
substituted: true,
wanted: wanted,
reason: wanted + ' needs two images (T0 and T1); only one was provided'
};
}
return { task: wanted, substituted: false, wanted: wanted, reason: '' };
}
```
(`frontend/assets/js/mission.js:184-213`.)
**Three steps, in order:**
1. **Map the reading to the server enum.** `ROUTER_TASK_TO_SERVER[intent.task] || 'caption'` β the
`|| 'caption'` is the default for an unmapped reading.
2. **The quantifier upgrade.** Only for `change`: a quantificational question becomes `change_vqa`.
The upgrade predicate is
`/\bhow (much|many|large)\b|\bhas\b|\bdid\b|percent|\barea\b|quantif/`.
3. **The pair guard.** If the wanted task requires a pair and fewer than two assets exist, substitute
`vqa` and **name the substitution**.
**The pair guard is checked *after* the quantifier upgrade**, so a quantificational change question
with one asset goes `change β change_vqa β vqa`. That is the path recorded in the live run
`run_62ca98d510be` β wait, no: that run dispatched `change_vqa` with a pair present. With one asset
it would be `vqa`. Both are correct under the rule.
**The upgrade predicate, term by term:**
| Term | Matches | Why |
|---|---|---|
| `\bhow (much\|many\|large)\b` | "how much", "how many", "how large" | quantity questions |
| `\bhas\b` | "has the coastline advanced" | present-perfect change questions |
| `\bdid\b` | "did the area grow" | past-tense change questions |
| `percent` | "what percent changed" | ratio questions |
| `\barea\b` | "how much area changed" | **word-bounded**, so `"areas"` does **not** match |
| `quantif` | "quantify", "quantification" | stem |
**`\barea\b` versus bare `area` is one of the four defect fixes** (Β§46), and the comment states the
failure precisely: without the boundary, the substring matched inside `"areas"`, so
`"Where are the built-up areas ..."` β *already* a mis-read change question β was **upgraded again**
to `change_vqa`. Two wrongs compounding: the reading was wrong, and the dispatch escalated the wrong
reading.
**The returned object's four fields:**
| Field | Meaning |
|---|---|
| `task` | the task actually dispatched |
| `substituted` | `true` when the pair guard fired |
| `wanted` | what the reading asked for, before the guard |
| `reason` | a human-readable explanation, empty when no substitution |
**`substituted` and `wanted` are what make the substitution visible rather than silent.** The
comment is explicit: *"Silently sending a pair-requiring task would produce a degraded envelope whose
reason β¦ the user never sees."* This is the same "absence must not be indistinguishable from a
non-event" principle the planner applies in Python (`core/planner.py:454`, Β§42).
### 37.3 `serverTaskFor()` β the test shim
```javascript
/* Back-compat shim for the unit tests that assert the pure question->task
mapping. Pair-awareness lives in `chooseTask`; this is the no-pair-limit
view of the same rule, exercised with a 2-asset assumption. */
function serverTaskFor(intent, query) {
return chooseTask(intent, query, 2).task;
}
```
(`frontend/assets/js/mission.js:215-220`.)
**A pure questionβtask mapping, obtained by assuming a pair.** This is how the unit tests assert the
*upgrade* rule (change β change_vqa) independently of the *pair* rule. Two behaviours, two entry
points, one implementation.
## 38. `assetsForTask()` and `validateOpticalSar()`
### 38.1 `assetsForTask()` β only the files the task consumes
```javascript
function assetsForTask(task, t1, t0) {
if (task === 'vqa' || task === 'grounding' || task === 'caption') {
/* Single-image tasks: the T0 frame is NOT part of the request. */
return [t1];
}
/* change / change_vqa / optical_sar all consume the pair. */
var pair = [t1];
if (t0) pair.push(t0);
return pair;
}
```
(`frontend/assets/js/mission.js:245-254`.)
**This is the third place the asset contract is enforced** β after `PAIRED_TASKS` and
`/api/capabilities` β and its comment records the live failure that motivated it:
> Sending the optional T0 frame to a single-image task makes the backend reject the whole request
> with `invalid_request` ("... requires exactly 2 assets ...; got 1") β **measured live, 2026-09-25**,
> when a pair was uploaded and a VQA question asked. Isolating the file set here is the fix: the pair
> is only ever sent to the tasks that declared it.
(`frontend/assets/js/mission.js:229-233`.)
**The failure mode is worth naming: the request was rejected because it sent *too many* assets.**
`vqa` declares `max_assets: 1`, so uploading a pair and asking a VQA question produced
`invalid_request`. The fix is not to drop the extra file silently but to send only the files the
dispatched task consumes β the single-image tasks get `[t1]` and the pair tasks get `[t1, t0]`.
**The optical-SAR slot semantics are stated in the comment:**
> For `optical_sar` the pair is read as optical (T1) + SAR/radar (T0) β the same two slots, a
> different meaning, and the input validator (`validateOpticalSar`) warns when they look wrong.
(`frontend/assets/js/mission.js:236-238`.)
So `t1` and `t0` carry **two different meanings** depending on the task: for `change` they are two
dates of the same sensor; for `optical_sar` they are two *modalities*. The slots are reused, and the
ambiguity is resolved by the task.
### 38.2 `validateOpticalSar()` β an advisory client-side check
```javascript
function validateOpticalSar(t1, t0) {
if (!t1 || !t0) {
return {
level: 'error',
message: 'Optical-SAR fusion needs two images: an optical scene (T1) and a SAR/radar image (T0). Add the radar image before running.'
};
}
var raster = { jpg: 1, jpeg: 1, png: 1 };
if (raster[_extOf(t0)] && raster[_extOf(t1)]) {
/* Both files are single-band photos β almost certainly NOT an optical +
SAR product. The model expects a real radar backscatter (usually a
GeoTIFF). We cannot prove the failure, but the pattern is wrong enough
to warn rather than let the round-trip fail opaquely. */
return {
level: 'warn',
message: 'Both files look like plain photos (JPEG/PNG). Optical-SAR fusion pairs an optical image with a radar/SAR backscatter β usually a GeoTIFF. The model may reject this pair.'
};
}
return { level: 'ok', message: '' };
}
```
(`frontend/assets/js/mission.js:273-292`.)
**The docstring above it states exactly what it is and is not:**
> Best-effort client-side check that an optical-SAR request has the right **SHAPE** of inputs. The
> browser cannot read sensor modality or band count from a File, so this is advisory, not a
> measurement: it flags the obviously-wrong case (two plain photos masquerading as an optical+SAR
> pair) so the user is warned BEFORE the backend returns `invalid_request`, and it names the
> missing-input case outright.
(`frontend/assets/js/mission.js:263-269`.)
| Level | Condition | Meaning |
|---|---|---|
| `error` | either file missing | the request cannot be formed |
| `warn` | both files are `.jpg` / `.jpeg` / `.png` | the pair is *probably* not an optical+SAR product |
| `ok` | otherwise | no client-side objection |
**The three levels are honest about their epistemics.** `error` is a fact the client can establish
(a file is absent). `warn` is a **heuristic** β the comment says *"We cannot prove the failure, but
the pattern is wrong enough to warn rather than let the round-trip fail opaquely."* `ok` means "no
objection", not "valid". The function never claims to validate the pair; it claims to catch the
obvious case before an opaque round-trip failure.
**Note the asymmetry: `.jpg`/`.jpeg`/`.png` produce a warning, but so would any non-GeoTIFF.** A
`.tif` file gets `ok` regardless of whether it is actually SAR. The check is a *shape* check on the
extension, and its docstring says so.
---
# Part I β The server-side decision: `PolicyPlanner`
## 39. Why the planner exists at all
`core/planner.py` opens with the reason, and it is the same reason `interpret()` and `chooseTask()`
are separate functions:
> A design where `intent.task` selects a specialist in one step has collapsed understand into
> decide. Three things break when that happens:
>
> 1. The router can never be overruled, downgraded, or widened. A confident-but-wrong prediction
> becomes unappealable.
> 2. A query needing two specialists can never get both. The router returns ONE task; *"what changed
> between these two images, and describe the scene"* routes to `change` alone, and the caption is
> lost. Section Β§3.5 is what recovers it.
> 3. There is no record of the decision. A plan step that carries a `reason` naming the rule that
> produced it is auditable; an `argmax` over a softmax is not a policy and cannot be inspected or
> tested.
(`core/planner.py:23-34`.)
**Three failure modes, each with a mechanism that prevents it.** Point 1 is what the closed refusal
list (Β§41) and the widening rules (Β§42) provide. Point 2 is the widening rules specifically. Point 3
is `PlanStep.reason`, which is *"a machine-generated RULE IDENTIFIER from a closed set, not a
narrative"* (`core/planner.py:174-176`) β the reason a free-text field would be a defect:
> A free-text field here would become a place to put chain-of-thought, and eventually would contain
> some (design Β§6).
### 39.1 The planner is pure, and the purity is load-bearing
```python
if TYPE_CHECKING: # pragma: no cover - typing only, never executed
# Imported for annotation only. `router.classifier` imports torch at module
# scope, and this module must stay importable and testable without torch β
# the planner is pure, and design Β§9.3 leans on that ("the planner is pure
# ... no models, no torch, no filesystem"). The only surface the planner
# actually uses is `prediction.intent` and `prediction.above_threshold`,
# so the dependency is structural and need not be a runtime import.
from router.classifier import RouterPrediction
```
(`core/planner.py:93-100`.)
**This is a real engineering constraint, not a style preference.** `router.classifier` imports torch
at module scope; importing it for real would make the planner un-importable in a torch-free
environment. Because the planner touches only `prediction.intent` and
`prediction.above_threshold`, the dependency is **structural** and the import can be annotation-only.
The class docstring restates the contract:
> Pure. Takes dataclasses, returns a dataclass. No models, no torch, no filesystem, no clock β every
> rule below is unit-testable with a hand-built `RouterPrediction`.
(`core/planner.py:255-258`.)
### 39.2 The two frozen constants
```python
#: Multiplier the planner applies to its *reading* of a lexical-fallback
#: confidence. Frozen by design Β§2.4. It exists so a matched keyword cannot
#: outrank a learned model when the planner decides whether to widen a plan.
#: The router's own reported confidence is never altered.
LEXICAL_FALLBACK_DISCOUNT: float = 0.75
#: Above this effective confidence the planner treats the route as settled and
#: does not add the discretionary explanation step. Below it, and when the
#: request also asks for language output, a `vlm` step is added so the run
#: explains itself. Also frozen by Β§2.3.
SETTLED_CONFIDENCE: float = 0.60
```
(`core/planner.py:106-116`.)
**`SETTLED_CONFIDENCE` is defined but not referenced in `plan()`.** The explanation step is added
when `not prediction.above_threshold and prediction.intent.language_output` (Β§40.4) β the gate is
`above_threshold`, the router's 0.70 threshold, not `SETTLED_CONFIDENCE`. So the constant documents a
design intent (Β§2.3) that the implementation expresses through `above_threshold` instead. This is
`IMPLEMENTED` with a constant that has no reader; it is named here because a reader will search for
its use and not find one.
### 39.3 The capability maps
```python
TASK_CAPABILITY: Mapping[Task, str] = {
Task.VQA: "vqa",
Task.CAPTION: "caption",
Task.GROUNDING: "grounding",
Task.CHANGE: "change",
Task.OPTICAL_SAR: "optical_sar",
Task.CHANGE_VQA: "change_vqa",
}
CAPABILITY_ASSETS: Mapping[str, int] = {
"vqa": 1,
"caption": 1,
"grounding": 1,
"change": 2,
"optical_sar": 2,
# Two assets, like `change` β the difference is the output, not the input:
# a short language answer rather than a change map.
"change_vqa": 2,
}
```
(`core/planner.py:118-142`.)
**`TASK_CAPABILITY` maps six `Task` members** β note it maps `Task.CHANGE_VQA`, so a forced
`change_vqa` is a first-class capability. `Task.UNSUPPORTED` is deliberately absent, because it is a
refusal, not a capability.
**`CAPABILITY_ASSETS` mirrors the specialists' own `validate_request`**, and the docstring names the
relationship:
> Mirrors the specialists' own `validate_request`, which stays authoritative; this is the planner's
> cheap precondition so it can refuse before construction is attempted.
(`core/planner.py:130-132`.)
**The authority is the specialist, not the planner.** `CAPABILITY_ASSETS` exists so the planner can
refuse cheaply; it does not replace the specialist's check, and the `_indices_for` helper (Β§40.5)
deliberately does not duplicate the pairing logic:
> A two-asset capability takes the first two in request order; a one-asset capability takes the
> first. Deliberately simple: the specialists validate their own asset count and pairing rules, and
> duplicating that logic here would give two places to disagree.
(`core/planner.py:530-534`.)
## 40. `plan()` β the single decision point
```python
def plan(
self,
prediction: "RouterPrediction",
request: AnalysisRequest,
*,
assets: list[AssetMetadata] | None = None,
) -> ExecutionPlan:
```
(`core/planner.py:276-282`.)
**`assets` is accepted and never read for policy.** The docstring says so: *"Not read for policy β
only the COUNT matters here, and that comes from `request.assets`. Accepted so the controller can
pass its already-resolved list without a second conversion."*
### 40.1 The `force_task` bypass
```python
asset_count = len(request.assets)
source = prediction.intent.source
# --- force_task bypasses the router entirely (Β§8) -----------------
# `Intent(source="forced")` already exists in the schema. When the
# caller forces a task, the router's opinion is not consulted at all,
# and that fact is recorded rather than hidden.
forced = request.force_task
task = forced if forced is not None else prediction.intent.task
if forced is not None:
source = "forced"
```
(`core/planner.py:296-306`.)
**The router is not consulted when a task is forced**, and the source is recorded as `"forced"` β
which is the third literal of `Intent.source` (`core/schemas.py:105`:
`Literal["learned", "lexical_fallback", "forced"]`). **A forced task still goes through the refusal
check and the widening rules**, so `force_task=Task.CHANGE` with one asset still refuses (Β§41).
### 40.2 The refusal check, first
```python
refusal = self._refusal_for(task, asset_count)
if refusal is not None:
return ExecutionPlan(
steps=(),
refused=True,
refusal=refusal,
uncertain=not prediction.above_threshold,
router_source=source,
effective_confidence=self._effective_confidence(prediction),
notes=(f"refused:{refusal.reason}",),
)
```
(`core/planner.py:309-319`.)
**A refusal is returned, never raised.** The module docstring states the contract:
> Refusing is not an error. It returns a valid plan with `refused=True` and a typed `PlanRefusal`, so
> the controller can assemble a normal envelope carrying the reason.
(`core/planner.py:75-79`.)
**The refusal's `uncertain` and `effective_confidence` are still populated**, so a refusal carries the
same router-provenance information a plan does.
### 40.3 The primary step
```python
primary = TASK_CAPABILITY[task]
self._append(
steps,
capability=primary,
indices=self._indices_for(primary, asset_count),
reason=f"task:{task.value}",
required=True,
params=self._params_for(primary, task),
)
```
(`core/planner.py:325-333`.)
**The primary step is always `required=True`** β a failure in it degrades the whole run, unlike the
widening steps which are `required=False`. **The `reason` is `task:<value>`**, e.g. `"task:change"` β
a rule identifier from the closed set, as the dataclass requires.
**`_params_for` exists because the VLM specialist branches on it:**
```python
@staticmethod
def _params_for(capability: str, task: Task) -> dict[str, Any]:
"""Specialist kwargs a step must carry.
The VLM specialist branches on `params["task"]` to decide between VQA
and captioning, so a `vlm`-backed step without it is unrunnable. The
mapping is by CAPABILITY, not by specialist, so `caption` passes
`"task": "caption"` even though the same object serves both.
"""
if capability in ("vqa", "caption"):
return {"task": capability}
return {}
```
(`core/planner.py:514-525`.) Only `vqa` and `caption` carry params, and both carry `{"task": <cap>}`
β because one specialist object serves both capabilities and needs to be told which.
### 40.4 The discretionary explanation step
```python
if (
not prediction.above_threshold
and prediction.intent.language_output
and "vqa" not in {s.capability for s in steps}
and self._capability_known("vqa")
):
self._append(
steps,
capability="vqa",
indices=(0,),
reason="uncertain_route:explain",
required=False,
)
notes.append("uncertain route: added an explanation step")
```
(`core/planner.py:339-352`.)
**Four conditions, all required.** The route must be uncertain; the request must want language; no
`vqa` step may already be planned (so a VQA request does not get a second VQA step); and the `vqa`
capability must be declared. The step is `required=False` and reads asset index `(0,)`.
**Note the ordering difference from the widening steps:** this check uses `_capability_known`
**at plan time**, whereas `_widening_steps` deliberately does **not** (Β§42.0). The comment above
`_widening_steps` explains the principle, and this block is a case where the availability gate is
applied inline β so the explanation step is never added *and then* dropped; it is simply not added.
Both approaches avoid the "silent omission" failure, because this one records
`"uncertain route: added an explanation step"` only when it fires and the drop path records its own
note when it drops.
### 40.5 `_indices_for` and `_append`
```python
@staticmethod
def _indices_for(capability: str, asset_count: int) -> tuple[int, ...]:
needed = CAPABILITY_ASSETS.get(capability, 1)
return tuple(range(min(needed, asset_count)))
```
(`core/planner.py:527-537`.) `range(min(needed, asset_count))` β so a two-asset capability on a
one-asset request yields `(0,)`, and the specialist's own validation is what ultimately refuses.
```python
@staticmethod
def _append(
steps: list[PlanStep],
*,
capability: str,
indices: tuple[int, ...],
reason: str,
required: bool,
params: Mapping[str, Any] | None = None,
) -> None:
steps.append(
PlanStep(
# Provisional id; `_renumber` assigns the canonical one after
# widening, availability filtering and any truncation, so the
# ids always match the final order.
step_id=f"step_{len(steps) + 1:03d}",
capability=capability,
specialist=capability,
asset_indices=indices,
params=params or {},
reason=reason,
required=required,
)
)
```
(`core/planner.py:539-562`.)
**`step_id` is provisional** β `_renumber` reassigns after widening, filtering and truncation. This
is why a truncated plan's step ids are still `step_001..step_N` with no gaps. **`specialist` is set
equal to `capability`** β the docstring on `PlanStep` notes they are kept distinct *"because the VQA
specialist answers to `"vqa"` AND `"caption"` while its own `name` is `"vlm"`"*
(`core/planner.py:167-171`); today the planner sets them equal and the registry resolves the rest.
### 40.6 The availability gate and the second refusal
```python
steps, avail_notes = self._drop_unavailable(steps)
notes.extend(avail_notes)
if not steps:
return ExecutionPlan(
steps=(),
refused=True,
refusal=self._refusal(
"model_unavailable",
"no_available_specialist",
"No specialist is available for this request in this "
"environment.",
),
uncertain=not prediction.above_threshold,
router_source=source,
effective_confidence=self._effective_confidence(prediction),
notes=tuple(notes),
)
```
(`core/planner.py:354-372`.)
**A plan that loses every step to availability becomes a refusal**, with code
`"model_unavailable"` and reason `"no_available_specialist"`. This is the **second** way a refusal
arises, and it is distinct from the closed list of Β§41 β which is why the reason string differs.
```python
def _drop_unavailable(
self, steps: list[PlanStep]
) -> tuple[list[PlanStep], list[str]]:
"""Remove steps whose capability the registry does not declare.
...
Note this filters only on *declaration*. Whether a declared capability
can actually be CONSTRUCTED is discovered by the controller, which
records an `UNAVAILABLE` registry entry β the planner must not attempt
construction, which would defeat the lazy-load design.
"""
known = {c for c in self.registry.available()} if self._registry_ok() else None
if known is None:
return steps, []
kept: list[PlanStep] = []
notes: list[str] = []
for step in steps:
if step.capability in known:
kept.append(step)
else:
notes.append(
f"dropped step {step.step_id} ({step.capability}): "
f"capability not registered"
)
return kept, notes
```
(`core/planner.py:572-602`.)
**Three design points, each stated in the source:**
1. **The planner filters on *declaration* only.** Whether a declared capability can be *constructed*
is the controller's discovery, which records an `UNAVAILABLE` registry entry. The planner must not
attempt construction, *"which would defeat the lazy-load design."*
2. **A broken registry is not a plan error.** `_registry_ok()` catches any exception and returns
`False`, in which case **no filtering happens** (`known is None` β return steps unchanged). A
registry that cannot be queried degrades to "plan everything", not "plan nothing".
3. **Every drop is recorded in `notes`**, so *"not planned because it cannot run"* stays
distinguishable from *"not planned because the planner chose not to"* (design Β§4.3).
### 40.7 Truncation, renumbering, and the mode
```python
if self.max_steps is not None and len(steps) > self.max_steps:
dropped = [s.capability for s in steps[self.max_steps:]]
notes.append(
f"plan truncated at {self.max_steps} steps; dropped {dropped}"
)
steps = steps[: self.max_steps]
steps = self._renumber(steps)
```
(`core/planner.py:375-382`.)
**Truncation is recorded with the names of the dropped capabilities** β never silent. The cap
mirrors `agent.max_specialists`, per the constructor docstring.
```python
@staticmethod
def _mode_for(steps: list[PlanStep]) -> PlanMode:
"""`PARALLEL_SAFE` only when the asset partitions are demonstrably safe.
Requires at least two steps, pairwise-disjoint asset indices, and no
step declaring a resource constraint. With today's four specialists
asset indices commonly overlap β a change plan and a caption plan can
both read asset 1 β so this legitimately returns SEQUENTIAL in normal
operation. That is correct, not a limitation (Β§3.4).
"""
if len(steps) < 2:
return PlanMode.SEQUENTIAL
seen: set[int] = set()
for step in steps:
indices = set(step.asset_indices)
if indices & seen:
return PlanMode.SEQUENTIAL
seen |= indices
return PlanMode.PARALLEL_SAFE
```
(`core/planner.py:613-632`.)
**`PARALLEL_SAFE` is a declaration, not a concurrency directive.** `PlanMode`'s docstring says so:
> `PARALLEL_SAFE` is a *declaration* that the steps share no mutable state, not a concurrency
> directive. v1 executes both modes sequentially β freeze section 5 forbids worker pools, queues and
> async frameworks. The marker records that a future `ThreadPoolExecutor` would be sound.
(`core/planner.py:149-155`.)
**In normal operation this returns `SEQUENTIAL`**, because a `change` + `caption` plan has both steps
reading asset 1 (`_indices_for("change", 2) = (0, 1)` and
`_indices_for("caption", 2) = (0,)` β index 0 in both). That overlap is *correct*, not a defect: the
two steps genuinely both read the same image.
## 41. The closed refusal list
```python
def _refusal_for(self, task: Task, asset_count: int) -> PlanRefusal | None:
"""The closed refusal list from Β§8. Order matters: query first."""
if task is Task.UNSUPPORTED:
return self._refusal(
UnsupportedQueryError.code,
"task_unsupported",
UnsupportedQueryError.user_message,
)
if asset_count < 1:
return self._refusal(
InvalidRequestError.code,
"zero_assets",
InvalidRequestError.user_message,
)
return None
```
(`core/planner.py:412-426`.)
**Two conditions, checked in a fixed order.** The docstring says *"Order matters: query first"* β
an `unsupported` task with zero assets reports `task_unsupported`, because the query is the more
fundamental problem.
| Order | Condition | Code | Reason | User message source |
|---|---|---|---|---|
| 1 | `task is Task.UNSUPPORTED` | `UnsupportedQueryError.code` | `"task_unsupported"` | `UnsupportedQueryError.user_message` |
| 2 | `asset_count < 1` | `InvalidRequestError.code` | `"zero_assets"` | `InvalidRequestError.user_message` |
**Both codes and user messages are imported from `core.errors`**, not written inline β so the
refusal's client-facing text and its error code come from the same definitions the exception classes
use. The module docstring's claim that the list is closed is what makes it testable:
> The refusal conditions are a closed list (Β§8) β no free-form logic, because a refusal rule
> invented at the call site is a policy that cannot be tested.
(`core/planner.py:80-81`.)
**Note what the list does *not* contain.** There is **no** refusal for "asset count too high for this
task" and **no** refusal for "asset count wrong for this specific capability". Those are the
specialists' `validate_request` responsibilities (Β§39.3) β so a `change` request with one asset
reaches the specialist and is refused there, not by the planner. That is the division the comment on
`CAPABILITY_ASSETS` states: the specialist stays authoritative.
**`_refusal()` is a thin constructor:**
```python
@staticmethod
def _refusal(code: str, reason: str, user_message: str) -> PlanRefusal:
return PlanRefusal(code=code, reason=reason, user_message=user_message)
```
(`core/planner.py:428-430`.)
## 42. The Β§3.5 widening rules β one task, several steps
```python
def _widening_steps(
self,
prediction: "RouterPrediction",
task: Task,
asset_count: int,
notes: list[str],
) -> list[PlanStep]:
"""The Β§3.5 rules that turn ONE router task into SEVERAL steps.
These are the reason the planner exists rather than a switch on
`intent.task`. The router's binary heads describe *aspects* of the
request that a single task label cannot carry.
"""
extra: list[PlanStep] = []
intent = prediction.intent
```
(`core/planner.py:434-448`.)
**This is the function that justifies the whole two-stage design.** The router returns *one* task;
the binary heads describe *aspects*; the widening rules are what recover the aspects the task label
cannot carry.
### 42.0 Why availability is not gated here
```python
# Availability is NOT gated here. The step is added whenever its rule
# fires, and `_drop_unavailable` removes it with a recorded note if the
# capability is unregistered. Gating here as well would create a second,
# silent availability check whose omission leaves no trace β exactly the
# "absence indistinguishable from a non-event" failure that Β§4.3 forbids.
```
(`core/planner.py:450-454`.)
**This is a deliberate choice, and it contradicts the pattern in Β§40.4.** The widening rules add
unconditionally; the explanation step checks availability inline. The comment's reasoning is sound
for the widening rules: a gate here plus `_drop_unavailable` would be two checks, and the inline one
would leave no trace when it silently declined. Recorded here as an internal inconsistency in
approach β not in outcome, because both paths produce a `notes` entry.
### 42.1 Rule one β `change` + `language_output`
```python
if task is Task.CHANGE and asset_count >= 2 and intent.language_output:
if self._capability_known("change_vqa"):
if not self._capability_planned(extra, "change_vqa"):
self._append(
extra,
capability="change_vqa",
indices=self._indices_for("change_vqa", asset_count),
reason="task:change+language_output:answer",
required=False,
)
notes.append(
"change + language request: added a change-VQA step"
)
elif not self._capability_planned(extra, "caption"):
self._append(
extra,
capability="caption",
indices=(asset_count - 1,),
reason="task:change+language_output",
required=False,
)
notes.append("change + language request: added a caption step")
```
(`core/planner.py:470-491`.)
**Three conditions:** the task is `change`; there are at least two assets (because `change` needs a
pair, and so does `change_vqa`); and the language head fired.
**The two branches are alternatives, not a sequence** β and the comment explains why:
> WHICH language output, though, depends on what the deployment has. With a change-VQA capability
> registered (R-02), the request is satisfied by an ANSWER to the change question β which is the
> specific thing that was asked. Without it, the fallback is to caption the later acquisition, which
> is the one a "what changed" answer describes.
>
> These are alternatives, not two things to do: captioning the post image is not a second
> requirement, it is the best available substitute when the answer capability is absent. Adding both
> would spend a step on a strictly weaker output.
(`core/planner.py:460-469`.)
**Two reasons, two rule ids, and a different asset index for the fallback:**
| Branch | Reason id | Asset indices | Meaning |
|---|---|---|---|
| `change_vqa` registered | `task:change+language_output:answer` | `(0, 1)` | answer the change question |
| `change_vqa` absent | `task:change+language_output` | `(asset_count - 1,)` | caption the **later** acquisition |
**`indices=(asset_count - 1,)` is the subtle one.** The caption step reads the **last** asset β the
post-change image β because that is the acquisition a "what changed" description describes. With two
assets, that is index 1.
### 42.2 Rule two β `spatial_output` + `language_output`
```python
# A request that wants BOTH coordinates and prose on one asset gets
# grounding and vqa. The router can only return one task, so without
# this rule the second aspect is silently lost.
if intent.spatial_output and intent.language_output and asset_count >= 1:
if task is not Task.GROUNDING:
if not self._capability_planned(extra, "grounding"):
self._append(
extra,
capability="grounding",
indices=(0,),
reason="aspect:spatial_output",
required=False,
)
notes.append("spatial+language request: added a grounding step")
```
(`core/planner.py:493-507`.)
**The `task is not Task.GROUNDING` guard** prevents adding a grounding step when grounding is
already the primary step. The step reads asset `(0,)` and is `required=False`.
**`_capability_planned` is the de-duplication helper:**
```python
@staticmethod
def _capability_planned(steps: list[PlanStep], capability: str) -> bool:
return any(step.capability == capability for step in steps)
```
(`core/planner.py:510-512`.) Note it checks the **`extra`** list, not the full step list β so a
primary step with the same capability does not suppress a widening step. The `task is not
Task.GROUNDING` guard is what prevents the duplicate in the one case that matters.
### 42.3 The widening rules, tabulated
| Rule | Fires when | Adds | Reason id | Indices | Required |
|---|---|---|---|---|---|
| 1a | `task=change` β§ `asset_count β₯ 2` β§ `language_output` β§ `change_vqa` known | `change_vqa` | `task:change+language_output:answer` | `(0, 1)` | F |
| 1b | same, but `change_vqa` **not** known β§ `caption` not already planned | `caption` | `task:change+language_output` | `(asset_count-1,)` | F |
| 2 | `spatial_output` β§ `language_output` β§ `asset_count β₯ 1` β§ `task β grounding` | `grounding` | `aspect:spatial_output` | `(0,)` | F |
**Worked example β the query the module docstring names.** *"what changed between these two images,
and describe the scene"* with two assets, a `change` reading and `language_output=True`:
| Step | Capability | Reason | Required | Indices |
|---|---|---|---|---|
| `step_001` | `change` | `task:change` | **T** | `(0, 1)` |
| `step_002` | `change_vqa` | `task:change+language_output:answer` | F | `(0, 1)` |
Mode: `SEQUENTIAL` β both steps read asset 0, so the partitions overlap. Notes:
`("change + language request: added a change-VQA step",)`.
**Without the widening rule, step 2 would not exist** and the caption/answer aspect would be lost β
which is precisely failure mode 2 in the module docstring.
## 43. The provenance discount
```python
@staticmethod
def _effective_confidence(prediction: "RouterPrediction") -> float:
"""The planner's reading of the router's confidence.
Applies the provenance discount for a lexical fallback (Β§2.4). The
router's own `Intent.confidence` is never modified β this is a separate
number used only for planning.
"""
raw = float(prediction.intent.confidence)
if prediction.intent.source == "lexical_fallback":
return raw * LEXICAL_FALLBACK_DISCOUNT
return raw
```
(`core/planner.py:397-408`.)
**The discount is 0.75, applied to a lexical-fallback confidence only.** The module docstring states
the reasoning in full:
> A lexical fallback at 0.9 is not the same evidence as a learned model at 0.9: one is a regex that
> matched, the other is a learned distribution. Treating them identically would let a matched keyword
> outrank the model it fell back from. So the planner applies a **provenance discount** β not a
> second numeric gate β to its own reading of the confidence, and never edits `Intent.confidence`
> itself. A measurement is the router's to report; the discount is the planner's reading.
(`core/planner.py:57-62`.)
**The discounted values, for the fallback's ten rule ids** (Β§27.10):
| Rule | Raw confidence | Effective (Γ 0.75) | Above `SETTLED_CONFIDENCE` 0.60? |
|---|---|---|---|
| `dual_modality` | 0.92 | **0.690** | yes |
| `temporal_spatial` | 0.90 | **0.675** | yes |
| `spatial` | 0.88 | **0.660** | yes |
| `temporal` | 0.85 | **0.6375** | yes |
| `caption` | 0.85 | **0.6375** | yes |
| `vqa` | 0.78 | **0.585** | **no** |
| `sar_single` | 0.75 | **0.5625** | **no** |
| `optical_single` | 0.72 | **0.540** | **no** |
| `no_match` | 0.30 | **0.225** | **no** |
| `empty_query` | 0.00 | **0.000** | no |
**Five of the ten rules discount below 0.60.** The three lowest-confidence task-producing rules β
`vqa`, `sar_single`, `optical_single` β all fall below `SETTLED_CONFIDENCE` after the discount, which
is exactly the intended effect: a bare `"what"` or `"optical image"` match is weak evidence and
should be treated as an unsettled route. The discount does not by itself change the plan (because
`SETTLED_CONFIDENCE` is not read by `plan()`), but it is the number `ExecutionPlan.effective_confidence`
carries into the trace β so a reader of the trace sees `0.585` for a fallback VQA, not `0.78`.
**`ExecutionPlan.effective_confidence` is rounded to 4 dp in the trace**
(`core/planner.py:245`), and the `router_source` field carries `"trained"`, `"lexical_fallback"` or
`"forced"` β so the discount and the provenance travel together.
---
# Part J β The router-defect case study
## 44. Symptom
`docs/RESEARCH_NOTES.md` Β§3.1 states it in three lines:
> The query *"Where are the built-up areas in this image?"* collapsed to **`vqa`** and answered
> **"River"** β instead of routing to `grounding`. A second query, *"Where is the new airport?"*,
> behaved the same way.
**Two queries, both legitimate location questions, both answered as if they were questions about a
single object.** The answer `"River"` is the tell: the VQA specialist was asked *where* something was
and answered with a *land-cover class*. The request was for coordinates and the response was a noun.
`README.md:556-562` gives the release-level summary:
> An earlier revision evaluated the temporal rule before the location rule, so *"Where are the
> built-up areas in this image?"* matched `\bbuilt\b` as a *change* marker and `area` inside
> *"areas"* as a quantifier. With one asset it collapsed to `vqa` and answered **"River"**. Fixed on
> 2026-09-25 in `frontend/assets/js/mission.js`; the fix is covered by regression tests and verified
> live. The same defect existed on a second surface (`SQ.policy` in `core.js`) and was fixed the same
> day.
**Three things in that paragraph matter:**
1. **Two compounding mis-reads.** `\bbuilt\b` fired the *temporal* predicate, and `area` matched
inside `"areas"` to fire the *quantifier* upgrade.
2. **The collapse was to `vqa`**, not to `change`. With one asset, a `change` reading would be
substituted to `vqa` by the pair guard (Β§37.2) β so the temporal mis-read *and* the pair guard
together produced the VQA answer. Two layers of the pipeline each did the locally-correct thing on
a wrong input.
3. **The defect existed on two surfaces** β `frontend/assets/js/mission.js` and `SQ.policy` in
`frontend/assets/js/core.js` β and both were fixed on 2026-09-25. This is the sibling-site pattern:
a fix applied to one implementation of a rule is not a fix until every implementation of that rule
is checked.
## 45. Root cause
`docs/RESEARCH_NOTES.md` Β§3.2:
> Two functions with different information:
>
> - **`interpret()`** β produces the console's *reading*; **asset-count-blind** (text only).
> - **`chooseTask()`** β performs *dispatch*; **asset-count-aware**.
>
> The defect was in the dispatch path's handling of spatial/lexical cues, so region queries fell
> through to the generic VQA specialist.
**The root cause is a lexical predicate, not a control-flow bug.** The `interpret()` ladder (Β§36.2)
is correct in structure β `sar` β `temporal` β `where` β `caption` β `vqa` β but its **inputs** were
wrong: `temporal` was `'required'` for a query that was not temporal, because `built` was in the
temporal term set.
**Trace the defective path, step by step:**
| Stage | Defective behaviour |
|---|---|
| `interpret("Where are the built-up areas in this image?")` | `where` = `true`; `changeStem` = `false`; `pairWording` = `false`; **`newAsChange`** = `false`; **temporal = `'required'`** because `built` was a temporal term |
| branch selection | `sar` false β **`temporal === 'required'` is TRUE** β `task = 'change'`, `spatial = 'box'` |
| `chooseTask(..., assetCount=1)` | `wanted = 'change'`; the upgrade predicate matched `area` inside `"areas"` β **`wanted = 'change_vqa'`**; `assetCount < 2` and `change_vqa` requires a pair β **substitute `vqa`** |
| server | `vqa` specialist answers the question about one image β **"River"** |
**Both mis-reads were in the *reading*, and the dispatch compounded them.** The `interpret()` output
already had `spatial_output = 'box'` β the reading *knew* this was a spatial request β but the
`task` was `change`, and the ladder's `change` branch is what carried the wrong task forward. The
spatial flag was never consulted for task selection because the temporal branch had already claimed
the query.
**This is why the fix had to be in the predicates, not the ladder.** Reordering the branches would
not help: with `temporal === 'required'` true, the temporal branch fires before `where` under any
ordering that puts temporal above spatial β and the correct reading requires `where` to win. The only
fix that makes `"Where are the built-up areas in this image?"` a `grounding` request is to stop
`built` from making it temporal.
## 46. The four lexical changes
`README.md:564-571` presents the fix as a table, and the source comments record each change
verbatim. All four are in `frontend/assets/js/mission.js`.
| # | Change | Source location | Why |
|---|---|---|---|
| 1 | **`built` removed from the temporal term set entirely** | `mission.js:87-91` | *"built-up areas"* is land-cover vocabulary, not a change marker. While it sat in the temporal set, the location question *"Where are the built-up areas in this image?"* was read as a change request and answered with the degenerate one-word *"River"*. **Measured live, 2026-09-25.** |
| 2 | **`\barea\b` instead of bare `area`** | `mission.js:189-194` | Without the boundary the substring matched inside *"areas"*, so the already-mis-read change question was upgraded **again** to `change_vqa`. The boundary keeps the quantifier reading for a real *"how much area changed"* while refusing the plural land-cover noun. |
| 3 | **`new` counts as a change marker only when the query is not a `where` question** | `mission.js:84-92` | The repo ships `eo/new-airport.jpg`, so *"Where is the new airport?"* is a real question, and `new` is a place descriptor as often as a change marker. |
| 4 | **the change stem is matched without a trailing `\b`** | `mission.js:73-79` | `\bchang\b` cannot match *"changed"*, *"changes"* or *"changing"* β there is no word boundary between the stem and its inflection. With the boundary, the page's own default question (*"What changed here?"*) fell through to the `vqa` branch, so the change path was **unreachable from the UI that exists to reach it**. |
### 46.1 Change 1 β `built` removed
The source comment, verbatim (`frontend/assets/js/mission.js:87-91`):
```
`built` was dropped outright -- "built-up areas" is land-cover
vocabulary, not a change marker. While it sat in the temporal set, the
location question "Where are the built-up areas in this image?" was read
as a change request and the server answered it with the degenerate
one-word "River". Measured live, 2026-09-25.
```
**Note the corroboration in the corpus.** The curated examples include
`Compare optical and SAR to identify built-up regions.` and
`Show built-up areas using radar and optical together.` β both labelled `optical_sar`, and both
containing "built-up" (Β§12.1). The corpus had already learned that "built-up" is a **land-cover
noun**; the frontend predicate had not. The fix aligns the predicate with the corpus.
### 46.2 Change 2 β `\barea\b`
The source comment, verbatim (`frontend/assets/js/mission.js:189-193`):
```
`\barea\b`, not bare `area`: without the boundary the substring matched
inside "areas", so "Where are the built-up areas ..." (already a
mis-read change question) was upgraded again to change_vqa. The
boundary keeps the quantifier reading for a real "how much area
changed" while refusing the plural land-cover noun.
```
**This is the fix that prevents the *second* error.** Even after change 1, the query would be
`grounding` β but without change 2, a *different* query would still be mis-dispatched. The comment's
parenthetical β *"(already a mis-read change question)"* β records that this boundary fix was applied
knowing the upstream reading was also being fixed, i.e. it is defence in depth rather than the
primary repair.
### 46.3 Change 3 β `new` gated on `where`
```javascript
var newAsChange = /\bnew\b/.test(q) && !where;
```
(`frontend/assets/js/mission.js:92`.)
**A single boolean AND, and it is the whole change.** `new` fires the temporal predicate only when
the query is **not** a `where` question. So:
| Query | `where` | `newAsChange` | temporal |
|---|---|---|---|
| `Where is the new airport?` | `true` | `false` | `'none'` β **grounding** |
| `What is the new runway?` | `false` | **`true`** | `'required'` β **change** |
| `Show me the new construction.` | `true` (`show me the`) | `false` | `'none'` β grounding |
**Row 2 is the documented residual** (Β§48): *"What is the new runway?"* reads `change` rather than
`vqa`. The gate is `where`-specific, so a non-`where` question containing `new` still reads temporal.
That is the *known and accepted* cost of the fix β the alternative would be to drop `new` entirely,
which would lose the temporal reading of `"What is new here?"`.
### 46.4 Change 4 β the `chang` stem
```javascript
var changeStem = /chang/.test(q);
```
(`frontend/assets/js/mission.js:79`.)
**The single most severe of the four, because it made a whole path unreachable.** The comment
explains the mechanism exactly:
```
The change stem is matched WITHOUT a trailing \b: `\bchang\b` cannot
match "changed", "changes" or "changing", because there is no word
boundary between the stem and its inflection. With the boundary, the
page's own default question ("What changed here?") fell through to the
vqa branch β so the change path was unreachable from the UI that exists
to reach it. The other terms keep their boundaries; they are whole words.
```
(`frontend/assets/js/mission.js:73-78`.)
**`\bchang\b` requires a word boundary after the `g`.** In `"changed"`, the character after `chang`
is `e`, which is a word character β so there is **no** boundary, and the pattern fails. The same
applies to `"changes"` and `"changing"`. The predicate matched only the bare word `"chang"`, which
appears in no English sentence.
**The consequence is stated in the comment and is worse than a single mis-routed query:** the page's
**own default question** is `'What changed here?'` (`mission.js:26`:
`var QUERY = params.get('q') || 'What changed here?';`). So the default question β the one the page
asks when no query is supplied β fell through to `vqa`. The change path was **unreachable from the
default UI state**.
**The final sentence is the generalisable lesson:** *"The other terms keep their boundaries; they are
whole words."* The `chang` stem is the one predicate that is deliberately **not** word-bounded,
because it is a stem rather than a word. Every other predicate in `interpret()` is either a whole
word or an explicit alternation.
### 46.5 The four changes, and which defect query each fixes
| Query | Fixed by | Result |
|---|---|---|
| `Where are the built-up areas in this image?` | **changes 1 + 2** | `grounding` |
| `Where is the new airport?` | **change 3** | `grounding` |
| `What changed here?` (the page default) | **change 4** | `change` |
| `How much area changed?` (regression guard) | change 2 (preserved) | still `change_vqa` |
**Change 2 is a two-sided fix:** it *fixes* the `"areas"` false positive **and** *preserves* the
`"how much area changed"` true positive. A fix that broke the quantifier reading would have been a
regression, which is why the boundary form is `\barea\b` rather than removing `area` from the
predicate.
## 47. Live verification
### 47.1 Three independent passes
`docs/RESEARCH_NOTES.md` Β§3.3 and `README.md:935-942` record the same table:
| Pass | Deployed HEAD | Result |
|---|---|---|
| 1 | `ff46eba42b18` + `d413d3672311` | **8/8** |
| 2 | `2d7ae53b482d` | **8/8** |
| 3 | `2d7ae53b482d` | **8/8** |
**8 cases per pass** β 6 regression plus 2 defect (`README.md:927`). **3 of 3 passes at 8/8**, for
**24 live runs** and **24 correct dispatches**, with **0 mock-node contamination** on every run
(`README.md:924-933`).
**Two of the three passes ran against the same HEAD (`2d7ae53b482d`).** That is not redundancy for
its own sake β it is what makes pass 3 an *independent* confirmation of pass 2 rather than a repeat
of a shared failure mode. `README.md:935` states the property that makes the passes genuinely
independent: **"Each pass produced fresh run identifiers β no run id is shared between passes."**
### 47.2 The two defect queries, with run ids
| Query | Run id | Dispatched |
|---|---|---|
| `Where are the built-up areas in this image?` | **`run_467ffa406f22`** | `grounding` |
| `Where is the new airport?` | **`run_46980ba55c62`** | `grounding` |
(`docs/RESEARCH_NOTES.md:139-142`; `README.md:970-971`.) `README.md:973` adds:
**"The last two are the router-defect queries. Both previously collapsed to `vqa` and answered
'River'."**
**Both now dispatch to `grounding`** β the correct task for a location question.
### 47.3 The full pass-3 run-id table
`README.md:962-971` records one complete pass, and the same pass the screenshots are drawn from. It is
reproduced here because it is the strongest available evidence that the router works end-to-end:
| Case | Query | Dispatched | Run ID |
|---|---|---|---|
| vqa | What type of terrain dominates this scene? | `vqa` | `run_fef26e91e7e6` |
| caption | Describe the main visual characteristics of this scene. | `caption` | `run_96281bdfcc08` |
| grounding | Where are the visible buildings in this image? | `grounding` | `run_e49adc8d319f` |
| change | What changed between the earlier and later image? | `change` | `run_aedc59cbcdc9` |
| change_vqa | Did the coastline advance between the two observations? | `change_vqa` | `run_62ca98d510be` |
| optical_sar | β¦combining the optical and SAR observations? | `optical_sar` | `run_beacf6aa4e21` |
| **grounding** | **Where are the built-up areas in this image?** | **`grounding`** | **`run_467ffa406f22`** |
| **grounding** | **Where is the new airport?** | **`grounding`** | **`run_46980ba55c62`** |
**All six tasks are exercised in one pass.** The `change_vqa` row is the documented quantifier
upgrade working as designed:
> Note the fifth row: *"Did the coastline advance between the two observations?"* is read as `change`
> and **dispatches `change_vqa`** β the documented quantifier upgrade, because the page's Answer block
> promises an answer and the server's `change` returns a spatial map with no language output.
(`README.md:975-979`.)
### 47.4 The harness false-positive β why the passes are trustworthy
`docs/RESEARCH_NOTES.md` Β§4 records a failure of the *harness*, not the code, and it is included
because it is the reason the passes can be believed:
> An earlier live-validation harness typed queries with **synthetic CDP key events**, which Chrome
> **silently drops when the window lacks OS focus**. The harness therefore dispatched the page's
> *default* query and still recorded a "result" β a **false pass**.
>
> **Fix:** the current harness **asserts form state before dispatch** (`q_ok`, `obs_ok`, `t0_ok`), and
> uses deterministic query entry (`js()` value-set + `type_text()` via CDP `Input.insertText`).
>
> **Independent check:** the earlier 8/8 run was re-examined and confirmed **not** infected β its
> answers were query-specific and the query text was embedded in the answers.
**The false-positive mechanism is worth naming precisely.** Chrome dropped the synthetic keystrokes,
so the query box still held the page's default (`'What changed here?'`), the harness pressed Run, and
a real result came back for the *wrong query*. Because a result came back, the harness recorded a
pass. **The default query is `'What changed here?'` β which is exactly the query change 4 was needed
to fix (Β§46.4).** So the harness false positive and the `chang`-stem defect interacted: a harness that
dispatched the default question would have been asking the one question the broken predicate
mis-routed.
**The independent re-examination is what makes the earlier pass usable**, and it used the right test:
*"its answers were query-specific and the query text was embedded in the answers."* A genuine pass
cannot be produced by a query the harness never entered.
### 47.5 The frontend regression suite
`README.md:933` records **106 passed** in `tests/unit/test_frontend_live_wiring.py`. This is the
suite that pins the routing behaviour, and it is the reason the four lexical changes cannot silently
regress.
**Status of the fix: `RESOLVED`.** Deployed to `SatQuery-Frontend`, verified across three independent
live passes, and covered by regression tests. The style guide's deployed-HEAD fact applies:
**Frontend `2d7ae53b482d`**.
---
# Part K β Residuals and honest boundaries
## 48. The known residuals
`README.md` limitation 8 states them, and they are the honest frontier of a lexical router:
> **Router lexical residuals.** *"What is the new runway?"* reads `change` rather than `vqa` (the
> `new`-as-change heuristic fires outside `where` questions), and *"How much built-up area was
> added?"* reads `vqa` (under-trigger). A lexical router cannot cleanly separate "the new X" from
> "what's new"; a trained intent router exists in `artifacts/router/` but is not attached.
**Two residuals, two different mechanisms:**
| Query | Reads as | Correct would be | Mechanism |
|---|---|---|---|
| `What is the new runway?` | `change` | `vqa` | **over-trigger** β `newAsChange` fires because the query is not a `where` question |
| `How much built-up area was added?` | `vqa` | `change_vqa` (or `change`) | **under-trigger** β no predicate claims it |
### 48.1 Residual 1 β the `new` over-trigger
`What is the new runway?` has `where = false` (no `where`/`locate`/`position`/`bound`/`outline`/
`coordinate` term), so `newAsChange = /\bnew\b/.test(q) && !where` is **`true`**, and the reading is
`change`.
**This is the accepted cost of change 3** (Β§46.3). The gate is `where`-specific: it fixes
`"Where is the new airport?"` without dropping `new` from the temporal vocabulary, which would lose
`"What is new here?"`. The residual is the queries in between β a non-`where` question about a
newly-built feature.
**The README's own diagnosis is the generalisable statement:** *"A lexical router cannot cleanly
separate 'the new X' from 'what's new'."* Both readings are defensible from the surface form; the
distinction requires knowing whether `new` modifies a **noun** (a new runway) or a **state** (what is
new).
### 48.2 Residual 2 β the `built-up area` under-trigger
`How much built-up area was added?` reads `vqa`. It has:
- `where = false` β no location term;
- `changeStem = false` β no `chang` stem;
- `pairWording = false` β none of the bounded pair terms match (`added` is not in the list);
- `newAsChange = false` β no `new`;
- `sar = false`;
- `caption = false`;
- so the ladder falls through to the final `else` β `vqa`.
**Note that change 2 is what prevents this query from being *upgraded*:** `\barea\b` does not match
`"area"` inside `"built-up area"`... actually it **does** β `"built-up area"` ends with the word
`area`, so `\barea\b` **matches**. But the upgrade predicate is only consulted inside
`if (wanted === 'change')`, and `wanted` is `'vqa'` here because the reading was `vqa`. So the
upgrade never runs. **The under-trigger is in `interpret()`, not in the upgrade predicate.**
**The residual is a missing predicate, not a wrong one.** `added` is a change verb, but it is not in
`pairWording`'s alternation (`between|versus|vs\.?|pair|temporal|before|after|differ\w*|expand\w*|
grow\w*|encroach\w*|lost|removed`). Adding `add\w*` would fix this query β and would risk new false
positives, which is why a lexical router has a frontier rather than a fixpoint.
### 48.3 The trained router exists and is not attached
The README's final sentence is the most important part of the residual entry: **"a trained intent
router exists in `artifacts/router/` but is not attached."**
**This is the state of the whole router subsystem at this release.** The trained adapter is real,
measured, and shipped (Β§30βΒ§33). The **frontend console does not use it** β the console runs
`interpret()` and `chooseTask()`, both lexical. And `IntentRouter.from_config` defaults
`adapter_path` to `None`, so **the server's default router is also lexical** unless a caller passes a
path (`router/classifier.py:143-153`).
**So at this release, in normal operation, the router is lexical on both surfaces, and the trained
adapter is a shipped artifact that is not wired in by default.** That is the single most important
fact a reader should take from this chapter, and it is `OPEN` β not a defect, but an explicit
deployment decision that the docs record rather than hide.
## 49. What is `NOT RUN`, `OPEN`, `BLOCKED` or `REJECTED` for this topic
Per `DOCS_STYLE_GUIDE.md` Β§4, every chapter ends with an explicit list. For the router:
| Item | Status | Detail |
|---|---|---|
| Router **test split**, at the release level | **`NOT RUN`** | `threshold_sweep_val.json` β `n_test_examples_scored: 0`, `test_split_touched: false`. The training-time test score (0.975, n=80) exists in `metadata.json`; the release-level test evaluation does not. |
| Router **calibration** | **`NOT RUN`** | The 0.70 threshold is uncalibrated by the phase's own admission: *"Calibration belongs in Phase 13 with real validation data, not here."* |
| Router **accuracy as a benchmark claim** | **`NOT RUN`** | The corpus is synthetic β 576 examples, 54 groups, `corpus_limited: true`. *"Treat the router as working, not as benchmarked."* |
| Plan minimum **val queries β₯ 500** | **`NOT RUN`** β missed | 86 scored against a plan minimum of 500 (`plan_min_val_queries`). |
| Plan minimum **hard negatives β₯ 100** | **`NOT RUN`** β missed | The sweep's val split carries **0** hard negatives (`hard_negatives_in_val: 0`) against a plan minimum of 100. The `hn_*` families are held out to test **by design**, so val can never satisfy this. |
| The two `hn_*` test misses (0.800 = 11/14) | **`OPEN`** | Which two examples were mis-classified is not recorded in any artifact. |
| The 50,822 vs 51,725 parameter discrepancy in three comments | **`OPEN`** | Documentation-only; the code and the artifact both give 51,725. Β§8. |
| Whether the training-time test score (0.975) counts as a test-set result | **`OPEN`** | `README.md` limitation 2 says the test set *"was never run"*; the artifact contains a test evaluation. Β§33.3. |
| Router **lexical residuals** (`"What is the new runway?"`, `"How much built-up area was added?"`) | **`OPEN`** | `README.md` limitation 8. Β§48. |
| The trained adapter is **not attached by default** on either surface | **`OPEN`** | `IntentRouter.from_config(adapter_path=None)` defaults to the fallback; the console is lexical. Β§48.3. |
| `adapter_source` returns `"trained"` when an adapter exists but no encoder is loaded | **`OPEN`** | The property answers "is an adapter attached?", while its docstring claims "which path `route()` will take". Β§20.2. |
| `SETTLED_CONFIDENCE = 0.60` has no reader in `plan()` | **`OPEN`** | The constant documents design Β§2.3; the implementation gates on `above_threshold`. Β§39.2. |
| The coherence repair fires silently β nothing records that it did | **`OPEN`** | The comment says *"we record that we did so"*; `to_trace()` has no such field. Β§21. |
| Per-family provenance of the shipped 576-example corpus | **`UNKNOWN`** | The current `dataset.py` declares 70 curated + 45 templates = 585 before dedupe; the recorded totals exceed the declared totals for four families. Β§13.3. |
| The exact list of duplicates `dedupe()` removed | **`UNKNOWN`** | Four within-`CURATED` duplicates are visible by inspection; the complete list requires running `build_corpus()` + `dedupe()`. Β§12.3. |
| The derived `class_weights` used in training | **`UNKNOWN`** | Not recorded in `metadata.json`; computable from `split.task_counts.train`. Β§15.3. |
| The output of `self_check()` on the shipped code | **`UNKNOWN`** | Not executed by this document and not recorded in any artifact. Β§29. |
| The shipped adapter's training corpus, byte-for-byte | **`UNKNOWN`** | The adapter is dated 2026-09-16; `router/dataset.py` has been edited since. Β§30.1. |
| Whether `interpret()`'s bare-`sar` rule (single modality word β `optical_sar`) is intended | **`OPEN`** | Differs from the Python fallback, which has separate SAR-only and optical-only rules. Β§36.2. |
| The frontend/backend routing parity | **`NOT RUN`** | No test compares `interpret()`/`chooseTask()` against `IntentRouter.route()`/`PolicyPlanner.plan()` on the same queries. |
## 50. Where the evidence lives
### 50.1 Source of truth (the implementation)
| Path | What it establishes |
|---|---|
| `router/encoder.py` | the frozen MiniLM contract, `max_length β€ 256`, `build_encoder` |
| `router/adapter.py` | the 5-head architecture, init, forward, serialisation |
| `router/label_space.py` | the six task classes, four modalities, three binary heads |
| `router/fallback.py` | the seven term tables, eight rules, `self_check()` |
| `router/classifier.py` | `IntentRouter`, the 0.70 gate, the coherence repair, persistence |
| `router/dataset.py` | `CURATED`, `TEMPLATES`, `SUBJECTS`, `split_by_group`, `validate` |
| `router/train.py` | `train_router`, the loss, the cache, `gate2_passed` |
| `core/planner.py` | `PolicyPlanner`, the refusal list, the widening rules, the discount |
| `core/schemas.py` | `Task` (7), `Modality` (4), `Intent` + `_consistency` |
| `configs/base.yaml` Β§`router` | every configured router value |
| `frontend/assets/js/mission.js` | `interpret()`, `chooseTask()`, `assetsForTask`, `validateOpticalSar` |
| `frontend/assets/js/core.js` | `SQ.policy` β the **second surface** of the router defect |
### 50.2 Artifacts (the measurements)
| Path | What it records |
|---|---|
| `artifacts/router/router_adapter_v001/metadata.json` | the shipped adapter: 51,725 params, corpus 576/54, all three split metric blocks, the full 60-epoch history, `duration_seconds: 4.92` |
| `artifacts/router/router_adapter_v001/adapter.pt` | the weights + the embedded architecture config |
| `artifacts/router/threshold_sweep_val.json` | the release-level sweep: val only, 50 thresholds, `test_split_touched: false`, `overall_ungated_accuracy: 0.965116` |
| `artifacts/router/cache/` | the cached corpus embeddings (`.npy` + `.json` per fingerprint) |
### 50.3 Phase records and notes
| Path | What it establishes |
|---|---|
| `docs/PHASE4_ROUTER_REPORT.md` | Gate 2 PASS, F4-1/F4-2/F4-3, four defects found by running, the standing caveats |
| `release/repo/docs/RESEARCH_NOTES.md` Β§3 | the router-defect case study, the live-pass table, the two run ids |
| `release/repo/docs/RESEARCH_NOTES.md` Β§4 | the harness false-positive and its fix |
| `release/repo/docs/RESEARCH_NOTES.md` Β§6 | the `interpret()`/`chooseTask()` asymmetry ruling (`RESOLVED`) |
| `release/repo/README.md` Β§"Routing and the execution trace" | the two-stage description, the five-head table, the confidence-gate diagram, the four lexical changes |
| `release/repo/README.md` Β§"Live validation" | the three passes, the 24 runs, the full pass-3 run-id table |
| `release/repo/README.md` Β§"Known limitations" rows 2, 8 | the validation-only status and the two residuals |
| `release/repo/DOCS_STYLE_GUIDE.md` Β§3 | the router fact that must never be stated wrongly: *"0.965116 is validation, ungated, n = 86; the test split was NOT RUN"* |
### 50.4 Tests
| Path | What it pins |
|---|---|
| `tests/routing/test_router.py` | **82 tests** (`docs/PHASE4_ROUTER_REPORT.md` Β§Files) β the label space, the encoder guards, the adapter shapes, the splitter, the fallback's `self_check()`, the gate |
| `tests/unit/test_frontend_live_wiring.py` | **106 passed** (`README.md:933`) β the frontend routing and the live-wiring assertions |
**Named guards recorded in the phase report:**
| Test | Guards against |
|---|---|
| `test_stratification_does_not_weaken_leakage_safety` | stratification breaking group-level leakage safety, across five seeds |
| `test_no_template_uses_another_tasks_characteristic_vocabulary` | a template whose surface form belongs to another task |
| `test_no_class_dominates_the_corpus` | any class exceeding 35 % of the corpus |
| `test_leakage_report_catches_a_deliberately_leaky_split` | the leakage audit itself β **this guard was once deleted by a patch that meant to append, and the suite still reported 258 passed** |
| `test_every_guard_in_the_suite_still_exists` | the meta-test that catches a deleted guard; **proven to fire** by renaming a guard and confirming the build breaks |
(`docs/PHASE4_ROUTER_REPORT.md` Β§"Four real defects", defect 3.)
**The deleted-guard defect is the most instructive item in the phase record**, and it is reproduced
here in the report's own words:
> A patch intended to *append* a test instead *replaced*
> `test_leakage_report_catches_a_deliberately_leaky_split`. The suite still reported green β 258
> passed β because nothing was asserting leakage detection any more.
>
> **Fix:** restored the guard, and added `test_every_guard_in_the_suite_still_exists` β a meta-test
> that imports each module and asserts the named controls are present. It was then **proven to fire**
> by renaming the guard, confirming the build breaks, and restoring.
>
> Its first run also failed, correctly: it listed `test_fusion_dim_mismatch_is_rejected` in the
> routing module, where it does not live. A meta-test making a false claim should fail, and it did.
**A green suite is not evidence that a guard exists.** The meta-test exists because the phase learned
that lesson by running the suite and finding 258 passing tests with the leakage guard gone.
---
## Chapter summary
The router is the first verb β it *understands*, and nothing more. It is two stages on both surfaces:
a **reading** that is blind to asset count (`interpret()`, `IntentRouter.route()`) and a **dispatch**
that is not (`chooseTask()`, `PolicyPlanner.plan()`).
The learned path is a frozen MiniLM encoder (`1110a243fdf4`, 384-d, `max_length` 128 inside a verified
256 ceiling) feeding a **51,725-parameter five-head adapter** trained on cached embeddings in
**4.92 seconds on CPU**. It scored **0.975 on its training-time test split** with all six classes
measured, and **0.800 hard-negative accuracy** β the most honest generalisation number in the
artifact. The release-level threshold sweep was **val only**, and its own note says *"the test split
was NOT touched."*
The lexical path is eight ordered rules over seven term tables, never inventing capability, refusing
at **0.30** rather than guessing. It is what runs by default, because `adapter_path` defaults to
`None`.
The router's most instructive episode is a **defect** that was found live, diagnosed to two
compounding lexical mis-reads, fixed with four documented changes, and verified across **three
independent live passes at 8/8** β with run ids `run_467ffa406f22` and `run_46980ba55c62` now
dispatching `grounding` where they once collapsed to `vqa` and answered *"River"*. Two residuals
remain, and a trained router sits in `artifacts/router/` not yet attached.
|