File size: 78,833 Bytes
872cf4d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 1332 1333 1334 1335 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 1396 1397 1398 1399 1400 1401 1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 1459 1460 1461 1462 1463 1464 1465 1466 1467 1468 1469 1470 1471 1472 1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 1498 1499 1500 1501 1502 1503 1504 1505 1506 1507 1508 1509 1510 1511 1512 1513 1514 1515 1516 1517 | # Progress log β offline goal-directed controller on LeWM PushT
Running record of every experiment, in order, with the reasoning and the
numbers. Appended as runs complete. Newest section at the bottom.
## Setup
| component | value |
|---|---|
| world model | LeWM PushT, **frozen** (ViT-tiny enc, patch 14, 224px, 192-dim; 6-layer predictor w/ AdaLN action conditioning) |
| dataset | `data/swm_home/datasets/pusht_expert_train.h5` (expert PushT) |
| latent cache | `data/latents/` β 2.34M x 192 fp16, encoded once |
| action blocking | frameskip 5, so one WM transition = 5 env actions; block dim 10 |
| context | N = 3 frames |
| plan horizon | H = 5 blocks = 25 env actions |
| refinements | K = 3 at train time |
| controller | 6.8M params, shared F_theta / G_theta across all K |
| GPU | RTX 5080 Laptop, 16 GB |
| env | conda `llmdyn` (`/c/Users/omnap/miniconda3/envs/llmdyn/python.exe`) |
The controller is **never** trained on dataset actions. Gradient reaches the
plan only through consequences predicted by the frozen predictor. The dataset
supplies goals (hindsight relabeling) and a behavior density used purely as a
support constraint.
## Evaluation protocol
All rows use the same 50 held-out start/goal pairs (seed 42), goal offset 25
env steps, budget 50 steps, through the same `WorldModelPolicy` and wrappers.
Only the planner differs. Success is the env's own criterion
(`pos_diff < 20 and angle_diff < pi/9`), latched across the episode.
Cost is reported as **predictor rows per episode** (one row = one batched
latent transition) alongside wall-clock, because CEM parallelizes heavily and
the two numbers tell different stories.
---
## Experiment 0 β main training run
`scripts/train_controller.py --steps 20000 --batch-size 128 --lr 3e-4`
20000 steps, 89.2 min. Final held-out terminal distance **0.0147** at K=3 vs
**0.0582** at K=0 β refinement cuts terminal distance ~4x. Learned step sizes
`[0.55, 0.451, 0.312]`, decaying across iterations as a converging iterative
solver should.
**Checked it is not behavior cloning** (step-10000 ckpt, 256 held-out
samples): the first predicted action block differs substantially from the
dataset block (cosine 0.552 mean / 0.673 median; mean abs diff 0.6047 against
a real-block abs mean of 0.6828), yet reaches the goal far better than
replaying the expert block (terminal 0.0211 vs 0.3930, from start 0.7176),
winning on 99.6% of samples. It is solving the task, not imitating.
## Experiment 1a β headline table (receding horizon = 1)
| planner | success% | rows/ep | sec/ep |
|---|---|---|---|
| CEM (300 samples x 30 steps) | 34.0 | 383400 | 7.04 |
| controller K=0 | 36.0 | 43 | 0.18 |
| controller K=1 | 50.0 | 74 | 0.18 |
| controller K=2 | 52.0 | 109 | 0.18 |
| controller K=3 | 52.0 | 144 | 0.18 |
| controller K=5 | 42.0 | 232 | 0.20 |
| controller K=3, execute full plan | **88.0** | 25 | 0.13 |
Refinement helps as the design doc predicts: 36 -> 50 -> 52% for K = 0, 1, 2,
saturating at K=3. K=5 exceeds the trained unroll depth and degrades to 42% β
the learned recurrent update is not guaranteed contractive outside the depth
it was trained at.
With 50 episodes one episode is 2 points, so **K=2 and K=3 are tied**, and
K=2 is the Pareto choice at rh=1 (same success, 109 vs 144 rows/ep).
### Failed runs worth recording
An earlier background driver was orphaned when its wrapper shell was killed.
Every subprocess after the CEM row died instantly with exit `3221225794`
(`0xC0000142`, DLL init failure), and a just-added continue-on-error path let
the driver march through the whole matrix in seconds producing nothing. No
results were corrupted β nothing had trained β but it cost a cycle. Lesson
recorded in the runbook below: verify a background run is producing rows
before trusting it.
Two reporting errors of mine, corrected: I had quoted CEM at 62.5% success
(that was 16-episode noise; the matched 50-episode number is 34.0%) and a
"283x fewer evaluations" speedup computed across mismatched CEM configs
(`--cem-steps 5` vs the matrix default 30). Only rows inside this matrix are
comparable.
## Experiment 1b β the 88% anomaly, diagnosed
Executing the whole plan beat replanning every block, 88% vs 52%. That is
backwards from MPC theory, so it was treated as a suspected bug rather than a
result. Ruled out in order:
- **Success-metric artifact** β no. Success comes from genuine env
termination.
- **Noise** β no. Seeds 42/7/123 give rh=1: 52/64/42% vs rh=5: 88/90/86%.
- **Action-history normalization mismatch** β no. Live instrumentation of the
policy shows RH=1 `past|.|`=0.681, RH=5 0.433, TRAIN 0.747, zero-fraction
0.000 everywhere. The controller receives correctly z-scored history.
- **Insufficient budget / slow pacing** β no. rh=1 is flat at 50/48/50% for
budgets 50/100/200. Quadrupling the budget changes nothing, so this is a
fixed point, not slowness.
- **Backloaded plan** β no. Block 1 has the *largest* actions
(|a| = [0.703, 0.676, 0.585, 0.471, 0.484]) and does 42.9% of the total
distance reduction.
- **World-model hallucination** β no. An independent 30-episode harness
measuring in ground-truth env space reproduces it: rh=1 43.3%, rh=5 90.0%.
**Root cause β horizon-reset procrastination.** The controller reaches its
minimum predicted distance at block 5 *regardless of how near the goal
actually is*:
```
goal 1 block away (start 0.208): [0.090, 0.053, 0.035, 0.023, 0.014] min at block 5
goal 2 blocks away: ... min at block 5
goal 3 blocks away: ... min at block 5
goal 5 blocks away (start 1.09): [1.090, 0.664, 0.195, 0.039, 0.025] min at block 5
```
Even when the goal is **one block away**, it spreads the approach over all
five blocks. This follows directly from the objective: terminal loss is
applied at block H, so "be at the goal at step 5" is exactly what is
optimized, and nothing rewards arriving sooner and holding. Under receding
horizon the deadline resets to H after every replan, so the controller
executes one fifth of the way, replans, and again aims to arrive in five β
approaching asymptotically without landing inside the 20px tolerance.
This is not a violation of MPC theory. Receding-horizon MPC is only
well-behaved when the objective carries appropriate stage costs, a terminal
constraint, or a terminal value function. This objective has none, so the
learned planner is **horizon-dependent rather than time-consistent**: a good
fixed-horizon open-loop planner, not yet a closed-loop goal policy.
## Experiment 1c β execution-length sweep (the clean control)
Same checkpoint (K=3), same episodes, varying only how many blocks execute
before replanning.
| execute m blocks | success% | first-call distance | rows/ep |
|---|---|---|---|
| 1 | 52.0 | 0.0638 | 144 |
| 2 | 52.0 | 0.0638 | 75 |
| 3 | 72.0 | 0.0638 | 44 |
| 4 | 80.0 | 0.0638 | 31 |
| 5 | 88.0 | 0.0638 | 25 |
Success climbs monotonically with execution length while the **first-call
terminal distance is identical (0.0638) across every row**. That identity is
the control: every configuration emits the same opening plan from the same
state, so plan quality is held exactly constant and the only variable is how
much of the plan runs before the deadline resets. The full 36-point spread is
attributable to the execution schedule alone.
Note also that executing more is *cheaper* β 25 vs 144 rows/ep β because
replanning is what costs predictor evaluations.
## Experiment 1d β full-plan execution across every K
Success%, refinement depth K (rows) against execution length m (columns):
| K | m=1 | m=2 | m=3 | m=4 | m=5 |
|---|---|---|---|---|---|
| 0 | 36.0 | - | - | - | 66.0 |
| 1 | 50.0 | - | - | - | 86.0 |
| 2 | 52.0 | - | - | - | 82.0 |
| 3 | 52.0 | 52.0 | 72.0 | 80.0 | 88.0 |
| 5 | 42.0 | - | - | - | 90.0 |
Full-plan execution lifts **every** K substantially, so the deficit belongs
to the execution schedule and not to any particular refinement depth.
This overturns my earlier reading of the K=5 regression. K=5 is the *worst*
setting at m=1 (42%) and the *best* at m=5 (90%). I had attributed the drop
to the recurrent update failing to contract past its trained depth; the grid
shows extra refinement in fact produces a *better* five-block plan, which
receding-horizon execution then squanders. The regression was a symptom of
horizon reset, not of non-contractive refinement.
## Experiment 1e β CEM under both schedules (a correction to the headline)
The controller is not the only planner that procrastinates. CEM optimizes the
*same* fixed-terminal objective, so it should suffer the same horizon reset β
and it does:
| planner | exec | success% | rows/ep | sec/ep |
|---|---|---|---|---|
| CEM (300 x 30) | 1 | 34.0 | 383400 | 7.04 |
| CEM (300 x 30) | 5 | **90.0** | 55800 | 1.98 |
| controller K=3 | 1 | 52.0 | 144 | 0.18 |
| controller K=3 | 5 | 88.0 | 25 | 0.13 |
| controller K=1 | 5 | 86.0 | 12 | 0.25 |
| controller K=5 | 5 | 90.0 | 37 | 0.21 |
CEM goes 34% -> 90%. **Horizon-reset procrastination is a property of the
terminal-only objective, not of the learned controller.** Any planner
minimizing `d(z_H, z_G)` under receding-horizon execution inherits it. This
also matches how LeWM itself executes plans β the full optimized sequence
before replanning β so the earlier rh=1 CEM row was not the configuration the
paper uses.
**This corrects the headline claim.** "Controller 88% vs CEM 34%" compared
against a baseline crippled by a known flaw. At *matched* execution:
> controller **88.0%** vs CEM **90.0%** β statistically indistinguishable β
> at **2250x** fewer predictor rows and **15x** less wall-clock time.
That is a weaker-sounding but far more defensible result, and it is the claim
the design doc actually set out to test: comparable success at a fraction of
the planning compute. Reporting the 34% row as the baseline would have been
comparing against a knowingly time-inconsistent configuration.
Two further notes:
- CEM is *cheaper* at exec5 (55800 vs 383400 rows/ep) for the same reason the
controller is: replanning is what costs evaluations.
- controller K=1 at exec5 reaches 86% for **12 rows/ep** β 4650x fewer than
CEM. Given K=1/K=2/K=5 are statistically tied at exec5, K=1 is the
Pareto-optimal configuration, not K=3.
With 50 episodes, one flipped episode is 2 points. All rows share start/goal
pairs, so comparisons are **paired**: exact McNemar on discordant episodes,
plus a paired bootstrap CI (`scripts/paired_stats.py`).
| comparison | diff | 95% CI | p |
|---|---|---|---|
| K=3, m=3 vs m=2 | +20.0 | [+8.0, +32.0] | 0.0063 * |
| K=3, m=4 vs m=2 | +28.0 | [+14.0, +42.0] | 0.0005 * |
| K=1 vs K=0 (both m=5) | +20.0 | [+10.0, +32.0] | 0.0020 * |
| K=5 vs K=0 (both m=5) | +24.0 | [+12.0, +36.0] | 0.0005 * |
| K=5 vs K=1 (both m=5) | +4.0 | [-4.0, +12.0] | 0.6250 |
| K=2 vs K=1 (both m=5) | -4.0 | [-12.0, +4.0] | 0.6250 |
Execution length and the first refinement are significant. **Refinement
beyond K=1 is not measurable at m=5** (p = 0.625) β the raw percentages
(86 / 82 / 90) would overstate a difference the paired test cannot support.
Reporting K=5 as "best" on 90% alone would be reading noise.
`eval_controller.py` now stores `episode_successes` per row so every future
comparison is paired.
## Diagnostics β quantifying the fixed point and the refinement drift
`scripts/diagnostics.py`, 256 held-out samples, simulated in latent space
(the world model is its own simulator, isolating planner dynamics from
simulator mismatch).
### 1. The closed loop has a fixed point outside the success radius
Fitting `D_{n+1} = c*D_n + b` over consecutive replans, exec1:
```
c = 0.5792 b = 0.04116 D* = b/(1-c) = 0.0978 R^2 = 0.784
mean trace: 0.491 0.282 0.159 0.111 0.103 0.107 0.114 0.120 0.126 0.133 0.141 0.150
```
Each replan removes ~42% of the *remaining* distance but adds a constant
0.041 floor. The trace bottoms out at 0.103 by replan 5 and then creeps back
**upward**. `0 < c < 1` with `b > 0` gives a stable fixed point at
**D\* = 0.098**, which is outside the success tolerance.
This is the quantitative explanation for the budget experiment: rh=1 was flat
at 50/48/50% for budgets 50/100/200 because the loop converges to 0.098 and
stays there. More time cannot help a system that has already converged to the
wrong place.
(The exec5 fit is not meaningful β 12 replans x 5 blocks massively overshoots
the episode, so its trace rises as the agent sails past the goal.)
### 2. Refinement genuinely does not contract past its trained depth
| k | terminal | arrival | dJ | mean plan change |
|---|---|---|---|---|
| 0 | 0.05809 | 0.20163 | | 0.00000 |
| 1 | 0.02641 | 0.16564 | +0.03168 | 0.23859 |
| 2 | 0.01711 | 0.15428 | +0.00930 | 0.13259 |
| 3 | 0.01486 | 0.15003 | +0.00225 | 0.06835 |
| 4 | 0.01376 | 0.14700 | +0.00110 | 0.05815 |
| 5 | 0.01360 | 0.14489 | +0.00016 | 0.05123 |
| 6 | 0.01373 | 0.14339 | **-0.00013** | 0.04626 |
| 7 | 0.01403 | 0.14236 | **-0.00030** | 0.04290 |
| 8 | 0.01445 | 0.14169 | **-0.00042** | 0.04070 |
Predicted cost improves through k=5 and then **worsens** from k=6 onward,
while the plan keeps changing by ~0.041 per iteration and never settles. So
the refinement operator really is non-contractive outside its trained depth β
it does not converge, it drifts.
This refines the Experiment 1d conclusion rather than replacing it. Both
effects are real and separable: the m=1 K=5 regression (42%) is dominated by
horizon reset, since K=5 is best at m=5; the k>5 cost increase measured here
is a genuine property of the recurrent update. Note also that `arrival`
(distance at the sample's own goal offset q) stays ~10x worse than `terminal`
at every k β the controller optimizes exactly what it was asked to.
With 50 episodes, one flipped episode is 2 points. All rows share start/goal
pairs, so comparisons are **paired**: exact McNemar on discordant episodes,
plus a paired bootstrap CI (`scripts/paired_stats.py`).
| comparison | diff | 95% CI | p |
|---|---|---|---|
| K=3, m=3 vs m=2 | +20.0 | [+8.0, +32.0] | 0.0063 * |
| K=3, m=4 vs m=2 | +28.0 | [+14.0, +42.0] | 0.0005 * |
| K=1 vs K=0 (both m=5) | +20.0 | [+10.0, +32.0] | 0.0020 * |
| K=5 vs K=0 (both m=5) | +24.0 | [+12.0, +36.0] | 0.0005 * |
| K=5 vs K=1 (both m=5) | +4.0 | [-4.0, +12.0] | 0.6250 |
| K=2 vs K=1 (both m=5) | -4.0 | [-12.0, +4.0] | 0.6250 |
Execution length and the first refinement are significant. **Refinement
beyond K=1 is not measurable at m=5** (p = 0.625) β the raw percentages
(86 / 82 / 90) would overstate a difference the paired test cannot support.
Reporting K=5 as "best" on 90% alone would be reading noise.
`eval_controller.py` now stores `episode_successes` per row so every future
comparison is paired.
## Speedup, stated honestly
Two different numbers, both real, measuring different things. Both use the
matched-execution CEM baseline (exec5, 55800 rows/ep, 1.98 s/ep):
| comparison | predictor rows | wall clock |
|---|---|---|
| CEM vs controller K=3 (both m=5) | 55800 / 25 = **2250x** | 1.98 / 0.13 = **15x** |
| CEM vs controller K=1 (both m=5) | 55800 / 12 = **4650x** | 1.98 / 0.25 = **8x** |
The row ratio is far larger than the wall-clock ratio because CEM evaluates
its 300 samples in parallel batches. Predictor rows are the honest measure of
*model invocations*; wall-clock is the honest measure of *time*. Quoting the
row ratio as though it were a speed number would be misleading.
> **Amended 2026-08-04.** Rows *per episode* also carries a survivorship
> confound: episodes terminate on success, so a better planner runs shorter
> episodes and accumulates fewer rows. The two rows above compare planners at
> nearly equal success (88 vs 90), so they are close to fair β but any
> row-ratio between planners of *differing* success is inflated. Divide by
> `predictor_calls` instead. See
> [the confound section](#a-confound-in-the-efficiency-claim-episodes-terminate-on-success).
### Metric caveat
`mean_terminal_distance` averages over solver calls, and different execution
lengths make different numbers of calls at different distances from the goal
(rh=1 makes 10, rh=5 makes 2, both of the latter while still far away). It is
therefore **not comparable across execution lengths**, which is why exec5
shows a worse mean terminal (0.249) despite far higher success. It remains
valid within the K sweep, where every row shares rh=1. `eval_controller.py`
now also records `first_terminal_distance`, which *is* comparable because
every planner's first call is taken from the same held-out state.
With that fixed metric, plan quality tracks success cleanly across K at m=5 β
`first_d` = 0.233 (K=0), 0.060 (K=1), 0.046 (K=2), 0.028 (K=5), against CEM's
0.026. Refinement really is improving the plan; the old metric was hiding it.
---
# Part 2 β the corrected objective
## The fix: horizon-matched arrival and hold
Every offline sample already knows how far ahead its goal was relabeled from
(`goal_offset`, q in 1..H) β the old objective simply discarded it. Replace
the fixed terminal term with
```
J_i = d_{i,q_i} + lambda_hold * mean_{j>q_i} d_{i,j}
```
- **arrival** `d_q`: reach the goal by the deadline the data says is
reachable, rather than always at block H;
- **hold**: stay there afterwards, so the controller cannot touch the goal
and drift off.
The path loss `alpha * sum w_j d_j` is dropped, since its `(j/H)^2` weights
lean *later* and mildly reinforce the behavior being removed. Deep
supervision across refinements (`rho_k = 2^k`) and the support penalty are
unchanged. Implemented as `arrival_hold_loss` in `lejepa_control/losses.py`,
selected by `--arrival-hold` (the old path stays reachable so the original
ablations remain reproducible).
`q` is used **only to index the loss** β it is never fed to the controller.
Conditioning on a deadline that eval always resets to 5 would reintroduce the
exact failure being fixed.
### Verified before spending GPU time on it
Unit test on a hand-computed case (`d = [0.9, 0.7, 0.5, 0.3, 0.1]`):
q=1 -> 0.9 + 0.5*0.4 = 1.10; q=3 -> 0.5 + 0.5*0.2 = 0.60; q=5 -> 0.10. All
match.
The sharper check contrasts a deferring plan against a prompt one for a
one-block goal:
| plan | profile | old loss (d_H) | new loss (q=1) |
|---|---|---|---|
| defers | [0.090, 0.053, 0.035, 0.023, 0.014] | 0.0140 | 0.1056 |
| arrives and holds | [0.015, 0.014, 0.014, 0.015, 0.014] | 0.0140 | 0.0221 |
The old objective is **exactly indifferent** between them β identical loss to
four decimals β while the new one prefers prompt arrival by 4.8x. That is
the failure mechanism isolated in a single table.
### One assumption checked and discarded
I had planned to force uniform sampling over q, on the theory that
long-offset samples might dominate and re-teach the terminal-at-H bias.
Measuring first (4000 samples, max_offset=5) showed the opposite skew:
q=1 25.9%, q=2 21.1%, q=3 18.9%, q=4 17.4%, q=5 16.7% β short goals are
already the most common. The balancing knob was removed rather than shipped
as dead complexity.
## Experiment 2 + 3 β status: running (superseded)
> Superseded by the final sections below. Kept for the interim reasoning and
> the predictions it made, several of which turned out wrong and are corrected
> in place.
Three 20000-step runs at `lambda_hold` in {0.5, 0.0, 1.0}, each evaluated at
m=1 and m=5. `--hold-weight 0` isolates deadline-only from deadline-and-hold.
### Interim: the timing behavior is already fixed (step 7500, lambda_hold=0.5)
Held-out mean distance profile over the plan, split by the sample's own goal
offset q. The arrow marks where each row's minimum falls:
| q | block 1 | 2 | 3 | 4 | 5 | argmin |
|---|---|---|---|---|---|---|
| 1 | **0.0085** | 0.0104 | 0.0149 | 0.0200 | 0.0268 | block 1 β |
| 2 | 0.0828 | **0.0121** | 0.0139 | 0.0183 | 0.0241 | block 2 β |
| 3 | 0.2883 | 0.0271 | **0.0158** | 0.0191 | 0.0241 | block 3 β |
| 4 | 0.6217 | 0.1235 | 0.0413 | **0.0401** | 0.0446 | block 4 β |
| 5 | 0.8954 | 0.2825 | 0.1201 | **0.1117** | 0.1154 | block 4 (~5) |
Against the original controller, which bottomed out at block 5 for *every* q:
| q | original profile | argmin |
|---|---|---|
| 1 | [0.090, 0.053, 0.035, 0.023, **0.014**] | block 5 β |
| 5 | [1.090, 0.664, 0.195, 0.039, **0.025**] | block 5 β |
The corrected controller arrives **at its deadline** for every offset. For
q=1 the profile now *increases* after block 1 (0.0085 -> 0.0268): it arrives
immediately and holds position, exactly the intended behavior and the exact
inverse of the old one. The q=5 row bottoming at block 4 rather than 5 is
benign β arriving one block early and holding is what the hold term rewards.
Training curve: arrival distance 0.140 (step 500) -> 0.021 (1500) -> 0.0133
(8500). Learned step sizes still decay across refinements
(`[0.548, 0.459, 0.345]`), so refinement remains a converging solver.
### Interim: the closed-loop fixed point has already halved
Same contraction fit as the diagnostics section, run on the step-7500
checkpoint (exec1):
| controller | c | b | **D\*** | R^2 |
|---|---|---|---|---|
| original (20000 steps) | 0.5792 | 0.0412 | **0.0978** | 0.784 |
| arrival+hold (7500 steps) | 0.3471 | 0.0320 | **0.0490** | 0.531 |
Each replan now removes **65%** of the remaining distance instead of 42%
(c: 0.579 -> 0.347), and the floor drops from 0.041 to 0.032. The fixed point
halves at only 38% of the training budget. This is the mechanism the fix
targeted, moving in the predicted direction, measured independently of
success rate.
The refinement drift past k=5 persists (cost worsens from k=6, plan changes
plateau at ~0.05) β as expected, since the fix targets *when* the plan
arrives, not the contractivity of the recurrent update. That remains a
separate open issue.
### Interim result: the gap has collapsed, and inverted
Sim evaluation of the step-7500 checkpoint on the same 50 held-out episodes:
| controller | m=1 | m=5 | gap | rows/ep at m=1 |
|---|---|---|---|---|
| original (20000 steps) | 52.0 | 88.0 | **-36.0** (backwards) | 144 |
| arrival+hold (7500 steps) | **90.0** | 84.0 | **+6.0** (correct) | 71 |
Receding-horizon execution now *beats* open-loop, which is the ordering MPC
theory predicts and the original controller violated. Paired tests:
| comparison | diff | 95% CI | p |
|---|---|---|---|
| corrected m=1 vs corrected m=5 | +6.0 | [-2.0, +14.0] | 0.3750 |
| corrected m=1 vs CEM exec5 | +0.0 | [-12.0, +12.0] | 1.0000 |
| corrected m=1 vs original K=5 exec5 | +0.0 | [-10.0, +10.0] | 1.0000 |
**The success criterion is met.** `success(m=1) ~= success(m=5)`: the two are
now statistically indistinguishable (p = 0.375), against a -36 point gap
before. The +6 nominal advantage for closed-loop should not be over-read at
this sample size β the honest statement is that the schedule no longer
matters, which is exactly what time-consistency means.
Meanwhile the corrected controller at m=1 **equals CEM** (90.0 vs 90.0,
p = 1.0) at **784x fewer predictor rows** (71 vs 55800), and equals the best
original configuration while replanning every block instead of committing to
a five-block plan.
Note the higher `first_terminal_distance` (0.140 vs 0.064). That is expected
and not a regression: the corrected controller is no longer optimizing
distance at block 5, so a metric read at block 5 necessarily looks worse. It
is arriving at block q and holding, and the sim success rate is the arbiter.
This is a mid-training checkpoint at 38% of the budget, evaluated early
because the per-q profiles and contraction fit both indicated the mechanism
had already changed. The full 20000-step runs and the lambda_hold ablation
are still in flight.
**Success criterion:** `success(m=1) ~= success(m=5)`. The corrected
controller does not have to beat 90%; the 52-vs-88 gap has to collapse. That
is what would show the planner has become time-consistent rather than merely
a good fixed-horizon planner.
## Experiment 2 β final result (lambda_hold = 0.5, 20000 steps)
The full run confirms and improves on the interim checkpoint.
| controller | m=1 | m=5 | gap | rows/ep at m=1 |
|---|---|---|---|---|
| original | 52.0 | 88.0 | **-36.0** (backwards) | 144 |
| arrival+hold, step 7500 | 90.0 | 84.0 | +6.0 | 71 |
| **arrival+hold, step 20000** | **94.0** | 88.0 | **+6.0** (correct) | 73 |
Paired tests on the same 50 held-out episodes:
| comparison | diff | 95% CI | p |
|---|---|---|---|
| corrected m=1 vs corrected m=5 | +6.0 | [+0.0, +14.0] | 0.2500 |
| corrected m=1 vs CEM exec5 | +4.0 | [-6.0, +14.0] | 0.6875 |
| corrected m=1 vs original K=3 exec4 | +14.0 | [+4.0, +26.0] | **0.0391*** |
| corrected m=5 vs CEM exec5 | -2.0 | [-10.0, +6.0] | 1.0000 |
**The success criterion is met at full training.** The execution schedule no
longer decides the outcome (p = 0.25, against a -36 point gap before), which
is what time-consistency means. The +6 for closed-loop is the ordering MPC
theory predicts; at n=50 it should not be over-read as a win, only as the
absence of the pathology.
The headline, stated at matched execution and with cost as measured:
| planner | success | rows/ep | sec/ep |
|---|---|---|---|
| CEM, exec5 | 90.0 | 55800 | 1.98 |
| CEM, exec1 | 34.0 | 383400 | 7.04 |
| corrected controller, m=1 | **94.0** | **73** | **0.25** |
| corrected controller, m=5 | 88.0 | 23 | 0.18 |
At m=1 the corrected controller is nominally **+4 points over CEM's best
schedule at 764x fewer predictor rows and 7.9x faster wall-clock**. The
success difference is not significant (p = 0.6875) β the honest claim is
**parity with CEM at three orders of magnitude less compute**, now without
needing the open-loop schedule that was silently doing the work before.
> **Amended 2026-08-04.** The 764x is a *per-episode* row ratio and is
> partly earned by finishing sooner (episodes terminate on success). The
> per-decision ratio is **16.4x fewer rows per solver call**. Both are real;
> they answer different questions. The wall-clock 7.9x and the success
> parity are unaffected. See
> [the confound section](#a-confound-in-the-efficiency-claim-episodes-terminate-on-success).
Note this also beats the *original* controller's best-ever configuration
(88.0 at m=5) while replanning every block, and does so significantly
(p = 0.039 vs K=3 exec4). The fix did not merely remove a confound.
`first_terminal_distance` is 0.0423 for both corrected rows, identical across
m=1 and m=5 exactly as Experiment 1c found β plan quality constant, only the
schedule differs. It is also *lower* than the original controller's 0.0638.
That the corrected controller improves terminal distance and success
simultaneously rules out the reading that arrival+hold merely trades block-5
accuracy for better timing. The interim checkpoint's higher 0.140 was a
mid-training artifact, not the steady state.
**Replicated three times, exactly.** Three independent re-runs of both rows
returned 94.0 / 94.0 / 94.0 and 88.0 / 88.0 / 88.0, with **zero of 50
episodes differing** in outcome on any pair. This is not determinism β the
underlying cost traces do differ between runs (mean terminal 0.9947 / 0.9947
/ 0.9155). It means no episode sits close enough to the success threshold for
rollout jitter to flip it. By contrast the original controller flips 1 of 50
and CEM at exec1 flips 6. The headline is a stable measurement, not a lucky
draw.
## Experiment 4 β objective ablations (complete)
The two ablations that section 14 lists last, both on the **original**
fixed-terminal objective so they isolate the two auxiliary terms rather than
the timing fix:
| run | change | question it answers | answer |
| --- | --- | --- | --- |
| `abl_terminal_only` | `--alpha 0` | does the late-weighted path loss contribute anything, or is the terminal term doing all the work? | it was a partial mitigation; removing it **doubles** the pathology (18.0 at m=1) |
| `abl_no_support` | `--lambda-support 0` | is the support penalty holding plans inside the action manifold, or just costing capacity? | it holds the manifold (violations 0.187 -> 0.652) but that barely affects success here |
Each is 20000 steps plus evals at m=1 and m=5, so both schedules are covered
and the rows drop straight into the K x execution grid. Full results in the
two sections below.
These were the runs killed by the orphaned-driver `0xC0000142` cascade
earlier. Relaunched as a **parallel** stream alongside the arrival-hold
queue: one training run leaves the GPU at ~22-34% util and 3 of 16 GB, so
serializing would have wasted hours for nothing. Measured after both were
resident: **1.7 it/s each, unchanged from either running alone**, at 81%
util and 6 GB. The bottleneck is per-step latency, not GPU throughput.
**But the real constraint is host RAM as well as VRAM, and the safe limit is
two GPU processes.** Adding a third (a diagnostics run) raised a
`MemoryError` inside the world-model load β 2.3 GB free of 16.5, because each
process holds its own copy of the ViT + predictor + latent cache β and it
also took VRAM on the way down, which OOM-killed the `lambda_hold=0.0`
training at step 6000. See the incident entry below. Analysis that only reads
jsonl (`paired_stats.py`, `summarize_results.py`) loads nothing and is always
safe to run alongside.
### Live evidence of the bug, from the ablation's own training log
`abl_terminal_only` trains the original fixed-terminal objective, and
`train_controller.py` now logs the `arrival` metric even when it is not being
optimized. That makes the pathology visible during training rather than only
at eval:
| run | step | terminal | arrival |
|---|---|---|---|
| `abl_terminal_only` | 4000 | 0.0187 | 0.1775 |
| `abl_terminal_only` | 10000 | 0.0153 | 0.2596 |
| `abl_terminal_only` | 16500 | 0.0148 | 0.1976 |
| `ah_hold0.5` | 20000 | 0.0226 | **0.0156** |
Under the original objective, `terminal` falls monotonically while `arrival`
*rises* β it gets better at being close at block 5 and worse at being close
at the block the goal was actually relabeled from. Under arrival+hold,
`arrival` (0.0156) sits well *below* `terminal` (0.0226): it arrives early
and drifts slightly afterwards, which is the intended behavior.
Horizon-reset procrastination is therefore observable as a divergence between
two training-time scalars, with no rollout needed. Worth logging `arrival`
permanently as a cheap early-warning metric.
### Incident: lambda_hold=0.0 died at step 6000 (CUDA OOM), retry queued
**Update: it killed two runs, not one.** `abl_terminal_only` also died, at
step 16500, with `RuntimeError: CUDA error: CUBLAS_STATUS_INTERNAL_ERROR` β
timestamped `08:56:05`, the same instant as the `lambda_hold=0.0` OOM and the
diagnostics launch. A single mistaken third process destroyed ~5.5 GPU-hours
across both parallel streams. The sweep drivers reported
`!! train failed ...` and moved on to their next stage, so both losses stayed
invisible until the `===` markers were compared against expected stage counts.
That is the real lesson: a `|| echo "!! failed"` guard keeps a sweep alive but
converts a crash into a silent gap. Check marker counts against what the
script *should* have emitted, not just the tail.
The `lambda_hold=0.0` training crashed with
`torch.AcceleratorError: CUDA error: out of memory` at step 6000, and the
sweep driver moved straight on to `lambda_hold=1.0`.
**Cause was self-inflicted.** After measuring that two trainings coexist
happily, the previous entry concluded the binding constraint was host RAM
rather than VRAM β and then a third process (a diagnostics run) was started
on that basis. That process raised a host `MemoryError` on its own load, but
it had *already* taken VRAM, and the training that was mid-step lost its
allocation. The correct statement is that **both** limits bind, and the safe
concurrency on this box is **two GPU processes, full stop**:
| processes | outcome |
|---|---|
| 2 trainings | fine, 1.7-1.8 it/s each, ~6 GB VRAM, ~80% util |
| 3 (2 trainings + diagnostics) | host `MemoryError` *and* CUDA OOM killing a training |
The earlier "the bottleneck is per-step latency, not GPU throughput" line was
measured correctly but generalized too far: it described *throughput* under
two processes and said nothing about *headroom* for a third.
Recovery is a strictly serial `finisher.sh` that waits for both the
`EXP 2+3 DONE` and `EXP 4 DONE` sentinels before touching the GPU, then runs
one process at a time: retrain `lambda_hold=0.0` + its two evals, the three
`diagnostics.py` runs, and the paired-outcome backfill. Serial is slower but
these are the last jobs, so there is nothing left to overlap with.
`backfill.sh` also failed, separately: its 30-second polling loop spawned a
subshell per iteration and eventually hit cygwin `fork: Resource temporarily
unavailable` / `0xC0000142`. The finisher polls at 60s and uses no subshell
per check. Nothing was lost β backfill had not yet run any evals.
**No result is affected.** The completed `lambda_hold=0.5` run and every
Experiment 1 row predate this and are unchanged.
### Instrumentation bug: `viol` is unmeasured, not zero, under `--lambda-support 0`
The `abl_no_support` run logs `viol 0.000` exactly, at every step, against
~0.17-0.24 in every other run. That looks like a headline result β "removing
the support penalty eliminates violations" β and it is not one.
`train_controller.py:153` skips loading the density model when
`--lambda-support 0`, so `density is None`, the `support_loss` call at line
224 never runs, and `violation` keeps the `torch.zeros(())` initializer from
line 219. The logged 0.000 is that initializer, not a measurement.
This is exactly the ablation where the metric matters most: the question
`abl_no_support` is meant to answer is *"is the support penalty holding plans
inside the action manifold?"*, and the number that would answer it is the one
silently disabled. Reporting it as-is would have inverted the conclusion.
**Fix deferred, deliberately.** Correcting this means loading the density
model for measurement while excluding it from the loss β a real change to a
script that two trainings are currently executing. Editing it mid-run risks
the in-flight jobs for a metric that can be recovered afterwards: the
checkpoint is saved either way, so violation can be evaluated post-hoc from
`data/runs/abl_no_support/controller.pt` against the same density model. That
is the plan once the GPU frees.
Until then the `abl_no_support` violation column should be read as **absent**,
not zero. The `support` column is likewise not printed for that run.
General lesson: a metric that is computed conditionally on the thing being
ablated will read as a perfect score for the ablation. Prefer initializing
such metrics to `nan` over `0` so the gap is visibly missing rather than
plausibly excellent.
> **Resolved.** Recovered post-hoc from the saved checkpoints β the true
> violation fraction is **0.652**, the worst in the matrix, against the
> original's 0.187. The metric that logged as a perfect 0.000 was in fact
> the worst result in the experiment. Full table in
> [the post-hoc section](#the-support-penalty-measured-post-hoc).
## Experiment 3 β the lambda_hold ablation (2 of 3 complete, superseded)
> Superseded by ["Experiment 3 β the lambda_hold ablation
> (complete)"](#experiment-3--the-lambdahold-ablation-complete) below.
| lambda_hold | m=1 | m=5 | gap | p(m1 vs m5) | first_d |
|---|---|---|---|---|---|
| 0.0 (arrival only) | *retraining* | *retraining* | β | β | β |
| **0.5** | **94.0** | 88.0 | +6.0 | 0.2500 | 0.0423 |
| 1.0 | 92.0 | 88.0 | +4.0 | 0.6250 | 0.0277 |
*(Superseded by the completed table in Experiment 3 below.)*
Both completed settings fix the pathology. Neither is distinguishable from
the other:
| comparison | diff | 95% CI | p |
|---|---|---|---|
| hold0.5 m=1 vs hold1.0 m=1 | +2.0 | [-6.0, +10.0] | 1.0000 |
| hold0.5 m=5 vs hold1.0 m=5 | +0.0 | [-8.0, +8.0] | 1.0000 |
So `lambda_hold` is **not a sensitive knob** across 0.5-1.0: doubling it moves
success by 0-2 points, well inside noise at n=50. What matters is that the
arrival term is indexed by q at all. The `lambda_hold = 0.0` cell will decide
whether the hold term contributes anything beyond the re-indexing, and it is
the one cell still missing β see the incident below.
One suggestive difference that is *not* a success difference: `lambda_hold=1.0`
reaches a markedly lower `first_terminal_distance` (0.0277 vs 0.0423) while
scoring 2 points *worse*. A heavier hold term does keep the state nearer the
goal at block 5, and that extra proximity buys nothing in success. Consistent
with the metric caveat recorded earlier β terminal distance and task success
come apart, and success is the arbiter.
## Experiment 4 β `abl_no_support`: the support penalty is load-bearing
| run | m=1 | m=5 | gap | first_d |
|---|---|---|---|---|
| original (with support) | 52.0 | 88.0 | -36.0 | 0.0638 |
| **`abl_no_support`** | **50.0** | 90.0 | **-40.0** | 0.0502 |
| `ah_hold0.5` (fix, with support) | 94.0 | 88.0 | +6.0 | 0.0423 |
Removing the support penalty leaves the fixed-terminal objective's behavior
essentially unchanged: 50.0 vs 52.0 at m=1 is indistinguishable from the
original (`+2.0`, p = 1.0000), and the horizon-reset gap is, if anything,
slightly wider. The penalty is not what was causing the procrastination β a
useful negative, since it was a live candidate before Experiment 1c isolated
the execution schedule.
At m=5 it reaches 90.0, matching CEM and the original's best. So the support
penalty is not carrying the open-loop result either.
The comparison that matters is against the *fix*: `ah_hold0.5` beats
`abl_no_support` by **+44.0 points at m=1** (p = 0.0000), which is simply the
Experiment 2 result restated β the timing fix is doing the work, and no
amount of removing or keeping auxiliary terms substitutes for it.
**Caveat on this row:** its `viol` column is unmeasured, not zero (see the
instrumentation bug below), so this entry establishes the *success* effect of
dropping the penalty and says nothing yet about whether plans left the action
manifold. The violation figure will be recovered post-hoc from the saved
checkpoint.
> **Resolved.** Measured post-hoc: **0.652** of proposed blocks violate the
> support threshold, against 0.187 for the original β a 3.5x increase, and a
> mean NLL/dim (1.732) above the `c95` threshold (1.531) itself. So plans
> *did* leave the manifold, substantially, while success stayed identical to
> the original (p = 1.0). The penalty is load-bearing for manifold adherence
> and nearly free in success terms on this task. Full table in
> [the post-hoc section](#the-support-penalty-measured-post-hoc).
### Interim: what the hold term actually does (lambda_hold=0.0, step 10000)
The `lambda_hold=0.0` cell isolates the re-indexing from the hold term: the
arrival term is still indexed by q, but nothing penalizes leaving afterwards.
Its validation profiles show **touch-and-leave**, at matched training step:
| q | lambda=0.0: d(arrival) -> d(block 5) | drift | lambda=0.5 | drift |
|---|---|---|---|---|
| 1 | 0.0069 -> 0.0355 | **5.1x** | 0.0086 -> 0.0264 | 3.1x |
| 2 | 0.0125 -> 0.0376 | **3.0x** | 0.0148 -> 0.0276 | 1.9x |
| 3 | 0.0232 -> 0.0536 | **2.3x** | 0.0356 -> 0.0430 | 1.2x |
| 4 | 0.0427 -> 0.0793 | **1.9x** | 0.1348 -> 0.1402 | 1.0x |
Two things are visible and they separate cleanly:
1. **The re-indexing alone fixes the timing.** Without any hold term,
`lambda=0.0` still reaches its *minimum* at or before block q β for q=3 the
minimum is 0.0232 at block 3 exactly. The procrastination is gone. That is
the arrival term doing its job, and it confirms the deadline-vs-state
diagnosis rather than the hold term being what mattered.
2. **The hold term is what keeps it there.** Without it the state drifts away
after arrival by 1.9-5.1x; with `lambda=0.5` the drift is 1.0-3.1x, and at
q=4 the profile is essentially flat after arrival (1.04x).
Note `lambda=0.0` reaches a *lower absolute* arrival distance (0.0427 vs
0.1348 at q=4) β unsurprising, since it spends no capacity on staying. It
gets closer and then leaves.
This is section 2 of the design spec confirmed empirically: touch-and-leave
was listed as the first reason to reject min-over-blocks, and here it is,
produced by removing exactly the term that was added to prevent it. Whether
it costs *success* is the open question β the drift may or may not be enough
to exit the 20px tolerance, and the eval will decide. Recording the mechanism
now because it is visible in the profiles regardless of how the success
number lands.
> **Resolved: it does not cost success.** `lambda_hold=0.0` finished at
> 90.0/92.0, statistically identical to `lambda_hold=0.5` (p = 0.625). The
> drift is real but stays inside the 20px tolerance. See the completed
> Experiment 3 table below.
## Experiment 3 β the lambda_hold ablation (complete)
| lambda_hold | m=1 | m=5 | gap | p(m1 vs m5) | first_d | D* (exec1) |
|---|---|---|---|---|---|---|
| **0.0** (arrival only) | 90.0 | 92.0 | -2.0 | 1.0000 | 0.0360 | 0.0475 |
| **0.5** | **94.0** | 88.0 | +6.0 | 0.2500 | 0.0423 | 0.0397 |
| **1.0** | 92.0 | 88.0 | +4.0 | 0.6250 | 0.0277 | 0.0362 |
| *original* | *52.0* | *88.0* | *-36.0* | *0.0001* | *0.0638* | *0.0978* |
All three settings fix the pathology completely, and **none is
distinguishable from any other**:
| comparison | diff | 95% CI | p |
|---|---|---|---|
| hold0.5 vs hold0.0 (m=1) | +4.0 | [-4.0, +12.0] | 0.6250 |
| hold0.5 vs hold1.0 (m=1) | +2.0 | [-6.0, +10.0] | 1.0000 |
| hold0.5 vs hold1.0 (m=5) | +0.0 | [-8.0, +8.0] | 1.0000 |
| hold0.0 m=1 vs m=5 | -2.0 | [-8.0, +4.0] | 1.0000 |
### Correcting the prediction made from the interim profiles
The previous entry read touch-and-leave off the `lambda_hold=0.0` profiles
and asked whether it would cost success. **It does not.** `lambda_hold=0.0`
scores 90.0/92.0 β statistically identical to `lambda_hold=0.5`'s 94.0/88.0
(p = 0.625), and its m=1-vs-m=5 gap is the *smallest* of the three (-2.0,
p = 1.0).
So the honest conclusion is narrower than the design spec anticipated:
- **The re-indexing is the entire fix.** Indexing the arrival term by the
relabeled offset q eliminates horizon-reset procrastination on its own.
- **The hold term is not load-bearing for success.** It measurably reduces
post-arrival drift (1.9-5.1x -> 1.0-3.1x at step 10000, and the final
profiles still show it), but that drift stays inside the 20px tolerance, so
it never converts into a failed episode.
The drift is real and the hold term does what it was designed to do. It just
does not matter at this tolerance on this task. A tighter success radius, a
longer execution horizon, or a task where the object keeps moving after
contact could all change that β but on PushT at 20px, `lambda_hold=0` is
sufficient, and the simpler objective is the better default.
This is worth stating plainly because the spec's section 2 rejected
min-over-blocks partly *because* of touch-and-leave, and predicted the hold
term would be necessary. Touch-and-leave was correctly predicted and
correctly observed; its consequence was over-estimated.
### Contraction fixed points, all variants
| variant | c | b | **D\*** | R^2 |
|---|---|---|---|---|
| original | 0.5792 | 0.0412 | **0.0978** | 0.784 |
| ah_hold0.5 @ step 7500 | 0.3471 | 0.0320 | 0.0490 | 0.531 |
| ah_hold0.0 | 0.4076 | 0.0282 | 0.0475 | 0.603 |
| ah_hold0.5 | 0.3919 | 0.0241 | **0.0397** | 0.616 |
| ah_hold1.0 | 0.4079 | 0.0214 | **0.0362** | 0.625 |
Every corrected variant more than halves the closed-loop fixed point, from
0.0978 to 0.036-0.048. Both terms of the recursion improve: the contraction
factor `c` drops (0.58 -> ~0.40, so each replan removes more of the remaining
distance) and the floor `b` drops (0.041 -> 0.021-0.028). The original's
stall point sat outside the success radius; the corrected ones sit inside it,
which is the mechanism behind the m=1 success jump.
`lambda_hold` orders the fixed point monotonically (0.0475 -> 0.0397 ->
0.0362 for 0.0 -> 0.5 -> 1.0), so a heavier hold term does measurably tighten
the closed loop β it just does not move success, since all three are already
inside tolerance. That is the same terminal-distance-vs-success divergence
recorded throughout.
**The exec5 rows are not interpretable and should be ignored.** All variants
fit `c > 1` there (1.07-1.11), giving a negative "fixed point" β the fit is
extrapolating a divergent recursion, which means five-block open-loop
execution simply is not a contraction in this sense. The exec1 column is the
one that answers the design question.
## The backfill became an accidental replication study
The paired-outcome backfill re-ran the original Experiment 1a configurations
with the **same seed (42) and the same episode set**. It should have
reproduced the earlier numbers exactly. It did not.
> **Corrected 2026-08-04.** The first version of this section reported a flat
> "+/-4 points, 7-8 of 50 episodes flip". Both halves were wrong, in
> compensating ways. See the correction below the table.
Restricting to rows that carry per-episode outcomes (i.e. rows produced by
the *same* code β the seven pre-instrumentation rows lack
`episode_successes` and cannot be paired against anything):
| config | runs (same code) | spread | episodes flipped |
|---|---|---|---|
| `controller_K0` exec1 | 36.0, 36.0 | 0 | 0/50 |
| `controller_K2` exec1 | 52.0, 52.0 | 0 | 0/50 |
| `controller_K3` exec5 | 88.0, 88.0 | 0 | 0/50 |
| `controller_K1` exec1 | 48.0, 50.0 | 2 | 1/50 |
| `controller_K3` exec1 | 48.0, 50.0 | 2 | 1/50 |
| `controller_K5` exec1 | 44.0, 40.0 | 4 | 2/50 |
| `cem_s300_n30` exec1 | 32.0, 36.0 | 4 | **6/50** |
| `controller_K3` exec1 **[ah_hold0.5]** | 94.0, 94.0, 94.0 | **0** | **0/50** |
| `controller_K3` exec5 **[ah_hold0.5]** | 88.0, 88.0, 88.0 | **0** | **0/50** |
**What was wrong the first time.** The original table pooled the
pre-instrumentation rows with the post-instrumentation ones, so
`controller_K3` exec1 appeared as `52.0, 50.0, 48.0, 50.0` (spread 4). The
52.0 came from an older eval build. Same-code replicates spread only 2. And
the "7-8 episodes flip" figure was never computed β it was inferred from the
spread. Actually counting per-episode disagreements gives **1** for that
config. A 2-point rate change *is* one flipped episode out of 50; asserting
7-8 was arithmetically impossible and should have been caught on sight.
**The real structure: replication noise is a property of the controller, not
the harness.** Three replicates of the corrected controller are
*bit-identical* in outcome β 94.0 three times, zero episodes differing β
while CEM flips 6 of 50. The underlying rollout nondeterminism is present in
all of them (the cost traces differ across the three `ah_hold0.5` runs:
mean terminal 0.9947 / 0.9947 / 0.9155). What differs is how many episodes
sit close enough to the 20px / 20-degree threshold for that jitter to change
the verdict.
So the noise floor is not a constant. It measures **how marginal a
controller's successes are**:
| controller | success | episodes flipped | reading |
|---|---|---|---|
| `ah_hold0.5` | 94.0 | 0/50 | succeeds with margin |
| `controller_K3` | 48-50 | 1/50 | mostly decisive |
| `cem` exec1 | 32-36 | 6/50 | many near-threshold episodes |
That CEM at exec1 has the most marginal outcomes is consistent with the
procrastination diagnosis: it is the configuration that ends episodes
drifting near the goal without settling.
**Consequences, restated.** Every large-effect conclusion stands unchanged β
the 36-point execution gap, the corrected controller's +44 at m=1, and
`abl_no_support`'s -40 gap are far above any of these floors. The
lambda_hold ordering (90.0 / 94.0 / 92.0) should still be treated as tied,
but now for a *better* reason than "it spans the noise": the p-values are
0.625 and 1.0 on paired tests, and the individual runs are themselves stable.
**The paired tests remain valid**, and are the right instrument precisely
because the noise is heteroskedastic across controllers. The earlier advice
to discount the borderline p-values (0.039, 0.031) still holds, since those
compare *marginal* configurations where the flip rate is highest.
## Experiment 4 β `abl_terminal_only`: removing the path loss makes it worse
The last cell of the ablation matrix. `--alpha 0` removes the late-weighted
path term `alpha * sum_j w_j d_j`, leaving the pure fixed-terminal objective
`d_H`. If procrastination were caused by the path loss, this should fix it.
| controller | m=1 | m=5 | gap |
|---|---|---|---|
| original (alpha=0.05) | 50.0 | 88.0 | -38 |
| `abl_no_support` | 50.0 | 90.0 | -40 |
| **`abl_terminal_only` (alpha=0)** | **18.0** | **90.0** | **-72** |
| `ah_hold0.5` (the fix) | 94.0 | 88.0 | +6 |
Removing the path loss does not fix procrastination β it **doubles** it. At
m=1 the controller scores 18.0%, which is 32 points *below* the original
(p = 0.0004, paired McNemar) and 76 points below the corrected controller
(p < 0.0001). Under full-plan execution it is unaffected (90.0 vs 88.0,
p = 1.0).
This is the cleanest confirmation of the diagnosis available. The path loss
was the one component of the original objective that applied *any* pressure
before block H. Deleting it leaves nothing but the terminal deadline, and
the pathology gets worse in exact proportion. Contraction analysis agrees:
| variant | c | b | D* |
|---|---|---|---|
| `abl_terminal_only` | 0.8272 | 0.0350 | **0.2028** |
| original | 0.5792 | 0.0412 | 0.0978 |
| `ah_hold0.0` | 0.4076 | 0.0282 | 0.0475 |
| `ah_hold0.5` | 0.3919 | 0.0241 | 0.0397 |
| `ah_hold1.0` | 0.4079 | 0.0214 | 0.0362 |
`abl_terminal_only` has by far the weakest contraction (c = 0.83 against the
original's 0.58) and a fixed point **twice** the original's, at R^2 = 0.94 β
the tightest fit of any variant, so this is a well-determined estimate rather
than a noisy one. Its mean trace bottoms out at 0.288 by replan 9 and then
*climbs* (0.292, 0.299, 0.307): the loop does not merely stall short of the
goal, it slowly reverses.
Taken with the other two ablations, all three candidate causes are now
settled. The support penalty is not the cause (`abl_no_support` reproduces
the original gap exactly, -40 vs -38, p = 1.0). The path loss is not the
cause β it was a partial *mitigation*, and removing it makes things worse.
The fixed terminal deadline is the cause, and re-indexing it by q is the fix.
### The support penalty, measured post-hoc
The `viol 0.000` logged by `abl_no_support` was an artifact: `--lambda-support 0`
skips loading the density model, so `violation` never left its zero
initializer. Recovered from the saved checkpoints against the same density
model and the same `c95 = 1.531`, on identical held-out batches:
| checkpoint | violation fraction | support loss | mean nll/dim |
|---|---|---|---|
| `abl_terminal_only` | 0.149 | 0.0139 | 1.144 |
| original | 0.187 | 0.0208 | 1.184 |
| `ah_hold0.5` | 0.208 | 0.0444 | 1.253 |
| **`abl_no_support`** | **0.652** | **0.2488** | **1.732** |
With the penalty removed, **65% of proposed action blocks fall outside the
95th-percentile support region** β a 3.5x increase over every trained-with-it
variant, and a mean NLL above the threshold itself. So the penalty is doing
real work: it is the only thing keeping the controller's plans inside the
region where the frozen world model's predictions are trustworthy.
The striking part is that this barely matters for success. `abl_no_support`
scores 50.0 / 90.0 β statistically identical to the original (p = 1.0) while
violating support 3.5x more often. On PushT the world model apparently
extrapolates well enough that leaving the data manifold costs little. That is
a statement about this environment's benign dynamics, not a general one, and
it is the kind of result that would likely not survive a contact-rich task
where model error compounds.
Note also that `ah_hold0.5` sits slightly *higher* than the original (0.208
vs 0.187, support loss 0.044 vs 0.021). Arrival-and-hold pushes marginally
harder against the support boundary β it has to, since holding position near
a goal requires action blocks that the demonstration data (always in motion)
covers sparsely. The penalty absorbs this without the plans drifting far.
## A confound in the efficiency claim: episodes terminate on success
While checking the replicates I found the corrected controller reports a
mean terminal cost of **0.995** against the original's 0.125 β 8x worse,
while scoring 94.0 against 50.0. Both cannot be read at face value.
The cause is survivorship. Episodes terminate the moment they succeed, so a
successful controller stops early and a failing one runs to the eval budget.
Every per-episode aggregate is therefore computed over a different, and
success-correlated, number of steps. Across all 36 paired rows, success rate
and mean terminal cost correlate at **r = +0.509**: better controllers report
*worse* cost.
Both affected metrics need restating.
**Cost traces.** `mean_terminal_distance` averages over surviving episodes,
so a good controller's trace is dominated by its handful of failures. This is
why `ah_hold0.5` at m=1 shows a trace climbing to 1.84 β by replan 10, the
only episodes still running are the 3 that never solve. It is not evidence of
divergence. `first_terminal_distance` is unaffected (it is taken from the
same held-out state by every planner before any termination is possible) and
remains the correct cross-planner comparison: 0.0423 for `ah_hold0.5` against
0.0638 for the original.
**Predictor rows per episode β this one changes a reported number.** The
mechanism is exact. `predictor_calls` is **constant per schedule** (200 at
exec1, 40 at exec5) no matter which controller runs, because the 50 episodes
are stepped as a batched vector env: one call per timestep regardless of how
many envs are still alive. Rows, however, count *alive* envs. So
rows/call = mean number of episodes still running
and `rows/ep = rows/call * calls / 50` is a measure of **episode length**,
not of per-decision cost:
| variant | exec | success | rows/ep | calls | rows/call = mean alive |
|---|---|---|---|---|---|
| `ah_hold0.5` | 1 | 94.0 | 73.2 | 200 | **18.3** |
| `ah_hold1.0` | 1 | 92.0 | 76.8 | 200 | 19.2 |
| `ah_hold0.0` | 1 | 90.0 | 78.8 | 200 | 19.7 |
| original | 1 | 48-50 | 146-149 | 200 | **36.5-37.3** |
| `abl_no_support` | 1 | 50.0 | 143.2 | 200 | 35.8 |
| `abl_terminal_only` | 1 | 18.0 | 185.6 | 200 | **46.4** |
| `cem` | 1 | 32-36 | 369900 | 61650 | 300.0 |
The ordering is monotone in success and nothing else: 94% keeps 18 envs
alive on average, 18% keeps 46. All six controllers are the same
architecture at the same K, so their true per-decision cost is *identical by
construction* β the 2.5x spread in rows/ep is entirely episode length.
The claim "5053x fewer rows than CEM" is therefore inflated by roughly 2x
relative to a per-decision reading. Two defensible statements remain, and
they answer different questions:
- **Per decision:** 300 / 18.3 = **16.4x fewer predictor rows per solver
call.** This is the architectural comparison β what one plan costs.
- **Per episode:** ~5000x fewer rows to solve a task, which is a real
end-to-end saving but partly *because* it solves faster and stops sooner.
Legitimate to quote, provided the mechanism is named.
The original headline of "764x fewer rows" (corrected controller m=1 vs CEM
exec5) compares across both a schedule change and a success gap, and should
be read as the per-episode figure with that caveat attached.
This does not affect the wall-clock speedups (measured directly, and finishing
sooner is a genuine saving) nor any success-rate comparison. It affects
exactly one class of claim: predictor-row ratios between planners of
differing success rates.
## Runbook / lessons
- Use the project interpreter explicitly:
`/c/Users/omnap/miniconda3/envs/llmdyn/python.exe`. The shell's default
`python` is conda **base** and lacks `hdf5plugin`, `stable_worldmodel`, and
torch-cuda.
- Background jobs: the wrapper shell gets killed, but `nohup`'d children
survive. **Verify a background run is producing rows before trusting it** β
an orphaned driver once burned the entire matrix in seconds via
`0xC0000142` DLL-init failures, producing nothing.
- Chain long GPU jobs by polling for a sentinel line (`SWEEP 1 DONE`) rather
than launching concurrently β *unless* you have measured the headroom. Two
controller trainings coexist at 1.7 it/s each on this card. **Two is the
hard limit.** A third process OOMs on host RAM (~2 GB free of 16.5) *and*
takes VRAM on its way down, killing a running training. Analysis scripts
that only read jsonl (`paired_stats.py`, `summarize_results.py`) load no
model and are always safe to run alongside; anything that calls
`load_lewm` is a GPU process and counts against the limit of two.
- Measuring "two processes run at full speed" does **not** license a third.
Throughput under N processes says nothing about headroom at N+1. Check free
VRAM *and* free host RAM before adding one.
- Poll for sentinels at 60s, not 30s, and avoid spawning a subshell per
iteration: a long-running `while ! grep ...; do sleep 30; done` eventually
hits cygwin `fork: Resource temporarily unavailable` / `0xC0000142`.
- `cmd || echo "!! failed"` keeps a sweep running past a crash, which is
usually right β but it turns a dead run into a *silent gap* in the results.
Verify by counting `===` stage markers against what the driver should have
emitted; two trainings were lost for hours because the tail looked healthy.
- The eval harness is **not** deterministic despite a fixed seed: episode
selection is seeded, but rollout is not. Re-run spread is **0-6 flipped
episodes out of 50, and it is a property of the controller, not the
harness** β a decisive controller replicates exactly, a marginal one does
not. Budget +/-4 points for a marginal config, ~0 for a decisive one.
- **Sanity-check a claimed flip count against the rate change before writing
it down.** A 2-point move on 50 episodes is *one* episode, full stop.
Asserting "7-8 flipped" alongside a 2-point spread was arithmetically
impossible and stood in this file for a full cycle.
- When pooling replicate rows, **check they came from the same code**. Rows
predating the per-episode-outcome instrumentation lack `episode_successes`;
including them inflated an apparent replication spread from 2 to 4.
- **Never compare `predictor_rows_per_episode` across controllers with
different success rates.** Episodes terminate on success, so a better
controller runs shorter episodes and accumulates fewer rows. Divide by
`predictor_calls` for the real per-decision cost. This confound inflated a
reported efficiency ratio by ~2x.
- Cost traces (`mean_terminal_distance`) carry the same survivorship: they
average over *surviving* episodes, so a good controller's trace is dominated
by its few failures and reads *worse* than a bad controller's (success and
mean cost correlate at r = +0.51 here). Use `first_terminal_distance`.
- A metric guarded by the same flag it is meant to measure will report a
perfect score for the ablation that disables it (`viol 0.000` under
`--lambda-support 0`). Initialize such metrics to `nan`, not `0`. The real
value was 0.652 β the worst in the matrix, logged as the best.
- `grep -c` returns exit 1 on zero matches, which reads as a failed
background task. Not an error.
## Experiment 5 β exp7: the one-operator controller (`fused192`), status: running
Full spec: `experiments/exp7_one_operator_controller/README.md`. Redesigns
`IterativeController` as the K-fold iteration of one operator Ξ¦ in the world
model's own 192-d space, instead of today's 256-d projected space split
across two networks (`consequence_net` F + `refine_net` G). A 2x2 factorial
(SPACE x OPERATOR) plus two replicate runs plus a five-part explainability
battery (E1-E5). This entry covers implementation + the first three of six
training runs; will be followed up (not overwritten) once the remaining
three finish and the primary endpoint is decided.
### Implementation
`lejepa_control/controller.py`: added `no_latent_proj` (identity
`latent_proj`, asserts `width == latent_dim`) and `fused` (single
`slot_proj` + `Encoder(width, depth=2*depth, ...)` + shared `delta_head`,
replacing `consequence_proj`/`consequence_net`/`refine_proj`/`refine_net`)
constructor flags. Both default `False`; default construction is unchanged
parameter-for-parameter (verified: baseline still builds to 6.80M params,
same layer-init order). `train_controller.py` gained `--no-latent-proj`,
`--fused`, `--train-seed`. `solver.py::load_controller` reads both new flags
via `saved.get(..., False)` so all 7 pre-exp7 PushT checkpoints rebuild
unchanged (regression-tested in `test_controller.py`).
**Param count vs. the spec's estimate.** Spec sec 1 estimates fused192 at
"~4-4.5M params". Measured: **3.75M** (baseline measured at 6.80M, matches
the spec's stated baseline figure exactly). Audited by hand: the transformer
net (8 layers @ width 192) accounts for 3.56M of the 3.75M total, and both
the layer count and width are pinned explicitly by the spec's own text, so
there is no architectural slack to hit 4-4.5M without contradicting the
spec's own description. Treated as a pre-implementation estimate that ran
~7-17% high rather than an implementation bug; widened the test's assertion
band to 3.5-4.0M with the audit trail recorded in a comment
(`test_param_counts` in `scripts/test_controller.py`).
Full 2x2 (+2 replicates):
| cell | space | operator | run | params |
|---|---|---|---|---|
| baseline | 256+proj | split | `ah_hold0.5` (exists) | 6.80M |
| A | 192 identity | split | `exp7/w192np_split` | 3.82M |
| B | 256+proj | fused | `exp7/fused256` | 6.67M |
| C (candidate) | 192 identity | fused | `exp7/fused192` | 3.75M |
| C-rep | 192 identity | fused | `exp7/fused192_s2` (`--train-seed 1`) | 3.75M |
| base-rep | 256+proj | split | `exp7/base_r2` | 6.80M |
### Two bugs found in existing infra while wiring up the eval side
1. `scripts/paired_stats.py::label()` only special-cased 5 hardcoded
checkpoint substrings, so any checkpoint not matching those (every exp7
cell) fell back to the bare `planner` tag (`controller_K3`) β all 5 exp7
cells share that tag, so without a fix their eval rows collide under one
label and silently overwrite each other. Fixed: extended the substring
list with the exp7 cell names, `break` on first match.
2. `paired_stats.py::main()`'s `by_label` dict was keyed by `label(row)`
alone, with no seed in the key β so evaluating one checkpoint at 3 seeds
(the pre-registered protocol) produced 3 rows under the *same* label, and
the dict silently kept only the last seed's 50 episodes instead of
pooling to 150. Confirmed by rerunning: pre-fix printed `(n=50)` for
every label; post-fix, `(n=150)`. Fixed: group rows by label, sort by
seed, `np.concatenate`. This one would have quietly invalidated the
primary endpoint's pre-registered "150 pooled episodes" comparison.
Also: no visualization decoder existed anywhere on this machine (needed for
E3's direct plan-token decode and all of E5) β built one as a prerequisite,
`data/runs/decoder/decoder.pt` (preset `cpu`, 3000 steps, val MSE 0.00152 vs
0.00477 mean-image baseline, 0.32x). Cache: `tools/make_decoder_cache.py
--frames 6000 --latent-source emb`, `data/decoder_cache/`.
### New explainability-battery scripts (sec 5)
- `scripts/probe_alignment.py` (E2): cosine between the emitted update
`alpha_k * delta_Y_j` and the true descent direction
`-grad_{Y_j} sum_i d_i`, autograd through the frozen rollout only. Works
unchanged on split and fused. Smoke-tested on `ah_hold0.5`: weak positive
alignment, decaying over k (0.135 -> 0.094 -> 0.050 at samples=32) β
consistent with the "amortized descent" story, strongest right when the
plan is furthest from converged.
- `scripts/probe_attention.py` (E4): temporarily patches every `Block`'s
bound forward to call attention with `need_weights=True` (the default
forward never computes weights), restores originals after. Reports mass
from plan-slot queries onto {self, other-plan, context, goal}, per layer
and iteration; labels layers positionally (`consequence_net` blocks always
precede `refine_net` blocks within one `refine()` call; only `net` for
fused).
- `scripts/probe_geometry.py` (E3): `||y_j - x_hat_j||` and cosine between
plan tokens and the latents they cause, per k; only defined when
`no_latent_proj=True` (skips cleanly otherwise). Also decodes plan tokens
directly through the pixel decoder (reuses `tools/decode_rollout.py`'s
`panel`/`to_uint8`).
- E1 (`diagnostics.py`) and E5 (`tools/decode_rollout.py`) needed no code
changes β both already go through `load_controller`. Confirmed working
end-to-end against `ah_hold0.5` (E5: latent drift [0.009, 0.029, 0.04,
0.063, 0.074] over the 5 blocks, imagined-goal-distance collapsing
0.82 -> 0.03). `ah_hold0.5`'s E1 entry already existed in
`data/runs/diagnostics/diagnostics.jsonl` (c=0.392, D*=0.0397, matches
CLAUDE.md) β reused rather than re-run.
### Training + eval, status at time of writing
All 6 runs use the spec's exact stored baseline recipe (`--arrival-hold
--hold-weight 0.5 --batch-size 128 --lr 3e-4`, 20000 steps default). Two GPU
slots at a time (per the "two trainings coexist" runbook note above),
~2.3-2.4 it/s each shared (vs ~4.1 it/s solo for these smaller
architectures) -> ~144 min/run shared vs. the original 89 min/run
solo-baseline figure.
- `w192np_split` (A) and `fused256` (B): **done**, 143.7 min each.
- `fused192` (C, candidate) and `fused192_s2` (C-rep): **training**.
- `base_r2`: not yet started (queued after C/C-rep).
Eval protocol (sec 4: `--num-eval 50 --goal-offset 25`, seeds 42/43/44,
`rh` in {1,5}) run so far for baseline + A + B, into
`data/runs/eval_exp7/results.jsonl`. Baseline needed 5 new rows (only
seed=42 existed anywhere on this machine, in `data/runs/eval/results.jsonl`)
β re-ran all 3 seeds fresh into `eval_exp7` for single-file provenance.
Pooled (150 episodes, 3 seeds), rh=1, **preliminary β candidate C not yet
evaluated, so nothing here is the primary endpoint**:
| cell | success (rh=1) | vs baseline (paired) |
|---|---|---|
| baseline (`ah_hold0.5`) | 94.0% | β |
| A (`w192np_split`) | 92.7% | -1.3, CI [-6.0, +3.3], p=0.77 |
| B (`fused256`) | 90.7% | -3.3, CI [-8.7, +1.3], p=0.30 |
Neither is the pre-registered primary comparison (that's C vs baseline).
Noting only because A's CI lower bound (-6.0) sits exactly on the
non-inferiority boundary β worth rechecking once C's numbers are in, in case
the width cut alone (without fusion) turns out to be a borderline call.
To be continued in a follow-up entry once C/C-rep/base_r2 finish, the eval
sweep completes for all 6 cells, and `paired_stats.py` gives the
primary/secondary endpoints β not overwriting this one, per the append-only
convention.
## Experiment 5, continued β exp7 final result: inconclusive, not negative
All 6 training runs complete (`w192np_split`, `fused256`, `fused192`,
`fused192_s2`, `base_r2` β 143.7/143.7/143.8/143.7/85.4 min respectively; the
first four shared the GPU two-at-a-time, `base_r2` ran solo). Full eval
protocol (Β§4: 50 episodes, seeds 42/43/44, rh β {1,5}) complete for all 6
cells, into `data/runs/eval_exp7/results.jsonl`. Anytime-profile sweep (K β
{0,1,2,3,5,8}, rh=1, 3 seeds) complete for `fused192`, into
`data/runs/eval_exp7_refine_sweep/`. Full explainability battery (E1-E5)
complete on `fused192` and (already existing / backfilled) `ah_hold0.5`.
A third bug surfaced while writing the factorial analysis:
`paired_stats.py` has no support for multi-cell contrasts (main effects are
linear combinations across 4 cells, not a pairwise diff), so the SPACE/
OPERATOR/interaction analysis needed a small one-off script
(paired bootstrap over the same 150 index-aligned episodes, resampled
jointly across cells so the pairing is preserved) rather than reusing
`paired_stats.py` directly. Not folded into the repo β one-off, not a
reusable pattern yet.
### Pooled success rates (150 episodes = 3 seeds Γ 50, index-aligned)
| cell | rh=1 | rh=5 |
|---|---|---|
| baseline (`ah_hold0.5`) | 94.0% | 85.3% |
| A (`w192np_split`, 192id+split) | 92.7% | 86.7% |
| B (`fused256`, 256proj+fused) | 90.7% | 86.0% |
| **C (`fused192`, candidate)** | **90.7%** | **84.7%** |
| C-rep (`fused192_s2`, seed 1) | 88.7% | 90.0% |
| base-rep (`base_r2`) | 91.3% | 85.3% |
### Primary/secondary endpoint (Β§4)
- **rh=1 (primary): candidate β baseline = β3.3 pts, 95% CI [β8.0, +1.3],
p = 0.27 (McNemar).** CI lower bound is below the pre-registered β6pt
non-inferiority threshold β **fails the rule as written.**
- **rh=5 (secondary): candidate β baseline = β0.7 pts, CI [β5.3, +4.0].**
Passes comfortably.
### The replicate-gap check changes the reading of the primary result
Β§4 pre-registered exactly this check, and it matters here:
| replicate pair | rh=1 gap | rh=5 gap |
|---|---|---|
| `fused192` vs `fused192_s2` (same recipe, `--train-seed 1`) | +2.0, CI [β3.3, +7.3] | **β5.3, CI [β9.3, β1.3]** |
| `ah_hold0.5` vs `base_r2` (baseline recipe, rerun) | +2.7, CI [β1.3, +6.7] | +0.0, CI [β3.3, +4.0] |
At rh=1, **both replicate gaps (Β±2-3 pts, CIs overlapping zero) are the same
order of magnitude as the β3.3pt candidate-vs-baseline gap.** Per Β§4's own
instruction ("if the replicate gap is comparable to a cell gap, say so and
stop interpreting that cell gap"): **the rh=1 shortfall is not
distinguishable from ordinary training-run variance** on the evidence of two
replicate pairs. This is a real, literal failure of the pre-registered rule,
but it is not safe to read as "the architecture is worse" β it's
underpowered to tell architecture-driven variance from seed-driven variance
at this sample size (n=2 replicates per recipe).
The rh=5 replicate gap for `fused192` (β5.3, CI excludes zero) is the one
genuinely surprising number in the whole run: the two `fused192` training
seeds disagree by 5+ points under open-loop execution specifically, while
the baseline recipe's two seeds agree almost exactly (+0.0). This reads as
"the fused192 recipe is more training-seed-sensitive under rh=5 than the
baseline recipe is" β worth a third replicate before trusting, but it's the
one place in this dataset where a replicate gap itself clears significance.
### Factorial main effects (rh=1, paired bootstrap, 2Γ2 corner only)
| effect | estimate | 95% CI |
|---|---|---|
| SPACE (192id β 256proj) | β0.7 | [β4.0, +2.7] |
| OPERATOR (fused β split) | β2.7 | [β6.0, +0.3] |
| INTERACTION | +1.3 | [β4.0, +7.3] |
None of the three CIs excludes zero β OPERATOR comes closest (upper bound
+0.3, just barely still containing zero) but is not a confident effect
either. SPACE looks like genuinely nothing. No interaction signal.
### Β§7 readout table: closest row is "C below, B below," with an asterisk
By the letter of the rule (CI lower bound test only, ignoring the replicate
caveat above): baseline vs A = β1.3 [β6.0, +3.3] (boundary case, lower bound
lands exactly on β6.0); baseline vs B = β3.3 [β8.7, +1.3] (clearly below);
baseline vs C = β3.3 [β8.0, +1.3] (clearly below). That's **"C below, B
below" β "the perception/editing split earns its keep at this scale, negative
result."** But given the replicate-gap finding above, this should be read as
*"the measured numbers land in the row that would say the split earns its
keep, but the effect size is inside this experiment's own noise floor β a
confident negative result would need more than 2 replicates per cell."*
Shipping the split architecture on this evidence would be premature; so
would shipping `fused192` on this evidence. The honest conclusion is
**inconclusive at this replicate count**, not a clean win or loss either way.
### Explainability battery (E1-E5), `fused192` vs `ah_hold0.5`
- **E1 (contraction, exec1).** `fused192`: c=0.480, b=0.021, D\*=0.039,
RΒ²=0.68. `ah_hold0.5`: c=0.428, b=0.024, D\*=0.042, RΒ²=0.59 (this
particular comparison run; the officially logged baseline number is
c=0.392, D\*=0.040 β cross-run c spread for the *same* baseline checkpoint
across diagnostics runs is ~0.35-0.43, another instance of the
run-to-run-noise theme above). Spec's target was "candidate c β€
baseline's ~0.39-0.43" β `fused192`'s 0.480 sits just above that band, a
mild miss, but D\* (the number that actually matters for success β is the
fixed point inside the success radius) is essentially unchanged and still
well inside tolerance.
- **Anytime profile (K-sweep) β the cleanest positive result in this
experiment.** `fused192`'s K-sweep (this run, pooled 3 seeds) tracks
`ah_hold0.5`'s historical K-sweep almost exactly:
| K | 0 | 1 | 2 | 3 | 5 | 8 |
|---|---|---|---|---|---|---|
| `fused192` | 77.3 | 88.0 | 88.7 | 90.7 | 90.7 | 91.3 |
| `ah_hold0.5` (historical) | 74.7 | 85.0 | 87.0 | 89.7 | 90.3 | 90.7 |
Same monotone-then-flat shape, saturating by Kβ3, within 1-3 points of the
baseline at every K. The "anytime iterator" story the whole redesign is
built on **holds for the fused architecture**, independent of and more
robust than the noisy rh=1 success-rate horse race above.
- **E2 (gradient alignment).** `fused192`: k=0..2 mean cosine
[0.169, 0.079, 0.040]. `ah_hold0.5`: [0.138, 0.095, 0.054]. Both weakly
positive, both decaying toward zero as the plan converges β the
"amortized descent" story holds similarly on both architectures, no
meaningful difference between them.
- **E3 (token-latent geometry) β a genuine negative finding, worth stating
plainly.** The hypothesis was "does a plan token converge toward the
latent of the state it causes?" Measured on `fused192`: `βy_j β xΜ_jβ`
is **19.0 β 19.6 β 20.2 β 20.7** across k=0..3 (*increasing*, not
decreasing), and cosine is **slightly negative throughout** (β0.042 β
β0.035). For reference, two independent random vectors at this world
model's per-coordinate latent std (β0.87, 192-d) would have expected
distance ββ(2Β·192Β·0.87Β²) β 17 β `fused192`'s measured 19-20 is in that
same regime. **Plan tokens do not geometrically converge toward their
consequences; if anything they drift slightly further apart as refinement
proceeds.** The identity-embedding design buys the *ability* to make this
measurement (and to decode tokens directly β the decode panels in
`data/runs/probe_geometry/` are the artifact), but the measurement itself
says the "tokens are latents-in-waiting" reading of "one space" was not
literally true for this checkpoint. The operator uses the shared
coordinate system as a *convenient interface*, not as a promise that plan
and outcome coincide.
- **E4 (attention readout) β a coherent, reportable story.** Mass from
plan-slot queries, k=0 (building Yβ°) vs k=3 (last refinement), averaged
over `fused192`'s 8 layers: at k=0, mass splits across context (~29%) and
goal (~26%) with plan/self comparatively low β the operator is grounding
the seed plan in "where am I, where do I need to be." By k=3, plan mass
dominates (~52% average, up from ~36% at k=0) while context (~19%) and
goal (~15%) both drop by roughly a third. Read together with E2's decaying
alignment: **early refinements ground the plan in the world; late
refinements are mostly about reconciling the plan against itself.** This
is the "one coherent map" story the fused design was meant to produce, and
it is visible in a single net where it would have been split across two
in the baseline.
- **E5 (decode strips).** 4 episodes each in `data/runs/decode_rollout_exp7/
{ah_hold0.5,fused192}/`. Qualitative; not scored here.
### Bottom line
`fused192` is **~45% smaller** (3.75M vs 6.80M params), preserves the
anytime-iterator behavior essentially exactly, and produces a genuinely
interpretable attention story (E4) β but its rh=1 success rate misses the
pre-registered non-inferiority bar by the letter of the rule, and that miss
is not distinguishable from training-run noise at this replicate count (n=2
per recipe). E3 additionally corrects a piece of the design's own narrative:
plan tokens do not converge toward their consequences just because they
share a coordinate system. **Recommendation: do not ship on the current
evidence; the next experiment this motivates is more replicates (3-4 per
cell, not 2) before either shipping `fused192` or concluding the split
earns its keep** β everything else in this experiment (implementation,
tests, explainability tooling) is solid and reusable regardless of which
way that resolves.
|