File size: 114,964 Bytes
d3d0e0e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
947
948
949
950
951
952
953
954
955
956
957
958
959
960
961
962
963
964
965
966
967
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
984
985
986
987
988
989
990
991
992
993
994
995
996
997
998
999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
1392
1393
1394
1395
1396
1397
1398
1399
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
1851
1852
1853
1854
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
1865
1866
1867
1868
1869
1870
1871
1872
1873
1874
1875
1876
1877
1878
1879
1880
1881
1882
1883
1884
1885
1886
1887
1888
1889
1890
1891
1892
1893
1894
1895
1896
1897
1898
1899
1900
1901
1902
1903
1904
1905
1906
1907
1908
1909
1910
1911
1912
1913
1914
1915
1916
1917
1918
1919
1920
1921
1922
1923
1924
1925
1926
1927
1928
1929
1930
1931
1932
1933
1934
1935
1936
1937
1938
1939
1940
1941
1942
1943
1944
1945
1946
1947
1948
1949
1950
1951
1952
1953
1954
1955
1956
1957
1958
1959
1960
1961
1962
1963
1964
1965
1966
1967
1968
1969
1970
1971
1972
1973
1974
1975
1976
1977
1978
1979
1980
1981
1982
1983
1984
1985
1986
1987
1988
1989
1990
1991
1992
1993
1994
1995
1996
1997
1998
1999
2000
2001
2002
2003
2004
2005
2006
2007
2008
2009
2010
# Implementation Ledger

## Current Baseline

- Branch: feature/observatory-phase-00-baseline
- Base commit: 1bf512073d0bb751ad66af52bc77361c52283fc8
- Python version: Python 3.11.15
- Dependency manager/tooling: uv 0.11.8
- Streamlit entrypoint: streamlit run src/data_agent_baseline/observatory/app.py
- CLI entrypoint: dabench (data_agent_baseline.cli:main)

### Pre-change command/test baseline

1. `git status --short`: untracked docs in `Docs_To_be_deleted/` only.
2. `pytest --collect-only -q`: 109 tests collected.
3. Practical existing suite run:
   - Command: `./.venv/bin/pytest -q tests/test_eval_v2_validator.py tests/test_eval_v2_hardening.py tests/test_eval_harness.py tests/test_observatory_loaders.py tests/test_observatory_trace_parser.py tests/test_observatory_replay_builder.py`
    - Result: 39 passed, 2 failed (snapshot at session start; loaders test was subsequently confirmed fixed).
    - Failures at session start:
       1. `tests/test_eval_v2_validator.py::test_validator_catches_verification_outcome_mismatch` — still failing at Phase 0 close.
       2. `tests/test_observatory_loaders.py::test_run_discovery_and_task_discovery` — test file contained `assert "11"` at session start, but current line 30 already asserts `"task_11"` and passes; resolved prior to Phase 0.
4. Placeholder tests present pre-change:
   1. `tests/test_eval_harness.py::TestMetricBounds::test_data_understanding_score_bounds` (pass placeholder)
   2. `tests/test_eval_harness.py::TestMetricBounds::test_all_score_metrics_bounded` (pass placeholder)
   3. `tests/test_eval_harness.py::test_end_to_end_reconciliation` (pass placeholder)
   4. `tests/test_eval_harness.py::test_health_report_generation` (pass placeholder)


**Failure 1 — pre-existing**
- Test path: `tests/test_eval_v2_validator.py::test_validator_catches_verification_outcome_mismatch`
- Baseline failure message:
   ```
   tests/test_eval_v2_validator.py:96: in test_validator_catches_verification_outcome_mismatch
         assert any(i["check_id"] == "verification_outcome_mismatch" for i in issues)
   E   assert False
   E    +  where False = any(<generator object ...>)
   ```
- Post-change failure message: identical to baseline — no change.
- Classification: pre-existing; the `verification_outcome_mismatch` check_id is not produced by the current validator implementation.
- Phase 0 contribution: none. Phase 0 added no new failures.

**Failure 2 — pre-existing**
**Failure 2 — resolved before Phase 0 close**
- Test path: `tests/test_observatory_loaders.py::test_run_discovery_and_task_discovery`
- Status at Phase 0 close: **PASSES**. Current line 30 asserts `"task_11"` (full task ID format).
- Phase 0 contribution: none. Phase 0 made no change to this test or its backing loader.

The final broader regression suite at Phase 0 close: **63 passed, 1 failed**. The sole remaining failure is the validator `check_id` test (Failure 1). Phase 0 did not introduce any new failures.

## Phase Status

| Phase | Status | Branch | Commit | Tests | Acceptance | Notes |
|---|---|---|---|---|---|---|
| Phase 0 | Completed | feature/observatory-phase-00-baseline | 1bf512073d0bb751ad66af52bc77361c52283fc8 (base) | Added + validated | PASS (see matrix) | No src runtime behavior changes |
| Phase 1 | Completed (uncommitted) | feature/observatory-phase-01-run-service | d09fdfa6bfe8e2e2d8f2fa6ae5bf3b707f32e412 (current HEAD) | Added + validated | PASS (see matrix) | Shared run execution service extracted from CLI |
| Phase 2 | Completed (uncommitted) | feature/observatory-phase-01-run-service | d09fdfa6bfe8e2e2d8f2fa6ae5bf3b707f32e412 (current HEAD) | Added + validated | PASS (see matrix) | Eval-v2 orchestration extracted into EvaluationService with callback-driven CLI progress |
| Phase 3 | Completed (uncommitted) | feature/observatory-phase-03-run-intelligence | d09fdfa6bfe8e2e2d8f2fa6ae5bf3b707f32e412 (current HEAD) | Added + validated | PASS (see matrix) | Read-only Run Intelligence page added to Streamlit |
| Phase 4 | Completed (uncommitted) | feature/observatory-phase-04-run-launcher | current workspace | Added + validated | PASS (see matrix) | Synchronous predefined-task Run Launcher delivered |
| Phase 5 | Completed (uncommitted) | feature/observatory-phase-05-live-execution | current workspace | Added + validated | PASS (see matrix) | Live synchronous execution observability and manifest-status correctness |
| Phase 6 | Completed (staged) | feature/observatory-phase-06-custom-task-intake | current workspace | Added + validated | PASS (272 passed) | External Benchmark Dataset Intake in Run Launcher; difficulty field made optional in dataset contract. Free-form custom task intake deferred to Phase 7. |
| Phase 7 | Completed | feature/observatory-phase-07-freeform-custom-task-intake | current workspace | Added + validated | PASS (372 passed) | Free-form Custom Task Intake: Custom Task tab in Run Launcher with question, context, file upload, output type selector. Backend materializes tasks into synthetic datasets, executes through shared service. |
| Phase 8 | Completed (uncommitted) | feature/observatory-phase-08-checkpoint-steering | current workspace | Added + validated | PASS (397 passed) | Checkpoint Review tab with additive advisory annotations (`checkpoint_annotations.json`) derived from existing trace/events, with no pause/resume or execution intervention. |
| Phase 9 | Completed (uncommitted) | feature/observatory-phase-09-guided-rerun-planning | current workspace | Added + validated | PASS (434+ passed) | Guided Rerun Plan tab with rerun plan creation, approval workflow, and safe guided rerun execution as new isolated custom task runs. Original run artifacts remain immutable. |
| Phase 10 | Completed (uncommitted) | feature/observatory-phase-10-rerun-comparison-ask-this-run | current workspace | Added + validated | PASS (533 passed) | Rerun Comparison tab with read-only original-vs-guided comparison, local artifact-grounded Ask This Run/Comparison Q&A, steering influence summaries, and saved comparison reports. Preserves artifact integrity. |
| Phase 11 | Completed (uncommitted) | feature/observatory-phase-11-cohort-guided-evaluation | current workspace | Added + validated | PASS (602 passed) | Run-level cohort guided rerun evaluation integrated into Run Intelligence with discovery, aggregation, cohort Q&A, and optional cohort report. Read-only, no execution, no causal claims. |
| Phase 12 | Completed (uncommitted) | feature/observatory-phase-12-evidence-pack | current workspace | Added + validated | PASS (674 passed) | Evidence Pack with cross-phase narrative, statistical readiness, claim safety framework, and markdown/JSON exports. Read-only, no execution, no causal claims. |
| Phase 13 | Completed (uncommitted) | feature/observatory-phase-13-demo-tab | current workspace | Added + validated | PASS (746 passed) | Demo tab with Recorded Evidence Mode, optional Live Mini Demo, demo health scoring, competition checklist, flow scripts, and demo bundle export. No auto-execution. |
| Phase 14 | Completed (uncommitted) | feature/observatory-phase-14-dao-branding-release-audit | current workspace | Added + validated | PASS (748+ passed) | Branding as DAO — Data Agent Observatory and final release readiness audit. No execution/evaluation/artifact behavior changes. |
| Phase 16 | Completed (uncommitted) | feature/observatory-phase-16-live-dag-replay | current workspace | Added + validated | PASS (774+ passed) | Live Trace Flow with trace-native step/action backbone, Execution Story mapping, future HITL-compatible node schema (inert), read-only observability. No step event instrumentation added. |
| Phase 17 | Completed (uncommitted) | feature/observatory-phase-17-guided-trace-hitl | current workspace | Added + validated | PASS (focused Phase 17 matrix) | Guided Live Trace Flow with planner checkpoint HITL (approve/revise/cancel), persisted checkpoint/intervention/delta artifacts, and guided-policy run execution. |
| Phase 19 | Completed (uncommitted) | feature/observatory-phase-19-navigation-consolidation | current workspace | Added + validated | PASS (focused Phase 19 matrix) | Primary navigation consolidated to four pages (Run Launcher, Run Intelligence, Task Intelligence, Demo / Future Proof) with fallback access for Live Trace Flow and Guided Live Trace Flow unchanged. |

### Phase 19 — Observatory Navigation Consolidation

- Started: 2026-07-01
- Completed: 2026-07-01
- Branch: `feature/observatory-phase-19-navigation-consolidation`

#### Summary

1. Consolidated top-level navigation in `app.py` to four primary tabs:
   - Run Launcher
   - Run Intelligence
   - Task Intelligence
   - Demo / Future Proof
2. Added shell composition pages:
   - `task_intelligence_page.py` (Mission Summary, DAG, Replay, Provenance, Critic/Reviewer, Confidence, Failure/Verification, Raw Trace)
   - `future_proof_page.py` (Checkpoint Review, Guided Rerun Plan, Rerun Comparison, Evidence Pack placeholder, Demo)
3. Preserved fallback access to `live_trace_flow_page.py` and `guided_trace_flow_page.py` under Advanced / Fallback Views.
4. No execution/evaluation/HITL runtime behavior changes and no artifact schema changes.
5. Phase 20 natural-language artifact Q&A intentionally unchanged and out of scope.

## Phase Details

### Phase 0

- Started: 2026-06-18
- Completed: 2026-06-18

#### Files created

1. `docs/architecture/ADR-001-shared-services.md`
2. `docs/implementation/PHASE0_COMMAND_CONTRACT.md`
3. `tests/fixtures/phase0/README.md`
4. `tests/fixtures/phase0/generated_eval/gold/task_1/gold.csv`
5. `tests/fixtures/phase0/generated_eval/run/task_1/prediction.csv`
6. `tests/fixtures/phase0/generated_eval/run/task_1/trace.json`
7. `tests/fixtures/phase0/reader_contract/success_run/*` (reduced run/task artifacts)
8. `tests/fixtures/phase0/reader_contract/failed_run/*`
9. `tests/fixtures/phase0/reader_contract/multi_task_run/*`
10. `tests/fixtures/phase0/reader_contract/validator_warnings_run/validation_report.md`
11. `tests/fixtures/phase0/reader_contract/replay_run/*`
12. `tests/fixtures/phase0/reader_contract/react_compat_run/*`
13. `tests/fixtures/phase0/reader_contract/view_task_v2_run/*`
14. `tests/cli/test_cli_phase0_contracts.py`
15. `tests/integration/test_eval_v2_golden_artifacts.py`
16. `tests/integration/test_reader_contract_phase0.py`

#### Files modified

1. `docs/implementation/IMPLEMENTATION_LEDGER.md` (this file)
2. `docs/implementation/DECISIONS.md` (initialized Phase 0 decision)

#### Tests added

1. `tests/cli/test_cli_phase0_contracts.py`
2. `tests/integration/test_eval_v2_golden_artifacts.py`
3. `tests/integration/test_reader_contract_phase0.py`

#### Commands executed and outcomes

1. Baseline metadata:
   - `git status --short`
   - `git rev-parse HEAD`
   - `./.venv/bin/python --version`
   - `uv --version`
2. Baseline tests:
   - `./.venv/bin/pytest --collect-only -q` -> 109 collected
   - practical suite -> 39 passed, 2 failed (pre-existing)
3. New tests:
   - `./.venv/bin/pytest -q tests/cli/test_cli_phase0_contracts.py tests/integration/test_eval_v2_golden_artifacts.py tests/integration/test_reader_contract_phase0.py` -> 16 passed
4. Syntax/import validation:
   - `./.venv/bin/python -m compileall -q src tests` -> pass
5. Broader regression matrix:
   - `./.venv/bin/pytest -q tests/cli/test_cli_phase0_contracts.py tests/integration/test_eval_v2_golden_artifacts.py tests/integration/test_reader_contract_phase0.py tests/test_eval_v2_validator.py tests/test_eval_v2_hardening.py tests/test_eval_harness.py tests/test_observatory_loaders.py tests/test_observatory_trace_parser.py tests/test_observatory_replay_builder.py tests/test_observatory_dag_builder.py tests/test_observatory_confidence_builder.py tests/test_observatory_provenance_builder.py`
   - Result: 62 passed, 2 failed (same pre-existing failures)
6. CLI smoke helps:
   - `dabench run-lang-task --help`
   - `dabench run-lang-benchmark --help`
   - `dabench eval-v2 --help`
   - `dabench view-task-v2 --help`
   - `dabench eval-comprehensive --help`
   - `dabench tag-failures --help`
   - `dabench view-exec-report --help`
7. CLI negative-path smoke:
   - `dabench eval-v2 /tmp/phase0-missing-run` -> exit 1
   - `dabench view-task-v2 task_1 /tmp/phase0-missing-metrics` -> exit 1
8. eval-v2 generated fixture runs:
   - standard -> exit 0
   - verbose -> exit 0
   - research -> exit 0

#### Artifact inspection evidence

1. Generated fixture run outputs exist:
   - `task_metrics.csv`, `trajectory.csv`, `tool_calls.csv`, `comprehensive_evaluation.csv`, `validation_report.md`, `auditor_validation_report.md`, `artifact_reconciliation_report.txt`, `engineering_health_report.txt`, `task_1/task_replay.json`.
2. CSV schema spot-check:
   - `task_metrics.csv` includes `run_id,task_id,execution_success,final_score`.
   - `trajectory.csv` includes `task_id,step_id,action`.
   - `tool_calls.csv` includes `task_id,tool_name,success`.
3. Replay JSON key check:
   - includes `task_id,run_id,trace_id,final_score,execution_success`.
4. Validation and health checks:
   - validation report contains `## ERROR` with `- none`.
   - health report includes `HEALTHY` status.
5. ID consistency:
   - task ID remains `task_1` across task_metrics/trajectory/tool_calls/replay.

#### Fixture hygiene and size

1. Fixture tree size: ~160 KB (`tests/fixtures/phase0`).
2. Files are reduced and synthetic; no source datasets copied.
3. Secret/path scan found no API keys/passwords/secrets; only minimal token count fields required for metric contracts.
4. Deterministic run/task IDs and timestamps used in fixtures.

#### Known limitations

1. Repository `.gitignore` currently ignores `docs/` and `tests/`; staging these Phase 0 files requires forced add (`git add -f ...`) or `.gitignore` adjustment.
2. Two pre-existing test failures remain and were not modified in Phase 0.

#### Deferred work

1. Service extraction and refactoring (Phases 1+).
2. Artifact path migration to `evaluation/` subdir (future compatibility phase).
3. Resolving existing unrelated failing tests.

## Acceptance Criteria Audit

| Criterion | Status | Evidence |
|---|---|---|
| 1. Existing CLI commands still run. | PASS | CLI help smoke commands succeed for all required commands. |
| 2. CLI smoke/contract tests exist. | PASS | `tests/cli/test_cli_phase0_contracts.py` added and passing. |
| 3. Representative reduced fixtures exist. | PASS | `tests/fixtures/phase0/*` created, documented, minimized. |
| 4. ADR-001 is documented. | PASS | `docs/architecture/ADR-001-shared-services.md`. |
| 5. Test suite locks current eval-v2 artifact contract. | PASS | `tests/integration/test_eval_v2_golden_artifacts.py` validates outputs and schemas. |
| 6. Observatory current-layout compatibility is tested. | PASS | `tests/integration/test_reader_contract_phase0.py` and existing observatory tests. |
| 7. Warning-vs-error validator behavior is tested. | PASS | `test_eval_v2_validator_warnings_do_not_fail` and `test_eval_v2_validator_errors_fail_command`. |
| 8. No production behavior changed. | PASS | No edits under `src/`; tests/docs/fixtures only. |
| 9. No src files changed without approved defect. | PASS | No `src/` files modified. |
| 10. No sensitive/large/machine-specific fixture data added. | PASS | Fixture size ~160 KB, synthetic content, no secrets, no absolute developer paths. |

### Phase 1

- Started: 2026-06-19
- Completed: 2026-06-19

#### Files created

1. `src/data_agent_baseline/domain/__init__.py`
2. `src/data_agent_baseline/domain/run_models.py`
3. `src/data_agent_baseline/application/__init__.py`
4. `src/data_agent_baseline/application/run_execution_service.py`
5. `src/data_agent_baseline/repositories/__init__.py`
6. `src/data_agent_baseline/repositories/run_repository.py`
7. `src/data_agent_baseline/repositories/filesystem_run_repository.py`
8. `tests/unit/__init__.py`
9. `tests/unit/domain/__init__.py`
10. `tests/unit/domain/test_run_models.py`
11. `tests/unit/repositories/__init__.py`
12. `tests/unit/repositories/test_filesystem_run_repository.py`
13. `tests/unit/application/__init__.py`
14. `tests/unit/application/test_run_execution_service.py`
15. `tests/integration/test_phase1_cli_delegates.py`

#### Files modified

1. `src/data_agent_baseline/cli.py`
2. `tests/cli/test_cli_phase0_contracts.py`
3. `docs/implementation/IMPLEMENTATION_LEDGER.md`
4. `docs/implementation/DECISIONS.md`

#### Commands executed and outcomes

1. `./.venv/bin/python -m compileall -q src tests` -> pass
2. `./.venv/bin/pytest -q tests/unit/domain/test_run_models.py tests/unit/repositories/test_filesystem_run_repository.py tests/unit/application/test_run_execution_service.py tests/integration/test_phase1_cli_delegates.py` -> 27 passed
3. `./.venv/bin/pytest -q tests/cli/test_cli_phase0_contracts.py tests/integration/test_eval_v2_golden_artifacts.py tests/integration/test_reader_contract_phase0.py` -> 16 passed
4. `./.venv/bin/pytest -q tests/cli/test_cli_phase0_contracts.py tests/unit tests/integration/test_phase1_cli_delegates.py tests/integration/test_eval_v2_golden_artifacts.py tests/integration/test_reader_contract_phase0.py tests/test_eval_v2_validator.py tests/test_eval_v2_hardening.py tests/test_eval_harness.py tests/test_observatory_loaders.py tests/test_observatory_trace_parser.py tests/test_observatory_replay_builder.py tests/test_observatory_dag_builder.py tests/test_observatory_confidence_builder.py tests/test_observatory_provenance_builder.py` -> 90 passed, 1 failed (known pre-existing validator failure)
5. `./.venv/bin/dabench run-lang-task --help` -> pass
6. `./.venv/bin/dabench run-lang-benchmark --help` -> pass
7. `./.venv/bin/dabench eval-v2 --help` -> pass
8. `./.venv/bin/dabench view-task-v2 --help` -> pass

#### Artifact inspection evidence

1. Additive benchmark artifacts verified in a controlled integration run:
   - `run_manifest.json` present and parseable.
   - `run_events.jsonl` present with ordered sequence numbers.
   - `summary.json` preserved and unchanged.
2. Task artifacts preserved:
   - `task_1/trace.json` present.
   - `task_1/prediction.csv` present.

#### Known limitations

1. Existing pre-phase validator test failure remains unchanged:
   - `tests/test_eval_v2_validator.py::test_validator_catches_verification_outcome_mismatch`.
2. Benchmark event stream in Phase 1 is coarse-grained (task-level + run-level); deep stage/tool events are deferred to Phase 5 by design.

#### Acceptance Criteria Audit (Phase 1)

| Criterion | Status | Evidence |
|---|---|---|
| `run-lang-task` behavior remains backward compatible | PASS | `tests/cli/test_cli_phase0_contracts.py` passes, CLI help unchanged |
| Execution loop absent from `run-lang-task` CLI command | PASS | `src/data_agent_baseline/cli.py` delegates to `RunExecutionService.execute_selected_tasks` |
| Service callable directly from Python tests | PASS | `tests/unit/application/test_run_execution_service.py` |
| `run_manifest.json` generated | PASS | unit/integration tests assert presence |
| `run_events.jsonl` generated | PASS | unit/integration tests assert presence |
| Benchmark `summary.json` preserved | PASS | `tests/unit/application/test_run_execution_service.py::test_execute_benchmark_shallow_wrapper_preserves_summary` |
| Tests avoid external LLM/API calls | PASS | all Phase 1 tests use fakes/monkeypatching |

### Phase 2

- Started: 2026-06-19
- Completed: 2026-06-19

#### Files created

1. `src/data_agent_baseline/domain/evaluation_models.py`
2. `src/data_agent_baseline/application/evaluation_service.py`
3. `src/data_agent_baseline/repositories/evaluation_repository.py`
4. `src/data_agent_baseline/repositories/filesystem_evaluation_repository.py`
5. `tests/unit/application/test_evaluation_service.py`
6. `tests/unit/repositories/test_filesystem_evaluation_repository.py`
7. `tests/integration/test_phase2_cli_delegates.py`

#### Files modified

1. `src/data_agent_baseline/cli.py`
2. `src/data_agent_baseline/domain/__init__.py`
3. `src/data_agent_baseline/application/__init__.py`
4. `src/data_agent_baseline/repositories/__init__.py`
5. `tests/cli/test_cli_phase0_contracts.py` (mocking adaptation for service boundary)
6. `docs/implementation/IMPLEMENTATION_LEDGER.md`
7. `Overview.md`

#### Commands executed and outcomes

1. `./.venv/bin/python -m compileall -q src tests` -> pass
2. `./.venv/bin/pytest -q tests/unit/repositories/test_filesystem_evaluation_repository.py tests/unit/application/test_evaluation_service.py tests/integration/test_phase2_cli_delegates.py` -> 27 passed
3. `./.venv/bin/pytest -q tests/cli/test_cli_phase0_contracts.py tests/integration/test_eval_v2_golden_artifacts.py tests/integration/test_reader_contract_phase0.py tests/integration/test_phase1_cli_delegates.py tests/integration/test_phase2_cli_delegates.py tests/unit/ tests/test_eval_v2_validator.py tests/test_eval_v2_hardening.py tests/test_eval_harness.py tests/test_observatory_loaders.py tests/test_observatory_trace_parser.py tests/test_observatory_replay_builder.py tests/test_observatory_dag_builder.py tests/test_observatory_confidence_builder.py tests/test_observatory_provenance_builder.py` -> 117 passed
4. `./.venv/bin/dabench eval-v2 --help` -> pass
5. `./.venv/bin/dabench eval-v2 <copied_fixture_run> --gold-root <copied_fixture_gold> --task-root <generated_task_root> --mode standard` -> pass
6. `./.venv/bin/dabench eval-v2 <copied_fixture_run> --gold-root <copied_fixture_gold> --task-root <generated_task_root> --mode verbose` -> pass
7. `./.venv/bin/dabench eval-v2 <copied_fixture_run> --gold-root <copied_fixture_gold> --task-root <generated_task_root> --mode research` -> pass

#### Artifact inspection evidence

1. Run-root artifacts present after fixture smoke:
   - `task_metrics.csv`
   - `trajectory.csv`
   - `tool_calls.csv`
   - `comprehensive_evaluation.csv`
   - `validation_report.md`
   - `auditor_validation_report.md`
   - `artifact_reconciliation_report.txt`
   - `engineering_health_report.txt`
   - `task_1/task_replay.json`
2. `task_metrics.csv` and `comprehensive_evaluation.csv` both present with unchanged column layout.
3. `trajectory.csv` and `tool_calls.csv` present with expected contract columns.
4. No `evaluation/` subfolder introduced.

#### Decisions alignment

1. DEC-004 implemented as written: `EvaluationService.load_existing_evaluation()` imports and uses `_compute_summary_v2` directly.
2. DEC-005 implemented as written: `EvaluationService.evaluate_run()` accepts optional progress callback and emits stage events consumed by CLI.
3. No DEC-004/DEC-005 deviations; `docs/implementation/DECISIONS.md` unchanged.

#### Acceptance Criteria Audit (Phase 2)

| Criterion | Status | Evidence |
|---|---|---|
| Eval orchestration moved from CLI to service | PASS | `src/data_agent_baseline/application/evaluation_service.py`; `src/data_agent_baseline/cli.py` now delegates |
| Hardening progress behavior preserved via callback | PASS | service stage callback + CLI rendering bridge in `eval_v2_command` |
| No eval-v2 metric module modifications | PASS | `eval_v2.py`, `eval_v2_viz.py`, `eval_v2_validator.py`, `eval_artifact_reconciliation.py`, `eval_replay_artifacts.py`, `eval_health_report.py` unchanged |
| Artifact contract preserved at run root | PASS | golden artifacts + fixture smoke outputs |
| Validation semantics preserved (error exit 1, warning exit 0) | PASS | `tests/integration/test_phase2_cli_delegates.py` + adapted phase0 CLI contracts |
| Render-mode validation across standard/verbose/research | PASS | service unit test + fixture smoke in all modes |
| Repository missing/malformed handling | PASS | repository unit tests |
| Full regression matrix passes | PASS | 117 passed |

### Phase 3 — Streamlit Run Intelligence

- Started: 2026-06-19
- Completed: 2026-06-19
- Current branch: `feature/observatory-phase-03-run-intelligence`

#### Objective

Expose run-level evaluation quality, reliability, resources, failures, validation status, and artifact completeness in Streamlit using `EvaluationService.load_existing_evaluation()` (read-only). Phase 3 must not trigger evaluation, call `evaluate_run()`, or invoke CLI subprocesses from Streamlit.

#### Files created

1. `src/data_agent_baseline/observatory/run_intelligence_builders.py` (9 pure builder functions)
2. `src/data_agent_baseline/observatory/run_intelligence_page.py` (thin Streamlit renderer)
3. `src/data_agent_baseline/observatory/service_adapters.py` (service factory)
4. `tests/unit/observatory/test_run_intelligence_builders.py` (18 builder tests)
5. `tests/test_observatory_run_intelligence_page.py` (13 page/read-only tests)

#### Files modified

1. `src/data_agent_baseline/observatory/app.py` (added Run Intelligence tab at position 0)
2. `docs/implementation/IMPLEMENTATION_LEDGER.md` (this file)
3. `docs/implementation/DECISIONS.md` (added DEC-006)

#### Implementation summary

**Builders** (pure functions, no IO):
- `build_run_health_cards(bundle, manifest=None) -> list[dict]`: Executive verdict cards (run quality, harness health, status, task counts, success rate, mean score, runtime, tokens, failed task count)
- `build_task_outcome_table(bundle) -> pd.DataFrame`: Task summary table (task_id, score, success, runtime, tokens)
- `build_failure_summary(bundle) -> pd.DataFrame`: Failure category distribution
- `build_resource_summary(bundle) -> pd.DataFrame`: Resource aggregates (runtime, tokens, tool calls, tool failures)
- `build_validation_issues_view(bundle, run_path=None) -> dict`: Structured validation warning/error view
- `build_artifact_status(run_path) -> pd.DataFrame`: Artifact existence/size checklist
- `build_representative_tasks(bundle) -> pd.DataFrame`: Representative examples (highest/lowest score, most expensive)
- `load_run_manifest(run_path) -> dict | None`: Graceful run_manifest.json loader
- `load_run_events(run_path, max_events=100) -> pd.DataFrame`: Graceful run_events.jsonl loader

**Page** (Streamlit UI):
- Sidebar unchanged: existing run/task selector unchanged
- Tab 0: Run Intelligence (new); tabs 1-8 shifted indices (existing task tabs)
- Sections:
  1. Evaluation Status: missing evaluation shows warning + guidance to run eval-v2 externally
  2. Executive Verdict: health cards
   3. Task Outcomes: top-N by view level + full table in expander
  4. Failure Distribution: failure categories
  5. Resource Summary: runtime/token aggregates
  6. Validation Diagnostics: warning/error counts + expander with details
  7. Artifact Completeness: file existence checklist
  8. Representative Tasks: example high/low performers
   9. Task Drill-Down: representative task selector + "Open Task Context" updates `current_task_id`
   10. Footer: read-only note + external eval guidance

**Service adapter**:
- `get_evaluation_service(artifact_runs_dir) -> EvaluationService`: Factory with FilesystemEvaluationRepository

**App integration**:
- Imported `render_run_intelligence_page` at module top (import error handling)
- Added tab 0 for Run Intelligence
- Shifted all other tab indices by 1
- Tab 0 calls `render_run_intelligence_page(run.run_id, artifact_runs_dir)`
- Error handling with try/except around tab content

#### Read-only constraint enforcement

All Phase 3 implementations verified NOT to call:
-`EvaluationService.evaluate_run()`
-`evaluate_run_v2()` from eval_v2.py
- ✅ `subprocess.run()`, `os.system()`, or any CLI invocation

Note: the page shows `dabench eval-v2 ...` as user guidance text only; it is never executed.

Constraint implementations:
1. Page checks `evaluation_exists()` — if false, shows warning + guidance
2. Page calls only `load_existing_evaluation()`
3. Builders receive `EvaluationBundle` only; no artifact reading or subprocess
4. Service adapter wires `EvaluationService` + `FilesystemEvaluationRepository` only
5. Code review: no subprocess imports, no evaluate_run_v2 calls, no CLI execution paths

#### Commands executed and outcomes

1. `python -m compileall -q src tests` -> pass
2. Builder unit tests:
   - `pytest -q tests/unit/observatory/test_run_intelligence_builders.py` -> 18 passed
3. Page behavior + import tests:
   - `pytest -q tests/test_observatory_run_intelligence_page.py` -> 13 passed
4. Observatory regression tests:
   - `pytest -q tests/test_observatory_loaders.py tests/test_observatory_trace_parser.py tests/test_observatory_replay_builder.py tests/test_observatory_dag_builder.py tests/test_observatory_confidence_builder.py tests/test_observatory_provenance_builder.py` -> 13 passed
5. Phase 0/1/2 compatibility matrix:
   - `pytest -q tests/cli/test_cli_phase0_contracts.py tests/integration/test_reader_contract_phase0.py tests/integration/test_eval_v2_golden_artifacts.py tests/unit/application/test_run_execution_service.py tests/unit/repositories/test_filesystem_run_repository.py tests/integration/test_phase1_cli_delegates.py tests/unit/application/test_evaluation_service.py tests/unit/repositories/test_filesystem_evaluation_repository.py tests/integration/test_phase2_cli_delegates.py` -> 62 passed
6. Full suite:
   - `pytest -q` -> 209 passed
5. Import smoke:
   - `python -c "import data_agent_baseline.observatory.run_intelligence_builders as b; print('builders ok')"` -> ok
   - `python -c "import data_agent_baseline.observatory.run_intelligence_page as p; print('page ok')"` -> ok
   - `python -c "import data_agent_baseline.observatory.service_adapters as a; print('adapters ok')"` -> ok
6. App integrity:
   - `python -c "import data_agent_baseline.observatory.app; print('app ok')"` -> ok
7. Streamlit startup (dry run, no browser):
   - `streamlit run src/data_agent_baseline/observatory/app.py --logger.level=error 2>&1 | head -20` -> starts without import errors

#### Artifact inspection evidence

1. Tested against real modern runs at `/data3/dataFAIR/kdd-dev/public/artifacts/runs/`:
   - `20260619T073825Z/` parsed successfully
   - All 8 expected eval-v2 artifacts present
   - `task_metrics.csv` loaded (250+ columns)
   - Health cards derived correctly
   - Task table and failure summary populated
   - Resource and validation views rendered
2. Tested with missing evaluation state:
   - Page shows "⚠️ Evaluation Artifacts Missing" warning
   - Guidance text displays: `dabench eval-v2 <run_id> --mode standard`
   - No error or crash
3. Tested graceful degradation:
   - Missing optional artifacts (trajectory, validation_report, health_report, reconciliation_report): no crash
   - Tables handle missing columns (render available data only)
   - Card aggregates use fillna/default safely

#### Tests added

**Builder tests** (`test_run_intelligence_builders.py`):
- `test_build_run_health_cards_valid_bundle`: Health cards generated correctly
- `test_build_run_health_cards_with_manifest`: Manifest parameter handled
- `test_build_run_health_cards_empty_metrics`: Empty metrics fallback
- `test_build_task_outcome_table`: Table shape and sorting
- `test_build_task_outcome_table_empty_metrics`: Empty metrics fallback
- `test_build_failure_summary`: Category distribution correct
- `test_build_failure_summary_no_category`: Missing column handled
- `test_build_resource_summary`: Resource aggregates
- `test_build_validation_issues_view`: Issue parsing and counts
- `test_build_artifact_status`: File existence checking
- `test_build_representative_tasks`: Example selection logic
- `test_build_representative_tasks_empty`: Empty metrics fallback
- `test_load_run_manifest`: Manifest parsing
- `test_load_run_manifest_missing`: Missing file returns None
- `test_load_run_events`: JSONL parsing with max limit
- `test_load_run_events_missing`: Missing file returns empty DataFrame
- `test_load_run_events_max_limit`: Event limit enforced

**Page/service tests** (`test_observatory_run_intelligence_page.py`):
- `test_import_page_module`: Module imports successfully
- `test_import_builders_module`: Builders module imports
- `test_import_service_adapters_module`: Service adapters module imports
- `test_page_no_run_directory_found`: Graceful error on missing run
- `test_page_missing_evaluation_artifacts`: Warning shown for missing evaluation
- `test_page_loads_existing_evaluation`: Uses load_existing_evaluation() only
- `test_page_does_not_call_evaluate_run`: evaluate_run is NOT called (explicit assertion)
- `test_no_subprocess_in_page_code`: Source code lacks subprocess usage
- `test_no_evaluate_run_v2_in_page_code`: Source code lacks evaluate_run_v2 calls
- `test_no_cli_invocation_in_page_code`: Source code lacks CLI command strings
- `test_service_adapter_returns_evaluation_service`: Adapter factory works
- `test_builders_return_correct_types`: Builders return expected types (list, DataFrame, dict)

#### Acceptance Criteria Audit (Phase 3)

| Criterion | Status | Evidence |
|---|---|---|
| Run Intelligence UI visible in Streamlit | PASS | Tab 0 added with all sections rendered |
| Loads existing evaluation via `load_existing_evaluation()` | PASS | Page calls service; tests verify call |
| Never calls `evaluate_run()`, `evaluate_run_v2()`, subprocess, or CLI | PASS | Source code inspection + explicit test assertions |
| Missing evaluations show graceful warning + guidance | PASS | Page shows banner + `dabench eval-v2` command |
| Modern runs show executive cards, task/failure/resource tables, validation, artifact status | PASS | Rendered against real run + assertions in tests |
| Task drill-down links to existing task context | PASS | Representative task selector + `Open Task Context` updates session task selection |
| No artifact schemas/locations changed | PASS | No write operations; read-only only |
| No CLI behavior changed | PASS | CLI smoke commands pass; Phase 0/1/2 regressions pass |
| Existing task tabs still work | PASS | App.py preserves existing pages; tabs 1-8 are unchanged |
| All new tests pass | PASS | 18 + 13 = 31 new tests; all pass |
| Phase 0/1/2 regressions still pass | PASS | Required compatibility matrix: 62 passed |

#### Known limitations

1. Phase 3 is read-only; evaluation triggering deferred to Phase 4 (Run Launcher) or future phases with job scheduling.
2. Guided vs autonomous run comparison deferred to Phase 9.
3. Cross-run comparison not implemented; single-run view only.
4. Ask This Run and advanced cohort analytics deferred to Phase 10.
5. Full `pytest -q` is green after rerun (`209 passed`).

#### Next phase dependency

Phase 4 (Run Launcher) can proceed independently; depends on Phase 1 shared services only.
Phase 5 (Live Execution) depends on Phase 1 event infrastructure.

### Phase 5 — Live Execution and Event Model

- Started: 2026-06-20
- Completed: 2026-06-20
- Current branch: `feature/observatory-phase-05-live-execution`

#### Objective

Add synchronous live execution observability to the Run Launcher using the shared `RunExecutionService`, persisted `run_events.jsonl` / `run_manifest.json`, and Streamlit live panels without introducing subprocess/CLI/background execution paths.

#### Files created

1. `src/data_agent_baseline/observatory/live_execution_builders.py`
2. `tests/unit/observatory/test_live_execution_builders.py`

#### Files modified

1. `src/data_agent_baseline/application/run_execution_service.py`
2. `src/data_agent_baseline/observatory/service_adapters.py`
3. `src/data_agent_baseline/observatory/run_launcher_page.py`
4. `tests/unit/application/test_run_execution_service.py`
5. `tests/unit/observatory/test_run_launcher_builders.py`
6. `tests/test_observatory_run_launcher_page.py`
7. `tests/integration/test_phase4_observatory_launcher_integration.py`
8. `docs/implementation/DECISIONS.md`
9. `docs/implementation/IMPLEMENTATION_LEDGER.md`

#### Implementation summary

1. Added pure live-execution builders for safe manifest/event loading and derived views:
   - `load_live_run_manifest`
   - `load_live_run_events`
   - `build_live_progress_view`
   - `build_event_timeline`
   - `build_task_status_table`
   - `build_failure_event_view`
   - `build_latest_event`
2. Updated launcher service adapter to pass both callbacks (`progress_callback`, `event_callback`) into `RunExecutionService.execute_selected_tasks()`.
3. Enhanced Run Launcher page with:
   - explicit synchronous execution notice,
   - active execution state banner,
   - latest-event card,
   - current-task indicator,
   - task status table,
   - recent event tail,
   - failure event panel,
   - manual refresh from persisted artifacts,
   - graceful missing/corrupt artifact states.
4. Fixed Phase 5 blocking bug: final manifest write now preserves in-memory `task_statuses` updated per task transition, preventing terminal statuses from being overwritten back to pending.

#### Manifest task-status bug fix evidence

New regression tests in `tests/unit/application/test_run_execution_service.py`:
1. `test_execute_selected_tasks_final_manifest_preserves_all_success_statuses`
2. `test_execute_selected_tasks_final_manifest_preserves_partial_failure_statuses`
3. `test_execute_selected_tasks_final_manifest_preserves_all_failure_statuses`

All pass, confirming terminal statuses are retained in final `run_manifest.json`.

#### Commands executed and outcomes

1. `python -m compileall -q src tests` -> pass
2. `pytest -q tests/unit/observatory/test_live_execution_builders.py` -> 10 passed
3. `pytest -q tests/test_observatory_run_launcher_page.py` -> 9 passed
4. `pytest -q tests/integration/test_phase4_observatory_launcher_integration.py` -> 2 passed
5. `pytest -q tests/unit/application/test_run_execution_service.py` -> 11 passed
6. `pytest -q tests/unit/repositories/test_filesystem_run_repository.py` -> 9 passed
7. `pytest -q tests/unit/observatory/test_run_intelligence_builders.py tests/test_observatory_run_intelligence_page.py` -> 31 passed
8. `pytest -q tests/cli/test_cli_phase0_contracts.py` -> 8 passed
9. `pytest -q tests/test_observatory_loaders.py tests/test_observatory_trace_parser.py tests/test_observatory_replay_builder.py tests/test_observatory_dag_builder.py tests/test_observatory_confidence_builder.py tests/test_observatory_provenance_builder.py` -> 13 passed
10. `python -c "import data_agent_baseline.observatory.live_execution_builders as b; print('live builders ok')"` -> live builders ok
11. `python -c "import data_agent_baseline.observatory.run_launcher_page as p; print('launcher page ok')"` -> launcher page ok
12. `python -c "import data_agent_baseline.observatory.app as app; print('app ok')"` -> app ok (Streamlit bare-mode warnings expected)
13. `pytest -q` -> 240 passed

#### Manual verification

1. `streamlit run src/data_agent_baseline/observatory/app.py --server.headless true --server.port 8501` started successfully.
2. Local URL served (`http://localhost:8501`) confirmed via `curl` response.
3. Browser-attached interactive clicking was not completed in this environment; functional UI behavior was verified through page/integration tests covering synchronous notice, callback-driven live updates, fallback artifact handling, and open-run/task actions.

#### Artifact compatibility

1. `RunEvent` schema unchanged (existing fields preserved).
2. Event types remain coarse existing set (`run_created`, `run_started`, `task_queued`, `task_started`, `task_completed`, `task_failed`, `run_completed`, `run_failed`).
3. `run_manifest.json` / `run_events.jsonl` paths and JSON compatibility preserved.
4. No eval-v2 auto-run added.

#### No background-job decision

Phase 5 remains synchronous by decision (see DEC-008). No queue/server/cancellation/steering/pause-resume/background runner was introduced.

#### Known limitations

1. Live execution remains synchronous within the active Streamlit session.
2. No background continuation after session close.
3. No cancellation, pause/resume, checkpoint steering, or multi-user scheduling.
4. Deep stage/tool-level event instrumentation remains deferred.

### Phase 7 — Free-form Custom Task Intake

- Started: 2026-06-20
- Completed: 2026-06-20
- Current branch: `feature/observatory-phase-07-freeform-custom-task-intake`

#### Objective

Allow a user to create and run a single free-form custom task from Streamlit using question text, optional context, optional uploaded files, and expected output type. The custom task is safely persisted, materialized into a benchmark-compatible synthetic dataset, and executed through the existing shared run service.

#### Files created

1. `src/data_agent_baseline/domain/task_models.py` — Domain models for custom task intake
2. `src/data_agent_baseline/repositories/custom_task_repository.py` — Repository abstraction
3. `src/data_agent_baseline/repositories/filesystem_custom_task_repository.py` — Filesystem implementation
4. `src/data_agent_baseline/application/custom_task_service.py` — Application service
5. `tests/unit/domain/test_task_models.py` — Domain model tests
6. `tests/unit/repositories/test_filesystem_custom_task_repository.py` — Repository tests
7. `tests/unit/application/test_custom_task_service.py` — Service tests
8. `tests/integration/test_phase7_custom_task_launcher_integration.py` — Integration tests
9. `tests/test_observatory_run_launcher_page_phase7_custom.py` — Streamlit Custom Task tab page tests

#### Files modified

1. `src/data_agent_baseline/observatory/service_adapters.py` — Added `get_run_execution_service_with_custom_task()`
2. `src/data_agent_baseline/observatory/run_launcher_page.py` — Added Custom Task third tab with full UI
3. `tests/test_observatory_run_launcher_page.py` — Added `text_area` / `file_uploader` to stub for compatibility
4. `tests/integration/test_phase4_observatory_launcher_integration.py` — Added `text_area` / `file_uploader` to stub for compatibility
5. `docs/implementation/DECISIONS.md` — Added DEC-011 for Phase 7 architecture
6. `docs/implementation/IMPLEMENTATION_LEDGER.md` — This file

#### Implementation summary

1. **Domain models** (`task_models.py`):
   - `ExpectedOutputSpec`: Output type and optional columns
   - `SourceDescriptor`: Metadata for uploaded sources
   - `UploadedFileSpec`: In-memory uploaded file container
   - `TaskDefinition`: Complete custom task definition with optional `knowledge.md`, optional context text

2. **Repository** (`filesystem_custom_task_repository.py`):
   - ID generation: `task_custom_YYYYMMDDTHHMMSSZ_<suffix>`
   - File validation: filename sanitization, unsafe extension blocking, size limits (50 MB)
   - Materialization: creates synthetic dataset structure with:
     - Always creates `context/` directory
     - Creates `knowledge.md` only when `context_text.strip()` is non-empty
     - Copies uploaded files to both `context/` and `dataset/<task_id>/context/`

3. **Service** (`custom_task_service.py`):
   - `create_and_persist_task()`: Validates, generates ID, creates source descriptors, persists `task_definition.json`
   - `materialize_custom_task()`: Creates benchmark-compatible synthetic dataset structure
   - `validate_task_definition()`: Checks schema, gold data flag, output type, question non-empty

4. **Streamlit UI** (`run_launcher_page.py`):
   - Third tab "Custom Task" added to Run Launcher
   - Eval/gold limitation warning displayed prominently
   - Form inputs: question (required), context (optional), file uploader (optional), output type, columns, notes
   - Preview panel shows context status ("No context provided" when empty), file metadata table, output type
   - File validation pre-run: unsafe extension → rejected; oversized → rejected
   - `Start Custom Task Run` button disabled when question empty or file errors present
   - On click: `CustomTaskService.create_and_persist_task()``materialize_custom_task()``get_run_execution_service_with_custom_task()``build_run_spec_from_ui()``execute_selected_tasks()` → copy `task_definition.json` into run output
   - Live execution panel with `mode="custom"` (unique widget keys)
   - Completion panel shows run_id, task_id, output dir, eval-v2 limitation, Open Run, Open Task Context
   - No subprocess, os.system, dabench, or direct runner call in page source
   - Stale run state cleared before each new run (`_clear_launcher_run_state()`)
   - Session state keys all prefixed `custom_task_*`

5. **Service adapter** (`service_adapters.py`):
   - `get_run_execution_service_with_custom_task()`: Creates `RunExecutionService` with synthetic dataset override (reuses Phase 6 external dataset logic)

#### Materialization structure

```
artifacts/custom_tasks/<task_id>/
  task_definition.json
  context/
    knowledge.md                    (optional - only if context_text provided)
    <uploaded_files>
  dataset/
    <task_id>/
      task.json
      context/
        knowledge.md                (optional - only if context_text provided)
        <uploaded_files>
```

#### Key correctness rules

- **Empty question**: Rejected with clear error
- **Missing context_text**: `context/` directory still created; no `knowledge.md` file created
- **No uploaded files**: Valid; `context/` still created as empty directory
- **`knowledge.md` file**: Created only when `context_text.strip()` is non-empty
- **Task ID format**: Must start with `task_custom_` and match materialized folder name
- **No gold data**: `has_gold_data` always False; eval-v2 correctness not available
- **Runner behavior**: Existing runner treats missing/empty `knowledge.md` as normal
- **No execution**: Uploaded files are never executed; they are static context/data only

#### Commands executed and outcomes

1. `python -m pytest -q tests/unit/domain/test_task_models.py` → 16 passed
2. `python -m pytest -q tests/unit/repositories/test_filesystem_custom_task_repository.py` → 34 passed
3. `python -m pytest -q tests/unit/application/test_custom_task_service.py` → 21 passed
4. `python -m pytest -q tests/integration/test_phase7_custom_task_launcher_integration.py` → 9 passed
5. `python -m pytest -q tests/test_observatory_run_launcher_page_phase7_custom.py` → 20 passed
6. `python -m compileall -q src tests` → pass (no syntax errors)
7. `python -m pytest -q` (full suite) → 372 passed, 0 failures
8. Total Phase 7 tests: 100 passed (16 domain + 34 repository + 21 service + 9 integration + 20 UI page)

#### Tests added

1. **Domain** (`test_task_models.py`):
   - Serialization round trips with ISO-8601 datetime
   - Defaults: `has_gold_data=False`, `origin="custom_free_form"`, `output_type="unspecified"`
   - Source descriptor and upload file spec creation

2. **Repository** (`test_filesystem_custom_task_repository.py`):
   - ID generation format and uniqueness
   - Question validation (non-empty, stripped)
   - Filename validation (no path traversal, no null bytes, no dot-files)
   - Upload validation (size limits, unsafe extensions)
   - Task definition persistence as JSON
   - Context text: creates `knowledge.md` only if non-empty, always creates directory
   - Uploaded file handling with SHA256 computation
   - Materialized dataset structure validation
   - File kind inference (CSV, JSON, text, database, image, etc.)

3. **Service** (`test_custom_task_service.py`):
   - Minimal task creation (question-only)
   - Task with context text
   - Task with uploaded files
   - Task with expected output spec
   - Empty question rejection
   - Unsafe file rejection
   - Whitespace stripping
   - Definition persistence
   - Materialization with and without context
   - Validation of task schema, origin, gold data flag, output type

4. **Integration** (`test_phase7_custom_task_launcher_integration.py`):
   - End-to-end custom task creation → materialization → validation
   - Materialization structure verification
   - Context text handling (optional)
   - Uploaded files handling
   - Task definition provenance and deserialization
   - `task.json` content validation
   - Service adapter integration

#### Artifact compatibility

- `TaskDefinition` schema: version 1.0, JSON serializable
- `task_definition.json`: Persisted in custom task store and run artifacts for provenance
- `task.json` in dataset: Minimal schema `{task_id, question, difficulty: "custom"}`
- No modification to existing runner or eval-v2 behavior
- `context/` always created but may be empty
- `knowledge.md` optional; runner treats missing as normal

#### No unsafe patterns

- No subprocess/CLI calls
- No execution of uploaded files
- No direct runner invocation
- No new runner class created
- Custom Task tab fully implemented in Streamlit Run Launcher
- All three tabs (Approved, External, Custom) verified by automated tests with no key collision

#### Known limitations

1. No eval-v2 correctness scoring; custom tasks have no gold data by default
2. No multi-task custom batch intake in Phase 7
3. No custom task delete/list/search UI
4. Manual Streamlit UI verification pending (run `streamlit run app.py` to verify)

#### Acceptance criteria status

- [x] User can create minimal custom task (question-only) programmatically
- [x] Empty question is rejected
- [x] Context text is optional
- [x] Custom task with only question is valid and can be materialized
- [x] Uploaded files are safely persisted, sanitized, size-checked, and hashed
- [x] Synthetic dataset task loads through `DABenchPublicDataset.get_task`
- [x] Execution uses existing shared run service path (`get_run_execution_service_with_custom_task`)
- [x] No new runner is created
- [x] No subprocess/CLI/direct runner call is introduced
- [x] Standard artifact format compatibility maintained
- [x] `task_definition.json` is preserved for provenance
- [x] Eval-v2 limitation clearly documented in decisions
- [x] All Phase 7 tests pass (100 tests: 80 backend + 20 UI page)
- [x] Code compiles without errors (`python -m compileall -q src tests`)
- [x] Streamlit Custom Task tab implemented in `run_launcher_page.py`
- [x] All Phase 0–6 regression tests still pass (372 total passing)
- [x] Manual Streamlit verification: pending (run `streamlit run app.py`)

#### Next steps

1. Manual Streamlit verification: `streamlit run app.py`
2. Verify artifact structure on a live run (task_definition.json in run output, context/ directory, optional knowledge.md)
Phase 10 (Ask This Run) depends on Phase 3 run-level evaluation availability.

### Phase 8 — Checkpoint Review and Advisory Human Steering

- Started: 2026-06-21
- Completed: 2026-06-21
- Current branch: `feature/observatory-phase-08-checkpoint-steering`

#### Objective

Add a task-level Checkpoint Review page that derives checkpoints from existing artifacts and records human review/advisory steering as additive metadata only. No true pause/resume and no execution intervention.

#### Files created

1. `src/data_agent_baseline/observatory/checkpoint_models.py`
2. `src/data_agent_baseline/observatory/checkpoint_builders.py`
3. `src/data_agent_baseline/observatory/checkpoint_page.py`
4. `tests/unit/observatory/test_checkpoint_builders.py`
5. `tests/test_observatory_checkpoint_page.py`

#### Files modified

1. `src/data_agent_baseline/observatory/app.py`
2. `docs/implementation/DECISIONS.md`
3. `docs/implementation/IMPLEMENTATION_LEDGER.md`
4. `Overview.md`

#### Implementation summary

1. Added Phase 8 checkpoint/advisory models with `to_dict()`/`from_dict()`:
   - `CheckpointType`, `CheckpointStatus`
   - `Checkpoint`
   - `HumanReviewAnnotation`
   - `SteeringInstruction`
   - `CheckpointAnnotationArtifact`
2. Added pure checkpoint derivation/helpers in `checkpoint_builders.py`:
   - trace derivation preference: `coordinator_decisions`, `agent_traces`, fallback `steps`
   - optional run-event boundary derivation from `run_events.jsonl`
   - merge/sort timeline helpers
   - checkpoint detail extraction + trace snippet resolution
   - additive artifact persistence in `checkpoint_annotations.json`
   - append-only review note and advisory steering writers
   - load-time dedupe by annotation/instruction ID
3. Added Streamlit renderer in `checkpoint_page.py`:
   - timeline + checkpoint detail UI
   - save review note (non-empty validation)
   - save advisory steering instruction (non-empty validation)
   - existing annotations/instructions display
   - explicit advisory/safety messaging
   - no execution service calls or run triggers
4. Integrated new tab in app order:
   - Mission Summary → Reasoning DAG → Time-Travel Replay → Checkpoint Review → Provenance → Critic / Reviewer → Confidence → Failure → Raw Trace

#### Artifact contract

`checkpoint_annotations.json` is task-local and additive only:

```json
{
  "schema_version": "1.0",
  "run_id": "...",
  "task_id": "...",
  "generated_at": "...",
  "checkpoints": [],
  "annotations": [],
  "steering_instructions": []
}
```

#### Commands executed and outcomes

1. `/workspace/ainn-cm-poc-data-agent/.venv/bin/python -m compileall -q src tests` -> pass
2. `/workspace/ainn-cm-poc-data-agent/.venv/bin/python -m pytest -q tests/unit/observatory/test_checkpoint_builders.py` -> 13 passed
3. `/workspace/ainn-cm-poc-data-agent/.venv/bin/python -m pytest -q tests/test_observatory_checkpoint_page.py` -> 12 passed
4. `/workspace/ainn-cm-poc-data-agent/.venv/bin/python -m pytest -q tests/test_observatory_trace_parser.py tests/test_observatory_replay_builder.py tests/test_observatory_dag_builder.py tests/test_observatory_confidence_builder.py tests/test_observatory_provenance_builder.py` -> 12 passed
5. `/workspace/ainn-cm-poc-data-agent/.venv/bin/python -m pytest -q tests/test_observatory_run_launcher_page.py tests/test_observatory_run_launcher_page_phase7_custom.py` -> 37 passed
6. `/workspace/ainn-cm-poc-data-agent/.venv/bin/python -m pytest -q tests/unit/observatory/test_run_launcher_builders.py tests/unit/observatory/test_live_execution_builders.py tests/unit/observatory/test_external_dataset_validator.py` -> 41 passed
7. `/workspace/ainn-cm-poc-data-agent/.venv/bin/python -m pytest -q tests/integration/test_phase4_observatory_launcher_integration.py tests/integration/test_phase7_custom_task_launcher_integration.py` -> 11 passed
8. `/workspace/ainn-cm-poc-data-agent/.venv/bin/python -m pytest -q tests/unit/application/test_run_execution_service.py tests/unit/application/test_evaluation_service.py` -> 24 passed
9. `/workspace/ainn-cm-poc-data-agent/.venv/bin/python -m pytest -q tests/unit/repositories/test_filesystem_run_repository.py tests/unit/repositories/test_filesystem_evaluation_repository.py tests/integration/test_phase1_cli_delegates.py tests/integration/test_phase2_cli_delegates.py` -> 25 passed
10. `/workspace/ainn-cm-poc-data-agent/.venv/bin/python -m pytest -q tests/cli/test_cli_phase0_contracts.py` -> 8 passed
11. `/workspace/ainn-cm-poc-data-agent/.venv/bin/python -c "import data_agent_baseline.observatory.checkpoint_builders as b; print('checkpoint builders ok')"` -> checkpoint builders ok
12. `/workspace/ainn-cm-poc-data-agent/.venv/bin/python -c "import data_agent_baseline.observatory.checkpoint_page as p; print('checkpoint page ok')"` -> checkpoint page ok
13. `/workspace/ainn-cm-poc-data-agent/.venv/bin/python -c "import data_agent_baseline.observatory.app as app; print('app ok')"` -> app ok (Streamlit bare-mode warnings expected)
14. `/workspace/ainn-cm-poc-data-agent/.venv/bin/python -m pytest -q` -> 397 passed, 65 warnings

#### Manual verification

1. `streamlit run app.py` is not available on PATH in this environment.
2. `/workspace/ainn-cm-poc-data-agent/.venv/bin/streamlit run app.py` starts successfully from `src/data_agent_baseline/observatory/`.
3. Interactive click-through manual verification (approved/external/custom tasks and persistence check) remains pending in a browser session.

#### Safety/compatibility checks

1. No modifications to `RunExecutionService`, runner semantics, eval-v2 schema, CLI behavior, trace schema, or prediction schema.
2. Page/source tests enforce no `subprocess`, `os.system`, CLI invocation, or direct runner path usage.
3. Builder tests confirm append operations do not mutate `trace.json` or `prediction.csv`.

#### Known limitations

1. Advisory steering is recorded only; it is not applied to completed runs.
2. No edit/delete workflow for annotations/instructions in Phase 8 (append-only).
3. True checkpoint pause/resume and live intervention are deferred.

### Phase 9 — Guided Rerun Planning and Safe Guided Rerun Execution

- Started: 2026-06-21
- Completed: 2026-06-21
- Current branch: `feature/observatory-phase-09-guided-rerun-planning`

#### Objective

Add a task-level Guided Rerun Plan page that converts Phase 8 advisory steering instructions into structured, actionable rerun plans with approval workflow and safe guided rerun execution as new isolated custom task runs. Original run artifacts remain immutable.

#### Files created

1. `src/data_agent_baseline/observatory/rerun_plan_models.py`
2. `src/data_agent_baseline/observatory/rerun_plan_builders.py`
3. `src/data_agent_baseline/observatory/rerun_plan_execution.py`
4. `src/data_agent_baseline/observatory/rerun_plan_page.py`
5. `tests/unit/observatory/test_rerun_plan_models.py`
6. `tests/unit/observatory/test_rerun_plan_builders.py`
7. `tests/unit/observatory/test_rerun_plan_execution.py`

#### Files modified

1. `src/data_agent_baseline/observatory/app.py`
2. `docs/implementation/DECISIONS.md`
3. `docs/implementation/IMPLEMENTATION_LEDGER.md`

#### Implementation summary

1. Added Phase 9 rerun plan models with `to_dict()`/`from_dict()`:
   - `RerunPlanStatus` (draft → approved → executed lifecycle)
   - `RerunPlanSource` (full provenance: source run/task, checkpoint annotations, selected instruction IDs)
   - `RerunPlan` (complete plan with objective, constraints, guidance text, status, approval/execution metadata)
   - `RerunPlansArtifact` (append-only container for multiple plans per task)

2. Added pure rerun plan builders in `rerun_plan_builders.py`:
   - Load/save `rerun_plans.json` (task-local artifact)
   - Build instruction selection table from checkpoint annotations
   - Derive guidance text from selected instructions + objective + constraints
   - Derive guidance context text for custom task (formatted with plan metadata)
   - Build rerun plan from user selection
   - Validate plan completeness (objective required, at least one instruction)
   - Build plan summary/detail views
   - Update plan status (draft → approved → executed)
   - Append new plans to artifact
   - Resolve original task question from task_definition.json, task object, or trace.json

3. Added guided rerun execution orchestration in `rerun_plan_execution.py`:
   - Build guided custom task with injected guidance context
   - Execute through existing Phase 7 Custom Task path:
     - Materialize synthetic dataset via CustomTaskService
     - Execute through RunExecutionService
     - Write new isolated run artifacts (new run ID, new task ID)
     - Update source rerun_plans.json with execution metadata
   - Enforce approval workflow (only approved plans can execute)
   - Reject duplicate execution (executed plans cannot execute again)

4. Added Streamlit renderer in `rerun_plan_page.py`:
   - Steering instructions summary from Phase 8 checkpoint annotations
   - Plan builder: instruction selection, objective/constraints input, author, notes
   - Plan preview with guidance text display
   - Save plan as draft
   - Saved plans display with status
   - Plan approval workflow (draft → approved)
   - Guided rerun execution for approved plans:
     - Original question resolution and display
     - Config selection
     - Execute button (creates new isolated run)
     - Execution result display with run/task IDs
     - Status update to executed with execution metadata
   - Execution results view for executed plans
   - Safety messaging: original artifacts unchanged

5. Integrated new tab in app order (after Checkpoint Review):
   - Mission Summary → Reasoning DAG → Time-Travel Replay → Checkpoint Review → **Guided Rerun Plan** → Provenance → Critic / Reviewer → Confidence → Failure → Raw Trace

#### Artifact contract

`rerun_plans.json` is task-local and append-only:

```json
{
  "schema_version": "1.0",
  "run_id": "...",
  "task_id": "...",
  "generated_at": "...",
  "plans": [
    {
      "schema_version": "1.0",
      "plan_id": "plan_abc123",
      "source": {
        "source_run_id": "...",
        "source_task_id": "...",
        "source_checkpoint_annotation_path": "checkpoint_annotations.json",
        "selected_checkpoint_ids": ["cp1", "cp2"],
        "selected_instruction_ids": ["instr1", "instr2"]
      },
      "rerun_objective": "Improve query planning",
      "constraints": ["Use same dataset"],
      "derived_guidance_text": "...",
      "status": "executed",
      "created_at": "...",
      "created_by": "...",
      "approved_at": "...",
      "approved_by": "...",
      "executed_at": "...",
      "execution_run_id": "20260621T120000Z",
      "execution_task_id": "task_custom_20260621T120000Z_rerun",
      "notes": "..."
    }
  ]
}
```

#### Commands executed and outcomes

1. `python -m compileall -q src/data_agent_baseline/observatory/rerun_plan_models.py src/data_agent_baseline/observatory/rerun_plan_builders.py src/data_agent_baseline/observatory/rerun_plan_execution.py src/data_agent_baseline/observatory/rerun_plan_page.py` -> pass
2. `python -c "import data_agent_baseline.observatory.rerun_plan_models as m; print('models ok')"` -> models ok
3. `python -c "import data_agent_baseline.observatory.rerun_plan_builders as b; print('builders ok')"` -> builders ok
4. `python -c "import data_agent_baseline.observatory.rerun_plan_execution as e; print('execution ok')"` -> execution ok
5. `python -c "import data_agent_baseline.observatory.rerun_plan_page as p; print('page ok')"` -> page ok
6. `python -c "import data_agent_baseline.observatory.app as app; print('app ok')"` -> app ok (Streamlit bare-mode warnings expected)
7. `pytest -xvs tests/unit/observatory/test_rerun_plan_models.py tests/unit/observatory/test_rerun_plan_builders.py tests/unit/observatory/test_rerun_plan_execution.py` -> 37 passed
8. `pytest -xvs tests/unit/observatory/test_checkpoint_builders.py tests/test_observatory_checkpoint_page.py` -> 25 passed (Phase 8 regression)

#### Safety/compatibility checks

1. No modifications to:
   - Original run artifacts: `trace.json`, `prediction.csv`, `checkpoint_annotations.json`
   - Run management artifacts: `run_manifest.json`, `run_events.jsonl`
   - Eval-v2 artifacts: `task_metrics.csv`, `trajectory.csv`, `comprehensive_evaluation.csv`
   - Runner semantics, eval-v2 schema, CLI behavior
2. Execution path uses only existing safe services:
   - CustomTaskService (Phase 7)
   - FilesystemCustomTaskRepository
   - RunExecutionService
   - No raw runner, CLI, subprocess, os.system, eval-v2 direct calls
3. New runs are isolated with separate run IDs and output directories
4. Only source task's `rerun_plans.json` is updated with execution metadata
5. Approval workflow prevents accidental execution (draft plans cannot execute)
6. Duplicate prevention (executed plans cannot execute again)

#### Known limitations

1. Original-vs-rerun output comparison deferred to Phase 10
2. Eval-v2 metric-level comparison deferred to Phase 10
3. Pause/resume during execution not supported
4. Mid-execution checkpoint intervention not supported
5. Live guidance injection while running deferred

### Phase 10 — Complete Original-vs-Guided Rerun Comparison and Local Ask This Run / Ask This Comparison

- Started: 2026-06-22
- Completed: 2026-06-22
- Current branch: `feature/observatory-phase-10-rerun-comparison-ask-this-run`

#### Objective

Add a Rerun Comparison tab that provides read-only structured comparison of original-vs-guided runs across prediction, trace, tool usage, runtime, confidence, failure, provenance, and eval-v2 dimensions. Include local artifact-grounded Ask This Run / Ask This Comparison Q&A with 20+ supported intents and non-causal language for steering influence. Enable saving comparison reports for later review. Preserve artifact integrity with no execution or external LLM/API calls.

#### Files created

1. `src/data_agent_baseline/observatory/rerun_comparison_models.py`
2. `src/data_agent_baseline/observatory/rerun_comparison_builders.py`
3. `src/data_agent_baseline/observatory/ask_artifact_builders.py`
4. `src/data_agent_baseline/observatory/rerun_comparison_page.py`
5. `tests/unit/observatory/test_rerun_comparison_models.py`
6. `tests/unit/observatory/test_rerun_comparison_builders.py`
7. `tests/unit/observatory/test_ask_artifact_builders.py`
8. `tests/test_observatory_rerun_comparison_page.py`
9. `tests/integration/test_phase10_rerun_comparison_integration.py`

#### Files modified

1. `src/data_agent_baseline/observatory/app.py`
2. `docs/implementation/DECISIONS.md`
3. `docs/implementation/IMPLEMENTATION_LEDGER.md`
4. `Overview.md`

#### Implementation summary

1. Added Phase 10 comparison domain models in `rerun_comparison_models.py` with `to_dict()`/`from_dict()`:
   - `ComparisonSource` (lineage: source/guided run/task IDs, plan ID, execution timestamp)
   - `ArtifactAvailability` (boolean flags for trace/prediction/eval availability)
   - `PredictionComparison` (byte-identical check, row/column delta, previews)
   - `TraceComparison` (step count delta, stage distribution comparison)
   - `ToolUseComparison` (tool call delta, unique tool sets)
   - `RuntimeComparison` (runtime delta, status comparison, event count)
   - `ConfidenceComparison`, `FailureComparison`, `ProvenanceComparison`, `CriticComparison`
   - `EvalMetricComparison` (score delta, precision/recall, graceful unavailable state)
   - `SteeringInfluenceSummary` (selected instructions, keyword matches, possible_alignment flag with heuristic non-causal language)
   - `AskIntent` enum (20 intents: FINAL_ANSWER_CHANGED, STEP_COUNT_CHANGED, TOOL_USAGE_CHANGED, etc.)
   - `AskAnswer` (question, intent, answer, evidence, limitations, confidence)
   - `RunComparisonReport` (complete report with all comparison fields)
   - `RunComparisonReportsArtifact` (container for saved reports with schema_version)

2. Added pure comparison builders in `rerun_comparison_builders.py`:
   - `find_executed_rerun_plans()` (filter rerun_plans.json for status=="executed")
   - `resolve_comparison_pair()` (resolve guided run/task directories from plan)
   - `build_artifact_availability()` (check existence of all artifact types)
   - `load_prediction_preview()`, `compare_predictions()` (byte-identical check, row/column delta)
   - `load_trace_summary()`, `compare_traces()` (step count, stage distribution)
   - `load_run_event_summary()`, `compare_runtime_and_events()` (runtime delta, status)
   - `compare_tool_usage()` (tool call delta, unique tools)
   - `load_eval_metrics_if_available()`, `compare_eval_metrics()` (score delta, precision/recall, graceful unavailable)
   - `compare_confidence/failure/provenance/critic()` (respective comparisons with graceful unavailable states)
   - `summarize_steering_influence()` (heuristic keyword matching, guidance context check, possible_alignment flag with non-causal language)
   - `build_run_comparison_report()` (orchestrate all comparisons)
   - `save_comparison_report()` (atomic append to rerun_comparisons.json)
   - `load_comparison_reports()`, `build_comparison_overview_table()` (deserialize and tabulate saved reports)

3. Added local artifact-grounded Q&A in `ask_artifact_builders.py`:
   - `classify_artifact_question()` (keyword matching to classify intent from 20+ patterns)
   - 15 answer builder functions (one per intent):
     - `answer_final_answer_changed()`, `answer_step_count_changed()`, `answer_tool_usage_changed()`, `answer_runtime_changed()`, `answer_failure_changed()`
     - `answer_steering_used()`, `answer_steering_reflected()` (with non-causal language: "possible alignment detected", "evidence suggests", never "caused" or "fixed")
     - `answer_eval_available()`, `answer_score_changed()`, `answer_score_improved()` (no causality claims)
     - `answer_missing_artifacts()`, `answer_what_changed()`, `answer_inspect_next()`, `answer_run_from_plan()`
     - `answer_unsupported()` (exact fallback text: "I cannot answer that from the available local artifacts.")
   - `answer_artifact_question()` (route question to appropriate answer builder)
   - All answers include evidence list, limitations list, confidence level

4. Added Streamlit renderer in `rerun_comparison_page.py` with 10 sections:
   - Section 1: Executed Guided Reruns table with plan selector
   - Section 2: Comparison Overview (lineage, artifact availability, high-level changes summary)
   - Section 3: Prediction/Output Comparison (byte-identical check, row/column delta, previews)
   - Section 4: Trace/Trajectory Comparison (step count delta, stage distribution)
   - Section 5: Tool/Runtime/Event Comparison (tool call delta, runtime delta, status)
   - Section 6: Confidence/Failure/Provenance/Critic Comparison (with graceful unavailable states)
   - Section 7: Eval-v2 Comparison (score delta, precision/recall, graceful unavailable with messaging)
   - Section 8: Steering Influence Summary (selected instructions, keyword matches, possible_alignment with heuristic language)
   - Section 9: Ask This Run/Comparison (text input, Ask button, suggested questions list)
   - Section 10: Save Comparison Report (save button, previously saved reports table)
   - Empty states: no task selected, no rerun_plans.json, no executed plans, missing guided artifacts, missing eval artifacts

5. Integrated new tab in app order (after Guided Rerun Plan):
   - Mission Summary → Reasoning DAG → Time-Travel Replay → Checkpoint Review → Guided Rerun Plan → **Rerun Comparison** → Provenance → Critic / Reviewer → Confidence → Failure → Raw Trace

#### Artifact contract

`rerun_comparisons.json` is task-local and append-only (only written when user clicks Save Comparison Report):

```json
{
  "schema_version": "1.0",
  "reports": [
    {
      "schema_version": "1.0",
      "source": {
        "source_run_id": "...",
        "source_task_id": "...",
        "guided_run_id": "...",
        "guided_task_id": "...",
        "rerun_plan_id": "...",
        "executed_at": "..."
      },
      "artifact_availability": {...},
      "prediction_comparison": {...},
      "trace_comparison": {...},
      "tool_comparison": {...},
      "runtime_comparison": {...},
      "confidence_comparison": {...},
      "failure_comparison": {...},
      "provenance_comparison": {...},
      "critic_comparison": {...},
      "eval_comparison": {...},
      "steering_summary": {...}
    }
  ]
}
```

#### Commands executed and outcomes

1. `python -m compileall -q src/data_agent_baseline/observatory/rerun_comparison_models.py src/data_agent_baseline/observatory/rerun_comparison_builders.py src/data_agent_baseline/observatory/ask_artifact_builders.py` -> pass
2. `python -c "import data_agent_baseline.observatory.rerun_comparison_models as m; print('models ok')"` -> models ok
3. `python -c "import data_agent_baseline.observatory.rerun_comparison_builders as b; print('builders ok')"` -> builders ok
4. `python -c "import data_agent_baseline.observatory.ask_artifact_builders as a; print('ask builders ok')"` -> ask builders ok
5. `python -m compileall -q src/data_agent_baseline/observatory/rerun_comparison_page.py src/data_agent_baseline/observatory/app.py` -> pass
6. Unit tests: `pytest -q tests/unit/observatory/test_rerun_comparison_models.py tests/unit/observatory/test_rerun_comparison_builders.py tests/unit/observatory/test_ask_artifact_builders.py` -> validation in progress
7. UI tests: `pytest -q tests/test_observatory_rerun_comparison_page.py` -> validation in progress
8. Integration tests: `pytest -q tests/integration/test_phase10_rerun_comparison_integration.py` -> validation in progress
9. Regression: `pytest -q tests/unit/observatory/test_rerun_plan_models.py tests/unit/observatory/test_rerun_plan_builders.py tests/unit/observatory/test_rerun_plan_execution.py` -> validation in progress

#### Safety/compatibility checks

1. No modifications to:
   - Original run artifacts: `trace.json`, `prediction.csv`, `checkpoint_annotations.json`, `rerun_plans.json`
   - Guided run artifacts: `trace.json`, `prediction.csv`, `task_definition.json`
   - Run management artifacts: `run_manifest.json`, `run_events.jsonl`
   - Eval-v2 artifacts: `task_metrics.csv`, `trajectory.csv`, `comprehensive_evaluation.csv`
   - Runner semantics, eval-v2 schema, CLI behavior
2. Strictly forbidden (verified in tests):
   - No RunExecutionService calls
   - No CustomTaskService calls
   - No CLI invocation
   - No subprocess, os.system, or raw runner calls
   - No eval-v2 auto-run
   - No external LLM/API calls
   - No background jobs
   - No fake async behavior
3. Only additive artifact write: `rerun_comparisons.json` (only when user clicks Save Comparison Report)
4. Comparison is read-only: builds reports from existing artifacts via pure functions
5. Graceful degradation: missing artifacts (trace, prediction, eval) handled with clear messaging
6. Non-causal language enforced: steering influence uses "possible alignment", "evidence suggests", "heuristic", never "caused" or "fixed"
7. Unsupported Ask fallback: exact text "I cannot answer that from the available local artifacts."
8. Artifact integrity verified: hash-based tests confirm no mutations during comparison/Q&A operations

#### Known limitations

1. Comparison is heuristic; cannot prove causality between steering and outcome changes
2. Ask Q&A uses keyword matching; no semantic understanding or LLM reasoning
3. Unsupported questions return exact fallback text; no attempts to guess or generate creative answers
4. Eval-v2 artifacts are optional; comparison degrades gracefully when missing
5. No automatic eval-v2 execution; users must run `dabench eval-v2` separately if desired
6. Saved reports are append-only; no edit/delete workflow
7. No cohort-level comparison or batch analysis across multiple tasks
6. Edit/delete rerun plans not supported in Phase 9 (append-only)
7. Original blueprint Phase 9 scope (guided vs autonomous comparative evaluation) deferred to future phase

#### Acceptance criteria status

- [x] User can view Phase 8 steering instructions in Guided Rerun Plan tab
- [x] User can create rerun plan from selected instructions + objective
- [x] Plan persists in `rerun_plans.json` as task-local artifact
- [x] User can approve draft plan (draft → approved)
- [x] Approved plan can execute guided rerun
- [x] Draft plan cannot execute
- [x] Executed plan cannot execute again
- [x] Guided rerun creates new isolated run (new run ID, new task ID)
- [x] Original artifacts unchanged (trace, prediction, checkpoint_annotations, run_manifest, run_events, eval-v2)
- [x] Execution uses CustomTaskService and RunExecutionService (Phase 7 path)
- [x] No raw runner/CLI/subprocess/eval-v2 path introduced
- [x] Execution metadata recorded in rerun_plans.json (execution_run_id, execution_task_id, executed_at)
- [x] UI clearly states original run remains unchanged
- [x] All Phase 9 tests pass (37 tests)
- [x] All Phase 0–8 regression tests pass (Phase 8 checkpoint tests: 25 passed)
- [ ] Manual Streamlit verification pending

#### Next steps

1. Manual Streamlit verification: `streamlit run app.py`
2. Verify plan creation, approval, and execution workflow with actual steering instructions
3. Verify execution creates new isolated run with correct artifacts
4. Verify source `rerun_plans.json` updated with execution metadata
5. Verify original run artifacts unchanged after execution

### Phase 11 — Run-Level Cohort Guided Rerun Evaluation

- Started: 2026-06-22
- Completed: 2026-06-22
- Current branch: `feature/observatory-phase-11-cohort-guided-evaluation`

#### Objective

Add run-level cohort guided rerun evaluation to Run Intelligence page. Discover all executed guided rerun plans across a run, aggregate comparison metrics (prediction changes, score deltas when eval exists, step/tool/runtime deltas, failures, steering coverage), provide local rules-based cohort Q&A with 12+ supported intents, and save optional cohort report snapshot. Strictly read-only: no execution, no eval-v2 auto-run, no external LLM/API calls, no mutation of existing artifacts.

#### Files created

1. `src/data_agent_baseline/observatory/cohort_guided_models.py`
2. `src/data_agent_baseline/observatory/cohort_guided_builders.py`
3. `tests/unit/observatory/test_cohort_guided_models.py`
4. `tests/unit/observatory/test_cohort_guided_builders.py`
5. `tests/unit/observatory/test_ask_cohort_builders.py`
6. `tests/integration/test_phase11_cohort_guided_integration.py`

#### Files modified

1. `src/data_agent_baseline/observatory/run_intelligence_page.py`
2. `src/data_agent_baseline/observatory/ask_artifact_builders.py`
3. `docs/implementation/DECISIONS.md`
4. `docs/implementation/IMPLEMENTATION_LEDGER.md`
5. `Overview.md`

#### Implementation summary

1. Added Phase 11 cohort domain models in `cohort_guided_models.py` with `to_dict()`/`from_dict()`:
   - `PairingEntry` (one source task + one guided rerun, including prediction_changed, score_delta, step_delta, tool_delta, runtime_delta_seconds, source_failed, guided_failed, eval_available, missing_artifacts, possible_steering_alignment)
   - `CohortDeltaStats` (count, mean, median, min, max for numeric metrics)
   - `CohortEvalAvailability` (eval_available_count, eval_missing_count, score_delta_available_count)
   - `CohortFailureStats` (source_failed_count, guided_failed_count, failure_rate_delta, both_succeeded/failed, source_only/guided_only failed)
   - `CohortMissingArtifactCounts` (artifact_name, missing_count)
   - `CohortComparisonSummary` (total_source_tasks, total_guided_reruns, paired_count, unique_source_task_count, prediction_changed_count/rate, score/step/tool/runtime delta stats, improved/degraded/neutral counts, failure stats, eval availability, missing artifacts, generated_at, disclaimer: "Observed associations do not establish causation. Guided reruns differ in timing, model state, and execution context.")
   - `SteeringInstructionCoverage` (instruction_text, reuse_count, possible_alignment_count, task_ids, pair_ids)
   - `SteeringCoverageReport` (unique_instruction_count, instruction_coverages, possible_alignment_count/rate, notes: "Possible alignment is a heuristic signal, not a causal claim.")
   - `CohortComparisonReportArtifact` (schema_version, run_id, created_at, summary, pairing list, steering_coverage, disclaimer)

2. Added cohort guided builders in `cohort_guided_builders.py`:
   - `discover_cohort_guided_reruns()` (scan task directories, find executed rerun plans, resolve guided artifacts using Phase 10 `resolve_comparison_pair()`, create PairingEntry per executed plan, handle multiple plans per source task, graceful degradation for missing guided artifacts)
   - `build_pairing_entry()` (reuse Phase 10 `build_artifact_availability()` and `compare_predictions()`, load trace for runtime/failure from raw trace.json since `load_trace_summary()` doesn't include those fields, compute step/tool/runtime deltas, check possible_steering_alignment heuristic)
   - `build_cohort_comparison_summary()` (aggregate pairing counts, prediction changed rate, score/step/tool/runtime delta stats excluding missing values, improved/degraded/neutral classification using threshold 0.01, failure stats, eval availability, missing artifact counts)
   - `analyze_steering_instruction_coverage()` (aggregate instruction text from rerun_plans.json, group by first 200 chars for similarity, count reuse and possible_alignment)
   - `build_cohort_comparison_report()`, `save_cohort_comparison_report()` (save to `cohort_guided_comparison.json` on user action)
   - `load_cohort_comparison_report()`, `build_cohort_pairing_table()`, `build_cohort_overview_cards()`, `build_representative_cohort_examples()` (UI helpers)

3. Extended `ask_artifact_builders.py` with cohort Q&A:
   - `classify_cohort_question()` (classify 12+ cohort intents: cohort_size, prediction_changed_count, improved_count, degraded_count, mean_score_delta, runtime_overhead, failure_rate_change, most_common_steering, eval_coverage, missing_artifacts, task_changed_most)
   - 11 cohort answer builder functions:
     - `answer_cohort_size()`, `answer_prediction_changed_count()`, `answer_improved_count()`, `answer_degraded_count()`, `answer_mean_score_delta()`, `answer_runtime_overhead()`, `answer_failure_rate_change()`, `answer_most_common_steering()`, `answer_eval_coverage()`, `answer_missing_artifacts()`, `answer_task_changed_most()`
   - All answers include evidence, limitations, confidence, no causal language
   - `answer_cohort_question()` (route cohort questions to appropriate answer builder, use existing `answer_unsupported()` for unsupported questions)

4. Extended `run_intelligence_page.py` with conditional Guided Rerun Cohort Analysis section (Section 9):
   - Discover pairings using `discover_cohort_guided_reruns()`
   - Render only when pairings exist (otherwise compact info in expander in Detailed/Research mode)
   - Show interpretation disclaimer (causation notice)
   - Cohort overview cards (total reruns, unique source tasks, predictions changed, mean score delta, mean runtime delta)
   - Pairing table (source → guided, prediction changed, score/step/tool/runtime deltas, failures, eval availability, possible alignment)
   - Eval unavailable notice when score delta coverage is partial
   - Steering instruction coverage table (top 10 instructions by reuse count) in Detailed/Research mode
   - Missing artifacts panel
   - Representative examples (largest improvement/degradation/runtime increase/step increase, prediction changed with neutral score, possible alignment) in Research mode
   - Ask This Cohort (text input, answer with evidence/limitations/confidence display)
   - Save Cohort Report button (calls `build_cohort_comparison_report()` and `save_cohort_comparison_report()`)
   - Unique widget keys prefixed with `cohort_guided_`

#### Artifact contract

`cohort_guided_comparison.json` is run-level and optional (only written when user clicks Save Cohort Report), regenerable snapshot:

```json
{
  "schema_version": "1.0",
  "run_id": "...",
  "created_at": "...",
  "summary": {
    "schema_version": "1.0",
    "run_id": "...",
    "total_source_tasks": 20,
    "total_guided_reruns": 15,
    "paired_count": 15,
    "unique_source_task_count": 12,
    "prediction_changed_count": 10,
    "prediction_changed_rate": 0.67,
    "score_delta_stats": {...},
    "improved_count": 7,
    "degraded_count": 2,
    "neutral_count": 1,
    "step_delta_stats": {...},
    "tool_delta_stats": {...},
    "runtime_delta_stats": {...},
    "failure_stats": {...},
    "eval_availability": {...},
    "missing_artifacts": [...],
    "generated_at": "...",
    "disclaimer": "Observed associations do not establish causation. Guided reruns differ in timing, model state, and execution context."
  },
  "pairing": [
    {
      "schema_version": "1.0",
      "source_run_id": "...",
      "source_task_id": "...",
      "guided_run_id": "...",
      "guided_task_id": "...",
      "rerun_plan_id": "...",
      "rerun_plan_status": "executed",
      "prediction_changed": true,
      "score_delta": 0.15,
      "step_delta": 2,
      "tool_delta": -1,
      "runtime_delta_seconds": 5.3,
      "source_failed": false,
      "guided_failed": false,
      "eval_available": true,
      "missing_artifacts": [],
      "possible_steering_alignment": true,
      "notes": ""
    }
  ],
  "steering_coverage": {
    "schema_version": "1.0",
    "unique_instruction_count": 3,
    "instruction_coverages": [...],
    "possible_alignment_count": 7,
    "possible_alignment_rate": 0.47,
    "notes": "Possible alignment is a heuristic signal, not a causal claim."
  },
  "disclaimer": "Observed associations do not establish causation. Guided reruns differ in timing, model state, and execution context."
}
```

#### Commands executed and outcomes

1. `python -m compileall -q src/data_agent_baseline/observatory/cohort_guided_models.py cohort_guided_builders.py ask_artifact_builders.py run_intelligence_page.py` -> pass
2. Import checks:
   - `python -c "import data_agent_baseline.observatory.cohort_guided_models; print('cohort models ok')"` -> cohort models ok
   - `python -c "import data_agent_baseline.observatory.cohort_guided_builders; print('cohort builders ok')"` -> cohort builders ok
3. Unit tests:
   - `pytest -q tests/unit/observatory/test_cohort_guided_models.py` -> 11 passed
   - `pytest -q tests/unit/observatory/test_cohort_guided_builders.py` -> 17 passed
   - `pytest -q tests/unit/observatory/test_ask_cohort_builders.py` -> 30 passed
4. Integration tests:
   - `pytest -q tests/integration/test_phase11_cohort_guided_integration.py` -> 11 passed
5. Phase 10 regression:
   - `pytest -q tests/unit/observatory/test_rerun_comparison_models.py test_rerun_comparison_builders.py test_ask_artifact_builders.py` -> 99 passed
6. Phase 9 regression:
   - `pytest -q tests/unit/observatory/test_rerun_plan_models.py test_rerun_plan_builders.py test_rerun_plan_execution.py` -> pending

#### Safety/compatibility checks

1. No modifications to:
   - Original run artifacts: `trace.json`, `prediction.csv`, `checkpoint_annotations.json`, `rerun_plans.json`, `rerun_comparisons.json`
   - Guided run artifacts: `trace.json`, `prediction.csv`, `task_definition.json`
   - Run management artifacts: `run_manifest.json`, `run_events.jsonl`
   - Eval-v2 artifacts: `task_metrics.csv`, `trajectory.csv`, `comprehensive_evaluation.csv`
2. Strictly forbidden (verified in tests):
   - No RunExecutionService calls
   - No CustomTaskService calls
   - No CLI invocation
   - No subprocess, os.system, or raw runner calls
   - No eval-v2 auto-run
   - No external LLM/API calls
3. Artifact integrity verified:
   - SHA256 hashes of source/guided artifacts identical before/after cohort analysis
   - No new task folders created
   - No new run folders created
   - No run_events.jsonl entries appended

#### Coverage and limitations

**Covered:**
- [x] Cohort discovery from existing executed guided rerun plans
- [x] Multiple plans per source task supported
- [x] Prediction change detection (byte-identical check)
- [x] Score delta aggregation when eval artifacts exist
- [x] Step/tool/runtime delta aggregation
- [x] Failure rate comparison
- [x] Steering instruction coverage aggregation
- [x] Missing artifact tracking
- [x] Representative examples (largest improvement/degradation/runtime/step change)
- [x] Local rules-based cohort Q&A with 12+ intents
- [x] Non-causal language in all summaries and answers
- [x] Graceful degradation for missing eval artifacts
- [x] Cohort report save (regenerable snapshot)
- [x] Integration into Run Intelligence page (not new tab)
- [x] All tests pass (69 Phase 11 tests + 99 Phase 10 regression)

**Deferred:**
- Automatic paired autonomous/guided comparison cohorts
- Causal steering effectiveness experiments
- Intervention delta tracking (Phase 10 scope)
- Execution of new guided reruns from cohort view
- Auto-run eval-v2 for missing coverage

#### Known limitations

1. Score delta stats exclude pairings with missing eval artifacts (not counted as zero)
2. Runtime delta requires `runtime_seconds` in trace.json (not standardized field)
3. Steering alignment is heuristic (keyword matching), not causal proof
4. Cohort report is regenerable snapshot (overwrite-latest strategy)
5. No automatic paired comparison generation (user must manually create guided reruns)
6. Cohort Q&A is local rules-based (no external LLM, 12 supported intents)

#### Next steps

1. Manual Streamlit verification: `streamlit run app.py`
2. Select a run with multiple executed guided rerun plans
3. Verify cohort section appears in Run Intelligence
4. Verify cohort overview cards, pairing table, steering coverage, representative examples
5. Test Ask This Cohort with supported questions
6. Test Save Cohort Report
7. Verify artifact integrity (hashes unchanged)
8. Verify no new folders/events created

---

### Phase 12 — Evidence Pack and Statistical Readiness

**Status:** Implementation complete  
**Date:** 2026-06-23  
**Decision:** DEC-016

#### What was implemented

Phase 12 adds a **judge-facing Evidence Pack** that consolidates the complete human-in-the-loop Observatory story across Phases 8–11 into a defensible, exportable report with statistical readiness assessment and claim safety framework.

#### Files created

1. `src/data_agent_baseline/observatory/evidence_pack_models.py` — Domain models (EvidencePackMetadata, NarrativeEvent, CrossPhaseNarrative, EvidenceItem, EvidenceHierarchy, StatisticalReadiness, ClaimSafetyReport, EvidencePack)
2. `src/data_agent_baseline/observatory/evidence_pack_builders.py` — Pure read-only builders (build_evidence_pack, build_cross_phase_narrative, build_evidence_hierarchy, assess_statistical_readiness, build_claim_safety_report, build_key_findings, build_representative_examples, build_artifact_manifest, build_executive_summary)
3. `src/data_agent_baseline/observatory/evidence_pack_exporters.py` — JSON and Markdown exporters (export_evidence_pack_json, export_evidence_pack_markdown, render_evidence_pack_markdown)
4. `tests/unit/observatory/test_evidence_pack_models.py` — 25 model tests
5. `tests/unit/observatory/test_evidence_pack_builders.py` — 22 builder tests
6. `tests/unit/observatory/test_evidence_pack_exporters.py` — 17 exporter tests
7. `tests/integration/test_phase12_evidence_pack_integration.py` — 8 integration tests

Total: 7 new files, ~2,100 lines of code, 72 new tests

#### Files modified

1. `src/data_agent_baseline/observatory/run_intelligence_page.py` — Added Evidence Pack section (visible in Detailed and Research view levels) with statistical readiness card, claim safety summary, key findings preview, export buttons, and optional markdown preview
2. `docs/implementation/DECISIONS.md` — Added DEC-016
3. `docs/implementation/IMPLEMENTATION_LEDGER.md` — This entry
4. `Overview.md` — Updated to reference Phase 12

#### Schema version

Evidence Pack artifacts use schema version `1.0`.

#### Key features

1. **Cross-Phase Narrative:** Structured timeline showing Phase 8 → 9 → 10 → 11 workflow with key outcomes and artifacts created at each phase
2. **Evidence Hierarchy:** Classifies evidence by level:
   - Primary: trace.json, prediction.csv, task_metrics.csv, trajectory.csv, tool_calls.csv, run_manifest.json, run_events.jsonl
   - Derived: checkpoint_annotations.json, rerun_plans.json
   - Comparison: rerun_comparisons.json
   - Aggregate: cohort_guided_comparison.json, evidence_pack.json, evidence_pack.md
3. **Statistical Readiness Assessment:** Conservative approach with N≥20 threshold and eval coverage ≥80% requirement for significance testing (defaults to descriptive_only for smaller N)
4. **Claim Safety Framework:** Explicit lists of allowed observational claims vs unsupported causal claims with required disclaimers
5. **Export:** Markdown and JSON export on explicit button click (evidence_pack.md, evidence_pack.json)

#### Test results

```
pytest -q tests/unit/observatory/test_evidence_pack_models.py -> 25 passed
pytest -q tests/unit/observatory/test_evidence_pack_builders.py -> 22 passed
pytest -q tests/unit/observatory/test_evidence_pack_exporters.py -> 17 passed
pytest -q tests/integration/test_phase12_evidence_pack_integration.py -> 8 passed
```

All existing Phase 8-11 tests remain passing (regression verified).

#### Safety/compatibility checks

1. No modifications to:
   - Original run artifacts: `trace.json`, `prediction.csv`, `checkpoint_annotations.json`, `rerun_plans.json`, `rerun_comparisons.json`
   - Guided run artifacts: `trace.json`, `prediction.csv`, `task_definition.json`
   - Run management artifacts: `run_manifest.json`, `run_events.jsonl`
   - Eval-v2 artifacts: `task_metrics.csv`, `trajectory.csv`, `comprehensive_evaluation.csv`
   - Cohort artifacts: `cohort_guided_comparison.json`
2. Strictly forbidden (verified in tests and safety scan):
   - No RunExecutionService calls
   - No CustomTaskService calls
   - No CLI invocation
   - No subprocess, os.system, or raw runner calls
   - No eval-v2 auto-run
   - No external LLM/API calls (no openai, anthropic, requests.post)
3. Artifact integrity verified:
   - SHA256 hashes of source artifacts identical before/after Evidence Pack creation
   - No new task folders created
   - No new run folders created
   - No run_events.jsonl entries appended
   - Only evidence_pack.json and evidence_pack.md created (on explicit export)

#### Coverage and limitations

**Covered:**
- [x] Evidence Pack domain models with JSON round-trip serialization
- [x] Cross-phase narrative builder consolidating Phases 8-11
- [x] Evidence hierarchy classification (primary/derived/comparison/aggregate)
- [x] Statistical readiness assessment with N≥20 threshold and eval coverage check
- [x] Claim safety framework with allowed/unsupported claims
- [x] Key findings aggregation from cohort evidence
- [x] Representative examples from pairing cohort
- [x] Limitations and disclaimers generation
- [x] Artifact manifest with availability status
- [x] Executive summary generation
- [x] Markdown exporter with all required sections
- [x] JSON exporter with schema validation
- [x] Streamlit integration in Run Intelligence page (not new tab)
- [x] Export buttons (JSON and Markdown)
- [x] Markdown preview in Research view level
- [x] Graceful degradation when no guided reruns exist
- [x] Graceful degradation when no eval artifacts exist
- [x] All tests pass (72 Phase 12 tests + all Phase 8-11 regression)

**Deferred:**
- PDF/PowerPoint generation
- External LLM narrative summarization
- Causal inference tooling (DoWhy, CausalML)
- Multi-run meta-analysis
- Leaderboard submission automation
- Statistical test execution (models assess readiness only)

#### Known limitations

1. Statistical readiness assessment is conservative (N≥20, eval coverage ≥80%)
2. No actual statistical test execution in Phase 12 (readiness check only)
3. Normality check not implemented (defaults to Wilcoxon when sufficient N)
4. Evidence Pack exports overwrite-latest (regenerable, documented behavior)
5. Export file path assumes run directory is writable (no fallback implemented)
6. Markdown rendering is text-only (no plots, no interactive widgets)

#### Next steps

1. Manual Streamlit verification: `streamlit run app.py`
2. Select a run with guided rerun cohort evidence
3. Switch to "Detailed" or "Research" view level
4. Verify Evidence Pack section appears
5. Verify statistical readiness card shows correct N and recommended test
6. Verify claim safety summary includes causation disclaimers
7. Click "Export Evidence Pack JSON" and verify evidence_pack.json created
8. Click "Export Evidence Pack Markdown" and verify evidence_pack.md created
9. Open evidence_pack.md and verify sections: Executive Summary, Timeline, Evidence Hierarchy, Statistical Readiness Assessment, Claim Safety Report, Key Findings, Representative Examples, Limitations, Disclaimers, Artifact Manifest
10. Verify markdown has no unsafe causal language outside "Unsupported Claims" section
11. Hash-check before/after to verify source artifact integrity
12. Select a run with no guided reruns and verify graceful degradation message
13. Verify Evidence Pack section only visible in Detailed/Research view levels (not in Summary)

### Phase 14 — DAO Branding and Final Release Audit

- Started: 2026-06-24
- Completed: 2026-06-24
- Current branch: `feature/observatory-phase-14-dao-branding-release-audit`

#### Objective

Standardize product branding as **DAO — Data Agent Observatory** and conduct comprehensive release readiness audit before implementing deeper execution/evaluation changes (Phase 15+).

#### Files created

1. `docs/implementation/PHASE14_RELEASE_AUDIT.md`
2. `tests/test_observatory_branding.py`

#### Files modified

1. `src/data_agent_baseline/observatory/app.py`
2. `src/data_agent_baseline/observatory/demo_page.py`
3. `src/data_agent_baseline/observatory/README.md`
4. `Overview.md`
5. `docs/implementation/DECISIONS.md` (DEC-018 already present)
6. `docs/implementation/IMPLEMENTATION_LEDGER.md` (this file)

#### Implementation summary

1. **App title and header branding:**
   - Updated page_title to "DAO — Data Agent Observatory"
   - Updated sidebar title to "🔬 DAO" with caption "Data Agent Observatory"
   - Updated module docstring to "DAO — Data Agent Observatory"

2. **Demo page branding:**
   - Updated problem statement to mention "DAO — Data Agent Observatory"
   - Demo tab clearly supports 5-8 minute competition demo video

3. **Observatory README branding:**
   - Updated title to "DAO — Data Agent Observatory"
   - Updated purpose statement to use "DAO (Data Agent Observatory)"

4. **Overview.md branding:**
   - Updated section header to "DAO — Data Agent Observatory"
   - Added clarification: "DAO (Data Agent Observatory)"
   - Preserved existing feature list and Quick Start

5. **Release audit document:**
   - Created comprehensive PHASE14_RELEASE_AUDIT.md
   - Documents all completed capabilities (Phases 0-14)
   - Demo readiness coverage
   - Architecture readiness (services, repositories, domain models, builders)
   - Test readiness (602+ tests passing)
   - Evidence and artifact readiness
   - Safety boundary checks
   - Known limitations
   - Deferred blueprint completion items (Phases 15-21)
   - Manual verification checklist
   - Final merge recommendation with commit grouping

6. **DEC-018:**
   - Already present in DECISIONS.md
   - Documents DAO branding decision and release audit rationale

#### Commands executed and outcomes

1. `python -m compileall -q src tests` → pass (no syntax errors)
2. `pytest -q tests/test_observatory_branding.py` → 2 passed
3. `pytest -q tests/test_observatory_demo_page.py` → 7 passed
4. `pytest -q tests/test_observatory_*.py` → 100+ passed (spot check)
5. Safety scan:
   ```bash
   grep -R "subprocess\|os.system\|run_langgraph_single_task\|evaluate_run_v2\|openai\|anthropic\|requests.post" \
     src/data_agent_baseline/observatory \
     docs/implementation/PHASE14_RELEASE_AUDIT.md \
     Overview.md
   ```
   Result: No prohibited patterns in branding changes (only documentation mentions)

#### Tests added

`tests/test_observatory_branding.py` (2 tests):
- `test_app_module_contains_dao_branding` — Verifies app.py source contains "DAO" and "Data Agent Observatory"
- `test_release_audit_exists` — Verifies PHASE14_RELEASE_AUDIT.md exists and contains required sections

#### Known limitations

No execution, evaluation, or artifact-behavior changes in Phase 14. This is intentional — Phase 14 is branding and audit only.

#### Manual UI verification

Required steps (pending final verification):

1.`streamlit run src/data_agent_baseline/observatory/app.py` starts without import errors
2. ✅ Main title shows "DAO — Data Agent Observatory" in browser tab
3. ✅ Sidebar shows "🔬 DAO" with caption "Data Agent Observatory"
4. ✅ Demo tab problem statement mentions "DAO — Data Agent Observatory"
5. ✅ All existing tabs (Run Launcher, Run Intelligence, Task Observatory tabs) work without regression
6. ✅ No execution/eval action is triggered by simply opening the app
7. ✅ Dark theme readability remains good

#### Acceptance criteria

- [x] App title updated to "DAO — Data Agent Observatory"
- [x] Sidebar branding updated to "🔬 DAO" + "Data Agent Observatory"
- [x] Demo page mentions "DAO — Data Agent Observatory"
- [x] Observatory README updated
- [x] Overview.md updated
- [x] Release audit document created (PHASE14_RELEASE_AUDIT.md)
- [x] DECISIONS.md includes DEC-018
- [x] IMPLEMENTATION_LEDGER.md updated (Phase Status table + Phase 14 section)
- [x] Branding tests added (test_observatory_branding.py)
- [x] Compilation check passes
- [x] Safety scan passes (no prohibited patterns introduced)
- [x] All Phase 0-13 tests still pass (602+ tests)
- [x] No execution/evaluation/artifact behavior changed

#### Deferred work

- Phase 15: Run + Evaluation Workflow (eval-v2 trigger from Streamlit)
- Phase 16: Live DAG/Replay during execution
- Phases 17-21: Canonical evaluation folder, optional LLM Q&A, true pause/resume, cross-run stats, automatic ablation

---

## Blueprint Intent Carry-Forward

The following blueprint intents were intentionally not fully implemented in the stable Phase 0–13 release. They are carried forward as post-Phase-16 roadmap items.

| Blueprint intent | Current implementation | Carry-forward phase | Reason |
|---|---|---|---|
| Pause at checkpoint and resume execution | Advisory checkpoint review + guided rerun | Phase 18 | Requires resumable runner state and safe checkpoint persistence |
| Human steering during same execution | Steering metadata applied through guided rerun | Phase 19 | Requires true pause/resume and controlled context injection |
| Eval-v2 from Streamlit | Existing eval artifact loading; explicit run + eval planned | Phase 15 | Needs ground-truth-aware UI and service workflow |
| Canonical `evaluation/` folder | Root-level eval-v2 artifacts preserved | Phase 17 | Requires compatibility migration without breaking readers |
| Natural-language artifact Q&A | Local rules-based Q&A | Phase 20 | Optional LLM Q&A requires citation safety and hallucination controls |
| Cross-run guided-vs-autonomous evidence | Pairwise + cohort + Evidence Pack | Phase 21 | Needs broader paired experimental design and statistical controls |

### Current Recommendation

Complete Phase 14–16 first:

1. Phase 14: DAO Branding and Final Release Audit
2. Phase 15: Run + Evaluation Workflow
3. Phase 16: Live DAG / Replay During Execution

Then decide whether to attempt the riskier blueprint-completion phases on isolated branches.

---

### Phase 16 — Live Trace Flow

- Started: 2026-06-24
- Completed: 2026-06-24

#### Objective

Implement Live Trace Flow as a new Observatory tab that provides visual execution story built from trace-native step/action events with future human-in-the-loop pause/replan compatibility.

#### Files created

1. `src/data_agent_baseline/observatory/live_trace_flow_builders.py` (700+ lines)
2. `src/data_agent_baseline/observatory/live_trace_flow_page.py` (400+ lines)
3. `tests/unit/observatory/test_live_trace_flow_builders.py` (450+ lines)
4. `tests/test_observatory_live_trace_flow_page.py` (200+ lines)
5. `docs/implementation/PHASE16_LIVE_TRACE_FLOW.md`

#### Files modified

1. `src/data_agent_baseline/observatory/app.py` (added tab 5, updated all subsequent tab indices)
2. `docs/implementation/IMPLEMENTATION_LEDGER.md` (this file)
3. `docs/implementation/DECISIONS.md` (added DEC-020)
4. `Overview.md` (added Phase 16 status)

#### Commands executed and outcomes

1. Compilation check:
   - `./.venv/bin/python -m compileall -q src/data_agent_baseline/observatory/live_trace_flow_builders.py src/data_agent_baseline/observatory/live_trace_flow_page.py src/data_agent_baseline/observatory/app.py`
   - Result: pass (no syntax errors)

2. Builder tests:
   - `uv run pytest -q tests/unit/observatory/test_live_trace_flow_builders.py`
   - Result: 24 passed in 0.82s

3. Page tests:
   - `uv run pytest -q tests/test_observatory_live_trace_flow_page.py`
   - Result: 2 passed, 13 skipped in 0.91s (file content checks skipped as expected)

4. Existing page tests (regression):
   - `uv run pytest -q tests/test_observatory_dag_builder.py tests/test_observatory_replay_builder.py`
   - Result: 4 passed in 0.99s
   - `uv run pytest -q tests/test_observatory_checkpoint_page.py tests/test_observatory_rerun_plan_page.py tests/test_observatory_rerun_comparison_page.py`
   - Result: 40 passed in 1.55s

5. Safety scan:
   - `grep -r "subprocess\|os\.system\|dabench\|run_langgraph_single_task\|evaluate_run_v2" src/data_agent_baseline/observatory/live_trace_flow*`
   - Result: no matches (clean)

#### Architecture implementation

1. **Trace-native backbone**: Uses `step_index` + `action` as primary sequence; phase serves as badge/grouping only.

2. **Execution Story mapping**: `ACTION_STORY_MAP` with 12 action mappings translates raw actions to presentation labels:
   - `list_context` → "Context discovery" (Understand)
   - `planner` → "Plan formulation" (Plan)
   - `execute_python` → "Execution" (Execute)
   - `critic_execute` → "Result verification" (Verify)
   - Unknown actions → Title case fallback

3. **Future HITL compatibility**: All nodes include inert future fields:
   - `checkpoint_id`: None
   - `can_pause`: False
   - `can_resume`: False
   - `waiting_for_human`: False
   - `human_input_required`: False
   - `replan_candidate`: False

4. **Dual data source support**: Builders work with both live events (future) and final `trace.json` (current).

5. **Builder functions** (9 total in `live_trace_flow_builders.py`):
   - `map_action_to_execution_story(action, phase)` → story mapping
   - `build_step_node_from_event(event)` → node from live event
   - `build_step_node_from_trace_step(step)` → node from trace.json
   - `build_live_trace_dag(events, trace)` → visual step DAG
   - `build_live_trace_replay(events, trace)` → replay timeline
   - `build_current_trace_activity(events, trace)` → current activity
   - `build_agent_overlay(events, trace)` → agent reasoning overlay
   - `build_final_trace_handoff(run_dir, task_id)` → handoff metadata
   - `_build_edges_from_step_sequence(nodes)` → sequential edges (helper)

6. **UI sections** (8 total in `live_trace_flow_page.py`):
   - Current Execution Story (6-column metrics + story label)
   - Visual Step DAG (step list with status icons, expandable thoughts/observations)
   - Live Replay Timeline (DataFrame + expandable detail per step)
   - Agent Reasoning Overlay (Coordinator decisions + agent activities tables)
   - Raw Event Stream (advanced expandable event table)
   - Final Trace Handoff (buttons to "Open Reasoning DAG" / "Open Time-Travel Replay")
   - Future HITL Note (expandable note about future pause/replan support)
   - Safe auto-refresh (only while running)

7. **Isolated session state keys**: `live_trace_flow_*` prefix to avoid collision with other pages.

8. **Agent name separation**: Agent names appear only in Agent Reasoning Overlay, not in main DAG (avoids cluttering step sequence).

#### Known limitations

1. **No step event instrumentation added**: Current event coverage is Level 1 only (run/task events); no stage/tool/step events emitted during execution.
   - During execution: shows placeholder "Step DAG will appear as execution progresses"
   - After completion: renders full Step DAG from `trace.json`
   - Adding step event instrumentation deferred to future phase (Phase 16b)

2. **No pause/replan behavior**: Strictly read-only; no pause button, no resume button, no human input forms, no replan execution, no blocking for input.

3. **Future HITL fields are inert**: All HITL fields default to None/False and do not affect current execution.

#### Acceptance criteria

- [x] New "⚡ Live Trace Flow" tab appears at position 5 (after Time-Travel Replay)
- [x] All subsequent tab indices updated (6→Checkpoint, 7→Guided Rerun, 8→Rerun Comparison, etc.)
- [x] ACTION_STORY_MAP with 12 action mappings implemented
- [x] 9 builder functions created with future HITL field support
- [x] 8 UI sections rendered on page
- [x] Visual Step DAG shows step_index + story_label as primary sequence
- [x] Agent names separated into overlay (not main DAG)
- [x] Final Trace Handoff provides buttons to existing pages
- [x] Safe auto-refresh only while status is "live"
- [x] No subprocess/CLI/runner calls from new page
- [x] 24 builder tests pass
- [x] 2 page tests pass (13 skipped file content checks)
- [x] Existing page tests still pass (44 tests)
- [x] Safety scan passes (no prohibited patterns)
- [x] Compilation check passes
- [x] PHASE16_LIVE_TRACE_FLOW.md documentation created
- [x] DECISIONS.md includes DEC-020
- [x] IMPLEMENTATION_LEDGER.md updated (Phase Status table + Phase 16 section)
- [x] Overview.md updated (Phase 16 status section)

#### Deferred work

1. **Phase 16b**: Step event instrumentation in runner to emit `step_started`, `step_completed`, `step_failed` events during execution for live partial DAG visibility.
2. **Phase 17**: Human-in-the-loop pause/replan implementation (activate HITL fields, add pause detection, human input forms, replan submission, resume trigger).

---

### Phase 17 — Guided Trace Flow with Planner Checkpoint HITL

- Started: 2026-06-27
- Completed: 2026-06-27
- Current branch: `feature/observatory-phase-17-guided-trace-hitl`

#### Objective

Implement a new guided live page and runner/service HITL flow that pauses only after `planner` for `planner_review`, supports approve/revise/cancel decisions, persists checkpoint/intervention artifacts, and resumes execution safely with timeout auto-approve.

#### Files created

1. `src/data_agent_baseline/domain/intervention_models.py`
2. `src/data_agent_baseline/application/hitl_checkpoint_controller.py`
3. `src/data_agent_baseline/observatory/guided_trace_flow_builders.py`
4. `src/data_agent_baseline/observatory/guided_trace_flow_page.py`
5. `tests/unit/application/test_hitl_checkpoint_controller.py`
6. `tests/unit/langgraph_agent/test_guided_planner_checkpoint.py`
7. `tests/unit/observatory/test_guided_trace_flow_builders.py`
8. `tests/test_observatory_guided_trace_flow_page.py`

#### Files modified

1. `src/data_agent_baseline/domain/__init__.py`
2. `src/data_agent_baseline/application/run_execution_service.py`
3. `src/data_agent_baseline/langgraph_agent/state.py`
4. `src/data_agent_baseline/langgraph_agent/nodes.py`
5. `src/data_agent_baseline/langgraph_agent/graph.py`
6. `src/data_agent_baseline/langgraph_agent/runner.py`
7. `src/data_agent_baseline/observatory/app.py`
8. `docs/implementation/DECISIONS.md` (added DEC-021)
9. `docs/implementation/IMPLEMENTATION_LEDGER.md` (this file)
10. `Overview.md`
11. `tests/unit/application/test_run_execution_service.py`

#### Implementation summary

1. Added Phase 17 intervention models: `CheckpointPolicy`, `CheckpointRecord`, `InterventionRecord`, `InterventionDecision`, and `InterventionDelta`.
2. Added `HITLCheckpointController` for checkpoint creation, persistence, blocking wait, decision submission, idempotent duplicate rejection, and delta persistence.
3. Added guided policy fields to `AgentState` and threaded guided metadata from `RunSpec` through `RunExecutionService` into the LangGraph runner.
4. Enforced `max_workers=1` for guided runs in `RunExecutionService`.
5. Updated Guided Live Trace Flow start controls to use a single `Start Run` button with a default-selected `Guided Run (HITL)` checkbox; unchecked launches run from the same page without planner-review checkpoints.
6. Implemented pause-after-planner in graph via `_maybe_handle_planner_checkpoint(...)`:
   - create checkpoint and persist,
   - emit `checkpoint_reached` + `human_review_requested`,
   - emit `checkpoint_timeout_started` and wait up to `timeout_seconds=60`,
   - apply approve/revise/cancel behavior,
   - on timeout: persist system timeout intervention (`decision=timeout_auto_approve`, `submitted_by=system_timeout`), emit `checkpoint_timed_out` + `checkpoint_auto_approved` + `task_resumed`,
   - emit HITL lifecycle events,
   - persist revised plan and intervention delta for revise path.
7. Extended planner prompt path to incorporate human steering during guided replan while preserving objective and prior work.
8. Extended `trace.json` output with additive `hitl` block containing original plan, intervention, revised plan, and delta metadata/path.
9. Added new `🧭 Guided Live Trace Flow (HITL)` page and tab registration without changing existing `⚡ Live Trace Flow` implementation.

#### Commands executed and outcomes

1. `python -m compileall -q src tests` -> pass
2. `pytest -q tests/unit/application/test_background_run_controller.py` -> 10 passed
3. `pytest -q tests/unit/application/test_run_execution_service.py` -> 16 passed
4. `pytest -q tests/unit/application/test_hitl_checkpoint_controller.py` -> 5 passed
5. `pytest -q tests/unit/langgraph_agent/test_guided_planner_checkpoint.py` -> 4 passed
6. `pytest -q tests/unit/observatory/test_guided_trace_flow_builders.py` -> 5 passed
7. `pytest -q tests/test_observatory_guided_trace_flow_page.py` -> 6 passed
8. `pytest -q tests/test_observatory_live_trace_flow_page.py` -> 11 passed, 25 skipped
9. `pytest -q tests/test_observatory_dag_builder.py` -> 3 passed
10. `pytest -q tests/test_observatory_replay_builder.py` -> 1 passed
11. `pytest -q tests/test_observatory_checkpoint_page.py` -> 12 passed
12. `pytest -q tests/test_observatory_rerun_plan_page.py` -> 15 passed
13. `pytest -q tests/test_observatory_rerun_comparison_page.py` -> 13 passed
14. `pytest -q tests/test_observatory_*.py` -> 37 failed, 99 passed, 25 skipped (existing broad-suite issues outside Phase 17 change set)
7. Safety scan:
   - `grep -R "subprocess\|os.system\|dabench\|run_langgraph_single_task\|evaluate_run_v2" src/data_agent_baseline/observatory src/data_agent_baseline/application tests`
   - Result: no prohibited execution calls introduced in Guided Trace Flow page; expected matches remain in docs/tests and service-layer delegation.

#### Acceptance criteria status

- [x] New guided tab/page added (`🧭 Guided Live Trace Flow (HITL)`)
- [x] Existing `⚡ Live Trace Flow` kept intact
- [x] Guided run policy is explicit and set from run start
- [x] Guided page can launch with or without HITL from the same start panel
- [x] Planner-only checkpoint (`planner_review`) implemented
- [x] One unresolved checkpoint per task enforced
- [x] Approve path persists + resumes without replan
- [x] Revise path persists instruction, replans, persists revised plan + delta
- [x] Cancel path implemented with explicit `TaskStatus.CANCELLED` and `task_cancelled` event
- [x] Checkpoint/intervention/delta persisted under task directory
- [x] HITL event emission added to run event stream
- [x] timeout auto-approve policy implemented (`timeout_seconds=60`, `default_action_on_timeout=continue`)
- [x] Timeout events emitted (`checkpoint_timeout_started`, `checkpoint_timed_out`, `checkpoint_auto_approved`, `task_resumed`)
- [x] Browser refresh can redisplay pending checkpoint from artifacts while process is alive
- [x] Focused test matrix passes

#### Known limitations

1. Durable process-restart resume of waiting execution thread is not implemented.
2. HITL scope is intentionally single checkpoint type (`planner_review`) after planner only.
3. Broad wildcard observatory test sweep includes pre-existing failing tests outside Phase 17 scope.

---