File size: 61,096 Bytes
03c7cdd
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
947
948
949
950
951
952
953
954
955
956
957
958
959
960
961
962
963
964
965
966
967
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
984
985
986
987
988
989
990
991
992
993
994
995
996
997
998
999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
# FATHOM β€” Antigravity Recovery Brief

**Project:** FATHOM (Meta Γ— PyTorch Γ— Hugging Face OpenEnv Hackathon Grand Finale, Bangalore, Apr 25–26 2026)
**Author of this brief:** triage handoff to Antigravity
**Today:** 2026-04-26 (final submission day)
**Repo root:** `C:\Users\prath\OneDrive\Desktop\Hackathons\Meta_finale\`
**Branch:** `master` (HF git remote is the source of truth; GitHub mirror not yet wired)
**HF user:** `Pratham-math`
**Live env Space:** <https://Pratham-math-fathom-env.hf.space> (CPU-basic, Docker SDK, RUNNING)
**Trained model repo:** <https://huggingface.co/Pratham-math/fathom-1.5b-grpo> (33 files, plots + adapters + merged_16bit live there)
**W&B run:** <https://wandb.ai/pratham-alwar05-indian-institute-of-information-technolo/huggingface/runs/sy1tqun0>

> **Recommended Antigravity model: Claude Sonnet 4.6** for bulk multi-file edits (deploy scripts, README, Streamlit, notebook, prompt-format alignment). **Escalate to Claude Opus 4.7** for the reward-design redesign in `rewards/compose.py` and the `train/grpo.py` prompt rewrite β€” those two need the hardest reasoning. **Gemini 2.5 Pro** is a fine substitute for either if you prefer Google models. Avoid running Haiku on this β€” the diagnosis below has too many interlocking pieces for a small model to keep coherent.

---

## 0. Read This First (Do Not Skip)

You are taking over a hackathon submission with **roughly 6–10 hours left** before judging. Code, training, and deployment are mostly built; **what's broken is small in surface area but huge in optics.** A judge clicking the README will see a flat-line reward curve, a JSON-only `/healthz` endpoint when they hit the Space URL, and a missing demo video link. Your job is to fix those three things in that order and stop. Do not refactor, do not rename, do not "clean up" β€” every change should be load-bearing for one of the four problems in Β§1.

**Hard rules for this session:**

1. **Do not retrain on a fresh A100 unless Β§2.A explicitly tells you to.** Cloud GPU minutes cost real money and the user has ~$10 of HF credits left. Every code change for Β§2.A must pass a **CPU dry-run** (no model weights, just shape + format checks) before you propose a re-train.
2. **Do not delete the existing failed run's evidence.** `outputs/plots/grpo_reward.png` is currently a flat line. Keep it, but reframe it in the README as a "v1, diagnosed" curve, then publish a v2 next to it once the fix lands. Honest evidence > deleted evidence.
3. **Stay on the existing stack.** No `pip install -U` of `unsloth`, `trl`, `transformers`, `vllm`, `peft`, or `bitsandbytes`. The pinned versions in `pyproject.toml` are the ones the smoke test passed on. If a library upgrade tempts you, the answer is no.
4. **Commit at every milestone with conventional-commit prefixes** (`fix(grpo):`, `feat(space):`, `docs(readme):`). The user reviews PRs by reading commits, not diffs.
5. **All HF pushes go to existing repos** (`Pratham-math/fathom-env`, `Pratham-math/fathom-1.5b-grpo`, `Pratham-math/fathom-code`). Do not create new HF repos.
6. **If you cannot reproduce a problem, say so.** Do not "fix" things by guessing. Every fix in Β§2 has a reproduction script you can run.

---

## 1. The Four Problems, Ranked

| # | Problem | Severity | Time to fix | Judges-visible? |
|---|---------|----------|-------------|----------------|
| **A** | GRPO reward curve is a flat line at 0.0 for every step | **CRITICAL** | 2–3 h (incl. retrain) | Yes β€” README plot |
| **B** | HF Space has no UI; judges see `{"detail":"Not Found"}` at root | **HIGH** | 1.5–2 h | Yes β€” first impression |
| **C** | Missing materials: mini-blog, video, slides, GitHub mirror | **HIGH** | 1.5 h | Yes β€” non-negotiable rubric items |
| **D** | The README implies the model "uses recursion in training"; it doesn't | **MEDIUM** | 30 min | Yes β€” judges may grep |

---

## 2. Problem A β€” Flat Reward Curve

### A.1 What the user sees

`outputs/plots/grpo_reward.png` (and the same file at <https://huggingface.co/Pratham-math/fathom-1.5b-grpo/blob/main/plots/grpo_reward.png>) is a horizontal line at y=0.0 across all logged GRPO steps. The reward never moves. Even though SFT loss looks healthy (3.2 β†’ 0.29) and SFT token accuracy hits 0.93, the GRPO phase teaches the model nothing.

### A.2 What is actually happening (root cause, verified)

I extracted training metrics from `job9b_full.log` (UTF-16-encoded; convert with `iconv -f UTF-16LE -t UTF-8`). **Every** logged GRPO step looks like this:

```python
{'loss': 0.0,
 'completions/mean_length': 3.25,        # ← model outputs ~3 tokens like "the man."
 'completions/min_length': 3.0,
 'completions/max_length': 4.0,
 'rewards/_instrumented_reward_fn/mean': 0.0,   # ← every generation scores 0
 'reward_std': 0.0,                              # ← all 8 generations identical reward
 'frac_reward_zero_std': 1.0,                    # ← GRPO advantage is 0 for 100% of examples
 'kl': 0.0,
 'clip_ratio/region_mean': 0.0,
 ...}
```

The chain of failure is:

1. The SFT-warm-started Qwen-1.5B is fine-tuned on `data/sft_traces.jsonl`, where each user message is shaped:

   ```
   Question: <q>

   [Document excerpt]:
   <ctx>
   ```

   The assistant target ends with `<answer>...</answer>`.

2. GRPO loads that SFT adapter, then `train/grpo.py:218-230` builds an entirely **different** user message shape:

   ```
   Context:
   <ctx_truncated>

   <example.prompt>
   ```

   Note the order is reversed (`Context:` first vs. `Question:` first), the field names differ (`Context:` vs. `[Document excerpt]:`), and the SFT-trained model has never seen this layout.

3. Confronted with an out-of-distribution prompt, the model collapses to the lowest-loss continuation it knows: bare 2–13 token answer spans like `the man.` or `silver.` β€” never wrapped in `<answer>...</answer>`.

4. `rewards/format_gate.py` is a **multiplicative** gate. Missing `<answer>` tag β†’ `format_gate=0.0` β†’ `compose.py:32` short-circuits the entire composite to `0.0`.

5. All 8 GRPO generations score exactly 0.0. **GRPO advantage = (reward βˆ’ group mean) / group std = 0 / 0 β†’ 0.** Gradient is therefore 0. The policy never moves. **Straight line forever.**

This is a single diagnosis with two compounding causes: **(i) prompt-shape drift between SFT and GRPO** and **(ii) a multiplicative gate with no soft floor**. Either one alone would degrade learning; together they zero it out. The recent commits (`fix(grpo): align sys_msg with SFT`, `flip Path 3 -> Path 1`) addressed the system message but not the user-content shape, and not the gate.

### A.3 Required fixes (apply all three; they are not redundant)

**Fix A.3.1 β€” Align the GRPO user-message shape with SFT.** Edit `train/grpo.py`:

```python
# train/grpo.py β€” replace lines 218-239 (the _to_prompt function)

def _to_prompt(example: dict) -> dict:
    ctx_full = example.get("context", "") or ""
    ctx_truncated = _truncate_to_tokens(ctx_full, ctx_budget_tok)
    # CRITICAL: must match data/sft_traces.jsonl user-message shape exactly.
    # SFT used "Question: <q>\n\n[Document excerpt]:\n<ctx>". Any deviation
    # puts the SFT-warm-started policy out-of-distribution and collapses
    # generation length to ~3 tokens (verified job9b_full.log).
    user_content = (
        f"{example.get('prompt', '')}\n\n"
        f"[Document excerpt]:\n{ctx_truncated}"
    )
    msgs = [
        {"role": "system", "content": sys_msg},
        {"role": "user", "content": user_content},
    ]
    prompt_str = tokenizer.apply_chat_template(
        msgs, tokenize=False, add_generation_prompt=True
    )
    return {
        "prompt": prompt_str,
        "gold_answer": str(example.get("gold_answer", "")),
        "prompt_token_count": int(example.get("context_length", 0)) // 4,
        "llm_call_count": 0,
    }
```

Also replace the `sys_msg` string (currently lines 187–192) with the **exact** system message that appears in `data/sft_traces.jsonl`, which is:

```
You are FATHOM, a recursive language model with a Python REPL sandbox. You can read a long document via the variable `ctx` and call `llm(prompt, chunk)` for sub-queries. Think step by step. Emit your final answer inside <answer>...</answer>.
```

(You can grep this from the first line of `data/sft_traces.jsonl` to confirm.) **Do not paraphrase.** Byte-identical or the SFT adapter will not transfer.

**Verification of A.3.1 (CPU-only, no GPU):**

```bash
python - <<'PY'
import json
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-Coder-1.5B-Instruct")
sft = json.loads(open("data/sft_traces.jsonl", encoding="utf-8").readline())
sft_msgs = sft["messages"][:2]  # system + user
sft_str = tok.apply_chat_template(sft_msgs, tokenize=False, add_generation_prompt=True)

# Now build the GRPO-side equivalent from data/train.jsonl row 0
row = json.loads(open("data/train.jsonl", encoding="utf-8").readline())
sys_msg = sft_msgs[0]["content"]
user = f"{row['prompt']}\n\n[Document excerpt]:\n{row['context'][:2000]}"
grpo_msgs = [{"role":"system","content":sys_msg},{"role":"user","content":user}]
grpo_str = tok.apply_chat_template(grpo_msgs, tokenize=False, add_generation_prompt=True)
# The system+user prefix should match byte-for-byte up to where the contexts differ
print("PREFIX_MATCH:", sft_str[:400] == grpo_str[:400])
PY
```

You should see `PREFIX_MATCH: True`. If False, the system message or chat template formatting still differs β€” keep iterating until True.

**Fix A.3.2 β€” Soft-format reward instead of a binary gate.** Edit `rewards/compose.py`. Replace `compose_reward_single` body so that missing `<answer>` no longer zeroes the composite; instead it loses a 0.10 bonus and gets bounded above by 0.05 (existing A-02 cap remains). This gives GRPO a non-zero gradient even when the policy is initially mis-formatted.

```python
# rewards/compose.py β€” replace compose_reward_single (lines 19-63)
def compose_reward_single(
    completion: str,
    gold_answer: str,
    prompt_token_count: int,
    llm_call_count: int,
    cfg_reward: Any,
) -> float:
    """Composite reward β€” soft format bonus instead of multiplicative gate.

    REW-02 v2: GRPO collapses when a multiplicative gate yields std=0 across
    a group (every generation scores 0). Replace with an additive 0.10
    format bonus so even malformed generations carry signal, then keep the
    A-02 correctness==0 cap at 0.05 to block format-only exploits.
    """
    has_format = format_gate(completion) == 1.0
    c = correctness(completion, gold_answer) if has_format else 0.0
    t = token_budget(
        completion,
        prompt_token_count,
        alpha=float(cfg_reward.alpha),
        variant=str(cfg_reward.token_budget_variant),
    )
    r = recursion_efficiency(
        int(llm_call_count),
        max_calls=int(cfg_reward.get("max_calls", 2))
        if hasattr(cfg_reward, "get")
        else int(getattr(cfg_reward, "max_calls", 2)),
    )
    w = cfg_reward.weights
    assert abs(float(w.correctness) + float(w.token_budget) + float(w.recursion_efficiency) - 1.0) < 1e-3

    composite = (
        float(w.correctness) * c
        + float(w.token_budget) * t
        + float(w.recursion_efficiency) * r
    )
    # NEW: small additive format bonus β€” the only signal when the model is
    # still learning the template. Keeps GRPO advantages non-zero.
    if has_format:
        composite += 0.10

    # Anti-hacking caps:
    # 1. correctness==0 β†’ at most 0.05 (blocks format-only exploit)
    if c == 0.0:
        return min(composite, 0.05 + (0.10 if has_format else 0.0))
    return composite
```

Update `REWARD_AUDIT.md` so the A-01 row reflects the new ceiling: `<answer></answer>` (empty answer with format) now scores at most **0.10** (the bonus alone, since correctness=0 and the cap is `0.05 + 0.10 = 0.15`; refine your cap math accordingly). Re-run `pytest tests/test_rewards.py -k "audit"` and confirm everything still passes; fix the assertions in the audit tests if their expected values shift.

**Fix A.3.3 β€” Add a regex-driven assertion before `trainer.train()` runs.** This is your insurance against the bug coming back silently. In `train/grpo.py`, just before `trainer.train()`:

```python
# Pre-flight: tokenize one example and confirm the chat-template prefix is
# the byte-identical match of an SFT trace prefix. If not, the SFT adapter
# is loaded but the policy will be out-of-distribution and reward will
# collapse (root cause of the v1 flat-line run).
import json as _json
_sft = _json.loads(open(str(cfg.data.sft_traces_path) if hasattr(cfg.data, "sft_traces_path") else "data/sft_traces.jsonl", encoding="utf-8").readline())
_sft_prefix = tokenizer.apply_chat_template(_sft["messages"][:2], tokenize=False, add_generation_prompt=True)[:200]
_grpo_first = train_dataset[0]["prompt"][:200]
assert _sft_prefix.split("Question:")[0] == _grpo_first.split("Question:")[0], (
    "SFT/GRPO chat-template prefix drift detected β€” see ANTIGRAVITY_BRIEF.md Β§A.3.1"
)
```

The exact split key may need tweaking depending on the tokenizer output; the goal is "if the system block diverges, raise loudly."

---

### A.4-bis β€” Real recursion-efficiency signal (the "actually teach the model to plan recursion" patch)

**Why this exists.** With only A.3.1–A.3.3, GRPO will start moving but it's optimizing for "produce a correctly-formatted exact-match answer." It is not optimizing for *when to recurse vs. when not to*. The current `recursion_efficiency` reward is dead β€” `train/grpo.py:239` hardcodes `llm_call_count=0` for every example, so that 5% weight is constant across all 8 generations and contributes zero variance. We're going to wake it up by:

1. **Parsing the model's completion** for `llm(` calls inside fenced Python code blocks
2. **Coupling the efficiency bonus to correctness** so the model can't farm reward by emitting `llm(` strings without solving the task
3. **Rebalancing weights** to give recursion behavior a real say (0.05 β†’ 0.15)

This converts FATHOM's GRPO from "single-turn QA training" to "plan-grading training." The model still doesn't *execute* recursion during the rollout (that requires the Β§D rewrite which is out of budget), but it learns to **predict good plans**: which task types deserve `llm()` calls and which don't. The trained model then drops into the inference-time recursion scaffold and executes those plans for real.

#### A.4-bis.1 The dataset β€” what the reward is actually shaping behavior across

`data/train.jsonl` contains 1000 rows with this distribution (verified):

| Task type | Count | Example prompt | Optimal recursion |
|-----------|-------|----------------|-------------------|
| `niah` | 400 (40%) | "What color is the mirror?" | **0 calls** β€” REPL grep finds the needle |
| `extractive` | 200 (20%) | "In which city was the 2000 agreement signed?" | **0 calls** β€” REPL regex |
| `multi_needle` | 300 (30%) | "Total cost of apple, lemon, plum?" | **0–2 calls** β€” REPL extracts; `llm()` if chunk too large |
| `counting` | 100 (10%) | "How many times does 'apple' appear?" | **0 calls** β€” REPL only (LLMs are bad at counting) |

Context lengths: min 4K, median 16K, p90 200K, max 200K. Larger contexts increasingly need `llm()` calls because they get tail-truncated to 4K in the prompt.

The reward function does **NOT** see `task_type`. The model has to *infer* recursion need from the prompt structure and context length β€” that's the right kind of generalization to teach.

#### A.4-bis.2 Create `rewards/recursion_extract.py` (NEW FILE)

Extracts `llm(` call counts from completion text, ignoring strings/comments using Python's `tokenize` module (with a regex fallback for syntactically-invalid blocks).

```python
"""Recursion call extractor β€” REW-04 v2.

Counts llm( function calls inside fenced ```python code blocks of a model
completion. Ignores occurrences in:
  - prose outside any code block
  - comments inside a code block (# llm(...) β†’ 0)
  - string literals inside a code block ("did llm(...)" β†’ 0)

Uses Python's tokenize module for accuracy; regex fallback when the code
block is syntactically invalid (the model writes broken Python sometimes
but we still want to count its intent).
"""
from __future__ import annotations

import io
import re
import tokenize

# Match ``` or ```python or ```py β€” case-insensitive, multi-line.
_CODE_BLOCK_RE = re.compile(
    r"```(?:python|py)?\s*\n(.*?)```",
    re.DOTALL | re.IGNORECASE,
)
_LLM_CALL_RE = re.compile(r"\bllm\s*\(")


def _count_in_block(code: str) -> int:
    """Count llm( calls in one code block. Tokenize-aware; regex fallback."""
    try:
        toks = list(tokenize.generate_tokens(io.StringIO(code).readline))
    except (tokenize.TokenizeError, IndentationError, SyntaxError):
        # Strip line comments, then regex. Conservative β€” does not strip
        # string literals, but the model rarely puts llm( in a string when
        # writing broken code.
        stripped = "\n".join(line.split("#", 1)[0] for line in code.splitlines())
        return len(_LLM_CALL_RE.findall(stripped))

    count = 0
    for i in range(len(toks) - 1):
        tok = toks[i]
        nxt = toks[i + 1]
        if (
            tok.type == tokenize.NAME
            and tok.string == "llm"
            and nxt.type == tokenize.OP
            and nxt.string == "("
        ):
            count += 1
    return count


def count_llm_calls(completion: str) -> int:
    """Total llm( calls inside all fenced code blocks of the completion.

    Returns 0 if completion is empty, has no code blocks, or only contains
    llm( in prose / comments / strings.
    """
    if not completion:
        return 0
    blocks = _CODE_BLOCK_RE.findall(completion)
    if not blocks:
        return 0
    return sum(_count_in_block(b) for b in blocks)


__all__ = ["count_llm_calls"]
```

#### A.4-bis.3 Replace `rewards/recursion_efficiency.py`

```python
"""Recursion efficiency reward component β€” REW-04 v2.

Linear decay on llm_call_count. Pure Python, stdlib-only.

Intended ranges:
  0 calls  β†’ 1.0 (best β€” task didn't need recursion)
  1 call   β†’ 0.75
  2 calls  β†’ 0.50
  3 calls  β†’ 0.25
  4+ calls β†’ 0.00 (recursion spam is wasteful)

This score is *coupled to correctness* in compose.py β€” wrong answers don't
earn an efficiency bonus, which prevents the model from learning to spam
`llm(` strings in code blocks for free reward.
"""

def recursion_efficiency(llm_call_count: int, **_) -> float:
    """Linear-decay efficiency on call count; gated to correctness in compose."""
    count = max(0, int(llm_call_count))
    return max(0.0, 1.0 - 0.25 * count)


__all__ = ["recursion_efficiency"]
```

#### A.4-bis.4 Replace `rewards/compose.py`

```python
"""Reward composition β€” REW-02 v3.

Changes from v2 (the Β§A.3.2 "soft format bonus" patch):
  - llm_call_count is now extracted from the completion's fenced Python
    code blocks (via rewards.recursion_extract.count_llm_calls), not
    hardcoded to 0 in train/grpo.py.
  - Recursion efficiency is gated on correctness β€” wrong answers cannot
    earn an efficiency bonus. Prevents the model from spamming `llm(`
    strings to harvest free reward.
  - Weights rebalanced: 0.70 correctness / 0.15 token_budget /
    0.15 recursion_efficiency. (Was 0.75 / 0.20 / 0.05.)
  - Per-component scalars are returned alongside the composite via the
    `_metrics` dict so the GRPOTrainer wrapper in train/grpo.py can
    log real per-component means to W&B (currently logs 0.0).

Anti-hacking caps preserved:
  - c == 0.0 β†’ composite ≀ 0.25 (was 0.05; raised to allow soft-format
    bonus to register, still well below any correct answer β‰₯ 0.80).
"""
from __future__ import annotations

from typing import Any, Callable

from .format_gate import format_gate
from .correctness import correctness
from .token_budget import token_budget
from .recursion_efficiency import recursion_efficiency
from .recursion_extract import count_llm_calls


def compose_reward_single(
    completion: str,
    gold_answer: str,
    prompt_token_count: int,
    cfg_reward: Any,
    llm_call_count: int | None = None,  # if None β†’ extract from completion
) -> tuple[float, dict[str, float]]:
    """Single-example composite reward + per-component metrics.

    Returns (composite_score, metrics_dict). The metrics dict has keys:
        format_pass, correctness, token_budget, recursion_eff_raw,
        recursion_eff_contribution, llm_call_count.
    """
    has_format = format_gate(completion) == 1.0
    c = correctness(completion, gold_answer) if has_format else 0.0
    t = token_budget(
        completion,
        prompt_token_count,
        alpha=float(cfg_reward.alpha),
        variant=str(cfg_reward.token_budget_variant),
    )

    if llm_call_count is None:
        llm_call_count = count_llm_calls(completion)
    eff_raw = recursion_efficiency(int(llm_call_count))
    # Couple efficiency to correctness β€” wrong answers earn 0 efficiency.
    eff_contribution = eff_raw if c == 1.0 else 0.0

    w = cfg_reward.weights
    assert abs(
        float(w.correctness) + float(w.token_budget) + float(w.recursion_efficiency) - 1.0
    ) < 1e-3, "REW-02 v3: composite weights must sum to 1.0"

    composite = (
        float(w.correctness) * c
        + float(w.token_budget) * t
        + float(w.recursion_efficiency) * eff_contribution
    )
    if has_format:
        composite += 0.10  # soft format bonus (Β§A.3.2)

    if c == 0.0:
        composite = min(composite, 0.25)

    metrics = {
        "format_pass": 1.0 if has_format else 0.0,
        "correctness": c,
        "token_budget": t,
        "recursion_eff_raw": eff_raw,
        "recursion_eff_contribution": eff_contribution,
        "llm_call_count": float(llm_call_count),
    }
    return composite, metrics


def compose_reward_fn(prompts: list, completions: list, **kwargs) -> list[float]:
    """TRL-compatible batched reward function. Returns scalars only.

    Per-component means are stashed under `kwargs['_component_means']` for
    the GRPOTrainer instrumentation wrapper to log to W&B. (TRL ignores
    extra kwargs.)
    """
    cfg_reward = kwargs.pop("cfg_reward")
    gold_answers = kwargs.get("gold_answer", [""] * len(completions))
    ptcs = kwargs.get("prompt_token_count", [1] * len(completions))

    pairs = [
        compose_reward_single(c, g, int(p), cfg_reward)
        for c, g, p in zip(completions, gold_answers, ptcs)
    ]
    rewards = [p[0] for p in pairs]
    metrics_list = [p[1] for p in pairs]

    # Aggregate component means for W&B logging via the wrapper.
    if metrics_list:
        keys = metrics_list[0].keys()
        means = {k: sum(m[k] for m in metrics_list) / len(metrics_list) for k in keys}
        kwargs["_component_means"] = means
    return rewards


def make_reward_fn(cfg_reward: Any) -> Callable:
    """Factory binding cfg_reward for GRPOTrainer.reward_funcs."""
    def _bound(prompts, completions, **kwargs):
        kwargs["cfg_reward"] = cfg_reward
        return compose_reward_fn(prompts, completions, **kwargs)
    return _bound


__all__ = ["compose_reward_fn", "compose_reward_single", "make_reward_fn"]
```

#### A.4-bis.5 Patch `train/grpo.py`

Two edits:

**(a)** In `_to_prompt` (lines 218–239 in current file, will shift after Β§A.3.1), **delete** the `"llm_call_count": 0` field from the returned dict β€” the extractor now computes it from each rollout's completion. Final return shape:

```python
return {
    "prompt": prompt_str,
    "gold_answer": str(example.get("gold_answer", "")),
    "prompt_token_count": int(example.get("context_length", 0)) // 4,
}
```

**(b)** Replace the `_instrumented_reward_fn` body (currently at lines 138–151) so it logs the **real** per-component means that `compose_reward_fn` now stashes under `kwargs['_component_means']`:

```python
def _instrumented_reward_fn(prompts, completions, **kwargs):
    rewards = reward_fn(prompts, completions, **kwargs)
    try:
        if wandb.run is not None:
            log_dict = {
                "reward/composite_mean": sum(rewards) / max(len(rewards), 1),
                "reward/composite_std": (
                    statistics.stdev(rewards) if len(rewards) > 1 else 0.0
                ),
            }
            cm = kwargs.get("_component_means", {})
            for k, v in cm.items():
                log_dict[f"reward/{k}_mean"] = float(v)
            wandb.log(log_dict)
    except Exception:
        pass
    return rewards
```

Add `import statistics` at the top of the file (already has `import inspect`, `import logging`, `import os`, etc., so just add the line).

#### A.4-bis.6 Update `configs/reward/v1.yaml`

```yaml
alpha: 0.2
weights:
  correctness: 0.70
  token_budget: 0.15
  recursion_efficiency: 0.15
token_budget_variant: "capped_linear"
answer_regex: "<answer>(.*?)</answer>"
max_calls: 4
```

#### A.4-bis.7 Add tests (`tests/test_recursion_extract.py`, NEW)

```python
"""Tests for rewards.recursion_extract β€” REW-04 v2."""
from rewards.recursion_extract import count_llm_calls


def test_empty_completion():
    assert count_llm_calls("") == 0


def test_no_code_block():
    assert count_llm_calls("The answer is <answer>silver</answer>") == 0


def test_single_call():
    c = "```python\nresult = llm('find', ctx[:5000])\n```\n<answer>silver</answer>"
    assert count_llm_calls(c) == 1


def test_two_calls_in_one_block():
    c = "```python\na = llm('q1', ctx[:1000])\nb = llm('q2', ctx[1000:])\n```"
    assert count_llm_calls(c) == 2


def test_calls_across_two_blocks():
    c = "```python\nx=llm('q', ctx)\n```\nthen\n```python\ny=llm('q2', ctx)\n```"
    assert count_llm_calls(c) == 2


def test_call_in_comment_not_counted():
    c = "```python\n# would call llm(stuff) but skipping\nprint('done')\n```"
    assert count_llm_calls(c) == 0


def test_call_in_string_literal_not_counted():
    c = '```python\nnote = "earlier code did llm(...)"\nprint(note)\n```'
    assert count_llm_calls(c) == 0


def test_call_outside_code_block_not_counted():
    c = "Maybe I should call llm(question, chunk) but I won't actually."
    assert count_llm_calls(c) == 0


def test_call_in_loop_counts_literal_occurrence():
    c = "```python\nfor chunk in chunks:\n    r = llm('find', chunk)\n```"
    assert count_llm_calls(c) == 1


def test_invalid_python_falls_back_to_regex():
    c = "```python\nthis is not valid python !!!\nresult = llm('q', ctx)\n```"
    assert count_llm_calls(c) >= 1  # fallback regex finds it


def test_bare_python_fence():
    c = "```\nans = llm('q', ctx)\n```"  # no language tag
    assert count_llm_calls(c) == 1
```

#### A.4-bis.8 Add tests to `tests/test_rewards.py`

Append a new test class at the end:

```python
class TestComposeV3:
    """REW-02 v3: soft format + recursion-extraction + correctness-gated efficiency."""

    @pytest.fixture
    def cfg_v3(self):
        return OmegaConf.create({
            "alpha": 0.2,
            "weights": {"correctness": 0.70, "token_budget": 0.15, "recursion_efficiency": 0.15},
            "token_budget_variant": "capped_linear",
            "answer_regex": "<answer>(.*?)</answer>",
            "max_calls": 4,
        })

    def test_correct_no_recursion_scores_high(self, cfg_v3):
        c = "```python\nimport re\nm=re.search('silver', ctx)\nprint(m.group())\n```\n<answer>silver</answer>"
        score, metrics = compose_reward_single(c, "silver", 100, cfg_v3)
        assert score >= 0.85, f"clean correct should score high, got {score}"
        assert metrics["llm_call_count"] == 0

    def test_zero_calls_beats_one_call_when_both_correct(self, cfg_v3):
        c0 = "```python\nimport re\nm=re.search('silver', ctx)\nprint(m.group())\n```\n<answer>silver</answer>"
        c1 = "```python\nans=llm('color', ctx[:5000])\nprint(ans)\n```\n<answer>silver</answer>"
        s0, _ = compose_reward_single(c0, "silver", 100, cfg_v3)
        s1, _ = compose_reward_single(c1, "silver", 100, cfg_v3)
        assert s0 > s1, f"0-call ({s0:.3f}) should beat 1-call ({s1:.3f}) when both correct"

    def test_efficiency_gated_on_correctness(self, cfg_v3):
        # Wrong answer with 0 calls β€” must NOT earn efficiency bonus.
        c = "```python\nprint('done')\n```\n<answer>gold</answer>"
        score, metrics = compose_reward_single(c, "silver", 100, cfg_v3)
        assert metrics["recursion_eff_contribution"] == 0.0
        assert score <= 0.25, f"wrong answer must be capped, got {score}"

    def test_recursion_spam_loses_to_minimal_recursion(self, cfg_v3):
        c2 = "```python\na=llm('q1',ctx[:1000])\nb=llm('q2',ctx[1000:2000])\n```\n<answer>silver</answer>"
        c5 = "```python\n" + "\n".join(f"x{i}=llm('q{i}',ctx)" for i in range(5)) + "\n```\n<answer>silver</answer>"
        s2, _ = compose_reward_single(c2, "silver", 100, cfg_v3)
        s5, _ = compose_reward_single(c5, "silver", 100, cfg_v3)
        assert s2 > s5, f"2-call ({s2:.3f}) should beat 5-call spam ({s5:.3f})"

    def test_format_only_capped(self, cfg_v3):
        c = "<answer>wrong</answer>"
        score, _ = compose_reward_single(c, "silver", 100, cfg_v3)
        assert 0.05 <= score <= 0.25, f"format-only wrong should be in [0.05, 0.25], got {score}"

    def test_no_format_gets_minimal_credit(self, cfg_v3):
        c = "silver"  # right text but no <answer> tag
        score, _ = compose_reward_single(c, "silver", 100, cfg_v3)
        assert score <= 0.20

    def test_group_variance_nonzero(self, cfg_v3):
        """Smoke check: a synthetic GRPO group of 8 must produce non-zero std.
        v1 had std=0.0 across all groups, which zeroed the GRPO advantage."""
        gens = [
            "```python\nimport re\nm=re.search('silver',ctx)\nprint(m.group())\n```\n<answer>silver</answer>",
            "```python\nans=llm('color',ctx[:5000])\nprint(ans)\n```\n<answer>silver</answer>",
            "<answer>silver</answer>",
            "<answer>gold</answer>",
            "the color is silver",
            "<answer></answer>",
            "```python\n" + "\n".join(f"x{i}=llm('q{i}',ctx)" for i in range(5)) + "\n```\n<answer>silver</answer>",
            "silver.",
        ]
        scores = [compose_reward_single(g, "silver", 200, cfg_v3)[0] for g in gens]
        import statistics as _st
        assert _st.stdev(scores) > 0.10, f"group std too low: {_st.stdev(scores)}"
```

Note: existing tests in `TestComposeReward` will break because `compose_reward_single` now returns `(score, metrics)` instead of just `score`. Either:
- (a) Update the existing tests to unpack `(score, _) = compose_reward_single(...)`, or
- (b) Keep backward compat by adding a `return_metrics: bool = False` flag with default False that returns just the float.

**Pick (a)** β€” explicit is better, and the v1 tests' expected values change anyway under the new weights. Find any existing call site of `compose_reward_single` and add `, _` to the unpacking. Update the old `TestComposeReward` cases to use new expected ranges (the cap moved from 0.05 to 0.25).

#### A.4-bis.9 Update `REWARD_AUDIT.md`

Add two new attack rows and revise A-05:

```markdown
## A-05: Recursion depth gaming (REVISED for v3)

**Vector:** Model uses 0 llm() calls on every task to maximize
recursion_efficiency, even on multi_needle / 200K tasks where recursion
would actually help correctness.

**Analysis (v3):**
- recursion_efficiency contributes only when correctness == 1.0 (gating
  in compose.py). On hard tasks where 0 calls fails to produce a correct
  answer, the efficiency bonus is forfeited entirely.
- Net incentive: use the *minimum* recursion that still produces a
  correct answer. Exactly the desired behavior.

**Status:** βœ… MITIGATED by correctness-gating.

---

## A-07: Comment-spam exploit (NEW)

**Vector:** Model emits `# llm(foo)` inside code blocks to inflate the
count regex without making real calls. (Inverted variant of A-05: spam
to make recursion_eff *lower*, useless because lower efficiency hurts.)

**Test:** `test_call_in_comment_not_counted` in
`tests/test_recursion_extract.py`.

**Result:** 0 calls counted βœ… β€” extractor uses tokenize, ignores comments.

**Status:** βœ… MITIGATED by tokenize-aware extraction.

---

## A-08: String-literal exploit (NEW)

**Vector:** Model writes `"earlier code did llm(...)"` in a string
literal to confuse a naive regex extractor.

**Test:** `test_call_in_string_literal_not_counted`.

**Result:** 0 calls counted βœ… β€” tokenize correctly identifies STRING
tokens and skips them.

**Status:** βœ… MITIGATED.
```

Update the bottom summary table accordingly. Bump verdict timestamp to today.

#### A.4-bis.10 CPU-only verification script (`scripts/verify_recursion_reward.py`, NEW)

This is the gate that must pass before any HF Job spend. Runs in <1 s on a laptop.

```python
"""CPU-only verification of REW-04 v2 reward design.

Confirms two GRPO-blocking properties:
  1. A synthetic 8-completion group produces non-zero std (v1's std was 0.0
     across every group, which is why the reward curve was flat).
  2. The ordering correct+0calls > correct+1call > correct+spam holds.

Run BEFORE spending any HF Jobs credits on a retrain.
"""
from __future__ import annotations

import statistics
import types

from rewards.compose import compose_reward_single

cfg = types.SimpleNamespace(
    alpha=0.2,
    weights=types.SimpleNamespace(
        correctness=0.70, token_budget=0.15, recursion_efficiency=0.15
    ),
    token_budget_variant="capped_linear",
    answer_regex="<answer>(.*?)</answer>",
    max_calls=4,
)

GOLD = "silver"
GENERATIONS = [
    ("correct + 0 llm calls (REPL grep)",
     "```python\nimport re\nm=re.search('silver', ctx)\nprint(m.group())\n```\n<answer>silver</answer>"),
    ("correct + 1 llm call",
     "```python\nans=llm('color', ctx[:5000])\nprint(ans)\n```\n<answer>silver</answer>"),
    ("correct + 3 llm calls (wasteful)",
     "```python\na=llm('q1',ctx[:1000])\nb=llm('q2',ctx[1000:2000])\nc=llm('q3',ctx[2000:3000])\n```\n<answer>silver</answer>"),
    ("correct + bare answer (no code, trivial-task path)",
     "<answer>silver</answer>"),
    ("wrong + format",
     "<answer>gold</answer>"),
    ("wrong + no format",
     "the color is gold"),
    ("right text + no format (v1 collapse mode)",
     "silver"),
    ("format-only spam",
     "<answer></answer>"),
]

print(f"{'idx':>3}  {'score':>6}  {'calls':>5}  description")
print("-" * 78)
scores = []
for i, (desc, gen) in enumerate(GENERATIONS):
    s, m = compose_reward_single(gen, GOLD, 200, cfg)
    scores.append(s)
    print(f"{i:>3}  {s:>6.3f}  {int(m['llm_call_count']):>5d}  {desc}")
print("-" * 78)
print(f"group mean: {statistics.mean(scores):.4f}")
print(f"group std:  {statistics.stdev(scores):.4f}  (must be > 0.10 for GRPO advantage)")
print(f"max - min:  {max(scores) - min(scores):.4f}")

# Hard gates β€” exit non-zero if any fail
assert statistics.stdev(scores) > 0.10, "FAIL: group std too low; GRPO will not learn"
assert scores[0] > scores[1] > scores[2], (
    f"FAIL: efficiency ordering broken (got {scores[0]:.3f} > {scores[1]:.3f} > {scores[2]:.3f})"
)
assert scores[0] > scores[4], "FAIL: correct must beat wrong"
assert scores[7] <= 0.25, "FAIL: format-only spam not capped"
print("\nPASS: REW-04 v2 produces learnable variance and correct orderings")
```

#### A.4-bis.11 Execution order β€” drop-in replacement for Β§6 steps 2–6

Replace steps 2–6 in the Β§6 table with:

| Step | Action | Time | Cost | Checkpoint |
|------|--------|------|------|------------|
| 2a | Apply A.3.1 (prompt alignment) + A.3.3 (assert) | 20 min | $0 | CPU dry-run prints `PREFIX_MATCH: True` |
| 2b | Apply A.4-bis: create `recursion_extract.py`, replace `recursion_efficiency.py`, replace `compose.py`, patch `train/grpo.py`, update `configs/reward/v1.yaml` | 40 min | $0 | All edits made, no test imports broken |
| 2c | Run new tests | 5 min | $0 | `pytest tests/test_recursion_extract.py tests/test_rewards.py -q` all green |
| 2d | Run CPU verifier | 1 min | $0 | `python scripts/verify_recursion_reward.py` prints `PASS:` |
| 3 | Commit: `fix(reward): align prompt with SFT, soft format, real recursion signal (A.3 + A.4-bis)` | 5 min | $0 | git log shows commit |
| 4 | Smoke on HF Jobs `a10g-large` | 5 min | ~$0.10 | `outputs/smoke/SMOKE_RESULT.md` GO |
| 5 | 50-step GRPO trial | 20 min | ~$2 | reward curve shows movement, group std > 0 |
| 6 | Decision gate (full retrain or honest fallback) | β€” | β€” | see A.4 below |

#### A.4-bis.12 Definition of done for this addendum

- [ ] `pytest tests/test_recursion_extract.py -q` β†’ 11 passed
- [ ] `pytest tests/test_rewards.py -q` β†’ all green (existing + new TestComposeV3 class)
- [ ] `python scripts/verify_recursion_reward.py` β†’ exits 0 with `PASS:` line
- [ ] On the 50-step trial run, W&B shows non-zero values for `reward/correctness_mean`, `reward/recursion_eff_contribution_mean`, `reward/llm_call_count_mean` β€” not just `reward/composite_mean`
- [ ] `frac_reward_zero_std` in the trainer logs is `< 0.5` for at least 80% of steps (v1 was `1.0` for 100% of steps)

If item 5 above fails (frac_reward_zero_std stays at 1.0), the prompt-alignment fix in A.3.1 didn't take. Re-check `PREFIX_MATCH: True` and that the SFT adapter is actually loading (look for the log line `TRN-03 SFT adapter loaded from .../sft_adapter (after unloading empty wrap)`).

---

### A.4 Re-train and republish

Once A.3.1–A.3.3 land and CPU dry-run prints `PREFIX_MATCH: True`:

1. **Smoke test on HF Jobs first** (1 min, ~$0.10 on `a10g-large`):

   ```bash
   bash scripts/job_smoke.sh
   ```

   Inspect `outputs/smoke/SMOKE_RESULT.md`. Expected: GO with 6/6 PASS.

2. **Short GRPO run** β€” 50 steps only, NOT 400. This is to verify the curve moves. Use `a10g-large` (β‰ˆ$2):

   ```bash
   # Override max_steps via Hydra
   bash scripts/job_train.sh -- train.max_steps=50
   ```

   Pull the resulting `trainer_state.json` and run `python scripts/make_plots.py`. The reward curve should now show **any non-zero variance** β€” even if it's only `0.05 β†’ 0.18`. That alone is a publishable curve.

3. **Decision gate:**
   - If 50-step curve moves: launch the full 400-step run on `a10g-large` (β‰ˆ$10–15) and replace the plots on `Pratham-math/fathom-1.5b-grpo/plots/*`.
   - If 50-step curve is still flat: stop. Do not spend more credits. Switch to the **honest fallback** in Β§A.5.

### A.5 Honest fallback (use only if A.4 step 3 still shows flat reward)

If the curve still doesn't move, do not fake it. Re-frame the README to claim what is actually true: **"the SFT phase taught the format; GRPO did not converge in our budget; the env, reward, and pipeline are nonetheless complete and reproducible."** This is genuinely a publishable result β€” most hackathon submissions don't even get SFT working. The judges' rubric awards points for "showing improvement in rewards" (20%); SFT loss `3.2 β†’ 0.29` and token accuracy `0.46 β†’ 0.93` are improvements. Lead with those plots; relegate the GRPO curve to a section titled "What we learned about reward design."

---

## 3. Problem B β€” No HF Space UI

### B.1 What the user sees

Hitting <https://Pratham-math-fathom-env.hf.space/> returns `{"detail":"Not Found"}`. There is no landing page. Judges who don't know to append `/healthz` or `/docs` see a blank 404. The Streamlit demo at `viz/app.py` exists locally but has never been deployed and is full of placeholder data anyway (sample tree literal at line 87, fake `[0.10, 0.18, 0.28, ...]` reward sparkline at line 191).

### B.2 Two-Space architecture (do this)

The OpenEnv contract requires the env Space to expose `/reset`, `/step`, etc. as JSON β€” that's correct, do not change it. But judges need a UI. **Solution: deploy a second Space (Streamlit SDK) that calls the env Space.** This is the canonical pattern in the OpenEnv hackathon submissions (the env Space is the "engine"; the demo Space is the "showroom").

| Space | URL | Purpose | SDK | What changes |
|-------|-----|---------|-----|--------------|
| `Pratham-math/fathom-env` | `Pratham-math-fathom-env.hf.space` | OpenEnv JSON server | Docker | Add a `GET /` HTML index page (B.3) |
| `Pratham-math/fathom-demo` (NEW) | `Pratham-math-fathom-demo.hf.space` | Streamlit UI for judges | Streamlit | New space, scaffolded from `viz/app.py` (B.4) |

### B.3 Patch the env Space β€” add a root index page

Edit `env/server/app.py` so a judge hitting the bare URL gets a useful HTML response, not a 404. Add this route **before** `@app.get("/healthz")`:

```python
from fastapi.responses import HTMLResponse

INDEX_HTML = """<!DOCTYPE html>
<html><head><meta charset="utf-8"><title>FATHOM Env Server</title>
<style>body{font-family:system-ui,sans-serif;max-width:760px;margin:40px auto;padding:0 20px;line-height:1.55;color:#111}
code{background:#f4f4f5;padding:2px 6px;border-radius:4px}
a{color:#4338ca}.tag{display:inline-block;padding:2px 8px;border-radius:999px;background:#eef2ff;color:#4338ca;font-size:12px;margin-right:6px}</style></head>
<body>
<h1>FATHOM Env Server <span class="tag">OpenEnv 0.2.3</span><span class="tag">Docker</span></h1>
<p><b>FATHOM</b> is the first openly-published OpenEnv RL environment that teaches a small language model to use a recursive-LM scaffold (Python REPL + recursive <code>llm()</code> calls) for long-context QA. Submission for the Meta Γ— PyTorch Γ— Hugging Face OpenEnv Hackathon Grand Finale, Bangalore, Apr 25–26 2026.</p>
<h2>Endpoints</h2>
<ul>
  <li><a href="/healthz"><code>GET /healthz</code></a> β€” liveness probe</li>
  <li><code>POST /reset</code> β€” start an episode (try via <a href="/docs">/docs</a>)</li>
  <li><code>POST /step</code> β€” execute REPL or llm() action</li>
  <li><a href="/state"><code>GET /state</code></a> β€” sanitized episode state</li>
  <li><a href="/docs"><code>GET /docs</code></a> β€” interactive OpenAPI</li>
</ul>
<h2>See also</h2>
<ul>
  <li><b>Demo UI:</b> <a href="https://Pratham-math-fathom-demo.hf.space">fathom-demo</a> (Streamlit)</li>
  <li><b>Trained model + plots:</b> <a href="https://huggingface.co/Pratham-math/fathom-1.5b-grpo">Pratham-math/fathom-1.5b-grpo</a></li>
  <li><b>Code repo:</b> <a href="https://huggingface.co/Pratham-math/fathom-code">Pratham-math/fathom-code</a></li>
  <li><b>Colab reproducer:</b> <code>notebooks/fathom_train.ipynb</code> in the code repo</li>
</ul>
</body></html>"""

@app.get("/", response_class=HTMLResponse)
def index() -> HTMLResponse:
    return HTMLResponse(content=INDEX_HTML)
```

Commit with `feat(space): add HTML index for judge first-impression`. Push to the env Space:

```bash
python scripts/deploy_space.py    # already wired to push env/ to fathom-env Space
```

Verify:

```bash
curl -sS https://Pratham-math-fathom-env.hf.space/ | head -20    # should be HTML, not 404
```

### B.4 Build and deploy the Streamlit demo Space

Create the demo Space programmatically:

```bash
mkdir -p space_demo
```

Files to create under `space_demo/`:

**`space_demo/README.md`** (Streamlit Space frontmatter):

```markdown
---
title: FATHOM Demo
emoji: 🧠
colorFrom: indigo
colorTo: purple
sdk: streamlit
sdk_version: 1.39.0
app_file: app.py
pinned: true
license: apache-2.0
---

# FATHOM Demo

Interactive UI for the FATHOM recursive-LM environment. Backed by [Pratham-math/fathom-env](https://huggingface.co/spaces/Pratham-math/fathom-env).
```

**`space_demo/requirements.txt`**:

```
streamlit>=1.39,<2.0
plotly>=5.24,<6.0
httpx>=0.27,<1.0
huggingface_hub>=0.28
pandas>=2.0
```

**`space_demo/app.py`** β€” port `viz/app.py` here, but **replace placeholder data with real artifacts**. Concretely:

1. **Reward composition pie** β€” keep, it's accurate.
2. **Recursion tree** (column 1) β€” replace the literal `sample_tree = {...}` with a live call to `https://Pratham-math-fathom-env.hf.space/reset` then `/step`, capturing the actual REPL trace from one episode. Cache it (`@st.cache_data(ttl=3600)`) so judges don't hammer the env. If the live call fails, fall back to a **clearly-labeled** "example trace" (do not pretend it's live).
3. **Pareto frontier** (column 2) β€” the current `[0.62, 0.61, 0.58, 0.52, 0.44]` numbers are fabricated. Either:
   - (a) Generate a real one by running the merged_16bit model from `Pratham-math/fathom-1.5b-grpo` against `data/eval.jsonl` at 5 different Ξ± values, or
   - (b) Remove this column and replace with an "Eval results" table reading `outputs/eval_*.json` if it exists, or
   - (c) Hide column 2 entirely and widen columns 1 + 3.

   **Pick (b) or (c) if you have <30 min.** Do not ship fabricated numbers.
4. **W&B iframe** (column 3) β€” set `WANDB_RUN_URL=https://wandb.ai/pratham-alwar05-indian-institute-of-information-technolo/huggingface/runs/sy1tqun0` in the Space's "Variables and secrets" panel so the iframe renders the real run.

Add a top banner cell:

```python
st.markdown(
    f"**Live env:** [Pratham-math/fathom-env]({os.environ.get('FATHOM_SPACE_URL','https://Pratham-math-fathom-env.hf.space')}) "
    f"Β· **Trained model:** [Pratham-math/fathom-1.5b-grpo](https://huggingface.co/Pratham-math/fathom-1.5b-grpo) "
    f"Β· **W&B:** [run sy1tqun0](https://wandb.ai/pratham-alwar05-indian-institute-of-information-technolo/huggingface/runs/sy1tqun0)"
)
```

Deploy:

```bash
cd space_demo
huggingface-cli login --token $HF_TOKEN
huggingface-cli repo create fathom-demo --type space --space_sdk streamlit
git init
git remote add origin https://Pratham-math:$HF_TOKEN@huggingface.co/spaces/Pratham-math/fathom-demo
git add -A && git commit -m "feat: initial fathom-demo Space"
git push -u origin main
```

Confirm at <https://Pratham-math-fathom-demo.hf.space> β€” you should see Streamlit boot in ~2 minutes. **Set the `FATHOM_SPACE_URL` and `WANDB_RUN_URL` Space variables** in the HF UI (Settings β†’ Variables and secrets).

### B.5 Update README submission table

In `README.md` lines 11–22 (the Submission Links table), add a row:

```markdown
| **Demo UI (Streamlit Space)** | <https://huggingface.co/spaces/Pratham-math/fathom-demo> |
| **Demo URL (live)** | <https://Pratham-math-fathom-demo.hf.space> |
```

---

## 4. Problem C β€” Missing Materials

The hackathon rubric explicitly lists these as **non-negotiable**. From the prompt:

> A short writeup: a mini-blog on Hugging Face or a < 2 minute video on YouTube explaining what your environment does and what you trained, or a short slide deck of presentation. Please make sure that all materials are linked from your README file so that judges can access them easily.

**Status:**

| Item | Required? | Status | Action |
|------|-----------|--------|--------|
| OpenEnv (latest) used | βœ… required | DONE (`openenv-core>=0.2.3`) | none |
| Working Unsloth/TRL training script | βœ… required | DONE (`train/grpo.py`) | none |
| Colab notebook | βœ… "ideally" | DONE-ish (`notebooks/fathom_train.ipynb`) | C.1 verify dataset paths |
| Loss + reward plots from a real run | βœ… required | DONE (10 PNGs on model repo) | A.4 will replace if curve moves |
| Mini-blog OR <2-min video OR slide deck | βœ… **NON-NEGOTIABLE** | **MISSING** | **C.2** |
| HF Space deployed | βœ… required | DONE (`fathom-env`) | B.3 + B.4 enrich |
| README motivation + env + results | βœ… required | DONE | C.3 polish |
| README links to Space + materials | βœ… required | PARTIAL | C.3 |
| **No big video files in env submission** | βœ… required | OK (no videos in repo) | none |

### C.1 Fix the Colab notebook dataset path

`notebooks/fathom_train.ipynb` cell-6 calls `hf_hub_download(repo_id='Pratham-math/fathom-code', filename='data/train.jsonl', ...)`. **Verify those files actually live on `Pratham-math/fathom-code`** β€” they may not. Run:

```bash
curl -sS "https://huggingface.co/api/models/Pratham-math/fathom-code/tree/main/data" | python -m json.tool
```

If `data/train.jsonl`, `data/eval.jsonl`, and `data/sft_traces.jsonl` are missing, push them:

```bash
huggingface-cli upload Pratham-math/fathom-code data/ data/ --repo-type=model
```

Re-run cell 6 in a Colab to confirm. (You can do this without a GPU β€” cells 1–5 only need CPU.)

### C.2 Create the mini-blog (fastest of the three options β€” do this)

A 600-word HF mini-blog beats a video for our time budget. Create it on the Hugging Face Hub:

```bash
huggingface-cli repo create fathom-blog --type space --space_sdk static
```

Then push a single `index.html` (or use the existing `assets/BLOG_DRAFT.md` if it's already drafted β€” check first with `cat assets/BLOG_DRAFT.md`). Required structure:

1. **Hook** (1 paragraph) β€” why teach a small model to recurse instead of buying a longer-context one
2. **Environment** (1 paragraph + screenshot of the demo Space) β€” REPL + `llm()` primitive, deterministic verifier, depth-2 cap
3. **Reward design** (1 paragraph + the reward-composition pie image) β€” 4 components, anti-hacking audit
4. **Training** (1 paragraph + the SFT loss curve and GRPO reward curve) β€” be **honest** about the GRPO curve. Frame the v1 flat-line as a finding ("our gate was multiplicative; this is what GRPO collapse looks like"); show the v2 curve underneath (after A.4 retrain) if it moved.
5. **Reproduce** (1 paragraph) β€” link to Colab notebook + HF Space + model repo
6. **Footer** β€” names, hackathon, license

Add to `README.md` Submission Links:

```markdown
| **Mini-blog** | <https://huggingface.co/spaces/Pratham-math/fathom-blog> |
```

If the user has already drafted `assets/BLOG_DRAFT.md`, port it into the Space's `index.html` with minimal styling β€” don't rewrite from scratch.

### C.3 GitHub mirror

```bash
cd C:/Users/prath/OneDrive/Desktop/Hackathons/Meta_finale
gh auth login    # if not already
gh repo create Pratham-math/fathom --public --source=. --remote=github --push
```

Then update README line 17 to:

```markdown
| **Code repo (GitHub mirror)** | <https://github.com/Pratham-math/fathom> |
```

Delete the stale `_to be added β€” see GITHUB_URL.txt once mirrored_` line.

### C.4 Submission-link block β€” final state

After C.1–C.3 + B.5, the README's "Submission Links (Judges Start Here)" table must contain **all** of:

- βœ… Environment Space (Hub page) + Endpoint URL + Health check
- βœ… Demo UI Space + Demo URL (NEW)
- βœ… Code repo (HF) + GitHub mirror (NEW)
- βœ… Trained model + plots
- βœ… Colab notebook
- βœ… Mini-blog (NEW)
- βœ… W&B run

Run `python scripts/submission_preflight.py` after the README edits and confirm `Submission package looks judge-ready.`

---

## 5. Problem D β€” Truth-in-Advertising

### D.1 What's misleading

The README, `CLAUDE.md`, and `assets/architecture.png` all imply that the model **uses** the REPL and recursive `llm()` calls **during GRPO training**. It does not. Read `train/grpo.py:153-169`:

```python
# Why no `env=` / `environment_url=` / `environment_factory` kwargs?
# - TRL 1.2.0's GRPOTrainer.__init__ only accepts env interaction via
#   `tools=` (needs transformers>=5.0), `environment_factory=`
#   (needs transformers>=5.2), or `rollout_func=`. We're on
#   transformers==4.56.2, so the first two raise. The third requires a
#   custom multi-turn rollout implementation we don't have time to
#   harden.
# - Our reward function (rewards.compose) operates on (prompt, completion,
#   gold_answer, prompt_token_count, llm_call_count) β€” zero env
#   interaction needed.
```

So:

- The **env exists** and is deployed (rubric requirement met).
- The **env is used at inference time** in the demo Space (judges can run a recursive episode).
- The **env is NOT used at training time**. GRPO is single-turn prompt β†’ completion β†’ deterministic reward.

A judge who reads code may flag this as inconsistent with the README's repeated claims about "teaching the model to use the REPL/recursion." That's a goodwill hit we can avoid with one paragraph of plain language.

### D.2 Required README edit

In `README.md`, just after the Architecture section (around line 33), insert this paragraph **verbatim**:

```markdown
### A note on the role of the env in training

TRL 1.2.0 with `transformers==4.56.2` does not yet expose multi-turn env-tool calls inside `GRPOTrainer.train()` (the `tools=` / `environment_factory=` kwargs require `transformers>=5.0`, and a custom `rollout_func=` was outside our time budget). FATHOM's GRPO phase is therefore single-turn: each step samples 8 generations from the policy on a chat-templated long-context QA prompt, scores them with our deterministic reward (format gate + correctness + token-budget + recursion-efficiency), and updates the policy with the standard GRPO advantage. **The env is exercised end-to-end at inference time** β€” the demo Space runs full multi-turn REPL + recursive `llm()` episodes against the trained model. Wiring the env directly into the training rollout is the natural next step once TRL 1.3 / transformers 5 ships.
```

This is honest, it preempts the obvious code-reading critique, and it actually **reframes our submission as forward-looking** rather than incomplete.

### D.3 Architecture image touch-up (optional, only if time)

`assets/architecture.png` shows arrows from "GRPOTrainer" to "REPL" and "llm()". Either:

- (a) Edit the source `assets/architecture.mmd` (Mermaid) so those arrows are dashed and labeled `inference-time only`, or
- (b) Skip this if you've already done D.2 β€” the README paragraph carries enough context.

---

## 6. Execution Order (with checkpoints)

| Step | Action | Time | Cost | Checkpoint |
|------|--------|------|------|------------|
| 1 | Read this brief, run `git status`, confirm clean working tree | 5 min | $0 | `git status` clean |
| 2 | Apply A.3.1 (prompt alignment) + A.3.2 (soft format) + A.3.3 (assert) | 30 min | $0 | CPU dry-run prints `PREFIX_MATCH: True` |
| 3 | Commit: `fix(grpo): align prompt with SFT, soft format reward` | 5 min | $0 | git log shows commit |
| 4 | Run smoke on HF Jobs `a10g-large` | 5 min | ~$0.10 | `outputs/smoke/SMOKE_RESULT.md` GO |
| 5 | 50-step GRPO trial run | 20 min | ~$2 | reward_curve.png shows movement |
| 6 | **Decision gate** β€” full retrain or honest fallback | β€” | β€” | see A.4 step 3 |
| 7 | Apply B.3 (env Space index) + push | 15 min | $0 | `curl /` returns HTML |
| 8 | Build B.4 (Streamlit demo Space) + push | 60 min | $0 | demo URL renders |
| 9 | Apply C.2 (mini-blog) | 30 min | $0 | blog Space live |
| 10 | Apply C.3 (GitHub mirror) | 5 min | $0 | GH repo public |
| 11 | Apply D.2 (README clarity paragraph) | 5 min | $0 | README diff |
| 12 | Run `python scripts/submission_preflight.py` | 1 min | $0 | "judge-ready" message |
| 13 | Refresh README submission table per C.4 | 10 min | $0 | all rows filled |
| 14 | Final commit + push to HF master + GH main | 5 min | $0 | both remotes in sync |
| 15 | Hit every link in the README from a fresh browser | 10 min | $0 | nothing 404s |

**Total: ~3.5 h of work + ~$2–15 of cloud GPU depending on retrain decision.**

---

## 7. Things You Might Be Tempted To Do β€” Don't

- ❌ **Upgrade `trl`/`transformers`/`unsloth` to enable env-tool calls in training.** This is a 2-day refactor with high failure risk. The D.2 paragraph defuses the critique without code changes.
- ❌ **Switch reward composition from weighted-sum to product or RLHF-style ranking.** The composition is fine; the format gate was the bug.
- ❌ **Train a 3B model "for better optics."** The CLAUDE.md is explicit that 1.5B was the deliberate choice. Sticking with 1.5B is part of the story (small-model recursion).
- ❌ **Increase `max_completion_length` past 2048.** STACK §10.3 calls this an anti-pattern. 2048 is fine.
- ❌ **Move from `vllm_mode='colocate'` to `'server'`.** STACK Β§10.4 + TRL #4543 β€” this breaks multi-turn. We don't even use multi-turn in training, but `colocate` is also the cheaper option memory-wise.
- ❌ **Refactor `env/server/environment.py`.** It's stable, audited, and shipped. Touch nothing inside `env/` except `app.py` for the index route.
- ❌ **Delete the v1 flat-line GRPO plot.** Honest evidence is part of the storytelling. Either replace with v2 (if A.4 succeeds) or annotate (if A.5 fallback).
- ❌ **Generate fake Pareto numbers because the demo Space looks empty.** Judges who notice are merciless. Either compute real numbers from `data/eval.jsonl` against the merged model or hide the column.
- ❌ **Run `pip install -U` of anything inside the venue venv.** The G10/G12 smoke gates passed on the current pin set; any upgrade voids that.

---

## 8. Reproduction Recipes (use when verifying)

### 8.1 Verify the flat-reward bug is real

```bash
# Convert the UTF-16 log Cursor wrote, then count how many steps had reward != 0
iconv -f UTF-16LE -t UTF-8 job9b_full.log 2>/dev/null \
  | grep -oE "'rewards/_instrumented_reward_fn/mean': [0-9.]+" \
  | sort -u
# Expected: only "'rewards/.../mean': 0.0" β€” confirms 100% flat
```

### 8.2 Verify the prompt-shape mismatch

```bash
python - <<'PY'
import json
sft = json.loads(open("data/sft_traces.jsonl", encoding="utf-8").readline())
print("SFT user content head:", repr(sft["messages"][1]["content"][:80]))
# Expected: starts with "Question: "

train = json.loads(open("data/train.jsonl", encoding="utf-8").readline())
# Simulate train/grpo.py's _to_prompt user content (PRE-FIX)
print("GRPO user content head (pre-fix):", repr(f"Context:\n{train['context'][:60]}\n\n{train['prompt']}"[:80]))
# Expected: starts with "Context:\n" β€” confirms the drift
PY
```

### 8.3 After A.3 fixes, dry-run the smoke test locally

```bash
python -m uvicorn env.server.app:app --host 0.0.0.0 --port 8001 &
sleep 5
python -m train.smoke_test --env-url http://localhost:8001
```

Expected: `outputs/smoke/SMOKE_RESULT.md` shows VERDICT: GO.

### 8.4 Verify Space deploy succeeded

```bash
for url in "/" "/healthz" "/state" "/docs"; do
  printf "GET $url β†’ "
  curl -s -o /dev/null -w "%{http_code}\n" "https://Pratham-math-fathom-env.hf.space$url"
done
# Expected: 200, 200, 200, 200 β€” currently / returns 404.
```

---

## 9. Glossary (Antigravity is cold; here's the cheat sheet)

- **OpenEnv** β€” Meta's standard for RL environments. Defines the JSON contract `/reset`, `/step`, `/state`, `/healthz`. We pin `openenv-core>=0.2.3,<0.3`.
- **GRPO** β€” Group Relative Policy Optimization. The trainer samples N completions (here 8), computes per-group advantages (relative to group mean/std), and updates the policy. **Critical:** if all N completions get the same reward, std=0 β†’ advantage=0 β†’ no update. That's exactly our v1 failure mode.
- **TRL** β€” Hugging Face's RLHF/GRPO trainer library. We're on `trl==1.2.0`.
- **Unsloth** β€” Memory-efficient LoRA + 4-bit loader. We use `Qwen2.5-Coder-1.5B-Instruct-bnb-4bit` + LoRA r=16.
- **RLM (Recursive Language Model)** β€” a scaffold where an LM can call itself recursively on document chunks. We cap depth at 2 in training, 4 at demo.
- **Format gate** β€” the `<answer>...</answer>` regex check. Was multiplicative (binary 0/1), now should be additive (+0.10 bonus).
- **vLLM colocate** β€” vLLM runs in the same process as TRL trainer, sharing the GPU. Required for multi-turn (we don't use multi-turn in training but colocate is also cheaper memory-wise).

---

## 10. Definition of Done

You are done when **all** of these are true at once:

- [ ] `python scripts/submission_preflight.py` prints `Preflight passed. Submission package looks judge-ready.`
- [ ] `curl -s -o /dev/null -w "%{http_code}" https://Pratham-math-fathom-env.hf.space/` returns **200** (not 404).
- [ ] `https://Pratham-math-fathom-demo.hf.space` renders a Streamlit page in <60 s with no placeholder/fake numbers visible.
- [ ] README's "Submission Links" table has zero `_to be added_` placeholders.
- [ ] Mini-blog space is live and linked from README.
- [ ] GitHub mirror is live and linked from README.
- [ ] `outputs/plots/grpo_reward.png` either (a) shows non-zero variance after A.4 retrain, or (b) is honestly framed in README as v1 with diagnosis (per A.5 fallback).
- [ ] D.2 paragraph appears in README β€” env's training-vs-inference role is explicit.
- [ ] `git status` is clean on `master`; remote HF master + GitHub main are pushed and in sync.
- [ ] Smoke test green on HF Jobs (`outputs/smoke/SMOKE_RESULT.md` GO timestamp within last 24 h).

If any item is unchecked, you are **not** done. Do not declare victory and stop.

---

## 11. Hand-back Format

When you finish, append a single block to the bottom of this file:

```
## ANTIGRAVITY SESSION RESULT β€” <YYYY-MM-DD HH:MM IST>

- A (reward curve): <fixed and retrained / honestly reframed / blocked because X>
- B (Space UI): <env index + demo space deployed / partial / blocked>
- C (materials): <blog + GH mirror linked / partial / blocked>
- D (truth-in-advertising): <paragraph in / not in>
- Final preflight: <PASS / FAIL+reason>
- Cloud spend: $<x>
- Commits: <list of commit short-shas>

Open risks for the user before submission:
- ...
```

Stop after that block. Do not push the brief itself; the user committed it.