fathom-code / train /grpo.py

Commit History

fix(train): accept wandb_v1_ key format in W&B gate check
9c646f0

23f2002275 Claude Sonnet 4.6 commited on

fix(grpo): allow FATHOM_USE_VLLM=0 to bypass vLLM rollout and avoid IS-ratio collapse under QLoRA
d92866e

23f2002275 Claude Sonnet 4.6 commited on

fix(train): tolerate missing/invalid WANDB_API_KEY, surface key length in smoke
1bf8189

23f2002275 commited on

fix(reward): align prompt with SFT, soft format, real recursion signal (A.3 + A.4-bis)
fa599d5

23f2002275 commited on

fix(grpo): align sys_msg with SFT - emit <answer>...</answer>
916512d

23f2002275 commited on

fix(grpo): tail-truncate context inside _to_prompt to satisfy vllm input-length gate
4e04091
verified

Pratham-math commited on

fix(grpo): TRL 1.2 stock path + LoRA double-wrap fix + smoke memory caps
75c8036
verified

Pratham-math commited on

feat: phase 1 complete — smoke green on HF Jobs, training scripts, plot generator, Colab notebook, submission preflight
8787bd3

23f2002275 commited on

clean repo without secrets or data
071ba6b

23f2002275 commited on