feat(train): flip Path 3 -> Path 1 — production 1.5B GRPO on a100-large 9e58a11 23f2002275 commited on Apr 25
docs(D): wire training evidence - link plots from HF model repo, add W&B run 707d9ee 23f2002275 commited on Apr 25
fix(grpo): TRL 1.2 stock path + LoRA double-wrap fix + smoke memory caps 7ad981e verified Pratham-math commited on Apr 25
fix(job_train): force CUDA_VISIBLE_DEVICES=0 to dodge TRL multi-GPU entropy bug 28d7ac6 verified Pratham-math commited on Apr 25
fix: heredocs use initialize_config_dir absolute path (no caller file = relative breaks) 83021ff 23f2002275 commited on Apr 25
fix(jobs): bump bitsandbytes to 0.47 (cu128 binary) so GPU dequant works in HF Job runtime image bcf5fdb 23f2002275 commited on Apr 25
fix(job_train): drop flash-attn install to avoid build-time OOM (R3 mitigation) 7c1bf5c 23f2002275 commited on Apr 25
fix(job_train): pip install huggingface_hub before downloading data files (R1 parity with job_smoke.sh) eae16b1 23f2002275 commited on Apr 25
config: scripts/job_train.sh \u2014 Path 3 (qwen_0_5b_smoke + GRPO max_steps=50) for sanity-check run before burning 1.5B credits d2b6a7e 23f2002275 commited on Apr 25
feat: phase 1 complete — smoke green on HF Jobs, training scripts, plot generator, Colab notebook, submission preflight 8787bd3 23f2002275 commited on Apr 25