# gpu-sft — Qwen3 self-instill SFT + eval (TPU→GPU port) Downstream of the SDG pipeline: **train Qwen3 on the self-instill datasets** (the `fzzhang/qwen3_8b_*_instill_...` sets our SDG runs produced) → **eval** on math/science/code → regenerate **`results.md`**. ## Provenance All scripts here were **extracted verbatim** from Tony's handoff doc `fangzhao.md` (Appendix A), via a fenced-code parser — not hand-typed. **They have NOT been run on a GPU** (written on macOS, no CUDA). Treat as a strong first draft to smoke-test on one node; the acceptance tests are the two parity gates in §10 of the handoff. Do not trust any experiment number until Gate 1 (training-loss parity) and Gate 2 (eval-baseline parity) pass. ## Pipeline ``` HF dataset@rev → scripts/gpu_sft/prepare_sft_data.py (CPU) → packed .npy + assistant mask → scripts/gpu_sft/train_sft_qwen3.py (GPU FSDP) → hf/step-N/ (bf16) → scripts/gpu_eval/gpu_eval_driver.py (vLLM+evalchemy) → eval JSON → claude/compile_results_local.py → results.md ``` ## Environment (two venvs — mandatory; see handoff §5.1) - `sft-train`: torch 2.7 cu124 + transformers 4.57 + flash-attn (this dir's `requirements-gpu.txt`). - `sft-eval`: vLLM (its own torch) + `datasets>=3.1,<4` + pinned evalchemy (via `scripts/gpu_eval/bootstrap_gpu_eval.sh`). - `HF_HOME` on the big local disk, NOT `$HOME`. Pin `datasets<4.0` in both (v4 breaks LiveCodeBench). - Gated licences: the `fzzhang` HF account must accept `cais/hle` AND `Idavidrein/gpqa` before any eval. ## Storage plan (this cluster: Azure ND96isr H100 v5, 4 nodes × 8×H100-80G, not preemptible) - **No HDFS, no cross-node FS.** 8B@32K fits on ONE node → run **4 independent single-node experiments**, one per node. - Working dir + `HF_HOME` + active checkpoints → **local NVMe** (`/`, ~3.3 TB free; one 8B experiment ~2 TB fits). - Durable artifact = **bf16 `hf/step-N` exports → private HF repo** `fzzhang/-stepN` (existing workflow; durable + eval-anywhere). - Tradeoff: HF export drops the fp32 optimizer state, so no bit-exact training resume. Given not-preemptible + ~1.5-day runs, accept re-run risk. Revisit HDFS write-perm only if bit-exact auto-resume is needed. ## Status - **Phase 0 (zero-GPU): DONE** — all 8 Python files compile; the `results.md` compiler is validated end-to-end via `claude/make_fixture_eval_tree.py` → `claude/compile_results_local.py` (correct §9 format). - **Phase 1 (1 node): pending** — two venvs + gated-licence preflight → `prepare_sft_data.py` on one dataset (with the §10 sanity asserts) → tiny smoke train → parity gate (`coding_depth4v3_n8_vr5`, 2000 steps) vs Tony's W&B (`tonyhlee/instill_ot4`). - **Phase 2:** GPU baselines (Qwen3-8B) into a fresh results root → experiments. ## Zero-GPU compiler self-test (reproduce anywhere) ``` python claude/make_fixture_eval_tree.py /tmp/fixture python claude/compile_results_local.py --root /tmp/fixture/evalout \ --models /tmp/fixture/evals.md --output /tmp/fixture/results.md \ --timestamp "2026-09-01 12:00" --no-cache cat /tmp/fixture/results.md ```