svd-code / gpu-sft /README.md
fzzhang's picture
Upload folder using huggingface_hub
58258b8 verified
|
Raw History Blame Contribute Delete
3.17 kB

gpu-sft — Qwen3 self-instill SFT + eval (TPU→GPU port)

Downstream of the SDG pipeline: train Qwen3 on the self-instill datasets (the fzzhang/qwen3_8b_*_instill_... sets our SDG runs produced) → eval on math/science/code → regenerate results.md.

Provenance

All scripts here were extracted verbatim from Tony's handoff doc fangzhao.md (Appendix A), via a fenced-code parser — not hand-typed. They have NOT been run on a GPU (written on macOS, no CUDA). Treat as a strong first draft to smoke-test on one node; the acceptance tests are the two parity gates in §10 of the handoff. Do not trust any experiment number until Gate 1 (training-loss parity) and Gate 2 (eval-baseline parity) pass.

Pipeline

HF dataset@rev → scripts/gpu_sft/prepare_sft_data.py  (CPU) → packed .npy + assistant mask
              → scripts/gpu_sft/train_sft_qwen3.py     (GPU FSDP) → hf/step-N/ (bf16)
              → scripts/gpu_eval/gpu_eval_driver.py    (vLLM+evalchemy) → eval JSON
              → claude/compile_results_local.py        → results.md

Environment (two venvs — mandatory; see handoff §5.1)

  • sft-train: torch 2.7 cu124 + transformers 4.57 + flash-attn (this dir's requirements-gpu.txt).
  • sft-eval: vLLM (its own torch) + datasets>=3.1,<4 + pinned evalchemy (via scripts/gpu_eval/bootstrap_gpu_eval.sh).
  • HF_HOME on the big local disk, NOT $HOME. Pin datasets<4.0 in both (v4 breaks LiveCodeBench).
  • Gated licences: the fzzhang HF account must accept cais/hle AND Idavidrein/gpqa before any eval.

Storage plan (this cluster: Azure ND96isr H100 v5, 4 nodes × 8×H100-80G, not preemptible)

  • No HDFS, no cross-node FS. 8B@32K fits on ONE node → run 4 independent single-node experiments, one per node.
  • Working dir + HF_HOME + active checkpoints → local NVMe (/, ~3.3 TB free; one 8B experiment ~2 TB fits).
  • Durable artifact = bf16 hf/step-N exports → private HF repo fzzhang/<exp>-stepN (existing workflow; durable + eval-anywhere).
  • Tradeoff: HF export drops the fp32 optimizer state, so no bit-exact training resume. Given not-preemptible + ~1.5-day runs, accept re-run risk. Revisit HDFS write-perm only if bit-exact auto-resume is needed.

Status

  • Phase 0 (zero-GPU): DONE — all 8 Python files compile; the results.md compiler is validated end-to-end via claude/make_fixture_eval_tree.py → claude/compile_results_local.py (correct §9 format).
  • Phase 1 (1 node): pending — two venvs + gated-licence preflight → prepare_sft_data.py on one dataset (with the §10 sanity asserts) → tiny smoke train → parity gate (coding_depth4v3_n8_vr5, 2000 steps) vs Tony's W&B (tonyhlee/instill_ot4).
  • Phase 2: GPU baselines (Qwen3-8B) into a fresh results root → experiments.

Zero-GPU compiler self-test (reproduce anywhere)

python claude/make_fixture_eval_tree.py /tmp/fixture
python claude/compile_results_local.py --root /tmp/fixture/evalout \
    --models /tmp/fixture/evals.md --output /tmp/fixture/results.md \
    --timestamp "2026-09-01 12:00" --no-cache
cat /tmp/fixture/results.md