|
Download gpu-sft/README.md from fzzhang/svd-code: direct link, hf CLI and curl.
- Browser
- Download file 3.17 kB
-
https://huggingface.co/fzzhang/svd-code/resolve/main/gpu-sft/README.md
- Command line
-
hf download hf://fzzhang/svd-code/gpu-sft/README.md
-
curl -L -o README.md https://huggingface.co/fzzhang/svd-code/resolve/main/gpu-sft/README.md
3.17 kB
gpu-sft — Qwen3 self-instill SFT + eval (TPU→GPU port)
Downstream of the SDG pipeline: train Qwen3 on the self-instill datasets (the
fzzhang/qwen3_8b_*_instill_... sets our SDG runs produced) → eval on
math/science/code → regenerate results.md.
Provenance
All scripts here were extracted verbatim from Tony's handoff doc
fangzhao.md (Appendix A), via a fenced-code parser — not hand-typed.
They have NOT been run on a GPU (written on macOS, no CUDA). Treat as a strong
first draft to smoke-test on one node; the acceptance tests are the two parity
gates in §10 of the handoff. Do not trust any experiment number until Gate 1
(training-loss parity) and Gate 2 (eval-baseline parity) pass.
Pipeline
HF dataset@rev → scripts/gpu_sft/prepare_sft_data.py (CPU) → packed .npy + assistant mask
→ scripts/gpu_sft/train_sft_qwen3.py (GPU FSDP) → hf/step-N/ (bf16)
→ scripts/gpu_eval/gpu_eval_driver.py (vLLM+evalchemy) → eval JSON
→ claude/compile_results_local.py → results.md
Environment (two venvs — mandatory; see handoff §5.1)
sft-train: torch 2.7 cu124 + transformers 4.57 + flash-attn (this dir'srequirements-gpu.txt).sft-eval: vLLM (its own torch) +datasets>=3.1,<4+ pinned evalchemy (viascripts/gpu_eval/bootstrap_gpu_eval.sh).HF_HOMEon the big local disk, NOT$HOME. Pindatasets<4.0in both (v4 breaks LiveCodeBench).- Gated licences: the
fzzhangHF account must acceptcais/hleANDIdavidrein/gpqabefore any eval.
Storage plan (this cluster: Azure ND96isr H100 v5, 4 nodes × 8×H100-80G, not preemptible)
- No HDFS, no cross-node FS. 8B@32K fits on ONE node → run 4 independent single-node experiments, one per node.
- Working dir +
HF_HOME+ active checkpoints → local NVMe (/, ~3.3 TB free; one 8B experiment ~2 TB fits). - Durable artifact = bf16
hf/step-Nexports → private HF repofzzhang/<exp>-stepN(existing workflow; durable + eval-anywhere). - Tradeoff: HF export drops the fp32 optimizer state, so no bit-exact training resume. Given not-preemptible + ~1.5-day runs, accept re-run risk. Revisit HDFS write-perm only if bit-exact auto-resume is needed.
Status
- Phase 0 (zero-GPU): DONE — all 8 Python files compile; the
results.mdcompiler is validated end-to-end viaclaude/make_fixture_eval_tree.py→claude/compile_results_local.py(correct §9 format). - Phase 1 (1 node): pending — two venvs + gated-licence preflight →
prepare_sft_data.pyon one dataset (with the §10 sanity asserts) → tiny smoke train → parity gate (coding_depth4v3_n8_vr5, 2000 steps) vs Tony's W&B (tonyhlee/instill_ot4). - Phase 2: GPU baselines (Qwen3-8B) into a fresh results root → experiments.
Zero-GPU compiler self-test (reproduce anywhere)
python claude/make_fixture_eval_tree.py /tmp/fixture
python claude/compile_results_local.py --root /tmp/fixture/evalout \
--models /tmp/fixture/evals.md --output /tmp/fixture/results.md \
--timestamp "2026-09-01 12:00" --no-cache
cat /tmp/fixture/results.md