|
Download gpu-sft/README.md from fzzhang/svd-code: direct link, hf CLI and curl.
- Browser
- Download file 3.17 kB
-
https://huggingface.co/fzzhang/svd-code/resolve/main/gpu-sft/README.md
- Command line
-
hf download hf://fzzhang/svd-code/gpu-sft/README.md
-
curl -L -o README.md https://huggingface.co/fzzhang/svd-code/resolve/main/gpu-sft/README.md
3.17 kB
| # gpu-sft β Qwen3 self-instill SFT + eval (TPUβGPU port) | |
| Downstream of the SDG pipeline: **train Qwen3 on the self-instill datasets** (the | |
| `fzzhang/qwen3_8b_*_instill_...` sets our SDG runs produced) β **eval** on | |
| math/science/code β regenerate **`results.md`**. | |
| ## Provenance | |
| All scripts here were **extracted verbatim** from Tony's handoff doc | |
| `fangzhao.md` (Appendix A), via a fenced-code parser β not hand-typed. | |
| **They have NOT been run on a GPU** (written on macOS, no CUDA). Treat as a strong | |
| first draft to smoke-test on one node; the acceptance tests are the two parity | |
| gates in Β§10 of the handoff. Do not trust any experiment number until Gate 1 | |
| (training-loss parity) and Gate 2 (eval-baseline parity) pass. | |
| ## Pipeline | |
| ``` | |
| HF dataset@rev β scripts/gpu_sft/prepare_sft_data.py (CPU) β packed .npy + assistant mask | |
| β scripts/gpu_sft/train_sft_qwen3.py (GPU FSDP) β hf/step-N/ (bf16) | |
| β scripts/gpu_eval/gpu_eval_driver.py (vLLM+evalchemy) β eval JSON | |
| β claude/compile_results_local.py β results.md | |
| ``` | |
| ## Environment (two venvs β mandatory; see handoff Β§5.1) | |
| - `sft-train`: torch 2.7 cu124 + transformers 4.57 + flash-attn (this dir's `requirements-gpu.txt`). | |
| - `sft-eval`: vLLM (its own torch) + `datasets>=3.1,<4` + pinned evalchemy (via `scripts/gpu_eval/bootstrap_gpu_eval.sh`). | |
| - `HF_HOME` on the big local disk, NOT `$HOME`. Pin `datasets<4.0` in both (v4 breaks LiveCodeBench). | |
| - Gated licences: the `fzzhang` HF account must accept `cais/hle` AND `Idavidrein/gpqa` before any eval. | |
| ## Storage plan (this cluster: Azure ND96isr H100 v5, 4 nodes Γ 8ΓH100-80G, not preemptible) | |
| - **No HDFS, no cross-node FS.** 8B@32K fits on ONE node β run **4 independent single-node experiments**, one per node. | |
| - Working dir + `HF_HOME` + active checkpoints β **local NVMe** (`/`, ~3.3 TB free; one 8B experiment ~2 TB fits). | |
| - Durable artifact = **bf16 `hf/step-N` exports β private HF repo** `fzzhang/<exp>-stepN` (existing workflow; durable + eval-anywhere). | |
| - Tradeoff: HF export drops the fp32 optimizer state, so no bit-exact training resume. Given not-preemptible + ~1.5-day runs, accept re-run risk. Revisit HDFS write-perm only if bit-exact auto-resume is needed. | |
| ## Status | |
| - **Phase 0 (zero-GPU): DONE** β all 8 Python files compile; the `results.md` compiler is validated end-to-end via `claude/make_fixture_eval_tree.py` β `claude/compile_results_local.py` (correct Β§9 format). | |
| - **Phase 1 (1 node): pending** β two venvs + gated-licence preflight β `prepare_sft_data.py` on one dataset (with the Β§10 sanity asserts) β tiny smoke train β parity gate (`coding_depth4v3_n8_vr5`, 2000 steps) vs Tony's W&B (`tonyhlee/instill_ot4`). | |
| - **Phase 2:** GPU baselines (Qwen3-8B) into a fresh results root β experiments. | |
| ## Zero-GPU compiler self-test (reproduce anywhere) | |
| ``` | |
| python claude/make_fixture_eval_tree.py /tmp/fixture | |
| python claude/compile_results_local.py --root /tmp/fixture/evalout \ | |
| --models /tmp/fixture/evals.md --output /tmp/fixture/results.md \ | |
| --timestamp "2026-09-01 12:00" --no-cache | |
| cat /tmp/fixture/results.md | |
| ``` | |