File size: 3,167 Bytes
58258b8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
# gpu-sft — Qwen3 self-instill SFT + eval (TPU→GPU port)

Downstream of the SDG pipeline: **train Qwen3 on the self-instill datasets** (the
`fzzhang/qwen3_8b_*_instill_...` sets our SDG runs produced) β†’ **eval** on
math/science/code β†’ regenerate **`results.md`**.

## Provenance
All scripts here were **extracted verbatim** from Tony's handoff doc
`fangzhao.md` (Appendix A), via a fenced-code parser β€” not hand-typed.
**They have NOT been run on a GPU** (written on macOS, no CUDA). Treat as a strong
first draft to smoke-test on one node; the acceptance tests are the two parity
gates in Β§10 of the handoff. Do not trust any experiment number until Gate 1
(training-loss parity) and Gate 2 (eval-baseline parity) pass.

## Pipeline
```
HF dataset@rev β†’ scripts/gpu_sft/prepare_sft_data.py  (CPU) β†’ packed .npy + assistant mask
              β†’ scripts/gpu_sft/train_sft_qwen3.py     (GPU FSDP) β†’ hf/step-N/ (bf16)
              β†’ scripts/gpu_eval/gpu_eval_driver.py    (vLLM+evalchemy) β†’ eval JSON
              β†’ claude/compile_results_local.py        β†’ results.md
```

## Environment (two venvs β€” mandatory; see handoff Β§5.1)
- `sft-train`: torch 2.7 cu124 + transformers 4.57 + flash-attn (this dir's `requirements-gpu.txt`).
- `sft-eval`: vLLM (its own torch) + `datasets>=3.1,<4` + pinned evalchemy (via `scripts/gpu_eval/bootstrap_gpu_eval.sh`).
- `HF_HOME` on the big local disk, NOT `$HOME`. Pin `datasets<4.0` in both (v4 breaks LiveCodeBench).
- Gated licences: the `fzzhang` HF account must accept `cais/hle` AND `Idavidrein/gpqa` before any eval.

## Storage plan (this cluster: Azure ND96isr H100 v5, 4 nodes Γ— 8Γ—H100-80G, not preemptible)
- **No HDFS, no cross-node FS.** 8B@32K fits on ONE node β†’ run **4 independent single-node experiments**, one per node.
- Working dir + `HF_HOME` + active checkpoints β†’ **local NVMe** (`/`, ~3.3 TB free; one 8B experiment ~2 TB fits).
- Durable artifact = **bf16 `hf/step-N` exports β†’ private HF repo** `fzzhang/<exp>-stepN` (existing workflow; durable + eval-anywhere).
- Tradeoff: HF export drops the fp32 optimizer state, so no bit-exact training resume. Given not-preemptible + ~1.5-day runs, accept re-run risk. Revisit HDFS write-perm only if bit-exact auto-resume is needed.

## Status
- **Phase 0 (zero-GPU): DONE** β€” all 8 Python files compile; the `results.md` compiler is validated end-to-end via `claude/make_fixture_eval_tree.py` β†’ `claude/compile_results_local.py` (correct Β§9 format).
- **Phase 1 (1 node): pending** β€” two venvs + gated-licence preflight β†’ `prepare_sft_data.py` on one dataset (with the Β§10 sanity asserts) β†’ tiny smoke train β†’ parity gate (`coding_depth4v3_n8_vr5`, 2000 steps) vs Tony's W&B (`tonyhlee/instill_ot4`).
- **Phase 2:** GPU baselines (Qwen3-8B) into a fresh results root β†’ experiments.

## Zero-GPU compiler self-test (reproduce anywhere)
```
python claude/make_fixture_eval_tree.py /tmp/fixture
python claude/compile_results_local.py --root /tmp/fixture/evalout \
    --models /tmp/fixture/evals.md --output /tmp/fixture/results.md \
    --timestamp "2026-09-01 12:00" --no-cache
cat /tmp/fixture/results.md
```