svd-code / gpu-sft /EVAL_RESULTS.md
fzzhang's picture
Upload folder using huggingface_hub
58258b8 verified
|
Raw History Blame Contribute Delete
2.5 kB
# GPU eval results (PARTIAL) β€” captured 2026-09-13 before a machine reorg
Raw per-task JSONs live on HDFS at `/mnt/hdfs/fangzhao_writable/marin-eval/results`
(durable across the reorg). This file is the accessible-anywhere (GitHub) copy of the
headline numbers so they're not gated behind HDFS access.
**Eval stack:** driver 535, cu124 **vLLM 0.8.5 + evalchemy@7f24168**, `gpu_eval_driver.py`,
tp=8, temp 0.7 / top_p 1.0 / max_gen 32768, seeds 42..51 (AIME*/AMC23/HMMT n=10; MATH500 &
OlympiadBench n=1; LiveCodeBench* n=6). Per the port's design, GPU numbers are **not
bit-comparable** to the TPU tables (different kernels/seeding) β€” deltas are computed vs the
`qwen3_8b_base` row below.
## qwen3_8b_base β€” baseline (the (+Ξ”) denominator)
### math
| task | mean | std | seeds |
|---|---|---|---|
| MATH500 | 0.9080 | – | 1 |
| OlympiadBench | 0.6795 | – | 1 |
| AIME24 | 0.7500 | 0.0478 | 10 |
| AIME25 | 0.6467 | 0.0358 | 10 |
| AIME26 | 0.6833 | 0.0572 | 10 |
| AMC23 | 0.9425 | 0.0290 | 10 |
| HMMT | **FAILED** (exited 1 @ 11s β€” bug, see `_raw/HMMT.evalchemy.log`) | – | – |
### code
| task | mean | std | seeds |
|---|---|---|---|
| LiveCodeBench | 0.6181 | 0.0230 | 6 |
| LiveCodeBenchv5_official | not finished (killed mid-run) | – | – |
| LiveCodeBenchv6_official | pending | – | – |
## Gate 2 (eval parity) β€” PASS
GPU AIME24 **0.750 [Β±0.048]** vs TPU reference **0.740 [Β±0.049]** β†’ within Β±1Οƒ. GPU baseline
runs slightly higher than the TPU reference across the board (MATH500 90.8 vs 89.2,
OlympiadBench 68.0 vs 62.5, AIME25 64.7 vs 60.0, AIME26 68.3 vs 62.0) β€” a consistent, small
upward shift consistent with the known kernel/seeding differences (not bit-comparable, as
designed). The eval stack is validated.
## Still pending (resume on the next machine)
- **baseline:** HMMT (fix the quick failure), LiveCodeBench v5/v6.
- **experiments (needed for the deltas):** `hs_competition_nofilter` (math),
`coding_depth4v3_nofilter` (code). A last-hour salvage of hs AIME24 may have banked β€”
append here if it did.
- then `claude/compile_results_local.py` over the HDFS results β†’ `results.md` with `(+Ξ”)`.
## Trained models (durable on HDFS, final step-1999)
- `/mnt/hdfs/fangzhao_writable/marin-sft/coding_depth4v3_nofilter/hf/step-1999` β€” DONE (2000 steps)
- `/mnt/hdfs/fangzhao_writable/marin-sft/hs_competition_nofilter/hf/step-1999` β€” DONE (2000 steps)
- `science_depth4v3_nofilter` β€” re-running fresh (`_v3`) on a separate machine