Download gpu-sft/EVAL_RESULTS.md from fzzhang/svd-code: direct link, hf CLI and curl.
- Browser
- Download file 2.5 kB
-
https://huggingface.co/fzzhang/svd-code/resolve/main/gpu-sft/EVAL_RESULTS.md
- Command line
-
hf download hf://fzzhang/svd-code/gpu-sft/EVAL_RESULTS.md
-
curl -L -o EVAL_RESULTS.md https://huggingface.co/fzzhang/svd-code/resolve/main/gpu-sft/EVAL_RESULTS.md
GPU eval results (PARTIAL) β captured 2026-09-13 before a machine reorg
Raw per-task JSONs live on HDFS at /mnt/hdfs/fangzhao_writable/marin-eval/results
(durable across the reorg). This file is the accessible-anywhere (GitHub) copy of the
headline numbers so they're not gated behind HDFS access.
Eval stack: driver 535, cu124 vLLM 0.8.5 + evalchemy@7f24168, gpu_eval_driver.py,
tp=8, temp 0.7 / top_p 1.0 / max_gen 32768, seeds 42..51 (AIME*/AMC23/HMMT n=10; MATH500 &
OlympiadBench n=1; LiveCodeBench* n=6). Per the port's design, GPU numbers are not
bit-comparable to the TPU tables (different kernels/seeding) β deltas are computed vs the
qwen3_8b_base row below.
qwen3_8b_base β baseline (the (+Ξ) denominator)
math
| task | mean | std | seeds |
|---|---|---|---|
| MATH500 | 0.9080 | β | 1 |
| OlympiadBench | 0.6795 | β | 1 |
| AIME24 | 0.7500 | 0.0478 | 10 |
| AIME25 | 0.6467 | 0.0358 | 10 |
| AIME26 | 0.6833 | 0.0572 | 10 |
| AMC23 | 0.9425 | 0.0290 | 10 |
| HMMT | FAILED (exited 1 @ 11s β bug, see _raw/HMMT.evalchemy.log) |
β | β |
code
| task | mean | std | seeds |
|---|---|---|---|
| LiveCodeBench | 0.6181 | 0.0230 | 6 |
| LiveCodeBenchv5_official | not finished (killed mid-run) | β | β |
| LiveCodeBenchv6_official | pending | β | β |
Gate 2 (eval parity) β PASS
GPU AIME24 0.750 [Β±0.048] vs TPU reference 0.740 [Β±0.049] β within Β±1Ο. GPU baseline runs slightly higher than the TPU reference across the board (MATH500 90.8 vs 89.2, OlympiadBench 68.0 vs 62.5, AIME25 64.7 vs 60.0, AIME26 68.3 vs 62.0) β a consistent, small upward shift consistent with the known kernel/seeding differences (not bit-comparable, as designed). The eval stack is validated.
Still pending (resume on the next machine)
- baseline: HMMT (fix the quick failure), LiveCodeBench v5/v6.
- experiments (needed for the deltas):
hs_competition_nofilter(math),coding_depth4v3_nofilter(code). A last-hour salvage of hs AIME24 may have banked β append here if it did. - then
claude/compile_results_local.pyover the HDFS results βresults.mdwith(+Ξ).
Trained models (durable on HDFS, final step-1999)
/mnt/hdfs/fangzhao_writable/marin-sft/coding_depth4v3_nofilter/hf/step-1999β DONE (2000 steps)/mnt/hdfs/fangzhao_writable/marin-sft/hs_competition_nofilter/hf/step-1999β DONE (2000 steps)science_depth4v3_nofilterβ re-running fresh (_v3) on a separate machine