|
Download gpu-sft/EVAL_RESULTS.md from fzzhang/svd-code: direct link, hf CLI and curl.
- Browser
- Download file 2.5 kB
-
https://huggingface.co/fzzhang/svd-code/resolve/main/gpu-sft/EVAL_RESULTS.md
- Command line
-
hf download hf://fzzhang/svd-code/gpu-sft/EVAL_RESULTS.md
-
curl -L -o EVAL_RESULTS.md https://huggingface.co/fzzhang/svd-code/resolve/main/gpu-sft/EVAL_RESULTS.md
2.5 kB
| # GPU eval results (PARTIAL) β captured 2026-09-13 before a machine reorg | |
| Raw per-task JSONs live on HDFS at `/mnt/hdfs/fangzhao_writable/marin-eval/results` | |
| (durable across the reorg). This file is the accessible-anywhere (GitHub) copy of the | |
| headline numbers so they're not gated behind HDFS access. | |
| **Eval stack:** driver 535, cu124 **vLLM 0.8.5 + evalchemy@7f24168**, `gpu_eval_driver.py`, | |
| tp=8, temp 0.7 / top_p 1.0 / max_gen 32768, seeds 42..51 (AIME*/AMC23/HMMT n=10; MATH500 & | |
| OlympiadBench n=1; LiveCodeBench* n=6). Per the port's design, GPU numbers are **not | |
| bit-comparable** to the TPU tables (different kernels/seeding) β deltas are computed vs the | |
| `qwen3_8b_base` row below. | |
| ## qwen3_8b_base β baseline (the (+Ξ) denominator) | |
| ### math | |
| | task | mean | std | seeds | | |
| |---|---|---|---| | |
| | MATH500 | 0.9080 | β | 1 | | |
| | OlympiadBench | 0.6795 | β | 1 | | |
| | AIME24 | 0.7500 | 0.0478 | 10 | | |
| | AIME25 | 0.6467 | 0.0358 | 10 | | |
| | AIME26 | 0.6833 | 0.0572 | 10 | | |
| | AMC23 | 0.9425 | 0.0290 | 10 | | |
| | HMMT | **FAILED** (exited 1 @ 11s β bug, see `_raw/HMMT.evalchemy.log`) | β | β | | |
| ### code | |
| | task | mean | std | seeds | | |
| |---|---|---|---| | |
| | LiveCodeBench | 0.6181 | 0.0230 | 6 | | |
| | LiveCodeBenchv5_official | not finished (killed mid-run) | β | β | | |
| | LiveCodeBenchv6_official | pending | β | β | | |
| ## Gate 2 (eval parity) β PASS | |
| GPU AIME24 **0.750 [Β±0.048]** vs TPU reference **0.740 [Β±0.049]** β within Β±1Ο. GPU baseline | |
| runs slightly higher than the TPU reference across the board (MATH500 90.8 vs 89.2, | |
| OlympiadBench 68.0 vs 62.5, AIME25 64.7 vs 60.0, AIME26 68.3 vs 62.0) β a consistent, small | |
| upward shift consistent with the known kernel/seeding differences (not bit-comparable, as | |
| designed). The eval stack is validated. | |
| ## Still pending (resume on the next machine) | |
| - **baseline:** HMMT (fix the quick failure), LiveCodeBench v5/v6. | |
| - **experiments (needed for the deltas):** `hs_competition_nofilter` (math), | |
| `coding_depth4v3_nofilter` (code). A last-hour salvage of hs AIME24 may have banked β | |
| append here if it did. | |
| - then `claude/compile_results_local.py` over the HDFS results β `results.md` with `(+Ξ)`. | |
| ## Trained models (durable on HDFS, final step-1999) | |
| - `/mnt/hdfs/fangzhao_writable/marin-sft/coding_depth4v3_nofilter/hf/step-1999` β DONE (2000 steps) | |
| - `/mnt/hdfs/fangzhao_writable/marin-sft/hs_competition_nofilter/hf/step-1999` β DONE (2000 steps) | |
| - `science_depth4v3_nofilter` β re-running fresh (`_v3`) on a separate machine | |