# GPU eval results (PARTIAL) — captured 2026-09-13 before a machine reorg Raw per-task JSONs live on HDFS at `/mnt/hdfs/fangzhao_writable/marin-eval/results` (durable across the reorg). This file is the accessible-anywhere (GitHub) copy of the headline numbers so they're not gated behind HDFS access. **Eval stack:** driver 535, cu124 **vLLM 0.8.5 + evalchemy@7f24168**, `gpu_eval_driver.py`, tp=8, temp 0.7 / top_p 1.0 / max_gen 32768, seeds 42..51 (AIME*/AMC23/HMMT n=10; MATH500 & OlympiadBench n=1; LiveCodeBench* n=6). Per the port's design, GPU numbers are **not bit-comparable** to the TPU tables (different kernels/seeding) — deltas are computed vs the `qwen3_8b_base` row below. ## qwen3_8b_base — baseline (the (+Δ) denominator) ### math | task | mean | std | seeds | |---|---|---|---| | MATH500 | 0.9080 | – | 1 | | OlympiadBench | 0.6795 | – | 1 | | AIME24 | 0.7500 | 0.0478 | 10 | | AIME25 | 0.6467 | 0.0358 | 10 | | AIME26 | 0.6833 | 0.0572 | 10 | | AMC23 | 0.9425 | 0.0290 | 10 | | HMMT | **FAILED** (exited 1 @ 11s — bug, see `_raw/HMMT.evalchemy.log`) | – | – | ### code | task | mean | std | seeds | |---|---|---|---| | LiveCodeBench | 0.6181 | 0.0230 | 6 | | LiveCodeBenchv5_official | not finished (killed mid-run) | – | – | | LiveCodeBenchv6_official | pending | – | – | ## Gate 2 (eval parity) — PASS GPU AIME24 **0.750 [±0.048]** vs TPU reference **0.740 [±0.049]** → within ±1σ. GPU baseline runs slightly higher than the TPU reference across the board (MATH500 90.8 vs 89.2, OlympiadBench 68.0 vs 62.5, AIME25 64.7 vs 60.0, AIME26 68.3 vs 62.0) — a consistent, small upward shift consistent with the known kernel/seeding differences (not bit-comparable, as designed). The eval stack is validated. ## Still pending (resume on the next machine) - **baseline:** HMMT (fix the quick failure), LiveCodeBench v5/v6. - **experiments (needed for the deltas):** `hs_competition_nofilter` (math), `coding_depth4v3_nofilter` (code). A last-hour salvage of hs AIME24 may have banked — append here if it did. - then `claude/compile_results_local.py` over the HDFS results → `results.md` with `(+Δ)`. ## Trained models (durable on HDFS, final step-1999) - `/mnt/hdfs/fangzhao_writable/marin-sft/coding_depth4v3_nofilter/hf/step-1999` — DONE (2000 steps) - `/mnt/hdfs/fangzhao_writable/marin-sft/hs_competition_nofilter/hf/step-1999` — DONE (2000 steps) - `science_depth4v3_nofilter` — re-running fresh (`_v3`) on a separate machine