svd-code / gpu-sft /EVAL_RESULTS.md
fzzhang's picture
Upload folder using huggingface_hub
58258b8 verified
|
Raw History Blame Contribute Delete
2.5 kB

GPU eval results (PARTIAL) β€” captured 2026-09-13 before a machine reorg

Raw per-task JSONs live on HDFS at /mnt/hdfs/fangzhao_writable/marin-eval/results (durable across the reorg). This file is the accessible-anywhere (GitHub) copy of the headline numbers so they're not gated behind HDFS access.

Eval stack: driver 535, cu124 vLLM 0.8.5 + evalchemy@7f24168, gpu_eval_driver.py, tp=8, temp 0.7 / top_p 1.0 / max_gen 32768, seeds 42..51 (AIME*/AMC23/HMMT n=10; MATH500 & OlympiadBench n=1; LiveCodeBench* n=6). Per the port's design, GPU numbers are not bit-comparable to the TPU tables (different kernels/seeding) β€” deltas are computed vs the qwen3_8b_base row below.

qwen3_8b_base β€” baseline (the (+Ξ”) denominator)

math

task mean std seeds
MATH500 0.9080 – 1
OlympiadBench 0.6795 – 1
AIME24 0.7500 0.0478 10
AIME25 0.6467 0.0358 10
AIME26 0.6833 0.0572 10
AMC23 0.9425 0.0290 10
HMMT FAILED (exited 1 @ 11s β€” bug, see _raw/HMMT.evalchemy.log) – –

code

task mean std seeds
LiveCodeBench 0.6181 0.0230 6
LiveCodeBenchv5_official not finished (killed mid-run) – –
LiveCodeBenchv6_official pending – –

Gate 2 (eval parity) β€” PASS

GPU AIME24 0.750 [Β±0.048] vs TPU reference 0.740 [Β±0.049] β†’ within Β±1Οƒ. GPU baseline runs slightly higher than the TPU reference across the board (MATH500 90.8 vs 89.2, OlympiadBench 68.0 vs 62.5, AIME25 64.7 vs 60.0, AIME26 68.3 vs 62.0) β€” a consistent, small upward shift consistent with the known kernel/seeding differences (not bit-comparable, as designed). The eval stack is validated.

Still pending (resume on the next machine)

  • baseline: HMMT (fix the quick failure), LiveCodeBench v5/v6.
  • experiments (needed for the deltas): hs_competition_nofilter (math), coding_depth4v3_nofilter (code). A last-hour salvage of hs AIME24 may have banked β€” append here if it did.
  • then claude/compile_results_local.py over the HDFS results β†’ results.md with (+Ξ”).

Trained models (durable on HDFS, final step-1999)

  • /mnt/hdfs/fangzhao_writable/marin-sft/coding_depth4v3_nofilter/hf/step-1999 β€” DONE (2000 steps)
  • /mnt/hdfs/fangzhao_writable/marin-sft/hs_competition_nofilter/hf/step-1999 β€” DONE (2000 steps)
  • science_depth4v3_nofilter β€” re-running fresh (_v3) on a separate machine