svd-code / gpu-sft /claude /evals.md
fzzhang's picture
Upload folder using huggingface_hub
58258b8 verified
|
Raw History Blame Contribute Delete
1.66 kB

Local Evals

The compiler (claude/compile_results_local.py) parses ONLY the ## Models to Evaluate section below. Everything else in this file is prose for humans.

Rules:

  • ### Baselines entries are the delta reference. One baseline per model size; the size token (0.6B, 1.7B, 4B, 8B, 32B) is read out of the name. A ## Qwen3 <SIZE> section is emitted only if a baseline for that size exists.
  • Every other ### ... heading is an experiment group. The size comes from the heading text, falling back to the experiment name.
  • Each - <name> under an experiment heading must be the EXACT experiment name, which is also the eval output directory prefix: <name>-step<N>/. Do not truncate, do not add a -step<N> suffix here.
  • Row order inside a suite table follows this file. The #### Best Checkpoints table re-sorts by sampling compute (nofilter first, then _n<N> asc, _vr<N> asc, allvalid last).
  • Suite routing is a substring match on the name: science -> Science; code / coding / rstarcoder -> Code; otherwise Math. The suite decides which benchmarks are read AND which table the row lands in.
  • Section order across sizes is int(re.sub(r"\D", "", size)), so 0.6B sorts as 6 and 1.7B as 17: 4B, 0.6B, 8B, 1.7B, 32B. Known quirk, kept for continuity with the existing results.md.

Models to Evaluate

Baselines

  • Qwen/Qwen3-4B
  • Qwen/Qwen3-8B

SFT 4B Experiments

  • exp_sft_qwen3_4b_selfinstill_ot3_math53k_n8_vr5_round1

SFT 8B Experiments

  • exp_sft_qwen3_8b_selfinstill_science_depth4v3_n8_vr5_2k_lr5e6_wd01
  • exp_sft_qwen3_8b_selfinstill_coding_depth4v3_n8_vr5_2k_lr5e6_wd01