pavlichenko's picture
Update model card and add evaluation results (#2)
1b68edc
Raw
History Blame Contribute Delete
1.03 kB
# Mellum 2 Base — evaluation results
# Self-reported by JetBrains. Numbers from Mellum 2 Technical Report (Table 5),
# measured on the pre-extension checkpoint Mellum2-12B-A2.5B-Base-Pretrain; the
# architecture is identical, only RoPE on global-attention layers was re-mapped
# via layer-selective YaRN to extend context from 8K to 128K.
# Only entries for benchmarks confirmed to be registered as HF Hub Benchmarks are listed.
- dataset:
id: Idavidrein/gpqa
task_id: diamond
value: 31.31
date: "2026-05-27"
notes: "pre-training eval (pre-YaRN), no-tools"
- dataset:
id: Idavidrein/gpqa
task_id: main
value: 35.04
date: "2026-05-27"
notes: "pre-training eval (pre-YaRN), no-tools"
- dataset:
id: TIGER-Lab/MMLU-Pro
task_id: mmlu_pro
value: 59.31
date: "2026-05-27"
notes: "pre-training eval (pre-YaRN), no-tools, exact match"
- dataset:
id: openai/gsm8k
task_id: gsm8k
value: 81.73
date: "2026-05-27"
notes: "pre-training eval (pre-YaRN), no-tools, exact match"