APEX-2

A from-scratch Mixture-of-Experts LLM: 3.87B total / 1.45B active parameters, pretrained on 86.5B tokens and instruction-tuned with SFT.

This checkpoint is the SFT release (pretrain → 2-stage SFT). A DPO stage was tried and dropped because it lowered code, math and instruction-following scores (see below).

Model details

Architecture Decoder-only MoE Transformer, every layer MoE. GQA (16Q/4KV, head_dim 128) + QK-RMSNorm + RoPE (θ=10K) + SwiGLU experts + RMSNorm, weight tying
Parameters 3,869.1M total · 1,453.2M active per token (1,142.0M non-embedding active), measured
Layers · d_model 32 · 2048
Experts 16 per layer, top-4 routing with renormalized probabilities, expert width 1024
Router losses load-balancing aux loss 0.01 (global-batch statistics) + router z-loss 1e-3
Context length 4096
Vocab 151,936 (Qwen3 tokenizer), tied embeddings
Weights stored as bfloat16 (~7.7GB), cast from the float32 training export
Pretrain 54,250 steps · 86.5B tokens (web · code in 11 languages · math · curated), 1M-token batches; the last 20.4B tokens are an LR decay phase (web 30 / code 35 / math 15 / curated 20); parts trained with DiLoCo on 2× GH200
Post-training SFT stage 1: 2.21B tokens (general 29 / code 44 / math 27). SFT stage 2: 0.28B tokens × 2 epochs (execution-verified code, competition math, precise instruction following)
HF format Weights map 1:1 onto Qwen3MoeForCausalLM (the training model uses fused expert tensors; the export splits them per expert). Export check: logits relative difference 1.0e-6 and 100% next-token agreement against the training checkpoint

Usage

The model uses the ChatML template (<|im_start|>role\n...<|im_end|>) shipped in tokenizer_config.json; generation ends at <|im_end|>.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("YOON1v/Apex-2", dtype=torch.bfloat16, device_map="cuda")
tokenizer = AutoTokenizer.from_pretrained("YOON1v/Apex-2")

messages = [{"role": "user", "content": "Write a Python function that checks if a number is prime."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt", return_dict=True).to("cuda")
out = model.generate(**inputs, max_new_tokens=512, do_sample=True, temperature=0.7, top_p=0.9)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

vLLM loads it directly as well: LLM(model="YOON1v/Apex-2", dtype="bfloat16", max_model_len=4096).

Benchmarks

Measured by us with vLLM greedy decoding and the chat template (prompt + answer ≤ 4096 tokens). Model-written code was executed in a sandbox.

Area Benchmark Apex-2 (SFT) Base
Code HumanEval 43.9 36.6
Code HumanEval+ 41.5 32.9
Code MBPP 56.3 54.8
Code MBPP+ 48.9 46.3
Code MultiPL-E HumanEval C++ 36.0 —
Code MultiPL-E MBPP C++ 41.6 —
Code LiveCodeBench v5–v6 3.2 —
Code └ LiveCodeBench easy (84) 11.9 —
Code └ LiveCodeBench medium (104) 1.0 —
Code └ LiveCodeBench hard (154) 0.0 —
Code CRUXEval-O 8.5 —
Code CRUXEval-I 3.2 —
Math GSM8K 32.4 (0-shot CoT) 15.1 (8-shot)
Math MATH-500 (0-shot CoT) 21.0 —
Instruction following IFEval prompt strict 44.7 —
Instruction following IFEval instruction strict 56.6 —
Knowledge MMLU (5-shot) 28.6 28.2
Commonsense HellaSwag (acc_norm) 62.2 60.8
Commonsense ARC-e (acc_norm) 64.1 66.7
Commonsense ARC-c (acc_norm) 38.4 39.7
Commonsense PIQA (acc_norm) 74.8 73.9
Commonsense WinoGrande (acc) 61.8 60.1
Commonsense LAMBADA (acc) 52.8 54.5

Compared with similar-size models (official numbers; protocols differ, so this is a rough comparison):

Model Pretraining tokens HumanEval HumanEval+ MBPP MBPP+ GSM8K IFEval MMLU
Apex-2 (3.87B · 1.45B active) 0.087T 43.9 41.5 56.3 48.9 32.4 44.7 28.6
Apex-1 DPO (1.1B dense) 0.02T 8.5 — 5.2 — 1.9 — 24.9
Qwen2.5-1.5B-Instruct 18T 61.6 — 63.2 — 73.2 42.5 50.7
Qwen2.5-Coder-1.5B-Instruct 5.5T 70.7 66.5 69.2 59.4 — — —
Qwen3-1.7B (non-thinking) 36T — — — — — 68.2 64.4
Llama-3.2-1B-Instruct 9T — — — — 44.4 59.5* 49.3
Gemma-3-1B-it 2T 41.5 — 35.2 — 62.8 80.2* 38.8
OLMoE-1B-7B (1.3B active) 5.1T 62.3 54.4 — — 72.4 66.4* 55.1
DeepSeek-Coder-1.3B 2T 65.9 60.4 65.3 54.8 — — —

DeepSeek-Coder numbers are the Qwen2.5-Coder report's re-evaluation. * Different IFEval metric: Apex-2 and Qwen report prompt-level strict; the others report an average of metrics or do not specify.

With 1/23 to 1/400 of the pretraining data of these models, Apex-2's base model matches Qwen2.5-1.5B (18T tokens) on HumanEval+ (32.9 vs 32.9). The big gaps are knowledge (MMLU) and math, which mainly track pretraining scale.

SFT data was 13-gram decontaminated against HumanEval(+), MBPP(+), GSM8K test, MATH, MMLU, ARC, HellaSwag, PIQA, WinoGrande, IFEval and LAMBADA. For LiveCodeBench only the newer v5–v6 problems (after 2024-08) were used, to reduce contamination risk.

Why no DPO: DPO on allenai/Dolci-Instruct-DPO (220K pairs) made answers 2.3× longer (320 → 733 tokens on average). Chosen-answer likelihood also fell during training. In evaluation HumanEval+ went 41.5 → 32.3, MBPP+ 48.9 → 37.3, GSM8K 32.4 → 14.3 and IFEval 44.7 → 35.7, so this SFT checkpoint is the release.

Known limitations

  • English-centric. Other languages, including Korean, barely work (multilingual data was excluded from SFT).
  • Limited world knowledge (MMLU ~29%), so it often states wrong facts with confidence.
  • Weak at competitive programming (LiveCodeBench medium ~1%) and at predicting code execution (CRUXEval).
  • 4096-token context. No tool use, no thinking mode.

Training data

  • Pretrain: 40 deduplicated sources, including:
    • web: FineWeb-Edu, DCLM, FinePDFs
    • code: StarCoder-family and The Stack-derived code in 11 languages, OpenCoder corpora, synthetic code
    • math: FineMath, InfiWebMath, Nemotron-CC-Math
    • curated: Wikipedia, StackExchange, peS2o, Cosmopedia, Gutenberg
  • SFT:
    • general: allenai/Dolci-Instruct-SFT (tool-use, multilingual and identity subsets removed)
    • code: OpenCoder opc-sft-stage1/2, nvidia/OpenCodeInstruct (unit-test pass rate ≥ 0.9), bigcode self-oss-instruct, m-a-p/Code-Feedback
    • math: nvidia/OpenMathInstruct-2, AI-MO/NuminaMath-CoT

Some SFT sources contain synthetic data generated by third-party models (e.g. GPT-4-class, Llama-3.1-405B, Qwen2.5-Coder); check each dataset's terms for your use case.

License

Apache 2.0.

Downloads last month
352
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for YOON1v/Apex-2

Finetuned
(1)
this model

Evaluation results