Apex-2-Base

The base checkpoint of a Mixture-of-Experts LLM pretrained from scratch: 3.87B total / 1.45B active parameters, trained on 86.5B tokens. Released as a starting point for continued pretraining (CPT) and fine-tuning.

Ready for continued pretraining

  • Standard Qwen3MoeForCausalLM format. It loads and trains in transformers, vLLM and other tools that support Qwen3-MoE, with no trust_remote_code.
  • Full-precision fp32 weights, exactly as in the training checkpoint (no bf16 rounding). Checked against the training checkpoint: logits relative difference 9.3e-7, 100% next-token agreement.
  • All training settings are public: optimizer, LR schedule and the exact data mix. See the Continued pretraining guide below.
  • A pre-decay checkpoint at peak LR (stable-44500) is also available for WSD-style continuation.

Model details

Tokens trained per dataset (86.5B total, click to expand)
Domain Dataset Stage 1 Main Decay Total Share
Web FineWeb-Edu 7.00 5.89 4.09 16.98 19.6%
Web DCLM-Baseline 4.40 5.24 1.02 10.66 12.3%
Web FinePDFs (English) 1.00 4.89 1.02 6.91 8.0%
Code StarCoderData source code (15 languages + Markdown) 3.20 11.77 3.68 18.65 21.6%
Code StarCoderData GitHub issues · Jupyter · commits 0.60 0.87 1.02 2.49 2.9%
Code OpenCoder FineWeb-Code 0.60 1.48 – 2.08 2.4%
Code OpenCoder Annealing (synthetic) – 1.12 2.04 3.16 3.7%
Code Dolma 3 CraneCode (synthetic) – 0.82 0.41 1.23 1.4%
Math FineMath 3+ 1.00 2.89 1.64 5.52 6.4%
Math InfiWebMath 3+ 0.60 1.67 0.82 3.08 3.6%
Math Nemotron-CC-Math 4+ – 0.98 0.61 1.59 1.8%
Curated Cosmopedia v2 (synthetic textbooks) 0.60 2.21 2.04 4.85 5.6%
Curated peS2o (scientific papers) 0.40 4.54 0.82 5.76 6.7%
Curated StackExchange – 0.22 0.61 0.84 1.0%
Curated Wikipedia (English) 0.60 0.73 0.61 1.94 2.2%
Curated Project Gutenberg (English books) – 0.75 – 0.75 0.9%
Total 20.00 46.06 20.45 86.51 100%

In billions of tokens, computed from each stage's mixture weights (see Training data).

Architecture Decoder-only MoE Transformer, every layer MoE. GQA (16Q/4KV, head_dim 128) + QK-RMSNorm + RoPE (θ=10,000) + SwiGLU experts + RMSNorm, tied embeddings
Parameters 3,869.1M total · 1,453.2M active per token (1,142.0M non-embedding active)
Layers · d_model 32 · 2048
Experts 16 per layer, top-4 routing with renormalized probabilities, expert width 1024
Router losses load-balancing aux loss 0.01 (global-batch statistics) + router z-loss 1e-3
Context length 4,096
Tokenizer Qwen3 tokenizer (vocab 151,936). eos <|endoftext|> (151643), no chat template
Weights fp32 safetensors (15.5GB), per-expert keys
Pretraining 54,250 steps · 86.5B tokens, 1M-token batches (4,096 × 256), attention masked at document boundaries. Parts trained with DiLoCo on 2× GH200

Usage (inference)

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("YOON1v/Apex-2-Base")
model = AutoModelForCausalLM.from_pretrained("YOON1v/Apex-2-Base", dtype=torch.bfloat16, device_map="cuda")

inputs = tokenizer("def is_prime(n: int) -> bool:\n", return_tensors="pt").to("cuda")
out = model.generate(**inputs, max_new_tokens=128, do_sample=False)
print(tokenizer.decode(out[0], skip_special_tokens=True))

vLLM: LLM(model="YOON1v/Apex-2-Base", dtype="bfloat16", max_model_len=4096)

This is a base model: it continues text and does not follow instructions or chat. For chat, use Apex-2.

Continued pretraining guide

Loading

import torch
from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained(
    "YOON1v/Apex-2-Base",
    dtype=torch.float32,        # keep the fp32 weights; train with bf16 autocast (mixed precision)
    output_router_logits=True,  # adds the expert load-balancing loss (off by default)
)
  • With output_router_logits=True, the load-balancing loss (router_aux_loss_coef=0.01) is added to the loss whenever labels are passed. Without it, routing can collapse onto a few experts.
  • The router z-loss (1e-3) used in pretraining is not part of the transformers implementation. Add it yourself if you need it.

Original training setup

Optimizer AdamW 8-bit (bitsandbytes, paged), β=(0.9, 0.95), eps 1e-8, weight decay 0.1
LR schedule WSD: 1,000 warmup steps → constant peak 4e-4 (stable) → linear decay to 0 over the last 9,750 steps (20.4B tokens)
Gradient clipping 1.0
Batch · sequence 1,048,576 tokens (4,096 × 256)
Loss CE + 0.01 × load-balancing + 1e-3 × router z-loss
Packing <|endoftext|> after every document, no attention across document boundaries

Recommendations

  • Re-warm the LR. This checkpoint has finished its decay (LR 0). Start CPT with a warmup and tune the peak LR at or below the original 4e-4.
  • Replay old data. Training only on new data can erode existing skills, code especially. Mixing in some of the original data (listed below) reduces forgetting.
  • Packing. If you pack several documents into one sequence, block attention across document boundaries (e.g. transformers DataCollatorWithFlattening + FlashAttention).
  • Context length. Trained at 4,096 tokens. Going longer needs changes such as a larger rope_theta.
  • Tokenizer. Use the default settings. transformers may suggest fix_mistral_regex=True; do not enable it, because it changes tokenization from what training used.
  • Special tokens. The embeddings of <|im_start|>, <|im_end|> and the other special tokens were barely trained during pretraining. Chat formatting is learned in SFT.

Intermediate checkpoint

revision step tokens state
main 54,250 86.5B decay finished (LR 0). Final model
stable-44500 44,500 66.1B before the decay, at peak LR (4e-4)

For WSD-style continuation, stable-44500 is the natural starting point: keep training on new data at the peak LR, then run your own decay whenever you choose.

model = AutoModelForCausalLM.from_pretrained("YOON1v/Apex-2-Base", revision="stable-44500", dtype=torch.float32)

Benchmarks

Area Benchmark stable 44,500 final 54,250
Code HumanEval (pass@1, greedy) 22.0 36.6
Code HumanEval+ 18.9 32.9
Code MBPP 41.0 54.8
Code MBPP+ 34.7 46.3
Math GSM8K (8-shot CoT, flexible) 10.2 15.1
Knowledge MMLU (5-shot) 25.3 28.2
Commonsense HellaSwag (0-shot, acc_norm) 56.9 60.8
Commonsense ARC-e (0-shot, acc_norm) 61.8 66.7
Commonsense ARC-c (0-shot, acc_norm) 35.5 39.7
Commonsense PIQA (0-shot, acc_norm) 73.2 73.9
Commonsense WinoGrande (0-shot) 57.7 60.1
Commonsense LAMBADA (0-shot) 53.4 54.5

Measured by us with lm-eval 0.4.13 (HF, bf16) and EvalPlus 0.3.1 (vLLM, base prompt, greedy).

Compared with similar-size base models (official numbers; protocols differ, so this is a rough comparison):

Model Pretraining tokens HumanEval / + MBPP / + MMLU HellaSwag
Apex-2-Base 0.087T 36.6 / 32.9 54.8 / 46.3 28.2 60.8
Qwen2.5-1.5B 18T 37.2 / 32.9 60.2 / 49.6 60.9 67.9
Qwen2.5-Coder-1.5B 5.5T 43.9 / 36.6 69.2 / 58.6 53.6 61.8
Gemma-2-2B 2T 17.7 / – 29.6 (3-shot) / – 51.3 73.0 (10-shot)

With about 1/200 of the pretraining data, Apex-2-Base matches Qwen2.5-1.5B (18T tokens) on HumanEval+ (32.9). Knowledge (MMLU) and math still show the gap in pretraining scale.

Training data

Public datasets only. 86.5B tokens in three stages:

Stage Steps Tokens Web / Code / Math / Curated (%)
Stage 1 0 → 19,074 20.0B 62 / 22 / 8 / 8
Main (stable) 19,074 → 44,500 46.1B 34.8 / 34.9 / 12.0 / 18.4
Decay 44,500 → 54,250 20.4B 30 / 35 / 15 / 20
Total 86.5B 39.9 / 31.9 / 11.8 / 16.3

Tokens trained per dataset: see the expandable table at the top of Model details.

StarCoderData by language (18.65B total)
Language Tokens (B) Share of all tokens
Python 4.64 5.4%
Java 1.65 1.9%
TypeScript 1.65 1.9%
Go 1.59 1.8%
JavaScript 1.58 1.8%
SQL 1.08 1.3%
Markdown 1.08 1.2%
C 1.05 1.2%
C++ 0.98 1.1%
PHP 0.81 0.9%
C# 0.74 0.9%
Rust 0.57 0.7%
Ruby 0.42 0.5%
Kotlin 0.33 0.4%
Scala 0.26 0.3%
Shell 0.22 0.3%
  • Token counts are computed from each stage's mixture weights. Every 4,096-token sequence was drawn from a source at random according to those weights, so actual amounts can differ by about 1%.
  • Documents that also appear in the stage 1 data were removed from the later corpus, and no source was used for more than one epoch, so no data was repeated.

Limitations

  • English-centric. Other languages, including Korean, barely work.
  • Limited world knowledge (MMLU ~28%).
  • As a base model it has no instruction following, chat ability or safety tuning.
  • 4,096-token context.

License

Apache 2.0. Each training dataset has its own license and terms (ODC-By, CC-BY, CC-BY-SA, the StarCoderData terms of use, NVIDIA data licenses, etc.); see the dataset pages.

Downloads last month
213
Safetensors
Model size
4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for YOON1v/Apex-2-Base

Finetunes
1 model
Quantizations
2 models

Datasets used to train YOON1v/Apex-2-Base

Evaluation results