Instructions to use YOON1v/Apex-2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use YOON1v/Apex-2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="YOON1v/Apex-2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("YOON1v/Apex-2") model = AutoModelForCausalLM.from_pretrained("YOON1v/Apex-2", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use YOON1v/Apex-2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "YOON1v/Apex-2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "YOON1v/Apex-2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/YOON1v/Apex-2
- SGLang
How to use YOON1v/Apex-2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "YOON1v/Apex-2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "YOON1v/Apex-2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "YOON1v/Apex-2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "YOON1v/Apex-2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use YOON1v/Apex-2 with Docker Model Runner:
docker model run hf.co/YOON1v/Apex-2
APEX-2
A from-scratch Mixture-of-Experts LLM: 3.87B total / 1.45B active parameters, pretrained on 86.5B tokens and instruction-tuned with SFT.
- Full code, architecture writeup and training diary: github.com/DW-dev-UE/LLM-from-scratch.
- The full benchmark writeup this card summarizes is BENCHMARK v3.
- Base model (for continued pretraining or your own fine-tuning): YOON1v/Apex-2-Base.
This checkpoint is the SFT release (pretrain → 2-stage SFT). A DPO stage was tried and dropped because it lowered code, math and instruction-following scores (see below).
Model details
| Architecture | Decoder-only MoE Transformer, every layer MoE. GQA (16Q/4KV, head_dim 128) + QK-RMSNorm + RoPE (θ=10K) + SwiGLU experts + RMSNorm, weight tying |
| Parameters | 3,869.1M total · 1,453.2M active per token (1,142.0M non-embedding active), measured |
| Layers · d_model | 32 · 2048 |
| Experts | 16 per layer, top-4 routing with renormalized probabilities, expert width 1024 |
| Router losses | load-balancing aux loss 0.01 (global-batch statistics) + router z-loss 1e-3 |
| Context length | 4096 |
| Vocab | 151,936 (Qwen3 tokenizer), tied embeddings |
| Weights | stored as bfloat16 (~7.7GB), cast from the float32 training export |
| Pretrain | 54,250 steps · 86.5B tokens (web · code in 11 languages · math · curated), 1M-token batches; the last 20.4B tokens are an LR decay phase (web 30 / code 35 / math 15 / curated 20); parts trained with DiLoCo on 2× GH200 |
| Post-training | SFT stage 1: 2.21B tokens (general 29 / code 44 / math 27). SFT stage 2: 0.28B tokens × 2 epochs (execution-verified code, competition math, precise instruction following) |
| HF format | Weights map 1:1 onto Qwen3MoeForCausalLM (the training model uses fused expert tensors; the export splits them per expert). Export check: logits relative difference 1.0e-6 and 100% next-token agreement against the training checkpoint |
Usage
The model uses the ChatML template (<|im_start|>role\n...<|im_end|>) shipped in tokenizer_config.json; generation ends at <|im_end|>.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("YOON1v/Apex-2", dtype=torch.bfloat16, device_map="cuda")
tokenizer = AutoTokenizer.from_pretrained("YOON1v/Apex-2")
messages = [{"role": "user", "content": "Write a Python function that checks if a number is prime."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt", return_dict=True).to("cuda")
out = model.generate(**inputs, max_new_tokens=512, do_sample=True, temperature=0.7, top_p=0.9)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
vLLM loads it directly as well: LLM(model="YOON1v/Apex-2", dtype="bfloat16", max_model_len=4096).
Benchmarks
Measured by us with vLLM greedy decoding and the chat template (prompt + answer ≤ 4096 tokens). Model-written code was executed in a sandbox.
| Area | Benchmark | Apex-2 (SFT) | Base |
|---|---|---|---|
| Code | HumanEval | 43.9 | 36.6 |
| Code | HumanEval+ | 41.5 | 32.9 |
| Code | MBPP | 56.3 | 54.8 |
| Code | MBPP+ | 48.9 | 46.3 |
| Code | MultiPL-E HumanEval C++ | 36.0 | — |
| Code | MultiPL-E MBPP C++ | 41.6 | — |
| Code | LiveCodeBench v5–v6 | 3.2 | — |
| Code | └ LiveCodeBench easy (84) | 11.9 | — |
| Code | └ LiveCodeBench medium (104) | 1.0 | — |
| Code | └ LiveCodeBench hard (154) | 0.0 | — |
| Code | CRUXEval-O | 8.5 | — |
| Code | CRUXEval-I | 3.2 | — |
| Math | GSM8K | 32.4 (0-shot CoT) | 15.1 (8-shot) |
| Math | MATH-500 (0-shot CoT) | 21.0 | — |
| Instruction following | IFEval prompt strict | 44.7 | — |
| Instruction following | IFEval instruction strict | 56.6 | — |
| Knowledge | MMLU (5-shot) | 28.6 | 28.2 |
| Commonsense | HellaSwag (acc_norm) | 62.2 | 60.8 |
| Commonsense | ARC-e (acc_norm) | 64.1 | 66.7 |
| Commonsense | ARC-c (acc_norm) | 38.4 | 39.7 |
| Commonsense | PIQA (acc_norm) | 74.8 | 73.9 |
| Commonsense | WinoGrande (acc) | 61.8 | 60.1 |
| Commonsense | LAMBADA (acc) | 52.8 | 54.5 |
Compared with similar-size models (official numbers; protocols differ, so this is a rough comparison):
| Model | Pretraining tokens | HumanEval | HumanEval+ | MBPP | MBPP+ | GSM8K | IFEval | MMLU |
|---|---|---|---|---|---|---|---|---|
| Apex-2 (3.87B · 1.45B active) | 0.087T | 43.9 | 41.5 | 56.3 | 48.9 | 32.4 | 44.7 | 28.6 |
| Apex-1 DPO (1.1B dense) | 0.02T | 8.5 | — | 5.2 | — | 1.9 | — | 24.9 |
| Qwen2.5-1.5B-Instruct | 18T | 61.6 | — | 63.2 | — | 73.2 | 42.5 | 50.7 |
| Qwen2.5-Coder-1.5B-Instruct | 5.5T | 70.7 | 66.5 | 69.2 | 59.4 | — | — | — |
| Qwen3-1.7B (non-thinking) | 36T | — | — | — | — | — | 68.2 | 64.4 |
| Llama-3.2-1B-Instruct | 9T | — | — | — | — | 44.4 | 59.5* | 49.3 |
| Gemma-3-1B-it | 2T | 41.5 | — | 35.2 | — | 62.8 | 80.2* | 38.8 |
| OLMoE-1B-7B (1.3B active) | 5.1T | 62.3 | 54.4 | — | — | 72.4 | 66.4* | 55.1 |
| DeepSeek-Coder-1.3B | 2T | 65.9 | 60.4 | 65.3 | 54.8 | — | — | — |
DeepSeek-Coder numbers are the Qwen2.5-Coder report's re-evaluation. * Different IFEval metric: Apex-2 and Qwen report prompt-level strict; the others report an average of metrics or do not specify.
With 1/23 to 1/400 of the pretraining data of these models, Apex-2's base model matches Qwen2.5-1.5B (18T tokens) on HumanEval+ (32.9 vs 32.9). The big gaps are knowledge (MMLU) and math, which mainly track pretraining scale.
SFT data was 13-gram decontaminated against HumanEval(+), MBPP(+), GSM8K test, MATH, MMLU, ARC, HellaSwag, PIQA, WinoGrande, IFEval and LAMBADA. For LiveCodeBench only the newer v5–v6 problems (after 2024-08) were used, to reduce contamination risk.
Why no DPO: DPO on allenai/Dolci-Instruct-DPO (220K pairs) made answers 2.3× longer (320 → 733 tokens on average). Chosen-answer likelihood also fell during training. In evaluation HumanEval+ went 41.5 → 32.3, MBPP+ 48.9 → 37.3, GSM8K 32.4 → 14.3 and IFEval 44.7 → 35.7, so this SFT checkpoint is the release.
Known limitations
- English-centric. Other languages, including Korean, barely work (multilingual data was excluded from SFT).
- Limited world knowledge (MMLU ~29%), so it often states wrong facts with confidence.
- Weak at competitive programming (LiveCodeBench medium ~1%) and at predicting code execution (CRUXEval).
- 4096-token context. No tool use, no thinking mode.
Training data
- Pretrain: 40 deduplicated sources, including:
- web: FineWeb-Edu, DCLM, FinePDFs
- code: StarCoder-family and The Stack-derived code in 11 languages, OpenCoder corpora, synthetic code
- math: FineMath, InfiWebMath, Nemotron-CC-Math
- curated: Wikipedia, StackExchange, peS2o, Cosmopedia, Gutenberg
- SFT:
- general: allenai/Dolci-Instruct-SFT (tool-use, multilingual and identity subsets removed)
- code: OpenCoder opc-sft-stage1/2, nvidia/OpenCodeInstruct (unit-test pass rate ≥ 0.9), bigcode self-oss-instruct, m-a-p/Code-Feedback
- math: nvidia/OpenMathInstruct-2, AI-MO/NuminaMath-CoT
Some SFT sources contain synthetic data generated by third-party models (e.g. GPT-4-class, Llama-3.1-405B, Qwen2.5-Coder); check each dataset's terms for your use case.
License
Apache 2.0.
- Downloads last month
- 352
Model tree for YOON1v/Apex-2
Base model
YOON1v/Apex-2-BaseEvaluation results
- pass@1 (greedy, chat) on HumanEval+self-reported41.500
- pass@1 (greedy, chat) on MBPP+self-reported48.900
- acc (0-shot CoT) on GSM8Kself-reported32.400
- prompt-level strict on IFEvalself-reported44.700
- acc (5-shot) on MMLUself-reported28.600
- acc_norm (0-shot) on HellaSwagself-reported62.200