Instructions to use YOON1v/Apex-2-Base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use YOON1v/Apex-2-Base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="YOON1v/Apex-2-Base")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("YOON1v/Apex-2-Base") model = AutoModelForCausalLM.from_pretrained("YOON1v/Apex-2-Base", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use YOON1v/Apex-2-Base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "YOON1v/Apex-2-Base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "YOON1v/Apex-2-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/YOON1v/Apex-2-Base
- SGLang
How to use YOON1v/Apex-2-Base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "YOON1v/Apex-2-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "YOON1v/Apex-2-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "YOON1v/Apex-2-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "YOON1v/Apex-2-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use YOON1v/Apex-2-Base with Docker Model Runner:
docker model run hf.co/YOON1v/Apex-2-Base
Apex-2-Base
The base checkpoint of a Mixture-of-Experts LLM pretrained from scratch: 3.87B total / 1.45B active parameters, trained on 86.5B tokens. Released as a starting point for continued pretraining (CPT) and fine-tuning.
Ready for continued pretraining
- Standard
Qwen3MoeForCausalLMformat. It loads and trains in transformers, vLLM and other tools that support Qwen3-MoE, with notrust_remote_code.- Full-precision fp32 weights, exactly as in the training checkpoint (no bf16 rounding). Checked against the training checkpoint: logits relative difference 9.3e-7, 100% next-token agreement.
- All training settings are public: optimizer, LR schedule and the exact data mix. See the Continued pretraining guide below.
- A pre-decay checkpoint at peak LR (
stable-44500) is also available for WSD-style continuation.
- Chat (SFT) version: YOON1v/Apex-2
- Code, architecture writeup and training diary: github.com/DW-dev-UE/LLM-from-scratch
Model details
Tokens trained per dataset (86.5B total, click to expand)
| Domain | Dataset | Stage 1 | Main | Decay | Total | Share |
|---|---|---|---|---|---|---|
| Web | FineWeb-Edu | 7.00 | 5.89 | 4.09 | 16.98 | 19.6% |
| Web | DCLM-Baseline | 4.40 | 5.24 | 1.02 | 10.66 | 12.3% |
| Web | FinePDFs (English) | 1.00 | 4.89 | 1.02 | 6.91 | 8.0% |
| Code | StarCoderData source code (15 languages + Markdown) | 3.20 | 11.77 | 3.68 | 18.65 | 21.6% |
| Code | StarCoderData GitHub issues · Jupyter · commits | 0.60 | 0.87 | 1.02 | 2.49 | 2.9% |
| Code | OpenCoder FineWeb-Code | 0.60 | 1.48 | – | 2.08 | 2.4% |
| Code | OpenCoder Annealing (synthetic) | – | 1.12 | 2.04 | 3.16 | 3.7% |
| Code | Dolma 3 CraneCode (synthetic) | – | 0.82 | 0.41 | 1.23 | 1.4% |
| Math | FineMath 3+ | 1.00 | 2.89 | 1.64 | 5.52 | 6.4% |
| Math | InfiWebMath 3+ | 0.60 | 1.67 | 0.82 | 3.08 | 3.6% |
| Math | Nemotron-CC-Math 4+ | – | 0.98 | 0.61 | 1.59 | 1.8% |
| Curated | Cosmopedia v2 (synthetic textbooks) | 0.60 | 2.21 | 2.04 | 4.85 | 5.6% |
| Curated | peS2o (scientific papers) | 0.40 | 4.54 | 0.82 | 5.76 | 6.7% |
| Curated | StackExchange | – | 0.22 | 0.61 | 0.84 | 1.0% |
| Curated | Wikipedia (English) | 0.60 | 0.73 | 0.61 | 1.94 | 2.2% |
| Curated | Project Gutenberg (English books) | – | 0.75 | – | 0.75 | 0.9% |
| Total | 20.00 | 46.06 | 20.45 | 86.51 | 100% |
In billions of tokens, computed from each stage's mixture weights (see Training data).
| Architecture | Decoder-only MoE Transformer, every layer MoE. GQA (16Q/4KV, head_dim 128) + QK-RMSNorm + RoPE (θ=10,000) + SwiGLU experts + RMSNorm, tied embeddings |
| Parameters | 3,869.1M total · 1,453.2M active per token (1,142.0M non-embedding active) |
| Layers · d_model | 32 · 2048 |
| Experts | 16 per layer, top-4 routing with renormalized probabilities, expert width 1024 |
| Router losses | load-balancing aux loss 0.01 (global-batch statistics) + router z-loss 1e-3 |
| Context length | 4,096 |
| Tokenizer | Qwen3 tokenizer (vocab 151,936). eos <|endoftext|> (151643), no chat template |
| Weights | fp32 safetensors (15.5GB), per-expert keys |
| Pretraining | 54,250 steps · 86.5B tokens, 1M-token batches (4,096 × 256), attention masked at document boundaries. Parts trained with DiLoCo on 2× GH200 |
Usage (inference)
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("YOON1v/Apex-2-Base")
model = AutoModelForCausalLM.from_pretrained("YOON1v/Apex-2-Base", dtype=torch.bfloat16, device_map="cuda")
inputs = tokenizer("def is_prime(n: int) -> bool:\n", return_tensors="pt").to("cuda")
out = model.generate(**inputs, max_new_tokens=128, do_sample=False)
print(tokenizer.decode(out[0], skip_special_tokens=True))
vLLM: LLM(model="YOON1v/Apex-2-Base", dtype="bfloat16", max_model_len=4096)
This is a base model: it continues text and does not follow instructions or chat. For chat, use Apex-2.
Continued pretraining guide
Loading
import torch
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained(
"YOON1v/Apex-2-Base",
dtype=torch.float32, # keep the fp32 weights; train with bf16 autocast (mixed precision)
output_router_logits=True, # adds the expert load-balancing loss (off by default)
)
- With
output_router_logits=True, the load-balancing loss (router_aux_loss_coef=0.01) is added to the loss wheneverlabelsare passed. Without it, routing can collapse onto a few experts. - The router z-loss (1e-3) used in pretraining is not part of the transformers implementation. Add it yourself if you need it.
Original training setup
| Optimizer | AdamW 8-bit (bitsandbytes, paged), β=(0.9, 0.95), eps 1e-8, weight decay 0.1 |
| LR schedule | WSD: 1,000 warmup steps → constant peak 4e-4 (stable) → linear decay to 0 over the last 9,750 steps (20.4B tokens) |
| Gradient clipping | 1.0 |
| Batch · sequence | 1,048,576 tokens (4,096 × 256) |
| Loss | CE + 0.01 × load-balancing + 1e-3 × router z-loss |
| Packing | <|endoftext|> after every document, no attention across document boundaries |
Recommendations
- Re-warm the LR. This checkpoint has finished its decay (LR 0). Start CPT with a warmup and tune the peak LR at or below the original 4e-4.
- Replay old data. Training only on new data can erode existing skills, code especially. Mixing in some of the original data (listed below) reduces forgetting.
- Packing. If you pack several documents into one sequence, block attention across document boundaries (e.g. transformers
DataCollatorWithFlattening+ FlashAttention). - Context length. Trained at 4,096 tokens. Going longer needs changes such as a larger
rope_theta. - Tokenizer. Use the default settings. transformers may suggest
fix_mistral_regex=True; do not enable it, because it changes tokenization from what training used. - Special tokens. The embeddings of
<|im_start|>,<|im_end|>and the other special tokens were barely trained during pretraining. Chat formatting is learned in SFT.
Intermediate checkpoint
| revision | step | tokens | state |
|---|---|---|---|
main |
54,250 | 86.5B | decay finished (LR 0). Final model |
stable-44500 |
44,500 | 66.1B | before the decay, at peak LR (4e-4) |
For WSD-style continuation, stable-44500 is the natural starting point: keep training on new data at the peak LR, then run your own decay whenever you choose.
model = AutoModelForCausalLM.from_pretrained("YOON1v/Apex-2-Base", revision="stable-44500", dtype=torch.float32)
Benchmarks
| Area | Benchmark | stable 44,500 | final 54,250 |
|---|---|---|---|
| Code | HumanEval (pass@1, greedy) | 22.0 | 36.6 |
| Code | HumanEval+ | 18.9 | 32.9 |
| Code | MBPP | 41.0 | 54.8 |
| Code | MBPP+ | 34.7 | 46.3 |
| Math | GSM8K (8-shot CoT, flexible) | 10.2 | 15.1 |
| Knowledge | MMLU (5-shot) | 25.3 | 28.2 |
| Commonsense | HellaSwag (0-shot, acc_norm) | 56.9 | 60.8 |
| Commonsense | ARC-e (0-shot, acc_norm) | 61.8 | 66.7 |
| Commonsense | ARC-c (0-shot, acc_norm) | 35.5 | 39.7 |
| Commonsense | PIQA (0-shot, acc_norm) | 73.2 | 73.9 |
| Commonsense | WinoGrande (0-shot) | 57.7 | 60.1 |
| Commonsense | LAMBADA (0-shot) | 53.4 | 54.5 |
Measured by us with lm-eval 0.4.13 (HF, bf16) and EvalPlus 0.3.1 (vLLM, base prompt, greedy).
Compared with similar-size base models (official numbers; protocols differ, so this is a rough comparison):
| Model | Pretraining tokens | HumanEval / + | MBPP / + | MMLU | HellaSwag |
|---|---|---|---|---|---|
| Apex-2-Base | 0.087T | 36.6 / 32.9 | 54.8 / 46.3 | 28.2 | 60.8 |
| Qwen2.5-1.5B | 18T | 37.2 / 32.9 | 60.2 / 49.6 | 60.9 | 67.9 |
| Qwen2.5-Coder-1.5B | 5.5T | 43.9 / 36.6 | 69.2 / 58.6 | 53.6 | 61.8 |
| Gemma-2-2B | 2T | 17.7 / – | 29.6 (3-shot) / – | 51.3 | 73.0 (10-shot) |
With about 1/200 of the pretraining data, Apex-2-Base matches Qwen2.5-1.5B (18T tokens) on HumanEval+ (32.9). Knowledge (MMLU) and math still show the gap in pretraining scale.
Training data
Public datasets only. 86.5B tokens in three stages:
| Stage | Steps | Tokens | Web / Code / Math / Curated (%) |
|---|---|---|---|
| Stage 1 | 0 → 19,074 | 20.0B | 62 / 22 / 8 / 8 |
| Main (stable) | 19,074 → 44,500 | 46.1B | 34.8 / 34.9 / 12.0 / 18.4 |
| Decay | 44,500 → 54,250 | 20.4B | 30 / 35 / 15 / 20 |
| Total | 86.5B | 39.9 / 31.9 / 11.8 / 16.3 |
Tokens trained per dataset: see the expandable table at the top of Model details.
StarCoderData by language (18.65B total)
| Language | Tokens (B) | Share of all tokens |
|---|---|---|
| Python | 4.64 | 5.4% |
| Java | 1.65 | 1.9% |
| TypeScript | 1.65 | 1.9% |
| Go | 1.59 | 1.8% |
| JavaScript | 1.58 | 1.8% |
| SQL | 1.08 | 1.3% |
| Markdown | 1.08 | 1.2% |
| C | 1.05 | 1.2% |
| C++ | 0.98 | 1.1% |
| PHP | 0.81 | 0.9% |
| C# | 0.74 | 0.9% |
| Rust | 0.57 | 0.7% |
| Ruby | 0.42 | 0.5% |
| Kotlin | 0.33 | 0.4% |
| Scala | 0.26 | 0.3% |
| Shell | 0.22 | 0.3% |
- Token counts are computed from each stage's mixture weights. Every 4,096-token sequence was drawn from a source at random according to those weights, so actual amounts can differ by about 1%.
- Documents that also appear in the stage 1 data were removed from the later corpus, and no source was used for more than one epoch, so no data was repeated.
Limitations
- English-centric. Other languages, including Korean, barely work.
- Limited world knowledge (MMLU ~28%).
- As a base model it has no instruction following, chat ability or safety tuning.
- 4,096-token context.
License
Apache 2.0. Each training dataset has its own license and terms (ODC-By, CC-BY, CC-BY-SA, the StarCoderData terms of use, NVIDIA data licenses, etc.); see the dataset pages.
- Downloads last month
- 213
Model tree for YOON1v/Apex-2-Base
Datasets used to train YOON1v/Apex-2-Base
wikimedia/wikipedia
HuggingFaceTB/smollm-corpus
Evaluation results
- pass@1 (greedy) on HumanEvalself-reported36.600
- pass@1 (greedy) on HumanEval+self-reported32.900
- pass@1 (greedy) on MBPP+self-reported46.300
- acc (8-shot CoT) on GSM8Kself-reported15.100
- acc (5-shot) on MMLUself-reported28.200
- acc_norm (0-shot) on HellaSwagself-reported60.800