tinymistral-276m

tinymistral-276m is a 276M-parameter dense decoder-only language model β€” the same-activation-parameter, same-tokens companion to the sparse Mixture-of-Experts model mikecovlee/tinymixtral.

It exists to answer one controlled question: at a fixed active parameter count and fixed FLOPs per token, what does MoE routing buy you? tinymistral-276m (dense, 276M active) is trained on the exact same data, tokenizer, 4-segment WSD schedule, batch size and LR ladder as the 477.5M-total / 276.1M-active MoE tinymixtral. Only the FFN differs.

This is a base (pretrained) model β€” it has no chat/instruction fine-tuning.

Model details

tinymistral-276m tinymixtral (MoE)
Architecture Dense SwiGLU FFN 4 routed experts, top-2 + aux-free
Total parameters 276,073,472 477,465,600
Active parameters 276,073,472 276,139,008
FFN intermediate 4096 2048 (per expert)
FFN MACs / token / layer 12,582,912 12,582,912 (top-2)
Hidden size 1024 1024
Layers 16 16
Attention GQA 16 Q / 4 KV heads, head_dim 64 same
Context length 2048 2048
Positional RoPE ΞΈ = 1e6, QK-Norm same
Norm / embeddings Pre-RMSNorm (eps 1e-6), tied embeddings same
Vocab 32,000 (TinyLlama tokenizer) same
Precision float32 checkpoint (bf16 training) same
License MIT MIT

The two models are iso-FLOPs per token and differ by exactly the 16 router matrices (16 Γ— 1024 Γ— 4 = 65,536 parameters, 0.024%).

Training

  • Tokens: 8.05B, split into 4 strictly disjoint segments (2.00 / 1.94 / 2.20 / 1.91B, zero repetition).
  • Schedule: WSD per segment β€” warmup 700 steps β†’ constant β†’ linear decay over the final 10%. LR ladder 5e-4 / 5e-4 / 4e-4 / 3e-4 across S1β†’S4.
  • Batch: 48 Γ— 1024 tokens = 49,152 tokens/step.
  • Optimizer: AdamW Ξ²(0.9, 0.95), weight decay 0.1 (no decay on norms/embeddings), grad clip 1.0.
  • Precision: bf16 autocast + bf16 optimizer states; gradient checkpointing.
  • Seed: 42. Segment boundaries are the anneal points; AdamW momentum carries across segments.
  • Hardware: single NVIDIA RTX PRO 4500 (Blackwell, 32 GB), ~23k tokens/s.

Data mix

FineWeb-Edu 44% Β· DCLM (web) 20% Β· Cosmopedia v2 12.5% Β· code (OpenCodeInstruct) 12.5% Β· math (OpenWebMath) 6% Β· Wikipedia 6%.

Results

Same-hardware evaluation (lm_eval 0.4.12, 0-shot, no chat template). Harness = mean of the 7 primary metrics (hellaswag acc_norm, piqa acc, winogrande acc, arc_easy acc, arc_challenge acc_norm, openbookqa acc_norm, lambada acc).

Metric tinymistral-276m (dense) tinymixtral (MoE)
val PPL (@8.05B, held-out) 16.33 15.59
7-task harness 0.3904 0.3992
MMLU 5-shot 0.2590 0.2480
TruthfulQA MC1 / MC2 0.2411 / 0.4304 0.2375 / 0.4169
GSM8K strict / flexible 0.0000 / 0.0136 0.0000 / 0.0159

Takeaway: at fixed active parameters and fixed FLOPs, the routed MoE is +0.88 pp on the 7-task harness and 4.7% lower val PPL β€” routing provides a real gain at this budget. MMLU / TruthfulQA are near chance for both models and are noise-dominated. This supports scaling the sparse-MoE path rather than reverting to dense.

Held-out val PPL across the four segments (dense vs MoE): 17.79/17.16 β†’ 16.65/16.22 β†’ 16.51/15.81 β†’ 16.33/15.59.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "mikecovlee/tinymistral-276m"
tok = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True)

inputs = tok("The capital of France is", return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=32, do_sample=False)
print(tok.decode(out[0], skip_special_tokens=True))

Requires transformers and trust_remote_code=True (custom tinymixtral architecture).

Limitations

  • Base model: no instruction tuning; not a chat model. Outputs should not be used as-is for assistant tasks.
  • Trained on only 8.05B tokens β€” far below modern small-model budgets (SmolLM2-360M / Qwen3-0.6B use 2–36T). Knowledge is capacity/budget-bound: MMLU β‰ˆ chance, GSM8K β‰ˆ 1–2%.
  • English-centric, no safety alignment.

Family

Citation

@misc{tinymistral-276m,
  title  = {tinymistral-276m: a 276M dense iso-active-parameter ablation of tinymixtral},
  author = {Mike Lee},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/mikecovlee/tinymistral-276m}}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for mikecovlee/tinymistral-276m

Finetuned
(2)
this model

Collection including mikecovlee/tinymistral-276m