Instructions to use mikecovlee/tinymistral-477m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mikecovlee/tinymistral-477m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="mikecovlee/tinymistral-477m", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("mikecovlee/tinymistral-477m", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use mikecovlee/tinymistral-477m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "mikecovlee/tinymistral-477m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mikecovlee/tinymistral-477m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/mikecovlee/tinymistral-477m
- SGLang
How to use mikecovlee/tinymistral-477m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "mikecovlee/tinymistral-477m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mikecovlee/tinymistral-477m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "mikecovlee/tinymistral-477m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mikecovlee/tinymistral-477m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use mikecovlee/tinymistral-477m with Docker Model Runner:
docker model run hf.co/mikecovlee/tinymistral-477m
tinymistral-477m — 477M dense base
tinymistral-477m is a 477M-parameter dense decoder-only language model — the
same-total-parameter, same-everything-else companion to the sparse Mixture-of-Experts model
mikecovlee/tinymixtral.
It exists to answer one controlled question: at a fixed total parameter count and fixed
data, what does MoE routing buy you? tinymistral-477m (dense, 477.4M total = active) is
trained on the exact same data, tokenizer, 4-segment WSD schedule, batch size and LR ladder
as the 477.5M-total / 276.1M-active MoE tinymixtral. Only the FFN differs: the four 2048-wide
expert FFNs are merged into one 8192-wide dense FFN per layer and the routers are removed.
This is a base (pretrained) model — it has no chat/instruction fine-tuning.
Model details
tinymistral-477m |
tinymixtral (MoE) |
|
|---|---|---|
| Architecture | Dense SwiGLU FFN | 4 routed experts, top-2, aux 1e-3 |
| Total parameters | 477,400,064 | 477,465,600 |
| Active parameters | 477,400,064 | 276,139,008 |
| FFN intermediate | 8192 (4 experts merged) | 2048 (per expert) |
| FFN MACs / token / layer | 25,165,824 | 12,582,912 (top-2) |
| Hidden size | 1024 | 1024 |
| Layers | 16 | 16 |
| Attention | GQA 16 Q / 4 KV heads, head_dim 64 | same |
| Context length | 2048 | 2048 |
| Positional | RoPE θ = 1e6, QK-Norm | same |
| Norm / embeddings | Pre-RMSNorm (eps 1e-6), tied embeddings | same |
| Vocab | 32,000 (TinyLlama tokenizer) | same |
| Precision | float32 checkpoint (bf16 training) | same |
| License | MIT | MIT |
The two models differ by exactly the 16 router matrices (16 × 1024 × 4 = 65,536 parameters,
0.014%); everything else is identical. The dense model pays ~1.7× the active FLOPs per
token (477.4M vs 276.1M active parameters) — this is a parameter-matched control, not a
compute-matched one. The compute-matched sibling is tinymistral-276m.
Training
- Tokens: 8.05B, split into 4 strictly disjoint segments (2.00 / 1.94 / 2.20 / 1.91B, zero repetition).
- Schedule: WSD per segment — warmup 700 steps → constant → linear decay over the final 10%. LR ladder
5e-4 / 5e-4 / 4e-4 / 3e-4across S1→S4. - Batch: 48 × 1024 tokens = 49,152 tokens/step.
- Optimizer: AdamW β(0.9, 0.95), weight decay 0.1 (no decay on norms/embeddings), grad clip 1.0.
- Precision: bf16 autocast + bf16 optimizer states; gradient checkpointing.
- Seed: 42. Segment boundaries are the anneal points; AdamW momentum carries across segments.
- Hardware: single NVIDIA RTX A5000 (Ampere, 24 GB), ~12k tokens/s.
Data mix
FineWeb-Edu 44% · DCLM (web) 20% · Cosmopedia v2 12.5% · code (OpenCodeInstruct) 12.5% · math (OpenWebMath) 6% · Wikipedia 6% (shares rounded, sum ≈ 100%).
Evaluation
Same-protocol evaluation (lm_eval 0.4.12, 0-shot, no chat template). Harness = mean of the
7 primary metrics: hellaswag acc_norm, piqa acc_norm, winogrande acc, arc_easy acc,
arc_challenge acc_norm, openbookqa acc_norm, lambada acc.
| Metric | tinymistral-477m (dense) | tinymixtral (MoE) |
|---|---|---|
| 7-task harness | 0.4013 | 0.3979 |
| MMLU 5-shot | 0.2399 | 0.2339 |
| TruthfulQA MC1 / MC2 | 0.2362 / 0.4090 | 0.2375 / 0.4171 |
Per-task scores (7-task suite):
| Task | tinymistral-477m (dense) | tinymixtral (MoE) |
|---|---|---|
| HellaSwag (acc_norm) | 0.3414 | 0.335 |
| PIQA (acc_norm) | 0.6355 | 0.638 |
| WinoGrande (acc) | 0.5114 | 0.515 |
| ARC-Easy (acc) | 0.4907 | 0.478 |
| ARC-Challenge (acc_norm) | 0.2568 | 0.255 |
| OpenBookQA (acc_norm) | 0.302 | 0.296 |
| LAMBADA (acc) | 0.2711 | 0.268 |
Dense per-segment trajectory (7-task harness / MMLU): 0.3931 / 0.2291 → 0.3918 / 0.2386 → 0.4016 / 0.2337 → 0.4013 / 0.2399.
Harness note. Every figure on this card is computed under the one formula stated above from the campaign artifacts. The
tinymistral-276mcard quotes its MoE column under a slightly different formula (piqa acc) and an earlier MMLU run, so cross-card means are not bitwise comparable. Here both columns are recomputed under the identical formula.
Takeaway: at fixed total parameters and identical data, removing MoE routing does not
hurt — the dense model is +0.34 pp on the 7-task harness and +0.6 pp on MMLU (six of
eleven metrics up; the losses are within noise for this scale). Read together with the
iso-FLOPs result (tinymistral-276m: the MoE is +0.88 pp at equal active parameters), the
conclusion is that the MoE advantage is compute efficiency, not parameter efficiency:
a dense model can match or beat the MoE at equal total parameters, but only by spending
~1.7× the FLOPs per token. MMLU / TruthfulQA are near chance for both models and are
noise-dominated.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "mikecovlee/tinymistral-477m"
tok = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True)
inputs = tok("The capital of France is", return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=32, do_sample=False)
print(tok.decode(out[0], skip_special_tokens=True))
Requires transformers and trust_remote_code=True (custom tinymixtral architecture).
Limitations
- Base model: no instruction tuning; not a chat model. Outputs should not be used as-is for assistant tasks.
- Trained on only 8.05B tokens — far below modern small-model budgets (SmolLM2-360M / Qwen3-0.6B use 2–36T). Knowledge is capacity/budget-bound: MMLU ≈ chance.
- English-centric, no safety alignment.
- The FFN is unusually wide (8× hidden, a merged-MoE shape); standard dense models at this
size use ~2.5–3.5× and more layers. This is deliberate — it keeps the comparison with
tinymixtralcontrolled — but a textbook-shaped 477M dense would likely score slightly differently.
Family
mikecovlee/tinymixtral— 477.5M MoE base (v3.0 flagship)mikecovlee/tinymixtral-it— MoE instruction-tuned (v3.0-it, 3M SFT)mikecovlee/tinymistral-276m— 276M dense base (iso-FLOPs ablation)mikecovlee/tinymistral-276m-it— 276M dense instruction-tuned (3M SFT)mikecovlee/tinymistral-477m— this model (477M dense base, iso-total-parameter ablation)mikecovlee/tinymixtral-v1.1-1b— 1B MoE (earlier flagship)mikecovlee/tinymixtral-v1.1-0.5b— 0.5B-class MoE (data-quality ablation)mikecovlee/tinymixtral-v2.0-beta— shared-expert experiment (beta)mikecovlee/tinymixtral-v1.0— legacy (C4)
Naming. The MoE family is published under
tinymixtral; the dense ablation companions use thetinymistralspelling. Both belong to the same project.
Citation
@misc{tinymistral477m2026,
title = {TinyMixtral: a small Mixture-of-Experts language-model family},
author = {Michael Lee},
year = {2026},
howpublished = {\url{https://huggingface.co/mikecovlee/tinymistral-477m}}
}
License
MIT (Copyright (C) 2026 Michael Lee).
- Downloads last month
- 239