Instructions to use mikecovlee/tinymistral-276m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mikecovlee/tinymistral-276m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="mikecovlee/tinymistral-276m", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("mikecovlee/tinymistral-276m", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use mikecovlee/tinymistral-276m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "mikecovlee/tinymistral-276m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mikecovlee/tinymistral-276m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/mikecovlee/tinymistral-276m
- SGLang
How to use mikecovlee/tinymistral-276m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "mikecovlee/tinymistral-276m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mikecovlee/tinymistral-276m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "mikecovlee/tinymistral-276m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mikecovlee/tinymistral-276m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use mikecovlee/tinymistral-276m with Docker Model Runner:
docker model run hf.co/mikecovlee/tinymistral-276m
tinymistral-276m
tinymistral-276m is a 276M-parameter dense decoder-only language model β the
same-activation-parameter, same-tokens companion to the sparse Mixture-of-Experts model
mikecovlee/tinymixtral.
It exists to answer one controlled question: at a fixed active parameter count and fixed
FLOPs per token, what does MoE routing buy you? tinymistral-276m (dense, 276M active) is
trained on the exact same data, tokenizer, 4-segment WSD schedule, batch size and LR ladder
as the 477.5M-total / 276.1M-active MoE tinymixtral. Only the FFN differs.
This is a base (pretrained) model β it has no chat/instruction fine-tuning.
Model details
tinymistral-276m |
tinymixtral (MoE) |
|
|---|---|---|
| Architecture | Dense SwiGLU FFN | 4 routed experts, top-2 + aux-free |
| Total parameters | 276,073,472 | 477,465,600 |
| Active parameters | 276,073,472 | 276,139,008 |
| FFN intermediate | 4096 | 2048 (per expert) |
| FFN MACs / token / layer | 12,582,912 | 12,582,912 (top-2) |
| Hidden size | 1024 | 1024 |
| Layers | 16 | 16 |
| Attention | GQA 16 Q / 4 KV heads, head_dim 64 | same |
| Context length | 2048 | 2048 |
| Positional | RoPE ΞΈ = 1e6, QK-Norm | same |
| Norm / embeddings | Pre-RMSNorm (eps 1e-6), tied embeddings | same |
| Vocab | 32,000 (TinyLlama tokenizer) | same |
| Precision | float32 checkpoint (bf16 training) | same |
| License | MIT | MIT |
The two models are iso-FLOPs per token and differ by exactly the 16 router matrices (16 Γ 1024 Γ 4 = 65,536 parameters, 0.024%).
Training
- Tokens: 8.05B, split into 4 strictly disjoint segments (2.00 / 1.94 / 2.20 / 1.91B, zero repetition).
- Schedule: WSD per segment β warmup 700 steps β constant β linear decay over the final 10%. LR ladder
5e-4 / 5e-4 / 4e-4 / 3e-4across S1βS4. - Batch: 48 Γ 1024 tokens = 49,152 tokens/step.
- Optimizer: AdamW Ξ²(0.9, 0.95), weight decay 0.1 (no decay on norms/embeddings), grad clip 1.0.
- Precision: bf16 autocast + bf16 optimizer states; gradient checkpointing.
- Seed: 42. Segment boundaries are the anneal points; AdamW momentum carries across segments.
- Hardware: single NVIDIA RTX PRO 4500 (Blackwell, 32 GB), ~23k tokens/s.
Data mix
FineWeb-Edu 44% Β· DCLM (web) 20% Β· Cosmopedia v2 12.5% Β· code (OpenCodeInstruct) 12.5% Β· math (OpenWebMath) 6% Β· Wikipedia 6%.
Results
Same-hardware evaluation (lm_eval 0.4.12, 0-shot, no chat template). Harness = mean of the
7 primary metrics (hellaswag acc_norm, piqa acc, winogrande acc, arc_easy acc,
arc_challenge acc_norm, openbookqa acc_norm, lambada acc).
| Metric | tinymistral-276m (dense) | tinymixtral (MoE) |
|---|---|---|
| val PPL (@8.05B, held-out) | 16.33 | 15.59 |
| 7-task harness | 0.3904 | 0.3992 |
| MMLU 5-shot | 0.2590 | 0.2480 |
| TruthfulQA MC1 / MC2 | 0.2411 / 0.4304 | 0.2375 / 0.4169 |
| GSM8K strict / flexible | 0.0000 / 0.0136 | 0.0000 / 0.0159 |
Takeaway: at fixed active parameters and fixed FLOPs, the routed MoE is +0.88 pp on the 7-task harness and 4.7% lower val PPL β routing provides a real gain at this budget. MMLU / TruthfulQA are near chance for both models and are noise-dominated. This supports scaling the sparse-MoE path rather than reverting to dense.
Held-out val PPL across the four segments (dense vs MoE): 17.79/17.16 β 16.65/16.22 β 16.51/15.81 β 16.33/15.59.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "mikecovlee/tinymistral-276m"
tok = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True)
inputs = tok("The capital of France is", return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=32, do_sample=False)
print(tok.decode(out[0], skip_special_tokens=True))
Requires transformers and trust_remote_code=True (custom tinymixtral architecture).
Limitations
- Base model: no instruction tuning; not a chat model. Outputs should not be used as-is for assistant tasks.
- Trained on only 8.05B tokens β far below modern small-model budgets (SmolLM2-360M / Qwen3-0.6B use 2β36T). Knowledge is capacity/budget-bound: MMLU β chance, GSM8K β 1β2%.
- English-centric, no safety alignment.
Family
mikecovlee/tinymixtralβ 477.5M MoE basemikecovlee/tinymixtral-itβ instruction-tunedmikecovlee/tinymistral-276mβ this model (dense iso-active ablation)
Citation
@misc{tinymistral-276m,
title = {tinymistral-276m: a 276M dense iso-active-parameter ablation of tinymixtral},
author = {Mike Lee},
year = {2026},
howpublished = {\url{https://huggingface.co/mikecovlee/tinymistral-276m}}
}
- Downloads last month
- -
Model tree for mikecovlee/tinymistral-276m
Base model
mikecovlee/tinymixtral