tinymistral-276m-it β€” instruction-tuned dense

tinymistral-276m-it is the instruction-tuned version of the 276M dense base mikecovlee/tinymistral-276m. It is the dense counterpart of the MoE flagship mikecovlee/tinymixtral-it: the exact same 3M-tier SFT recipe (same data, hyper-parameters and schedule) is applied to the dense base, so the two instruction-tuned models can be compared at the same active-parameter count and the same SFT budget.

Model details

tinymistral-276m-it tinymixtral-it (MoE)
Architecture Dense SwiGLU FFN 4 routed experts, top-2, aux 1e-3
Total parameters 276,073,472 477,465,600
Active parameters 276,073,472 276,139,008
FFN intermediate 4096 2048 (per expert)
Hidden size 1024 1024
Layers 16 16
Attention GQA 16 Q / 4 KV heads, head_dim 64 same
Context length 2048 2048
Positional RoPE ΞΈ = 1e6, QK-Norm same
Norm / embeddings Pre-RMSNorm (eps 1e-6), tied embeddings same
Vocab 32,000 (TinyLlama tokenizer) same
Precision float32 checkpoint (bf16 training) same
License MIT MIT

Training

  • Initialization: the 8.05B-token final checkpoint of the dense base tinymistral-276m.
  • SFT recipe: identical to the MoE flagship tinymixtral-it (3M tier) β€” 2,168,835 deduplicated, eval-decontaminated English instruction conversations blended from 10 public sources (Tulu3, OpenHermes, SlimOrca, OpenOrca, UltraChat, MetaMath, OrcaMath, OpenMathInstruct-2, SQuAD2, TriviaQA), 1 epoch, sequence packing to 1024 tokens, batch 24, AdamW lr 2e-5 (cosine, 100-step warmup), weight decay 0.1, seed 42 β†’ 56,793 steps.
  • Hardware: single NVIDIA RTX PRO 4500 (Blackwell, 32 GB), 1036.5 min (17.3 h).

Evaluation

Same-hardware evaluation (lm_eval 0.4.12, 0-shot, no chat template). Harness = mean of the 7 primary metrics (hellaswag acc_norm, piqa acc, winogrande acc, arc_easy acc, arc_challenge acc_norm, openbookqa acc_norm, lambada acc).

Model 7-task harness 7-task all-acc GSM8K strict / flex IFEval prompt / inst
tinymixtral-it (MoE + 3M SFT) 0.3994 0.3691 0.0182 / 0.0205 0.1756 / 0.2782
tinymixtral (MoE base) 0.3992 β€” 0.0000 / 0.0159 β€”
tinymistral-276m (dense base) 0.3904 0.3890 0.0000 / 0.0136 β€”
tinymistral-276m-it 0.3892 0.3631 0.0174 / 0.0205 0.1460 / 0.2602

Harness note. Means quoted here use the 7-task harness (excludes BoolQ); the base v1.0/v3.0 cards report an 8-task mean (includes BoolQ). The two are not directly comparable.

Takeaways:

  • SFT transfers to the dense base: GSM8K flexible rises from 0.0136 (dense base) to 0.0205, matching the MoE flagship under the same recipe.
  • The 7-task harness barely moves (0.3904 β†’ 0.3892), the same pattern as the MoE flagship β€” SFT mainly buys instruction-following and math, not general tasks.
  • Instruction following is weaker on the dense base: IFEval 0.1460 / 0.2602 vs 0.1756 / 0.2782 for tinymixtral-it β€” the dense base is less steerable than the MoE base.
  • Absolute numbers remain bounded by the 276M scale and the 8.05B-token pretraining budget.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "mikecovlee/tinymistral-276m-it"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True, dtype="bfloat16", device_map="auto")

messages = [{"role": "user", "content": "What is 12% of 250?"}]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))

Requires transformers and trust_remote_code=True (custom tinymixtral architecture).

Limitations

  • The base was pretrained on only 8.05B tokens β€” far below modern small-model budgets; knowledge tasks (MMLU / TruthfulQA) are near chance.
  • Math/reasoning is bounded by the 276M scale and the pretraining budget; SFT only helps marginally.
  • English-centric, no safety alignment.

Family

Naming. The MoE family is published under tinymixtral; the dense 276M iso-active ablation companions use the tinymistral spelling. Both belong to the same project.

Citation

@misc{tinymistral276mit2026,
  title  = {TinyMixtral: a small Mixture-of-Experts language-model family},
  author = {Michael Lee},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/mikecovlee/tinymistral-276m-it}}
}

License

MIT (Copyright (C) 2026 Michael Lee).

Downloads last month
409
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for mikecovlee/tinymistral-276m-it

Finetuned
(1)
this model

Datasets used to train mikecovlee/tinymistral-276m-it

Collection including mikecovlee/tinymistral-276m-it