tinymixtral-it

Instruction-tuned TinyMixtral: 477.5M total / 276.1M active parameters (MoE, top-2 of 4 experts), 2048 context, 32k SentencePiece vocab.

Fine-tuned for one epoch from the public mikecovlee/tinymixtral base on 2,168,835 deduplicated and eval-decontaminated English instruction conversations blended from public sources (Tulu3, OpenHermes, SlimOrca, OpenOrca, UltraChat, MetaMath, OrcaMath, OMI-Llama-2, SQuAD2, TriviaQA). Sequence packing to 1024 tokens, bf16, AdamW, lr 2e-5 cosine with 3% warmup.

Load

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "mikecovlee/tinymixtral-it", trust_remote_code=True, dtype="bfloat16", device_map="cuda"
)
tokenizer = AutoTokenizer.from_pretrained("mikecovlee/tinymixtral-it")

Conversational use via the bundled chat template (tokenizer.apply_chat_template); greedy decoding works well at this scale.

Benchmarks (0-shot)

metric score
Instruction following (IFEval, prompt-strict / inst-strict) 0.170 / 0.280
Open-ended answer quality (LLM rubric, 4,955 held-out prompts, 0-100) 15.0 ± 0.3
GSM8K (strict / flexible) 2.1% / 2.3%
8-task harness (canonical mean: hellaswag/piqa/arc_c/obqa acc_norm, rest acc) 0.403

Per-task harness: hellaswag 0.340, piqa 0.634, winogrande 0.528, arc_easy 0.448, arc_challenge 0.260, openbookqa 0.302, boolq 0.426, lambada 0.289.

Note: relative to the base model, instruction following and open-ended answer quality improve substantially while multiple-choice common-sense accuracy drops (boolq is the most sensitive task). Arithmetic remains far below practical use.

License: MIT (Copyright (C) 2026 Michael Lee).

Downloads last month
18
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mikecovlee/tinymixtral-it

Finetuned
(1)
this model

Collection including mikecovlee/tinymixtral-it