ShallowSeek-mini-it

ShallowSeek-mini-it is the instruction-tuned (chat) version of ShallowSeek-mini-base, a tiny DeepSeek-V3-style Mixture-of-Experts model trained from scratch on a laptop CPU. It has 11.0M parameters, of which 6.3M are active per token.

Part of the ShallowSeek family. Not affiliated with DeepSeek.

It answers in a chat format and stops when it's done, but it's still an 11M-parameter model: answers are fluent and on-topic but often wrong or vague, and it invents facts confidently. It's an educational and research artifact; don't use it for anything that matters.

Chat format

<|system|>You are a helpful assistant.<|end|><|user|>What causes the seasons?<|end|><|assistant|>

The model writes its reply and ends it with <|end|>. The GGUF embeds this template, and the GGUF repo includes Ollama template and params files.

Fine-tuning (SFT)

Base model ShallowSeek-mini-base (pretrained on 286M FineWeb-Edu tokens)
Conversations 16,347 (15,530 train / 817 validation), 1.62M assistant tokens trained on
Grounded Q&A (~36% of assistant tokens) ~10.6K single-turn question/answer pairs generated by Qwen3.8-9B from FineWeb-Edu documents in ShallowSeek's own pretraining data, so answers concern text the model saw. About 2% of pairs still refer to "the document"; these were left in
General chat (~64%) 5,684 conversations from HuggingFaceTB/smol-smoltalk: all 1,994 everyday-conversations (small talk, multi-turn), plus openhermes, smol-constraints and smol-magpie-ultra-short (capped at 15% of assistant tokens), filtered to remove code and to fit 1,024 tokens
Loss On assistant replies and their closing `<
Schedule 1942 steps, batch 16, 2 epochs, AdamW, LR 1.5e-4 → 1.5e-5 cosine (about 10% of the pretraining peak)
MoE routing Load-balancing biases frozen at their pretrained values during SFT, with padding excluded from the expert-load statistics

Validation loss on 817 held-out conversations (assistant tokens only): 3.779 → 2.949.

Evaluation: base vs IT

Both are scored on the same raw-text prompts, without the chat template, so the numbers are directly comparable. Chance on MMLU is 25%.

Metric Base IT
Probes top-5 28.6% 35.7%
Mean answer log-prob -5.68 -5.40
True-vs-false pairs won 58.3% 66.7%
MMLU (letter) 24.9% 25.6%
MMLU (cloze) 25.5% 25.6%
MMLU biology & medicine (cloze) 27.8% 27.7%

SFT mainly teaches format, turn-taking and when to stop, and MMLU is essentially unchanged. The factual probes improved slightly, plausibly because the grounded Q&A restates facts from documents the model pretrained on. With 14 probes and 12 pairs, treat that as a modest effect rather than a large one.

Usage

Ollama (uses the bundled chat template):

ollama run hf.co/edededdy/ShallowSeek-mini-it-GGUF:Q8_0

llama.cpp:

llama-cli -m ShallowSeek-mini-it-q8_0.gguf -cnv --temp 0.7 --top-p 0.9 --repeat-penalty 1.15

PyTorch: see chat.py in the training code:

python chat.py --ckpt mini-sft.pt --temperature 0.7 --rep-penalty 1.15

Files

File Contents
model.safetensors PyTorch weights (including the MTP module), our own parameter names
config.json ModelConfig for deepseek_moe
tokenizer.json Byte-level BPE tokenizer (Hugging Face tokenizers format)
chat_template.jinja Chat template (Jinja), the same one embedded in the GGUF
tokenizer_config.json Special tokens (`<
ShallowSeek-mini-it-f16.gguf GGUF (f16) for llama.cpp / Ollama / LM Studio
ShallowSeek-mini-it-q8_0.gguf GGUF (q8_0) for llama.cpp / Ollama / LM Studio

Acknowledgements

  • Architecture after DeepSeek-V3 (DeepSeek-AI, 2024).
  • Pretraining data: FineWeb-Edu (Hugging Face, ODC-BY 1.0).
  • Chat data: smol-smoltalk (Hugging Face, Apache-2.0).
  • Synthetic Q&A written by Qwen3.8-9B.

Card generated 2026-10-01 by prepare_release.py.

Downloads last month
198
Safetensors
Model size
12.1M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for edededdy/ShallowSeek-mini-it

Finetuned
(1)
this model
Quantizations
1 model

Datasets used to train edededdy/ShallowSeek-mini-it

Collection including edededdy/ShallowSeek-mini-it