ShallowSeek-mini-base

ShallowSeek-mini-base is a tiny DeepSeek-V3-style Mixture-of-Experts language model trained from scratch on a laptop CPU. It has 11.0M parameters in total, of which 6.3M are active per token. It was pretrained on 286M tokens of FineWeb-Edu.

Part of the ShallowSeek family. Not affiliated with DeepSeek. The name is a nod to the architecture it borrows, at a fraction of the depth.

This is the base model, before any instruction tuning. It continues text; it doesn't follow instructions. It's an educational and research artifact: it writes fluent, on-topic English but knows very few facts and confidently makes things up. Don't use it for anything that matters.

Architecture

It follows the main ideas of DeepSeek-V3, scaled down:

Component This model
Layers / hidden size 8 / 192 (layer 0 has a dense FFN of 512; layers 1–7 are MoE)
Attention Multi-head Latent Attention (MLA): 6 heads, KV latent 48, decoupled RoPE dim 16, head dims 24+16 (q/k) and 24 (v). The KV cache stores only the 48-dim latent plus a 16-dim RoPE key per token
MoE DeepSeekMoE: 12 fine-grained routed experts (SwiGLU, 128 hidden), top-3, plus 1 shared expert; sigmoid gating; group-limited routing (3 groups, top-2)
Load balancing Auxiliary-loss-free bias balancing (update speed 0.001) plus a small sequence-wise auxiliary loss (α = 0.0001). Expert load stayed at about 1.14× the mean (1.28× in the busiest layer) for most of training
Multi-Token Prediction 1 MTP module (λ = 0.3), used in training only. It's in model.safetensors and left out of the GGUF (1.14M extra params)
Context 1024 tokens
Tokenizer Byte-level BPE, 8,192 tokens, trained on the same FineWeb-Edu slice (~3.9 bytes/token)

Training

Data FineWeb-Edu sample/10BT, 236,544 documents, 286M tokens (~26 tokens per parameter), one pass
Steps 34,960 × 8,192 tokens (batch 8 × 1,024)
Optimizer AdamW (β = 0.9, 0.95), weight decay 0.1, gradient clipping 1.0
Learning rate 1.5e-3 peak, 700 warm-up steps, cosine decay to 1.5e-4
Precision fp32
Hardware One 8-core Intel Core i9-9880H (2019 MacBook Pro), CPU only: about 2 days of compute at ~1.9K tokens/s

Validation loss: 5.12 (step 1,000) → 4.28 (step 5,000) → 4.06 (step 10,000) → 3.97 (step 15,000) → 3.87 (step 20,000) → 3.78 (step 25,000) → 3.70 (step 30,000) → 3.67 (step 35,000)

Evaluation

Metric Value
Validation loss (nats/token) 3.667
Perplexity 39.1
Bits per byte 1.354

Factual probes (benchmark.py)

These are 14 fill-in-the-blank probes and 12 true-versus-false sentence pairs, scored on raw text.

Metric Value
Probes answered at top-1 / top-5 0% / 29%
Mean log-probability of the correct answer -5.68
True-vs-false pairs won 7/12 (58%)
Prompt Expected Rank of expected Model's top-3
The Earth revolves around the Sun 18 earth, world, Earth
Water is made of hydrogen and oxygen 2 hydrogen, oxygen, nitrogen
The heart pumps blood 27 are, can, have
Plants make their food through a process called photosynthesis 96 “, the, "
The capital of France is Paris 1,381 a, the, an
The largest planet in our solar system is Jupiter 29 the, a, solar
World War II ended in 1945 5 the, 19, 18
HIV is a virus 7 common, very, disease
The American flag is red, white and blue 4 white, red, black
The sun rises in the east 65 middle, morning, mid
Fish breathe using their gills 41 skin, mouth, l
2 + 2 = 4 4 2, 1, 0
3 + 5 = 8 19 1, 5, 2
7 + 6 = 13 30 5, 1, 4

MMLU (mmlu_eval.py)

Format Accuracy Questions scored Average few-shot examples
letter 24.9% 14,037 (skipped 5) 4.5
cloze 25.5% 14,037 (skipped 5) 0.0

By category (accuracy weighted by question count; chance is 25%):

Category Subjects Questions Letter Cloze
STEM 19 3,153 26.8% 24.0%
Humanities 13 4,705 23.8% 25.3%
Social Sciences 12 3,077 23.9% 26.1%
Other 13 3,107 25.5% 26.7%

Biology & medicine (cloze): 27.8% on 2,094 questions, against 25% chance. That's +2.8 points, about 3.0 standard errors above chance, so it's statistically significant. 9 of the 10 subjects are at or above chance. The overall cloze score is close to chance, so the signal sits mostly in the life-science and health subjects, where FineWeb-Edu's educational web text is richest. Maths-heavy subjects come out below chance, since cloze can't guess numeric answers. This group isn't an official MMLU category; it pools these 10 subjects:

Subject Questions Cloze accuracy
clinical knowledge 265 32.5%
human aging 223 29.6%
professional medicine 272 29.0%
virology 166 28.3%
medical genetics 100 28.0%
college medicine 173 27.4%
high school biology 310 26.8%
college biology 144 25.7%
nutrition 306 25.2%
anatomy 135 23.7%

Chance is 25%. In the letter format the model compares " A", " B", " C" and " D", with up to 5 same-subject examples included only while they fit in the 1,024-token window. In the cloze format each answer's text is scored by log-probability per byte. Small models in the letter format tend to answer the same letter within a subject, so per-subject letter scores mostly reflect how the answer key happens to be distributed. The cloze results are the meaningful ones here.

What it can and can't do

  • ✅ Grammatical, topic-consistent English for a few sentences, in the register of educational web text.
  • ✅ Coarse associations: it knows HIV is a virus rather than a mineral, ice is frozen water, and a week has seven days.
  • ❌ Specific facts and relations (capital cities, which body orbits which), arithmetic beyond memorised cases like "2 + 2 = 4", and following instructions.
  • ❌ It invents names, numbers and "facts" fluently, and falls into repetition loops at low temperature. Use top-p ≈ 0.9 and a repetition penalty of about 1.1–1.2.

Usage

llama.cpp, Ollama or LM Studio (GGUF, llama.cpp's built-in deepseek2 architecture):

llama-completion -m ShallowSeek-mini-base-q8_0.gguf -p "Photosynthesis is" -n 100 --temp 0.7 --top-p 0.9 --repeat-penalty 1.15

PyTorch (needs the deepseek_moe package from the training code):

import json, torch
from safetensors.torch import load_file
from deepseek_moe import DeepSeekMoEModel, ModelConfig
from deepseek_moe.tokenizer import BPETokenizer

model = DeepSeekMoEModel(ModelConfig.from_dict(json.load(open("config.json")))).eval()
model.load_state_dict(load_file("model.safetensors"))
tok = BPETokenizer.from_file("tokenizer.json")
ids = torch.tensor([[tok.eot_id] + tok.encode("Photosynthesis is")])
print(tok.decode(model.generate(ids, 60, temperature=0.7, top_p=0.9, repetition_penalty=1.15)[0].tolist()))

Files

File Contents
model.safetensors PyTorch weights (including the MTP module), our own parameter names
config.json ModelConfig for deepseek_moe
tokenizer.json Byte-level BPE tokenizer (Hugging Face tokenizers format)
ShallowSeek-mini-base-f16.gguf GGUF (f16) for llama.cpp / Ollama / LM Studio
ShallowSeek-mini-base-q8_0.gguf GGUF (q8_0) for llama.cpp / Ollama / LM Studio

Acknowledgements

  • The architecture follows DeepSeek-V3 (DeepSeek-AI, 2024): MLA, DeepSeekMoE, auxiliary-loss-free balancing and multi-token prediction.
  • Pretraining data: FineWeb-Edu (Hugging Face, ODC-BY 1.0).

Card generated 2026-10-01 by prepare_release.py.

Downloads last month
90
Safetensors
Model size
12.1M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for edededdy/ShallowSeek-mini-base

Finetunes
1 model
Quantizations
1 model

Dataset used to train edededdy/ShallowSeek-mini-base

Collection including edededdy/ShallowSeek-mini-base