Instructions to use edededdy/ShallowSeek-mini-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use edededdy/ShallowSeek-mini-base with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf edededdy/ShallowSeek-mini-base:F16 # Run inference directly in the terminal: llama cli -hf edededdy/ShallowSeek-mini-base:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf edededdy/ShallowSeek-mini-base:F16 # Run inference directly in the terminal: llama cli -hf edededdy/ShallowSeek-mini-base:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf edededdy/ShallowSeek-mini-base:F16 # Run inference directly in the terminal: ./llama-cli -hf edededdy/ShallowSeek-mini-base:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf edededdy/ShallowSeek-mini-base:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf edededdy/ShallowSeek-mini-base:F16
Use Docker
docker model run hf.co/edededdy/ShallowSeek-mini-base:F16
- LM Studio
- Jan
- vLLM
How to use edededdy/ShallowSeek-mini-base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "edededdy/ShallowSeek-mini-base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "edededdy/ShallowSeek-mini-base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/edededdy/ShallowSeek-mini-base:F16
- Ollama
How to use edededdy/ShallowSeek-mini-base with Ollama:
ollama run hf.co/edededdy/ShallowSeek-mini-base:F16
- Unsloth Desktop
- Docker Model Runner
How to use edededdy/ShallowSeek-mini-base with Docker Model Runner:
docker model run hf.co/edededdy/ShallowSeek-mini-base:F16
- Lemonade
How to use edededdy/ShallowSeek-mini-base with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull edededdy/ShallowSeek-mini-base:F16
Run and chat with the model
lemonade run user.ShallowSeek-mini-base-F16
List all available models
lemonade list
- Atomic Chat
ShallowSeek-mini-base
ShallowSeek-mini-base is a tiny DeepSeek-V3-style Mixture-of-Experts language model trained from scratch on a laptop CPU. It has 11.0M parameters in total, of which 6.3M are active per token. It was pretrained on 286M tokens of FineWeb-Edu.
Part of the ShallowSeek family. Not affiliated with DeepSeek. The name is a nod to the architecture it borrows, at a fraction of the depth.
This is the base model, before any instruction tuning. It continues text; it doesn't follow instructions. It's an educational and research artifact: it writes fluent, on-topic English but knows very few facts and confidently makes things up. Don't use it for anything that matters.
Architecture
It follows the main ideas of DeepSeek-V3, scaled down:
| Component | This model |
|---|---|
| Layers / hidden size | 8 / 192 (layer 0 has a dense FFN of 512; layers 1–7 are MoE) |
| Attention | Multi-head Latent Attention (MLA): 6 heads, KV latent 48, decoupled RoPE dim 16, head dims 24+16 (q/k) and 24 (v). The KV cache stores only the 48-dim latent plus a 16-dim RoPE key per token |
| MoE | DeepSeekMoE: 12 fine-grained routed experts (SwiGLU, 128 hidden), top-3, plus 1 shared expert; sigmoid gating; group-limited routing (3 groups, top-2) |
| Load balancing | Auxiliary-loss-free bias balancing (update speed 0.001) plus a small sequence-wise auxiliary loss (α = 0.0001). Expert load stayed at about 1.14× the mean (1.28× in the busiest layer) for most of training |
| Multi-Token Prediction | 1 MTP module (λ = 0.3), used in training only. It's in model.safetensors and left out of the GGUF (1.14M extra params) |
| Context | 1024 tokens |
| Tokenizer | Byte-level BPE, 8,192 tokens, trained on the same FineWeb-Edu slice (~3.9 bytes/token) |
Training
| Data | FineWeb-Edu sample/10BT, 236,544 documents, 286M tokens (~26 tokens per parameter), one pass |
| Steps | 34,960 × 8,192 tokens (batch 8 × 1,024) |
| Optimizer | AdamW (β = 0.9, 0.95), weight decay 0.1, gradient clipping 1.0 |
| Learning rate | 1.5e-3 peak, 700 warm-up steps, cosine decay to 1.5e-4 |
| Precision | fp32 |
| Hardware | One 8-core Intel Core i9-9880H (2019 MacBook Pro), CPU only: about 2 days of compute at ~1.9K tokens/s |
Validation loss: 5.12 (step 1,000) → 4.28 (step 5,000) → 4.06 (step 10,000) → 3.97 (step 15,000) → 3.87 (step 20,000) → 3.78 (step 25,000) → 3.70 (step 30,000) → 3.67 (step 35,000)
Evaluation
| Metric | Value |
|---|---|
| Validation loss (nats/token) | 3.667 |
| Perplexity | 39.1 |
| Bits per byte | 1.354 |
Factual probes (benchmark.py)
These are 14 fill-in-the-blank probes and 12 true-versus-false sentence pairs, scored on raw text.
| Metric | Value |
|---|---|
| Probes answered at top-1 / top-5 | 0% / 29% |
| Mean log-probability of the correct answer | -5.68 |
| True-vs-false pairs won | 7/12 (58%) |
| Prompt | Expected | Rank of expected | Model's top-3 |
|---|---|---|---|
| The Earth revolves around the | Sun | 18 | earth, world, Earth |
| Water is made of hydrogen and | oxygen | 2 | hydrogen, oxygen, nitrogen |
| The heart pumps | blood | 27 | are, can, have |
| Plants make their food through a process called | photosynthesis | 96 | “, the, " |
| The capital of France is | Paris | 1,381 | a, the, an |
| The largest planet in our solar system is | Jupiter | 29 | the, a, solar |
| World War II ended in | 1945 | 5 | the, 19, 18 |
| HIV is a | virus | 7 | common, very, disease |
| The American flag is red, white and | blue | 4 | white, red, black |
| The sun rises in the | east | 65 | middle, morning, mid |
| Fish breathe using their | gills | 41 | skin, mouth, l |
| 2 + 2 = | 4 | 4 | 2, 1, 0 |
| 3 + 5 = | 8 | 19 | 1, 5, 2 |
| 7 + 6 = | 13 | 30 | 5, 1, 4 |
MMLU (mmlu_eval.py)
| Format | Accuracy | Questions scored | Average few-shot examples |
|---|---|---|---|
| letter | 24.9% | 14,037 (skipped 5) | 4.5 |
| cloze | 25.5% | 14,037 (skipped 5) | 0.0 |
By category (accuracy weighted by question count; chance is 25%):
| Category | Subjects | Questions | Letter | Cloze |
|---|---|---|---|---|
| STEM | 19 | 3,153 | 26.8% | 24.0% |
| Humanities | 13 | 4,705 | 23.8% | 25.3% |
| Social Sciences | 12 | 3,077 | 23.9% | 26.1% |
| Other | 13 | 3,107 | 25.5% | 26.7% |
Biology & medicine (cloze): 27.8% on 2,094 questions, against 25% chance. That's +2.8 points, about 3.0 standard errors above chance, so it's statistically significant. 9 of the 10 subjects are at or above chance. The overall cloze score is close to chance, so the signal sits mostly in the life-science and health subjects, where FineWeb-Edu's educational web text is richest. Maths-heavy subjects come out below chance, since cloze can't guess numeric answers. This group isn't an official MMLU category; it pools these 10 subjects:
| Subject | Questions | Cloze accuracy |
|---|---|---|
| clinical knowledge | 265 | 32.5% |
| human aging | 223 | 29.6% |
| professional medicine | 272 | 29.0% |
| virology | 166 | 28.3% |
| medical genetics | 100 | 28.0% |
| college medicine | 173 | 27.4% |
| high school biology | 310 | 26.8% |
| college biology | 144 | 25.7% |
| nutrition | 306 | 25.2% |
| anatomy | 135 | 23.7% |
Chance is 25%. In the letter format the model compares " A", " B", " C" and " D", with up to 5 same-subject examples included only while they fit in the 1,024-token window. In the cloze format each answer's text is scored by log-probability per byte. Small models in the letter format tend to answer the same letter within a subject, so per-subject letter scores mostly reflect how the answer key happens to be distributed. The cloze results are the meaningful ones here.
What it can and can't do
- ✅ Grammatical, topic-consistent English for a few sentences, in the register of educational web text.
- ✅ Coarse associations: it knows HIV is a virus rather than a mineral, ice is frozen water, and a week has seven days.
- ❌ Specific facts and relations (capital cities, which body orbits which), arithmetic beyond memorised cases like "2 + 2 = 4", and following instructions.
- ❌ It invents names, numbers and "facts" fluently, and falls into repetition loops at low temperature. Use top-p ≈ 0.9 and a repetition penalty of about 1.1–1.2.
Usage
llama.cpp, Ollama or LM Studio (GGUF, llama.cpp's built-in deepseek2 architecture):
llama-completion -m ShallowSeek-mini-base-q8_0.gguf -p "Photosynthesis is" -n 100 --temp 0.7 --top-p 0.9 --repeat-penalty 1.15
PyTorch (needs the deepseek_moe package from the training code):
import json, torch
from safetensors.torch import load_file
from deepseek_moe import DeepSeekMoEModel, ModelConfig
from deepseek_moe.tokenizer import BPETokenizer
model = DeepSeekMoEModel(ModelConfig.from_dict(json.load(open("config.json")))).eval()
model.load_state_dict(load_file("model.safetensors"))
tok = BPETokenizer.from_file("tokenizer.json")
ids = torch.tensor([[tok.eot_id] + tok.encode("Photosynthesis is")])
print(tok.decode(model.generate(ids, 60, temperature=0.7, top_p=0.9, repetition_penalty=1.15)[0].tolist()))
Files
| File | Contents |
|---|---|
model.safetensors |
PyTorch weights (including the MTP module), our own parameter names |
config.json |
ModelConfig for deepseek_moe |
tokenizer.json |
Byte-level BPE tokenizer (Hugging Face tokenizers format) |
ShallowSeek-mini-base-f16.gguf |
GGUF (f16) for llama.cpp / Ollama / LM Studio |
ShallowSeek-mini-base-q8_0.gguf |
GGUF (q8_0) for llama.cpp / Ollama / LM Studio |
Acknowledgements
- The architecture follows DeepSeek-V3 (DeepSeek-AI, 2024): MLA, DeepSeekMoE, auxiliary-loss-free balancing and multi-token prediction.
- Pretraining data: FineWeb-Edu (Hugging Face, ODC-BY 1.0).
Card generated 2026-10-01 by prepare_release.py.
- Downloads last month
- 90