File size: 14,687 Bytes
277acf1 26bd9b9 277acf1 26bd9b9 277acf1 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 | ---
license: apache-2.0
language:
- en
pipeline_tag: text-generation
datasets:
- HuggingFaceFW/fineweb-edu
- HuggingFaceTB/cosmopedia
- HuggingFaceTB/finemath
- HuggingFaceTB/smoltalk
- codeparrot/github-code-clean
tags:
- mixture-of-experts
- moe
- deepseek
- multi-head-latent-attention
- mla
- from-scratch
- tiny
- littlekedi
- base
---
<p align="center"><img src="little_kedi.jpg" alt="LittleKedi logo" width="640"></p>
# LittleKedi-tiny-base
`LittleKedi-tiny-base` is a **small DeepSeek-V3-style Mixture-of-Experts language model trained from scratch**. It has
**2.0M parameters** in total, of which **1.5M are active per token**
(0.50M excluding the embedding table). It was pretrained on a 1,000M-token mix (FineWeb-Edu, Cosmopedia (Stanford), Cosmopedia (OpenStax), Cosmopedia (WikiHow), Cosmopedia (Khan Academy), FineMath 4+, SmolTalk, github-code-clean).
*Part of the LittleKedi family. Not affiliated with DeepSeek.*
This is the **base model**, before any instruction tuning. It continues text; it doesn't follow instructions. It's an educational and research artifact: it writes fluent, on-topic English but **knows few facts and confidently makes things up**. Don't use it for anything that matters.
## The LittleKedi family
LittleKedi is a family of small DeepSeek-V3-style Mixture-of-Experts language models trained from
scratch. Every size comes in three variants that differ in **how sparse** they are: how many routed experts
they have, and what share of their parameters a token actually passes through.
| Variant | Routed experts |
|---|---|
| **standard** (`<size>`) | a few wide experts; the densest variant |
| **sparse** (`<size>-sparse`) | many narrower, fine-grained experts, more of them active per token |
| **wide-sparse** (`<size>-wide-sparse`) | more experts again; the largest total and smallest active share |
More experts means more total capacity to store knowledge, but each expert
sees fewer tokens (tokens Γ top-k Γ· experts per layer), so sparser variants need more training data before
their experts are well trained. The variants of a size share a tokenizer and dataset, so they can be
compared directly. Shapes change as the family scales; the table below is generated from the current
presets.
This card is for **`tiny`**, the **standard** tiny model: 1.5M of its 2.0M parameters (78%) are active per token. Its siblings are `tiny-sparse` and `tiny-wide-sparse`.
| Model | Variant | Vocab | Total | Active | Active non-emb. | Active % | Routed experts |
|---|---|---|---|---|---|---|---|
| **`tiny` β this model** | **standard** | **8K** | **2.0M** | **1.5M** | **0.5M** | **78%** | **8 Γ 64, top-2** |
| `tiny-sparse` | sparse | 8K | 2.6M | 1.6M | 0.5M | 60% | 32 Γ 32, top-4 |
| `tiny-wide-sparse` | wide-sparse | 8K | 3.8M | 1.6M | 0.5M | 42% | 64 Γ 32, top-4 |
| `mini` | standard | 8K | 9.4M | 4.8M | 3.2M | 51% | 12 Γ 128, top-3 |
| `mini-sparse` | sparse | 8K | 19.8M | 4.8M | 3.2M | 24% | 64 Γ 64, top-6 |
| `mini-wide-sparse` | wide-sparse | 8K | 36.4M | 4.9M | 3.3M | 13% | 128 Γ 64, top-6 |
| `small` | standard | 8K | 12.2M | 6.3M | 4.2M | 52% | 16 Γ 128, top-4 |
| `small-sparse` | sparse | 8K | 24.1M | 6.4M | 4.3M | 27% | 80 Γ 64, top-8 |
| `small-wide-sparse` | wide-sparse | 8K | 43.9M | 6.5M | 4.4M | 15% | 160 Γ 64, top-8 |
| `moderate` | standard | 16K | 49.8M | 18.8M | 12.5M | 38% | 24 Γ 192, top-4 |
| `moderate-sparse` | sparse | 16K | 105.8M | 19.1M | 12.8M | 18% | 120 Γ 96, top-8 |
| `moderate-wide-sparse` | wide-sparse | 16K | 199.0M | 19.4M | 13.1M | 10% | 240 Γ 96, top-8 |
| `medium` | standard | 16K | 85.9M | 30.8M | 22.4M | 36% | 24 Γ 256, top-4 |
| `medium-sparse` | sparse | 16K | 185.3M | 31.2M | 22.8M | 17% | 120 Γ 128, top-8 |
| `medium-wide-sparse` | wide-sparse | 16K | 350.9M | 31.6M | 23.2M | 9% | 240 Γ 128, top-8 |
| `base` | standard | 32K | 258.7M | 105.4M | 80.2M | 41% | 32 Γ 256, top-6 |
| `base-sparse` | sparse | 32K | 542.8M | 106.3M | 81.2M | 20% | 160 Γ 128, top-12 |
| `base-wide-sparse` | wide-sparse | 32K | 1015.9M | 107.6M | 82.4M | 11% | 320 Γ 128, top-12 |
"Active" counts everything a token passes through: embeddings, attention, the dense layer, shared
experts and its top-k routed experts. "Active non-embedding" leaves out the embedding table (tied with the
output head), which costs memory but almost no compute; it's the fair number for comparing compute
between models.
## Architecture
| Component | This model |
|---|---|
| Parameters | 1.99M total, 1.55M active per token (78%), 0.50M active non-embedding |
| Layers / hidden size | 4 / 128 (layer 0 has a dense SwiGLU FFN of 352; layers 1β3 are MoE) |
| Attention | **Multi-head Latent Attention (MLA)**: 4 heads, KV latent 32, decoupled RoPE dim 16, head dims 16+16 (q/k) and 16 (v). The KV cache stores only the 32-dim latent plus a 16-dim RoPE key per token |
| MoE | **DeepSeekMoE**: 8 routed experts (SwiGLU, 64 hidden), top-2, plus 1 shared expert; sigmoid gating |
| Load balancing | **Auxiliary-loss-free** bias balancing (update speed 0.001) plus a small sequence-wise loss (Ξ± = 0.0001). Over the second half of training the median batch put 1.07Γ the mean load on the busiest expert (1.08Γ in the worst layer), and 0.00% of expert assignments were dropped for capacity. |
| Expert dispatch | Training: fixed capacity (capacity factor 2.0; overflow assignments are dropped). Inference is always dropless |
| Embeddings | Input embedding and output head **tied** |
| Context | 1,024 tokens |
## Training
| | |
|---|---|
| Data | 1,000M tokens, 801,792 documents (~502 tokens per parameter, ~646 per active parameter), one pass, sources interleaved throughout (mix below) |
| Tokenizer | Byte-level BPE, 8,192 tokens (~3.75 bytes/token on this mix); digits not split (v1 tokenizer); no chat tokens, so chat data is plain `User:` / `Assistant:` text |
| Steps | 15,258 Γ 65,536 tokens (batch 16 Γ 1024 Γ 4 accumulation) |
| Optimizer | **Muon** (momentum 0.95, Nesterov, 5 Newton-Schulz steps; each expert's matrix orthogonalized separately) for hidden matrices, AdamW (Ξ² = 0.9, 0.95) for embeddings, norms and the router; weight decay 0.1, gradient clipping 1.0 |
| Learning rate | 2e-3 peak; **WSD**: 200 warm-up steps, constant, then linear decay to 0 over the last 20% of steps (schedule set for 15,258 steps) |
| Precision | bf16 autocast, `torch.compile` |
| Compute | ~3e15 FLOPs (6 Γ active non-embedding parameters Γ tokens) |
| Hardware | One AMD Radeon 890M integrated GPU (Ryzen AI 300 laptop), ROCm on Windows: about 8.9 hours at a median ~31.1K tokens/s |
| Source | What it is | Tokens | Share |
|---|---|---|---|
| FineWeb-Edu `sample/10BT` | educational web pages | 700M | 70% |
| Cosmopedia (Stanford) | synthetic textbooks | 70M | 7% |
| Cosmopedia (OpenStax) | synthetic textbooks | 30M | 3% |
| Cosmopedia (WikiHow) | synthetic how-to articles | 30M | 3% |
| Cosmopedia (Khan Academy) | synthetic lessons | 20M | 2% |
| FineMath 4+ | maths web pages | 80M | 8% |
| SmolTalk | chat conversations | 50M | 5% |
| github-code-clean | source code | 20M | 2% |
## Evaluation
| Metric | Value |
|---|---|
| Validation loss (nats/token, held-out mix) | **3.869** |
| Perplexity | 47.9 |
| Bits per byte | 1.438 |
| Experts never chosen on the validation sample | 0 |
Validation loss over training: 5.60 (16M tok) β 4.17 (164M tok) β 4.08 (295M tok) β 4.04 (442M tok) β 4.02 (590M tok) β 4.01 (737M tok) β 3.96 (868M tok) β 3.87 (1,000M tok)
### Zero-shot benchmarks (`evals.py`)
| Task | acc_norm | acc | Items | Chance |
|---|---|---|---|---|
| SciQ | 56.3% Β± 1.6% | 59.8% | 1,000 | 25% |
| ARC-Easy | 32.7% Β± 1.0% | 32.3% | 2,376 | 25% |
| HellaSwag | 26.7% Β± 0.4% | 26.4% | 10,042 | 25% |
Each answer choice is scored by the log-probability of its text after the prompt, using lm-evaluation-harness
prompts (SciQ includes its supporting passage). **acc_norm** divides by the answer's length in characters;
**acc** doesn't. Splits: SciQ test, ARC-Easy test, HellaSwag validation.
### MMLU (`mmlu_eval.py`)
| Format | Accuracy | Questions scored | Average few-shot examples |
|---|---|---|---|
| letter | 25.4% | 14,037 (skipped 5) | 4.5 |
| cloze | 25.8% | 14,037 (skipped 5) | 0.0 |
| Category | Subjects | Questions | Letter | Cloze |
|---|---|---|---|---|
| STEM | 19 | 3,153 | 27.6% | 23.5% |
| Humanities | 13 | 4,705 | 24.4% | 25.5% |
| Social Sciences | 12 | 3,077 | 23.6% | 26.6% |
| Other | 13 | 3,102 | 26.4% | 27.9% |
**Biology & medicine (cloze): 28.4% on 2,089 questions**, +3.4 points against 25% chance (about 3.6 standard errors, significant). This pools the 10 life-science and health subjects; it isn't an official MMLU category.
In the **letter** format the model compares " A"β" D", with same-subject examples added while they fit the
context. In the **cloze** format each answer's text is scored by log-probability per byte; at this size the
cloze results are the meaningful ones. Chance is 25%.
### Compared with its siblings and other small models
| | **This model** | [Pythia-70M](https://huggingface.co/EleutherAI/pythia-70m) | [GPT-2](https://huggingface.co/openai-community/gpt2) | [Supra-50M](https://huggingface.co/SupraLabs/Supra-50M-Base) |
|---|---|---|---|---|
| Parameters (total) | 2.0M | 70.4M | 124.4M | 51.8M |
| Active non-embedding parameters | 0.5M | 18.9M | 85.1M | 35.4M |
| Training tokens | 1,000M | 300B (the Pile) | undisclosed (WebText) | 20B (FineWeb-Edu) |
| SciQ (acc_norm) | 56.3% | 55.2% | 53.2% | 77.2% |
| ARC-Easy (acc_norm) | 32.7% | 35.0% | 42.0% | 52.2% |
| HellaSwag (acc_norm) | 26.7% | β | 29.5% | 31.8% |
| MMLU (cloze) | 25.8% | β | β | β |
| MMLU biology & medicine (cloze) | 28.4% | β | β | β |
| Validation loss (held-out mix) | 3.869 | β | β | β |
LittleKedi columns are measured with this repo's `evals.py` and `mmlu_eval.py` on the checkpoints named. Validation
losses are only comparable between models that share a tokenizer and validation set. The other columns are
**published figures**: Pythia-70M from EleutherAI's evaluations as quoted on the
[Wisp-15M card](https://huggingface.co/DedeProGames/Wisp-15M), GPT-2 and Supra-50M from the
[Supra-50M-Base card](https://huggingface.co/SupraLabs/Supra-50M-Base). Their prompts may differ from ours,
so treat gaps of a few points as noise.
## What it can and can't do
At 0.5M active non-embedding parameters this model has learned **the form of English much better than its
content**. It's a research artifact for studying small MoE models, not a source of information.
**What it does**
- β
Writes grammatical, fluent sentences for a paragraph or two, in the register of its training data
(educational web text and textbooks): *"Photosynthesis is the process by which"* β *"β¦it is used to monitor
and measure the concentration of nutrients."*
- β
Stays roughly on topic: a prompt about the heart leads to blood vessels, the body and disease; a story
opening leads to characters and a plot.
- β
Picks up document structure from its data mix: markdown headings and worked "Example 1:" sections
after maths prompts, dialogue turns after `User:` / `Assistant:` text.
- β
Shows some benchmark signal above chance (SciQ 56.3%, ARC-Easy 32.7%, MMLU biology & medicine 28.4% in
the cloze format), mostly from matching answers to a passage or recognising which option sounds right,
not from recalling facts.
**What it can't do**
- β **Facts.** It reliably names the right kind of thing but the wrong thing: *"The French Revolution began
in"* β *"1906β¦"*; *"Albert Einstein was"* β *"a co-director of the French Marxi Palm."* Treat every name,
date and number it writes as made up.
- β **Arithmetic.** *"2 + 2 ="* β *"3"*, *"7 + 6 ="* β *"6"*.
- β **Code.** Code was 2% of its training data; given a Python function signature it produces empty
docstrings and whitespace.
- β **Instructions or questions.** It's a base model: it continues text rather than answering it, and with
only 0.5M non-embedding parameters an instruction-tuned version won't be a reliable assistant either.
- β **Long-range coherence.** Topics drift after a few sentences, and invented names appear
(*"the Grovi Palace"*).
**Decoding**
Greedy or low-temperature decoding falls into loops almost immediately (*"The process of the plant is formed
by the plant. The process of the plant is formed by the plantβ¦"*). Use sampling with **temperature β 0.7,
top-p β 0.9 and a repetition penalty of about 1.15**, which is what produced the examples above.
## Usage
**PyTorch** (needs the `littlekedi` package from [the training code](https://github.com/EdmundMartin/LittleKedi)):
```python
import json, torch
from safetensors.torch import load_file
from littlekedi import LittleKediModel, ModelConfig
from littlekedi.tokenizer import BPETokenizer
model = LittleKediModel(ModelConfig.from_dict(json.load(open("config.json")))).eval()
model.load_state_dict(load_file("model.safetensors"))
tok = BPETokenizer.from_file("tokenizer.json")
ids = torch.tensor([[tok.eot_id] + tok.encode("Photosynthesis is")])
print(tok.decode(model.generate(ids, 60, temperature=0.7, top_p=0.9, repetition_penalty=1.15)[0].tolist()))
```
Or interactively: `python sample.py --ckpt <checkpoint>.pt`. Base models are released as PyTorch weights only;
GGUF builds come with the instruct models.
## Files
| File | Contents |
|---|---|
| `model.safetensors` | PyTorch weights (embedding tied with the output head), `littlekedi` parameter names |
| `config.json` | `ModelConfig` for `littlekedi` |
| `tokenizer.json` | Byte-level BPE tokenizer (Hugging Face `tokenizers` format) |
| `little_kedi.jpg` | The LittleKedi logo |
## Acknowledgements
- The architecture follows **DeepSeek-V3** (DeepSeek-AI, 2024): MLA, DeepSeekMoE and auxiliary-loss-free balancing.
- **Muon** (Keller Jordan et al., 2024) with Moonlight's update scaling (Liu et al., 2025).
- Pretraining data: **FineWeb-Edu** (ODC-BY 1.0).
- Pretraining data: **Cosmopedia** (Apache 2.0).
- Pretraining data: **FineMath 4+** (ODC-BY 1.0).
- Pretraining data: **SmolTalk** (Apache 2.0 for its new subsets, with the datasets it includes under their own licenses).
- Pretraining data: **github-code-clean** (Apache 2.0, with each file under its own open-source license).
*Card generated 2026-10-09 by `prepare_release.py` from `tiny.pt` (step 15257).*
|