File size: 14,687 Bytes
277acf1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
26bd9b9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
277acf1
 
 
 
26bd9b9
277acf1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
---

license: apache-2.0
language:
- en
pipeline_tag: text-generation
datasets:
- HuggingFaceFW/fineweb-edu
- HuggingFaceTB/cosmopedia
- HuggingFaceTB/finemath
- HuggingFaceTB/smoltalk
- codeparrot/github-code-clean
tags:
- mixture-of-experts
- moe
- deepseek
- multi-head-latent-attention
- mla
- from-scratch
- tiny
- littlekedi
- base
---


<p align="center"><img src="little_kedi.jpg" alt="LittleKedi logo" width="640"></p>

# LittleKedi-tiny-base

`LittleKedi-tiny-base` is a **small DeepSeek-V3-style Mixture-of-Experts language model trained from scratch**. It has
**2.0M parameters** in total, of which **1.5M are active per token**
(0.50M excluding the embedding table). It was pretrained on a 1,000M-token mix (FineWeb-Edu, Cosmopedia (Stanford), Cosmopedia (OpenStax), Cosmopedia (WikiHow), Cosmopedia (Khan Academy), FineMath 4+, SmolTalk, github-code-clean).

*Part of the LittleKedi family. Not affiliated with DeepSeek.*

This is the **base model**, before any instruction tuning. It continues text; it doesn't follow instructions. It's an educational and research artifact: it writes fluent, on-topic English but **knows few facts and confidently makes things up**. Don't use it for anything that matters.

## The LittleKedi family

LittleKedi is a family of small DeepSeek-V3-style Mixture-of-Experts language models trained from
scratch. Every size comes in three variants that differ in **how sparse** they are: how many routed experts
they have, and what share of their parameters a token actually passes through.

| Variant | Routed experts |
|---|---|
| **standard** (`<size>`) | a few wide experts; the densest variant |
| **sparse** (`<size>-sparse`) | many narrower, fine-grained experts, more of them active per token |
| **wide-sparse** (`<size>-wide-sparse`) | more experts again; the largest total and smallest active share |

More experts means more total capacity to store knowledge, but each expert
sees fewer tokens (tokens Γ— top-k Γ· experts per layer), so sparser variants need more training data before
their experts are well trained. The variants of a size share a tokenizer and dataset, so they can be
compared directly. Shapes change as the family scales; the table below is generated from the current
presets.

This card is for **`tiny`**, the **standard** tiny model: 1.5M of its 2.0M parameters (78%) are active per token. Its siblings are `tiny-sparse` and `tiny-wide-sparse`.

| Model | Variant | Vocab | Total | Active | Active non-emb. | Active % | Routed experts |
|---|---|---|---|---|---|---|---|
| **`tiny` ← this model** | **standard** | **8K** | **2.0M** | **1.5M** | **0.5M** | **78%** | **8 Γ— 64, top-2** |
| `tiny-sparse` | sparse | 8K | 2.6M | 1.6M | 0.5M | 60% | 32 Γ— 32, top-4 |
| `tiny-wide-sparse` | wide-sparse | 8K | 3.8M | 1.6M | 0.5M | 42% | 64 Γ— 32, top-4 |
| `mini` | standard | 8K | 9.4M | 4.8M | 3.2M | 51% | 12 Γ— 128, top-3 |
| `mini-sparse` | sparse | 8K | 19.8M | 4.8M | 3.2M | 24% | 64 Γ— 64, top-6 |
| `mini-wide-sparse` | wide-sparse | 8K | 36.4M | 4.9M | 3.3M | 13% | 128 Γ— 64, top-6 |
| `small` | standard | 8K | 12.2M | 6.3M | 4.2M | 52% | 16 Γ— 128, top-4 |
| `small-sparse` | sparse | 8K | 24.1M | 6.4M | 4.3M | 27% | 80 Γ— 64, top-8 |
| `small-wide-sparse` | wide-sparse | 8K | 43.9M | 6.5M | 4.4M | 15% | 160 Γ— 64, top-8 |
| `moderate` | standard | 16K | 49.8M | 18.8M | 12.5M | 38% | 24 Γ— 192, top-4 |
| `moderate-sparse` | sparse | 16K | 105.8M | 19.1M | 12.8M | 18% | 120 Γ— 96, top-8 |
| `moderate-wide-sparse` | wide-sparse | 16K | 199.0M | 19.4M | 13.1M | 10% | 240 Γ— 96, top-8 |
| `medium` | standard | 16K | 85.9M | 30.8M | 22.4M | 36% | 24 Γ— 256, top-4 |
| `medium-sparse` | sparse | 16K | 185.3M | 31.2M | 22.8M | 17% | 120 Γ— 128, top-8 |
| `medium-wide-sparse` | wide-sparse | 16K | 350.9M | 31.6M | 23.2M | 9% | 240 Γ— 128, top-8 |
| `base` | standard | 32K | 258.7M | 105.4M | 80.2M | 41% | 32 Γ— 256, top-6 |
| `base-sparse` | sparse | 32K | 542.8M | 106.3M | 81.2M | 20% | 160 Γ— 128, top-12 |
| `base-wide-sparse` | wide-sparse | 32K | 1015.9M | 107.6M | 82.4M | 11% | 320 Γ— 128, top-12 |

"Active" counts everything a token passes through: embeddings, attention, the dense layer, shared
experts and its top-k routed experts. "Active non-embedding" leaves out the embedding table (tied with the
output head), which costs memory but almost no compute; it's the fair number for comparing compute
between models.

## Architecture

| Component | This model |
|---|---|
| Parameters | 1.99M total, 1.55M active per token (78%), 0.50M active non-embedding |
| Layers / hidden size | 4 / 128 (layer 0 has a dense SwiGLU FFN of 352; layers 1–3 are MoE) |
| Attention | **Multi-head Latent Attention (MLA)**: 4 heads, KV latent 32, decoupled RoPE dim 16, head dims 16+16 (q/k) and 16 (v). The KV cache stores only the 32-dim latent plus a 16-dim RoPE key per token |
| MoE | **DeepSeekMoE**: 8 routed experts (SwiGLU, 64 hidden), top-2, plus 1 shared expert; sigmoid gating |
| Load balancing | **Auxiliary-loss-free** bias balancing (update speed 0.001) plus a small sequence-wise loss (Ξ± = 0.0001). Over the second half of training the median batch put 1.07Γ— the mean load on the busiest expert (1.08Γ— in the worst layer), and 0.00% of expert assignments were dropped for capacity. |
| Expert dispatch | Training: fixed capacity (capacity factor 2.0; overflow assignments are dropped). Inference is always dropless |
| Embeddings | Input embedding and output head **tied** |
| Context | 1,024 tokens |

## Training

| | |
|---|---|
| Data | 1,000M tokens, 801,792 documents (~502 tokens per parameter, ~646 per active parameter), one pass, sources interleaved throughout (mix below) |
| Tokenizer | Byte-level BPE, 8,192 tokens (~3.75 bytes/token on this mix); digits not split (v1 tokenizer); no chat tokens, so chat data is plain `User:` / `Assistant:` text |
| Steps | 15,258 Γ— 65,536 tokens (batch 16 Γ— 1024 Γ— 4 accumulation) |
| Optimizer | **Muon** (momentum 0.95, Nesterov, 5 Newton-Schulz steps; each expert's matrix orthogonalized separately) for hidden matrices, AdamW (Ξ² = 0.9, 0.95) for embeddings, norms and the router; weight decay 0.1, gradient clipping 1.0 |
| Learning rate | 2e-3 peak; **WSD**: 200 warm-up steps, constant, then linear decay to 0 over the last 20% of steps (schedule set for 15,258 steps) |
| Precision | bf16 autocast, `torch.compile` |
| Compute | ~3e15 FLOPs (6 Γ— active non-embedding parameters Γ— tokens) |
| Hardware | One AMD Radeon 890M integrated GPU (Ryzen AI 300 laptop), ROCm on Windows: about 8.9 hours at a median ~31.1K tokens/s |

| Source | What it is | Tokens | Share |
|---|---|---|---|
| FineWeb-Edu `sample/10BT` | educational web pages | 700M | 70% |
| Cosmopedia (Stanford) | synthetic textbooks | 70M | 7% |
| Cosmopedia (OpenStax) | synthetic textbooks | 30M | 3% |
| Cosmopedia (WikiHow) | synthetic how-to articles | 30M | 3% |
| Cosmopedia (Khan Academy) | synthetic lessons | 20M | 2% |
| FineMath 4+ | maths web pages | 80M | 8% |
| SmolTalk | chat conversations | 50M | 5% |
| github-code-clean | source code | 20M | 2% |

## Evaluation

| Metric | Value |
|---|---|
| Validation loss (nats/token, held-out mix) | **3.869** |
| Perplexity | 47.9 |
| Bits per byte | 1.438 |
| Experts never chosen on the validation sample | 0 |

Validation loss over training: 5.60 (16M tok) β†’ 4.17 (164M tok) β†’ 4.08 (295M tok) β†’ 4.04 (442M tok) β†’ 4.02 (590M tok) β†’ 4.01 (737M tok) β†’ 3.96 (868M tok) β†’ 3.87 (1,000M tok)

### Zero-shot benchmarks (`evals.py`)

| Task | acc_norm | acc | Items | Chance |

|---|---|---|---|---|

| SciQ | 56.3% Β± 1.6% | 59.8% | 1,000 | 25% |

| ARC-Easy | 32.7% Β± 1.0% | 32.3% | 2,376 | 25% |

| HellaSwag | 26.7% Β± 0.4% | 26.4% | 10,042 | 25% |



Each answer choice is scored by the log-probability of its text after the prompt, using lm-evaluation-harness

prompts (SciQ includes its supporting passage). **acc_norm** divides by the answer's length in characters;

**acc** doesn't. Splits: SciQ test, ARC-Easy test, HellaSwag validation.



### MMLU (`mmlu_eval.py`)

| Format | Accuracy | Questions scored | Average few-shot examples |
|---|---|---|---|
| letter | 25.4% | 14,037 (skipped 5) | 4.5 |
| cloze | 25.8% | 14,037 (skipped 5) | 0.0 |

| Category | Subjects | Questions | Letter | Cloze |
|---|---|---|---|---|
| STEM | 19 | 3,153 | 27.6% | 23.5% |
| Humanities | 13 | 4,705 | 24.4% | 25.5% |
| Social Sciences | 12 | 3,077 | 23.6% | 26.6% |
| Other | 13 | 3,102 | 26.4% | 27.9% |

**Biology & medicine (cloze): 28.4% on 2,089 questions**, +3.4 points against 25% chance (about 3.6 standard errors, significant). This pools the 10 life-science and health subjects; it isn't an official MMLU category.

In the **letter** format the model compares " A"–" D", with same-subject examples added while they fit the
context. In the **cloze** format each answer's text is scored by log-probability per byte; at this size the
cloze results are the meaningful ones. Chance is 25%.

### Compared with its siblings and other small models

| | **This model** | [Pythia-70M](https://huggingface.co/EleutherAI/pythia-70m) | [GPT-2](https://huggingface.co/openai-community/gpt2) | [Supra-50M](https://huggingface.co/SupraLabs/Supra-50M-Base) |
|---|---|---|---|---|
| Parameters (total) | 2.0M | 70.4M | 124.4M | 51.8M |
| Active non-embedding parameters | 0.5M | 18.9M | 85.1M | 35.4M |
| Training tokens | 1,000M | 300B (the Pile) | undisclosed (WebText) | 20B (FineWeb-Edu) |
| SciQ (acc_norm) | 56.3% | 55.2% | 53.2% | 77.2% |

| ARC-Easy (acc_norm) | 32.7% | 35.0% | 42.0% | 52.2% |
| HellaSwag (acc_norm) | 26.7% | – | 29.5% | 31.8% |

| MMLU (cloze) | 25.8% | – | – | – |

| MMLU biology & medicine (cloze) | 28.4% | – | – | – |

| Validation loss (held-out mix) | 3.869 | – | – | – |



LittleKedi columns are measured with this repo's `evals.py` and `mmlu_eval.py` on the checkpoints named. Validation
losses are only comparable between models that share a tokenizer and validation set. The other columns are
**published figures**: Pythia-70M from EleutherAI's evaluations as quoted on the
[Wisp-15M card](https://huggingface.co/DedeProGames/Wisp-15M), GPT-2 and Supra-50M from the
[Supra-50M-Base card](https://huggingface.co/SupraLabs/Supra-50M-Base). Their prompts may differ from ours,
so treat gaps of a few points as noise.

## What it can and can't do

At 0.5M active non-embedding parameters this model has learned **the form of English much better than its

content**. It's a research artifact for studying small MoE models, not a source of information.

**What it does**

- βœ… Writes grammatical, fluent sentences for a paragraph or two, in the register of its training data
  (educational web text and textbooks): *"Photosynthesis is the process by which"* β†’ *"…it is used to monitor

  and measure the concentration of nutrients."*
- βœ… Stays roughly on topic: a prompt about the heart leads to blood vessels, the body and disease; a story
  opening leads to characters and a plot.
- βœ… Picks up document structure from its data mix: markdown headings and worked "Example 1:" sections
  after maths prompts, dialogue turns after `User:` / `Assistant:` text.
- βœ… Shows some benchmark signal above chance (SciQ 56.3%, ARC-Easy 32.7%, MMLU biology & medicine 28.4% in
  the cloze format), mostly from matching answers to a passage or recognising which option sounds right,
  not from recalling facts.

**What it can't do**

- ❌ **Facts.** It reliably names the right kind of thing but the wrong thing: *"The French Revolution began

  in"* β†’ *"1906…"*; *"Albert Einstein was"* β†’ *"a co-director of the French Marxi Palm."* Treat every name,
  date and number it writes as made up.
- ❌ **Arithmetic.** *"2 + 2 ="* β†’ *"3"*, *"7 + 6 ="* β†’ *"6"*.
- ❌ **Code.** Code was 2% of its training data; given a Python function signature it produces empty
  docstrings and whitespace.
- ❌ **Instructions or questions.** It's a base model: it continues text rather than answering it, and with
  only 0.5M non-embedding parameters an instruction-tuned version won't be a reliable assistant either.
- ❌ **Long-range coherence.** Topics drift after a few sentences, and invented names appear
  (*"the Grovi Palace"*).

**Decoding**

Greedy or low-temperature decoding falls into loops almost immediately (*"The process of the plant is formed

by the plant. The process of the plant is formed by the plant…"*). Use sampling with **temperature β‰ˆ 0.7,

top-p β‰ˆ 0.9 and a repetition penalty of about 1.15**, which is what produced the examples above.

## Usage

**PyTorch** (needs the `littlekedi` package from [the training code](https://github.com/EdmundMartin/LittleKedi)):

```python

import json, torch

from safetensors.torch import load_file

from littlekedi import LittleKediModel, ModelConfig

from littlekedi.tokenizer import BPETokenizer



model = LittleKediModel(ModelConfig.from_dict(json.load(open("config.json")))).eval()

model.load_state_dict(load_file("model.safetensors"))

tok = BPETokenizer.from_file("tokenizer.json")

ids = torch.tensor([[tok.eot_id] + tok.encode("Photosynthesis is")])

print(tok.decode(model.generate(ids, 60, temperature=0.7, top_p=0.9, repetition_penalty=1.15)[0].tolist()))

```

Or interactively: `python sample.py --ckpt <checkpoint>.pt`. Base models are released as PyTorch weights only;
GGUF builds come with the instruct models.

## Files

| File | Contents |
|---|---|
| `model.safetensors` | PyTorch weights (embedding tied with the output head), `littlekedi` parameter names |
| `config.json` | `ModelConfig` for `littlekedi` |
| `tokenizer.json` | Byte-level BPE tokenizer (Hugging Face `tokenizers` format) |
| `little_kedi.jpg` | The LittleKedi logo |

## Acknowledgements

- The architecture follows **DeepSeek-V3** (DeepSeek-AI, 2024): MLA, DeepSeekMoE and auxiliary-loss-free balancing.
- **Muon** (Keller Jordan et al., 2024) with Moonlight's update scaling (Liu et al., 2025).
- Pretraining data: **FineWeb-Edu** (ODC-BY 1.0).
- Pretraining data: **Cosmopedia** (Apache 2.0).
- Pretraining data: **FineMath 4+** (ODC-BY 1.0).
- Pretraining data: **SmolTalk** (Apache 2.0 for its new subsets, with the datasets it includes under their own licenses).
- Pretraining data: **github-code-clean** (Apache 2.0, with each file under its own open-source license).

*Card generated 2026-10-09 by `prepare_release.py` from `tiny.pt` (step 15257).*