---
license: apache-2.0
language:
- en
pipeline_tag: text-generation
datasets:
- HuggingFaceFW/fineweb-edu
- HuggingFaceTB/cosmopedia
- HuggingFaceTB/finemath
- HuggingFaceTB/smoltalk
- codeparrot/github-code-clean
tags:
- mixture-of-experts
- moe
- deepseek
- multi-head-latent-attention
- mla
- from-scratch
- tiny
- littlekedi
- base
---

# LittleKedi-tiny-base
`LittleKedi-tiny-base` is a **small DeepSeek-V3-style Mixture-of-Experts language model trained from scratch**. It has
**2.0M parameters** in total, of which **1.5M are active per token**
(0.50M excluding the embedding table). It was pretrained on a 1,000M-token mix (FineWeb-Edu, Cosmopedia (Stanford), Cosmopedia (OpenStax), Cosmopedia (WikiHow), Cosmopedia (Khan Academy), FineMath 4+, SmolTalk, github-code-clean).
*Part of the LittleKedi family. Not affiliated with DeepSeek.*
This is the **base model**, before any instruction tuning. It continues text; it doesn't follow instructions. It's an educational and research artifact: it writes fluent, on-topic English but **knows few facts and confidently makes things up**. Don't use it for anything that matters.
## The LittleKedi family
LittleKedi is a family of small DeepSeek-V3-style Mixture-of-Experts language models trained from
scratch. Every size comes in three variants that differ in **how sparse** they are: how many routed experts
they have, and what share of their parameters a token actually passes through.
| Variant | Routed experts |
|---|---|
| **standard** (``) | a few wide experts; the densest variant |
| **sparse** (`-sparse`) | many narrower, fine-grained experts, more of them active per token |
| **wide-sparse** (`-wide-sparse`) | more experts again; the largest total and smallest active share |
More experts means more total capacity to store knowledge, but each expert
sees fewer tokens (tokens × top-k ÷ experts per layer), so sparser variants need more training data before
their experts are well trained. The variants of a size share a tokenizer and dataset, so they can be
compared directly. Shapes change as the family scales; the table below is generated from the current
presets.
This card is for **`tiny`**, the **standard** tiny model: 1.5M of its 2.0M parameters (78%) are active per token. Its siblings are `tiny-sparse` and `tiny-wide-sparse`.
| Model | Variant | Vocab | Total | Active | Active non-emb. | Active % | Routed experts |
|---|---|---|---|---|---|---|---|
| **`tiny` ← this model** | **standard** | **8K** | **2.0M** | **1.5M** | **0.5M** | **78%** | **8 × 64, top-2** |
| `tiny-sparse` | sparse | 8K | 2.6M | 1.6M | 0.5M | 60% | 32 × 32, top-4 |
| `tiny-wide-sparse` | wide-sparse | 8K | 3.8M | 1.6M | 0.5M | 42% | 64 × 32, top-4 |
| `mini` | standard | 8K | 9.4M | 4.8M | 3.2M | 51% | 12 × 128, top-3 |
| `mini-sparse` | sparse | 8K | 19.8M | 4.8M | 3.2M | 24% | 64 × 64, top-6 |
| `mini-wide-sparse` | wide-sparse | 8K | 36.4M | 4.9M | 3.3M | 13% | 128 × 64, top-6 |
| `small` | standard | 8K | 12.2M | 6.3M | 4.2M | 52% | 16 × 128, top-4 |
| `small-sparse` | sparse | 8K | 24.1M | 6.4M | 4.3M | 27% | 80 × 64, top-8 |
| `small-wide-sparse` | wide-sparse | 8K | 43.9M | 6.5M | 4.4M | 15% | 160 × 64, top-8 |
| `moderate` | standard | 16K | 49.8M | 18.8M | 12.5M | 38% | 24 × 192, top-4 |
| `moderate-sparse` | sparse | 16K | 105.8M | 19.1M | 12.8M | 18% | 120 × 96, top-8 |
| `moderate-wide-sparse` | wide-sparse | 16K | 199.0M | 19.4M | 13.1M | 10% | 240 × 96, top-8 |
| `medium` | standard | 16K | 85.9M | 30.8M | 22.4M | 36% | 24 × 256, top-4 |
| `medium-sparse` | sparse | 16K | 185.3M | 31.2M | 22.8M | 17% | 120 × 128, top-8 |
| `medium-wide-sparse` | wide-sparse | 16K | 350.9M | 31.6M | 23.2M | 9% | 240 × 128, top-8 |
| `base` | standard | 32K | 258.7M | 105.4M | 80.2M | 41% | 32 × 256, top-6 |
| `base-sparse` | sparse | 32K | 542.8M | 106.3M | 81.2M | 20% | 160 × 128, top-12 |
| `base-wide-sparse` | wide-sparse | 32K | 1015.9M | 107.6M | 82.4M | 11% | 320 × 128, top-12 |
"Active" counts everything a token passes through: embeddings, attention, the dense layer, shared
experts and its top-k routed experts. "Active non-embedding" leaves out the embedding table (tied with the
output head), which costs memory but almost no compute; it's the fair number for comparing compute
between models.
## Architecture
| Component | This model |
|---|---|
| Parameters | 1.99M total, 1.55M active per token (78%), 0.50M active non-embedding |
| Layers / hidden size | 4 / 128 (layer 0 has a dense SwiGLU FFN of 352; layers 1–3 are MoE) |
| Attention | **Multi-head Latent Attention (MLA)**: 4 heads, KV latent 32, decoupled RoPE dim 16, head dims 16+16 (q/k) and 16 (v). The KV cache stores only the 32-dim latent plus a 16-dim RoPE key per token |
| MoE | **DeepSeekMoE**: 8 routed experts (SwiGLU, 64 hidden), top-2, plus 1 shared expert; sigmoid gating |
| Load balancing | **Auxiliary-loss-free** bias balancing (update speed 0.001) plus a small sequence-wise loss (α = 0.0001). Over the second half of training the median batch put 1.07× the mean load on the busiest expert (1.08× in the worst layer), and 0.00% of expert assignments were dropped for capacity. |
| Expert dispatch | Training: fixed capacity (capacity factor 2.0; overflow assignments are dropped). Inference is always dropless |
| Embeddings | Input embedding and output head **tied** |
| Context | 1,024 tokens |
## Training
| | |
|---|---|
| Data | 1,000M tokens, 801,792 documents (~502 tokens per parameter, ~646 per active parameter), one pass, sources interleaved throughout (mix below) |
| Tokenizer | Byte-level BPE, 8,192 tokens (~3.75 bytes/token on this mix); digits not split (v1 tokenizer); no chat tokens, so chat data is plain `User:` / `Assistant:` text |
| Steps | 15,258 × 65,536 tokens (batch 16 × 1024 × 4 accumulation) |
| Optimizer | **Muon** (momentum 0.95, Nesterov, 5 Newton-Schulz steps; each expert's matrix orthogonalized separately) for hidden matrices, AdamW (β = 0.9, 0.95) for embeddings, norms and the router; weight decay 0.1, gradient clipping 1.0 |
| Learning rate | 2e-3 peak; **WSD**: 200 warm-up steps, constant, then linear decay to 0 over the last 20% of steps (schedule set for 15,258 steps) |
| Precision | bf16 autocast, `torch.compile` |
| Compute | ~3e15 FLOPs (6 × active non-embedding parameters × tokens) |
| Hardware | One AMD Radeon 890M integrated GPU (Ryzen AI 300 laptop), ROCm on Windows: about 8.9 hours at a median ~31.1K tokens/s |
| Source | What it is | Tokens | Share |
|---|---|---|---|
| FineWeb-Edu `sample/10BT` | educational web pages | 700M | 70% |
| Cosmopedia (Stanford) | synthetic textbooks | 70M | 7% |
| Cosmopedia (OpenStax) | synthetic textbooks | 30M | 3% |
| Cosmopedia (WikiHow) | synthetic how-to articles | 30M | 3% |
| Cosmopedia (Khan Academy) | synthetic lessons | 20M | 2% |
| FineMath 4+ | maths web pages | 80M | 8% |
| SmolTalk | chat conversations | 50M | 5% |
| github-code-clean | source code | 20M | 2% |
## Evaluation
| Metric | Value |
|---|---|
| Validation loss (nats/token, held-out mix) | **3.869** |
| Perplexity | 47.9 |
| Bits per byte | 1.438 |
| Experts never chosen on the validation sample | 0 |
Validation loss over training: 5.60 (16M tok) → 4.17 (164M tok) → 4.08 (295M tok) → 4.04 (442M tok) → 4.02 (590M tok) → 4.01 (737M tok) → 3.96 (868M tok) → 3.87 (1,000M tok)
### Zero-shot benchmarks (`evals.py`)
| Task | acc_norm | acc | Items | Chance |
|---|---|---|---|---|
| SciQ | 56.3% ± 1.6% | 59.8% | 1,000 | 25% |
| ARC-Easy | 32.7% ± 1.0% | 32.3% | 2,376 | 25% |
| HellaSwag | 26.7% ± 0.4% | 26.4% | 10,042 | 25% |
Each answer choice is scored by the log-probability of its text after the prompt, using lm-evaluation-harness
prompts (SciQ includes its supporting passage). **acc_norm** divides by the answer's length in characters;
**acc** doesn't. Splits: SciQ test, ARC-Easy test, HellaSwag validation.
### MMLU (`mmlu_eval.py`)
| Format | Accuracy | Questions scored | Average few-shot examples |
|---|---|---|---|
| letter | 25.4% | 14,037 (skipped 5) | 4.5 |
| cloze | 25.8% | 14,037 (skipped 5) | 0.0 |
| Category | Subjects | Questions | Letter | Cloze |
|---|---|---|---|---|
| STEM | 19 | 3,153 | 27.6% | 23.5% |
| Humanities | 13 | 4,705 | 24.4% | 25.5% |
| Social Sciences | 12 | 3,077 | 23.6% | 26.6% |
| Other | 13 | 3,102 | 26.4% | 27.9% |
**Biology & medicine (cloze): 28.4% on 2,089 questions**, +3.4 points against 25% chance (about 3.6 standard errors, significant). This pools the 10 life-science and health subjects; it isn't an official MMLU category.
In the **letter** format the model compares " A"–" D", with same-subject examples added while they fit the
context. In the **cloze** format each answer's text is scored by log-probability per byte; at this size the
cloze results are the meaningful ones. Chance is 25%.
### Compared with its siblings and other small models
| | **This model** | [Pythia-70M](https://huggingface.co/EleutherAI/pythia-70m) | [GPT-2](https://huggingface.co/openai-community/gpt2) | [Supra-50M](https://huggingface.co/SupraLabs/Supra-50M-Base) |
|---|---|---|---|---|
| Parameters (total) | 2.0M | 70.4M | 124.4M | 51.8M |
| Active non-embedding parameters | 0.5M | 18.9M | 85.1M | 35.4M |
| Training tokens | 1,000M | 300B (the Pile) | undisclosed (WebText) | 20B (FineWeb-Edu) |
| SciQ (acc_norm) | 56.3% | 55.2% | 53.2% | 77.2% |
| ARC-Easy (acc_norm) | 32.7% | 35.0% | 42.0% | 52.2% |
| HellaSwag (acc_norm) | 26.7% | – | 29.5% | 31.8% |
| MMLU (cloze) | 25.8% | – | – | – |
| MMLU biology & medicine (cloze) | 28.4% | – | – | – |
| Validation loss (held-out mix) | 3.869 | – | – | – |
LittleKedi columns are measured with this repo's `evals.py` and `mmlu_eval.py` on the checkpoints named. Validation
losses are only comparable between models that share a tokenizer and validation set. The other columns are
**published figures**: Pythia-70M from EleutherAI's evaluations as quoted on the
[Wisp-15M card](https://huggingface.co/DedeProGames/Wisp-15M), GPT-2 and Supra-50M from the
[Supra-50M-Base card](https://huggingface.co/SupraLabs/Supra-50M-Base). Their prompts may differ from ours,
so treat gaps of a few points as noise.
## What it can and can't do
At 0.5M active non-embedding parameters this model has learned **the form of English much better than its
content**. It's a research artifact for studying small MoE models, not a source of information.
**What it does**
- ✅ Writes grammatical, fluent sentences for a paragraph or two, in the register of its training data
(educational web text and textbooks): *"Photosynthesis is the process by which"* → *"…it is used to monitor
and measure the concentration of nutrients."*
- ✅ Stays roughly on topic: a prompt about the heart leads to blood vessels, the body and disease; a story
opening leads to characters and a plot.
- ✅ Picks up document structure from its data mix: markdown headings and worked "Example 1:" sections
after maths prompts, dialogue turns after `User:` / `Assistant:` text.
- ✅ Shows some benchmark signal above chance (SciQ 56.3%, ARC-Easy 32.7%, MMLU biology & medicine 28.4% in
the cloze format), mostly from matching answers to a passage or recognising which option sounds right,
not from recalling facts.
**What it can't do**
- ❌ **Facts.** It reliably names the right kind of thing but the wrong thing: *"The French Revolution began
in"* → *"1906…"*; *"Albert Einstein was"* → *"a co-director of the French Marxi Palm."* Treat every name,
date and number it writes as made up.
- ❌ **Arithmetic.** *"2 + 2 ="* → *"3"*, *"7 + 6 ="* → *"6"*.
- ❌ **Code.** Code was 2% of its training data; given a Python function signature it produces empty
docstrings and whitespace.
- ❌ **Instructions or questions.** It's a base model: it continues text rather than answering it, and with
only 0.5M non-embedding parameters an instruction-tuned version won't be a reliable assistant either.
- ❌ **Long-range coherence.** Topics drift after a few sentences, and invented names appear
(*"the Grovi Palace"*).
**Decoding**
Greedy or low-temperature decoding falls into loops almost immediately (*"The process of the plant is formed
by the plant. The process of the plant is formed by the plant…"*). Use sampling with **temperature ≈ 0.7,
top-p ≈ 0.9 and a repetition penalty of about 1.15**, which is what produced the examples above.
## Usage
**PyTorch** (needs the `littlekedi` package from [the training code](https://github.com/EdmundMartin/LittleKedi)):
```python
import json, torch
from safetensors.torch import load_file
from littlekedi import LittleKediModel, ModelConfig
from littlekedi.tokenizer import BPETokenizer
model = LittleKediModel(ModelConfig.from_dict(json.load(open("config.json")))).eval()
model.load_state_dict(load_file("model.safetensors"))
tok = BPETokenizer.from_file("tokenizer.json")
ids = torch.tensor([[tok.eot_id] + tok.encode("Photosynthesis is")])
print(tok.decode(model.generate(ids, 60, temperature=0.7, top_p=0.9, repetition_penalty=1.15)[0].tolist()))
```
Or interactively: `python sample.py --ckpt .pt`. Base models are released as PyTorch weights only;
GGUF builds come with the instruct models.
## Files
| File | Contents |
|---|---|
| `model.safetensors` | PyTorch weights (embedding tied with the output head), `littlekedi` parameter names |
| `config.json` | `ModelConfig` for `littlekedi` |
| `tokenizer.json` | Byte-level BPE tokenizer (Hugging Face `tokenizers` format) |
| `little_kedi.jpg` | The LittleKedi logo |
## Acknowledgements
- The architecture follows **DeepSeek-V3** (DeepSeek-AI, 2024): MLA, DeepSeekMoE and auxiliary-loss-free balancing.
- **Muon** (Keller Jordan et al., 2024) with Moonlight's update scaling (Liu et al., 2025).
- Pretraining data: **FineWeb-Edu** (ODC-BY 1.0).
- Pretraining data: **Cosmopedia** (Apache 2.0).
- Pretraining data: **FineMath 4+** (ODC-BY 1.0).
- Pretraining data: **SmolTalk** (Apache 2.0 for its new subsets, with the datasets it includes under their own licenses).
- Pretraining data: **github-code-clean** (Apache 2.0, with each file under its own open-source license).
*Card generated 2026-10-09 by `prepare_release.py` from `tiny.pt` (step 15257).*