
LittleKedi-tiny-base
LittleKedi-tiny-base is a small DeepSeek-V3-style Mixture-of-Experts language model trained from scratch. It has
2.0M parameters in total, of which 1.5M are active per token
(0.50M excluding the embedding table). It was pretrained on a 1,000M-token mix (FineWeb-Edu, Cosmopedia (Stanford), Cosmopedia (OpenStax), Cosmopedia (WikiHow), Cosmopedia (Khan Academy), FineMath 4+, SmolTalk, github-code-clean).
Part of the LittleKedi family. Not affiliated with DeepSeek.
This is the base model, before any instruction tuning. It continues text; it doesn't follow instructions. It's an educational and research artifact: it writes fluent, on-topic English but knows few facts and confidently makes things up. Don't use it for anything that matters.
The LittleKedi family
LittleKedi is a family of small DeepSeek-V3-style Mixture-of-Experts language models trained from scratch. Every size comes in three variants that differ in how sparse they are: how many routed experts they have, and what share of their parameters a token actually passes through.
| Variant | Routed experts |
|---|---|
standard (<size>) |
a few wide experts; the densest variant |
sparse (<size>-sparse) |
many narrower, fine-grained experts, more of them active per token |
wide-sparse (<size>-wide-sparse) |
more experts again; the largest total and smallest active share |
More experts means more total capacity to store knowledge, but each expert sees fewer tokens (tokens × top-k ÷ experts per layer), so sparser variants need more training data before their experts are well trained. The variants of a size share a tokenizer and dataset, so they can be compared directly. Shapes change as the family scales; the table below is generated from the current presets.
This card is for tiny, the standard tiny model: 1.5M of its 2.0M parameters (78%) are active per token. Its siblings are tiny-sparse and tiny-wide-sparse.
| Model | Variant | Vocab | Total | Active | Active non-emb. | Active % | Routed experts |
|---|---|---|---|---|---|---|---|
tiny ← this model |
standard | 8K | 2.0M | 1.5M | 0.5M | 78% | 8 × 64, top-2 |
tiny-sparse |
sparse | 8K | 2.6M | 1.6M | 0.5M | 60% | 32 × 32, top-4 |
tiny-wide-sparse |
wide-sparse | 8K | 3.8M | 1.6M | 0.5M | 42% | 64 × 32, top-4 |
mini |
standard | 8K | 9.4M | 4.8M | 3.2M | 51% | 12 × 128, top-3 |
mini-sparse |
sparse | 8K | 19.8M | 4.8M | 3.2M | 24% | 64 × 64, top-6 |
mini-wide-sparse |
wide-sparse | 8K | 36.4M | 4.9M | 3.3M | 13% | 128 × 64, top-6 |
small |
standard | 8K | 12.2M | 6.3M | 4.2M | 52% | 16 × 128, top-4 |
small-sparse |
sparse | 8K | 24.1M | 6.4M | 4.3M | 27% | 80 × 64, top-8 |
small-wide-sparse |
wide-sparse | 8K | 43.9M | 6.5M | 4.4M | 15% | 160 × 64, top-8 |
moderate |
standard | 16K | 49.8M | 18.8M | 12.5M | 38% | 24 × 192, top-4 |
moderate-sparse |
sparse | 16K | 105.8M | 19.1M | 12.8M | 18% | 120 × 96, top-8 |
moderate-wide-sparse |
wide-sparse | 16K | 199.0M | 19.4M | 13.1M | 10% | 240 × 96, top-8 |
medium |
standard | 16K | 85.9M | 30.8M | 22.4M | 36% | 24 × 256, top-4 |
medium-sparse |
sparse | 16K | 185.3M | 31.2M | 22.8M | 17% | 120 × 128, top-8 |
medium-wide-sparse |
wide-sparse | 16K | 350.9M | 31.6M | 23.2M | 9% | 240 × 128, top-8 |
base |
standard | 32K | 258.7M | 105.4M | 80.2M | 41% | 32 × 256, top-6 |
base-sparse |
sparse | 32K | 542.8M | 106.3M | 81.2M | 20% | 160 × 128, top-12 |
base-wide-sparse |
wide-sparse | 32K | 1015.9M | 107.6M | 82.4M | 11% | 320 × 128, top-12 |
"Active" counts everything a token passes through: embeddings, attention, the dense layer, shared experts and its top-k routed experts. "Active non-embedding" leaves out the embedding table (tied with the output head), which costs memory but almost no compute; it's the fair number for comparing compute between models.
Architecture
| Component | This model |
|---|---|
| Parameters | 1.99M total, 1.55M active per token (78%), 0.50M active non-embedding |
| Layers / hidden size | 4 / 128 (layer 0 has a dense SwiGLU FFN of 352; layers 1–3 are MoE) |
| Attention | Multi-head Latent Attention (MLA): 4 heads, KV latent 32, decoupled RoPE dim 16, head dims 16+16 (q/k) and 16 (v). The KV cache stores only the 32-dim latent plus a 16-dim RoPE key per token |
| MoE | DeepSeekMoE: 8 routed experts (SwiGLU, 64 hidden), top-2, plus 1 shared expert; sigmoid gating |
| Load balancing | Auxiliary-loss-free bias balancing (update speed 0.001) plus a small sequence-wise loss (α = 0.0001). Over the second half of training the median batch put 1.07× the mean load on the busiest expert (1.08× in the worst layer), and 0.00% of expert assignments were dropped for capacity. |
| Expert dispatch | Training: fixed capacity (capacity factor 2.0; overflow assignments are dropped). Inference is always dropless |
| Embeddings | Input embedding and output head tied |
| Context | 1,024 tokens |
Training
| Data | 1,000M tokens, 801,792 documents (~502 tokens per parameter, ~646 per active parameter), one pass, sources interleaved throughout (mix below) |
| Tokenizer | Byte-level BPE, 8,192 tokens (~3.75 bytes/token on this mix); digits not split (v1 tokenizer); no chat tokens, so chat data is plain User: / Assistant: text |
| Steps | 15,258 × 65,536 tokens (batch 16 × 1024 × 4 accumulation) |
| Optimizer | Muon (momentum 0.95, Nesterov, 5 Newton-Schulz steps; each expert's matrix orthogonalized separately) for hidden matrices, AdamW (β = 0.9, 0.95) for embeddings, norms and the router; weight decay 0.1, gradient clipping 1.0 |
| Learning rate | 2e-3 peak; WSD: 200 warm-up steps, constant, then linear decay to 0 over the last 20% of steps (schedule set for 15,258 steps) |
| Precision | bf16 autocast, torch.compile |
| Compute | ~3e15 FLOPs (6 × active non-embedding parameters × tokens) |
| Hardware | One AMD Radeon 890M integrated GPU (Ryzen AI 300 laptop), ROCm on Windows: about 8.9 hours at a median ~31.1K tokens/s |
| Source | What it is | Tokens | Share |
|---|---|---|---|
FineWeb-Edu sample/10BT |
educational web pages | 700M | 70% |
| Cosmopedia (Stanford) | synthetic textbooks | 70M | 7% |
| Cosmopedia (OpenStax) | synthetic textbooks | 30M | 3% |
| Cosmopedia (WikiHow) | synthetic how-to articles | 30M | 3% |
| Cosmopedia (Khan Academy) | synthetic lessons | 20M | 2% |
| FineMath 4+ | maths web pages | 80M | 8% |
| SmolTalk | chat conversations | 50M | 5% |
| github-code-clean | source code | 20M | 2% |
Evaluation
| Metric | Value |
|---|---|
| Validation loss (nats/token, held-out mix) | 3.869 |
| Perplexity | 47.9 |
| Bits per byte | 1.438 |
| Experts never chosen on the validation sample | 0 |
Validation loss over training: 5.60 (16M tok) → 4.17 (164M tok) → 4.08 (295M tok) → 4.04 (442M tok) → 4.02 (590M tok) → 4.01 (737M tok) → 3.96 (868M tok) → 3.87 (1,000M tok)
Zero-shot benchmarks (evals.py)
| Task | acc_norm | acc | Items | Chance |
|---|---|---|---|---|
| SciQ | 56.3% ± 1.6% | 59.8% | 1,000 | 25% |
| ARC-Easy | 32.7% ± 1.0% | 32.3% | 2,376 | 25% |
| HellaSwag | 26.7% ± 0.4% | 26.4% | 10,042 | 25% |
Each answer choice is scored by the log-probability of its text after the prompt, using lm-evaluation-harness prompts (SciQ includes its supporting passage). acc_norm divides by the answer's length in characters; acc doesn't. Splits: SciQ test, ARC-Easy test, HellaSwag validation.
MMLU (mmlu_eval.py)
| Format | Accuracy | Questions scored | Average few-shot examples |
|---|---|---|---|
| letter | 25.4% | 14,037 (skipped 5) | 4.5 |
| cloze | 25.8% | 14,037 (skipped 5) | 0.0 |
| Category | Subjects | Questions | Letter | Cloze |
|---|---|---|---|---|
| STEM | 19 | 3,153 | 27.6% | 23.5% |
| Humanities | 13 | 4,705 | 24.4% | 25.5% |
| Social Sciences | 12 | 3,077 | 23.6% | 26.6% |
| Other | 13 | 3,102 | 26.4% | 27.9% |
Biology & medicine (cloze): 28.4% on 2,089 questions, +3.4 points against 25% chance (about 3.6 standard errors, significant). This pools the 10 life-science and health subjects; it isn't an official MMLU category.
In the letter format the model compares " A"–" D", with same-subject examples added while they fit the context. In the cloze format each answer's text is scored by log-probability per byte; at this size the cloze results are the meaningful ones. Chance is 25%.
Compared with its siblings and other small models
| This model | Pythia-70M | GPT-2 | Supra-50M | |
|---|---|---|---|---|
| Parameters (total) | 2.0M | 70.4M | 124.4M | 51.8M |
| Active non-embedding parameters | 0.5M | 18.9M | 85.1M | 35.4M |
| Training tokens | 1,000M | 300B (the Pile) | undisclosed (WebText) | 20B (FineWeb-Edu) |
| SciQ (acc_norm) | 56.3% | 55.2% | 53.2% | 77.2% |
| ARC-Easy (acc_norm) | 32.7% | 35.0% | 42.0% | 52.2% |
| HellaSwag (acc_norm) | 26.7% | – | 29.5% | 31.8% |
| MMLU (cloze) | 25.8% | – | – | – |
| MMLU biology & medicine (cloze) | 28.4% | – | – | – |
| Validation loss (held-out mix) | 3.869 | – | – | – |
LittleKedi columns are measured with this repo's evals.py and mmlu_eval.py on the checkpoints named. Validation
losses are only comparable between models that share a tokenizer and validation set. The other columns are
published figures: Pythia-70M from EleutherAI's evaluations as quoted on the
Wisp-15M card, GPT-2 and Supra-50M from the
Supra-50M-Base card. Their prompts may differ from ours,
so treat gaps of a few points as noise.
What it can and can't do
At 0.5M active non-embedding parameters this model has learned the form of English much better than its content. It's a research artifact for studying small MoE models, not a source of information.
What it does
- ✅ Writes grammatical, fluent sentences for a paragraph or two, in the register of its training data (educational web text and textbooks): "Photosynthesis is the process by which" → "…it is used to monitor and measure the concentration of nutrients."
- ✅ Stays roughly on topic: a prompt about the heart leads to blood vessels, the body and disease; a story opening leads to characters and a plot.
- ✅ Picks up document structure from its data mix: markdown headings and worked "Example 1:" sections
after maths prompts, dialogue turns after
User:/Assistant:text. - ✅ Shows some benchmark signal above chance (SciQ 56.3%, ARC-Easy 32.7%, MMLU biology & medicine 28.4% in the cloze format), mostly from matching answers to a passage or recognising which option sounds right, not from recalling facts.
What it can't do
- ❌ Facts. It reliably names the right kind of thing but the wrong thing: "The French Revolution began in" → "1906…"; "Albert Einstein was" → "a co-director of the French Marxi Palm." Treat every name, date and number it writes as made up.
- ❌ Arithmetic. "2 + 2 =" → "3", "7 + 6 =" → "6".
- ❌ Code. Code was 2% of its training data; given a Python function signature it produces empty docstrings and whitespace.
- ❌ Instructions or questions. It's a base model: it continues text rather than answering it, and with only 0.5M non-embedding parameters an instruction-tuned version won't be a reliable assistant either.
- ❌ Long-range coherence. Topics drift after a few sentences, and invented names appear ("the Grovi Palace").
Decoding
Greedy or low-temperature decoding falls into loops almost immediately ("The process of the plant is formed by the plant. The process of the plant is formed by the plant…"). Use sampling with temperature ≈ 0.7, top-p ≈ 0.9 and a repetition penalty of about 1.15, which is what produced the examples above.
Usage
PyTorch (needs the littlekedi package from the training code):
import json, torch
from safetensors.torch import load_file
from littlekedi import LittleKediModel, ModelConfig
from littlekedi.tokenizer import BPETokenizer
model = LittleKediModel(ModelConfig.from_dict(json.load(open("config.json")))).eval()
model.load_state_dict(load_file("model.safetensors"))
tok = BPETokenizer.from_file("tokenizer.json")
ids = torch.tensor([[tok.eot_id] + tok.encode("Photosynthesis is")])
print(tok.decode(model.generate(ids, 60, temperature=0.7, top_p=0.9, repetition_penalty=1.15)[0].tolist()))
Or interactively: python sample.py --ckpt <checkpoint>.pt. Base models are released as PyTorch weights only;
GGUF builds come with the instruct models.
Files
| File | Contents |
|---|---|
model.safetensors |
PyTorch weights (embedding tied with the output head), littlekedi parameter names |
config.json |
ModelConfig for littlekedi |
tokenizer.json |
Byte-level BPE tokenizer (Hugging Face tokenizers format) |
little_kedi.jpg |
The LittleKedi logo |
Acknowledgements
- The architecture follows DeepSeek-V3 (DeepSeek-AI, 2024): MLA, DeepSeekMoE and auxiliary-loss-free balancing.
- Muon (Keller Jordan et al., 2024) with Moonlight's update scaling (Liu et al., 2025).
- Pretraining data: FineWeb-Edu (ODC-BY 1.0).
- Pretraining data: Cosmopedia (Apache 2.0).
- Pretraining data: FineMath 4+ (ODC-BY 1.0).
- Pretraining data: SmolTalk (Apache 2.0 for its new subsets, with the datasets it includes under their own licenses).
- Pretraining data: github-code-clean (Apache 2.0, with each file under its own open-source license).
Card generated 2026-10-09 by prepare_release.py from tiny.pt (step 15257).
- Downloads last month
- 32