
LittleKedi-tiny-wide-sparse-base
LittleKedi-tiny-wide-sparse-base is a small DeepSeek-V3-style Mixture-of-Experts language model trained from scratch. It has
3.8M parameters in total, of which 1.6M are active per token
(0.52M excluding the embedding table). It was pretrained on a 1,000M-token mix (FineWeb-Edu, Cosmopedia (Stanford), Cosmopedia (OpenStax), Cosmopedia (WikiHow), Cosmopedia (Khan Academy), FineMath 4+, SmolTalk, github-code-clean).
Part of the LittleKedi family. Not affiliated with DeepSeek.
This is the base model, before any instruction tuning. It continues text; it doesn't follow instructions. It's an educational and research artifact: it writes fluent, on-topic English but knows few facts and confidently makes things up. Don't use it for anything that matters.
The LittleKedi family
LittleKedi is a family of small DeepSeek-V3-style Mixture-of-Experts language models trained from scratch. Every size comes in three variants that differ in how sparse they are: how many routed experts they have, and what share of their parameters a token actually passes through.
| Variant | Routed experts |
|---|---|
standard (<size>) |
a few wide experts; the densest variant |
sparse (<size>-sparse) |
many narrower, fine-grained experts, more of them active per token |
wide-sparse (<size>-wide-sparse) |
more experts again; the largest total and smallest active share |
More experts means more total capacity to store knowledge, but each expert sees fewer tokens (tokens × top-k ÷ experts per layer), so sparser variants need more training data before their experts are well trained. The variants of a size share a tokenizer and dataset, so they can be compared directly. Shapes change as the family scales; the table below is generated from the current presets.
This card is for tiny-wide-sparse, the wide-sparse tiny model: 1.6M of its 3.8M parameters (42%) are active per token. Its siblings are tiny and tiny-sparse.
| Model | Variant | Vocab | Total | Active | Active non-emb. | Active % | Routed experts |
|---|---|---|---|---|---|---|---|
tiny |
standard | 8K | 2.0M | 1.5M | 0.5M | 78% | 8 × 64, top-2 |
tiny-sparse |
sparse | 8K | 2.6M | 1.6M | 0.5M | 60% | 32 × 32, top-4 |
tiny-wide-sparse ← this model |
wide-sparse | 8K | 3.8M | 1.6M | 0.5M | 42% | 64 × 32, top-4 |
mini |
standard | 8K | 9.4M | 4.8M | 3.2M | 51% | 12 × 128, top-3 |
mini-sparse |
sparse | 8K | 19.8M | 4.8M | 3.2M | 24% | 64 × 64, top-6 |
mini-wide-sparse |
wide-sparse | 8K | 36.4M | 4.9M | 3.3M | 13% | 128 × 64, top-6 |
small |
standard | 8K | 12.2M | 6.3M | 4.2M | 52% | 16 × 128, top-4 |
small-sparse |
sparse | 8K | 24.1M | 6.4M | 4.3M | 27% | 80 × 64, top-8 |
small-wide-sparse |
wide-sparse | 8K | 43.9M | 6.5M | 4.4M | 15% | 160 × 64, top-8 |
moderate |
standard | 16K | 49.8M | 18.8M | 12.5M | 38% | 24 × 192, top-4 |
moderate-sparse |
sparse | 16K | 105.8M | 19.1M | 12.8M | 18% | 120 × 96, top-8 |
moderate-wide-sparse |
wide-sparse | 16K | 199.0M | 19.4M | 13.1M | 10% | 240 × 96, top-8 |
medium |
standard | 16K | 85.9M | 30.8M | 22.4M | 36% | 24 × 256, top-4 |
medium-sparse |
sparse | 16K | 185.3M | 31.2M | 22.8M | 17% | 120 × 128, top-8 |
medium-wide-sparse |
wide-sparse | 16K | 350.9M | 31.6M | 23.2M | 9% | 240 × 128, top-8 |
base |
standard | 32K | 258.7M | 105.4M | 80.2M | 41% | 32 × 256, top-6 |
base-sparse |
sparse | 32K | 542.8M | 106.3M | 81.2M | 20% | 160 × 128, top-12 |
base-wide-sparse |
wide-sparse | 32K | 1015.9M | 107.6M | 82.4M | 11% | 320 × 128, top-12 |
"Active" counts everything a token passes through: embeddings, attention, the dense layer, shared experts and its top-k routed experts. "Active non-embedding" leaves out the embedding table (tied with the output head), which costs memory but almost no compute; it's the fair number for comparing compute between models.
Architecture
| Component | This model |
|---|---|
| Parameters | 3.78M total, 1.57M active per token (42%), 0.52M active non-embedding |
| Layers / hidden size | 4 / 128 (layer 0 has a dense SwiGLU FFN of 352; layers 1–3 are MoE) |
| Attention | Multi-head Latent Attention (MLA): 4 heads, KV latent 32, decoupled RoPE dim 16, head dims 16+16 (q/k) and 16 (v). The KV cache stores only the 32-dim latent plus a 16-dim RoPE key per token |
| MoE | DeepSeekMoE: 64 routed experts (SwiGLU, 32 hidden), top-4, plus 2 shared experts; sigmoid gating; group-limited routing (4 groups, top-2) |
| Load balancing | Auxiliary-loss-free bias balancing (update speed 0.001) plus a small sequence-wise loss (α = 0.0001). Over the second half of training the median batch put 1.23× the mean load on the busiest expert (1.28× in the worst layer), and 0.00% of expert assignments were dropped for capacity. |
| Expert dispatch | Training: fixed capacity (capacity factor 2.0; overflow assignments are dropped). Inference is always dropless |
| Embeddings | Input embedding and output head tied |
| Context | 1,024 tokens |
Training
| Data | 1,000M tokens, 801,792 documents (~264 tokens per parameter, ~637 per active parameter), one pass, sources interleaved throughout (mix below) |
| Tokenizer | Byte-level BPE, 8,192 tokens (~3.75 bytes/token on this mix); digits not split (v1 tokenizer); no chat tokens, so chat data is plain User: / Assistant: text |
| Steps | 15,258 × 65,536 tokens (batch 16 × 1024 × 4 accumulation) |
| Optimizer | Muon (momentum 0.95, Nesterov, 5 Newton-Schulz steps; each expert's matrix orthogonalized separately) for hidden matrices, AdamW (β = 0.9, 0.95) for embeddings, norms and the router; weight decay 0.1, gradient clipping 1.0 |
| Learning rate | 2e-3 peak; WSD: 200 warm-up steps, constant, then linear decay to 0 over the last 20% of steps (schedule set for 15,258 steps) |
| Precision | bf16 autocast, torch.compile |
| Compute | ~3.1e15 FLOPs (6 × active non-embedding parameters × tokens) |
| Hardware | One AMD Radeon 890M integrated GPU (Ryzen AI 300 laptop), ROCm on Windows: about 10.4 hours at a median ~26.7K tokens/s |
| Source | What it is | Tokens | Share |
|---|---|---|---|
FineWeb-Edu sample/10BT |
educational web pages | 700M | 70% |
| Cosmopedia (Stanford) | synthetic textbooks | 70M | 7% |
| Cosmopedia (OpenStax) | synthetic textbooks | 30M | 3% |
| Cosmopedia (WikiHow) | synthetic how-to articles | 30M | 3% |
| Cosmopedia (Khan Academy) | synthetic lessons | 20M | 2% |
| FineMath 4+ | maths web pages | 80M | 8% |
| SmolTalk | chat conversations | 50M | 5% |
| github-code-clean | source code | 20M | 2% |
Evaluation
| Metric | Value |
|---|---|
| Validation loss (nats/token, held-out mix) | 3.736 |
| Perplexity | 41.9 |
| Bits per byte | 1.389 |
| Experts never chosen on the validation sample | 0 |
Validation loss over training: 5.59 (16M tok) → 4.07 (164M tok) → 3.98 (295M tok) → 3.93 (442M tok) → 3.91 (590M tok) → 3.89 (737M tok) → 3.84 (868M tok) → 3.74 (1,000M tok)
Zero-shot benchmarks (evals.py)
| Task | acc_norm | acc | Items | Chance |
|---|---|---|---|---|
| SciQ | 53.9% ± 1.6% | 60.4% | 1,000 | 25% |
| ARC-Easy | 32.4% ± 1.0% | 32.4% | 2,376 | 25% |
| HellaSwag | 26.3% ± 0.4% | 26.3% | 10,042 | 25% |
Each answer choice is scored by the log-probability of its text after the prompt, using lm-evaluation-harness prompts (SciQ includes its supporting passage). acc_norm divides by the answer's length in characters; acc doesn't. Splits: SciQ test, ARC-Easy test, HellaSwag validation.
MMLU (mmlu_eval.py)
| Format | Accuracy | Questions scored | Average few-shot examples |
|---|---|---|---|
| letter | 25.2% | 14,037 (skipped 5) | 4.5 |
| cloze | 25.7% | 14,037 (skipped 5) | 0.0 |
| Category | Subjects | Questions | Letter | Cloze |
|---|---|---|---|---|
| STEM | 19 | 3,153 | 27.3% | 24.2% |
| Humanities | 13 | 4,705 | 23.5% | 26.0% |
| Social Sciences | 12 | 3,077 | 24.0% | 26.2% |
| Other | 13 | 3,102 | 26.6% | 26.2% |
Biology & medicine (cloze): 27.8% on 2,089 questions, +2.8 points against 25% chance (about 2.9 standard errors, significant). This pools the 10 life-science and health subjects; it isn't an official MMLU category.
In the letter format the model compares " A"–" D", with same-subject examples added while they fit the context. In the cloze format each answer's text is scored by log-probability per byte; at this size the cloze results are the meaningful ones. Chance is 25%.
Compared with its siblings and other small models
| This model | tiny |
Pythia-70M | GPT-2 | Supra-50M | |
|---|---|---|---|---|---|
| Parameters (total) | 3.8M | 2.0M | 70.4M | 124.4M | 51.8M |
| Active non-embedding parameters | 0.5M | 0.5M | 18.9M | 85.1M | 35.4M |
| Training tokens | 1,000M | 1,000M | 300B (the Pile) | undisclosed (WebText) | 20B (FineWeb-Edu) |
| SciQ (acc_norm) | 53.9% | 56.3% | 55.2% | 53.2% | 77.2% |
| ARC-Easy (acc_norm) | 32.4% | 32.7% | 35.0% | 42.0% | 52.2% |
| HellaSwag (acc_norm) | 26.3% | 26.7% | – | 29.5% | 31.8% |
| MMLU (cloze) | 25.7% | 25.8% | – | – | – |
| MMLU biology & medicine (cloze) | 27.8% | 28.4% | – | – | – |
| Validation loss (held-out mix) | 3.736 | 3.869 | – | – | – |
LittleKedi columns are measured with this repo's evals.py and mmlu_eval.py on the checkpoints named. Validation
losses are only comparable between models that share a tokenizer and validation set. The other columns are
published figures: Pythia-70M from EleutherAI's evaluations as quoted on the
Wisp-15M card, GPT-2 and Supra-50M from the
Supra-50M-Base card. Their prompts may differ from ours,
so treat gaps of a few points as noise.
What it can and can't do
This is the sparsest tiny model: the same 0.5M active non-embedding parameters per token as tiny, but
with 64 narrower experts and ~3× tiny's expert capacity. It models text measurably better than tiny (validation
loss 3.736 against 3.869 on the same data, matching tiny's loss with roughly a third of the tokens), but in
use it's a model of the same kind: it has learned the form of English much better than its content. It's a
research artifact for studying sparsity in small MoE models, not a source of information.
What it does
- ✅ Writes grammatical, fluent sentences for a paragraph or two, in the register of educational web text and textbooks.
- ✅ Stays roughly on topic, and sometimes gets the broad frame right where tiny didn't: "The French Revolution began in" → "…the late 18th century" (with sampling; tiny said 1906). Treat this as an occasional improvement, not a reliable one: the same prompt decoded greedily gives "the late 19th century".
- ✅ Picks up document structure from its data mix: markdown headings and
# math:blocks after maths prompts, bullet lists, dialogue turns afterUser:/Assistant:text. - ✅ Benchmark signal above chance, on a par with
tiny(SciQ 53.9%, ARC-Easy 32.4%, MMLU biology & medicine 28% in the cloze format). Its lower loss hasn't turned into higher multiple-choice accuracy at this size.
What it can't do
- ❌ Facts. Names and dates are plausible in kind but usually wrong: "Albert Einstein was" → "a professor of medicine at the University of California"; "DNA is made of" → "a coin, a piece of paper". Treat every name, date and number it writes as made up.
- ❌ Arithmetic. "2 + 2 =" → "6", "7 + 6 =" → "6".
- ❌ Code. Given a Python function signature it produces docstring fragments and whitespace.
- ❌ Instructions or questions. It's a base model: it continues text rather than answering it.
- ❌ Long-range coherence. Topics drift after a sentence or two ("Photosynthesis is the process by which a person's body is able to move in a controlled way").
Decoding
Greedy decoding falls into loops even faster than tiny's ("the heart, the heart, the heart…"). Use sampling with temperature ≈ 0.7, top-p ≈ 0.9 and a repetition penalty of about 1.15.
Usage
PyTorch (needs the littlekedi package from the training code):
import json, torch
from safetensors.torch import load_file
from littlekedi import LittleKediModel, ModelConfig
from littlekedi.tokenizer import BPETokenizer
model = LittleKediModel(ModelConfig.from_dict(json.load(open("config.json")))).eval()
model.load_state_dict(load_file("model.safetensors"))
tok = BPETokenizer.from_file("tokenizer.json")
ids = torch.tensor([[tok.eot_id] + tok.encode("Photosynthesis is")])
print(tok.decode(model.generate(ids, 60, temperature=0.7, top_p=0.9, repetition_penalty=1.15)[0].tolist()))
Or interactively: python sample.py --ckpt <checkpoint>.pt. Base models are released as PyTorch weights only;
GGUF builds come with the instruct models.
Files
| File | Contents |
|---|---|
model.safetensors |
PyTorch weights (embedding tied with the output head), littlekedi parameter names |
config.json |
ModelConfig for littlekedi |
tokenizer.json |
Byte-level BPE tokenizer (Hugging Face tokenizers format) |
little_kedi.jpg |
The LittleKedi logo |
Acknowledgements
- The architecture follows DeepSeek-V3 (DeepSeek-AI, 2024): MLA, DeepSeekMoE and auxiliary-loss-free balancing.
- Muon (Keller Jordan et al., 2024) with Moonlight's update scaling (Liu et al., 2025).
- Pretraining data: FineWeb-Edu (ODC-BY 1.0).
- Pretraining data: Cosmopedia (Apache 2.0).
- Pretraining data: FineMath 4+ (ODC-BY 1.0).
- Pretraining data: SmolTalk (Apache 2.0 for its new subsets, with the datasets it includes under their own licenses).
- Pretraining data: github-code-clean (Apache 2.0, with each file under its own open-source license).
Card generated 2026-10-10 by prepare_release.py from tiny-wide-sparse.pt (step 15257).
- Downloads last month
- 21