LittleKedi logo

LittleKedi-tiny-wide-sparse-base

LittleKedi-tiny-wide-sparse-base is a small DeepSeek-V3-style Mixture-of-Experts language model trained from scratch. It has 3.8M parameters in total, of which 1.6M are active per token (0.52M excluding the embedding table). It was pretrained on a 1,000M-token mix (FineWeb-Edu, Cosmopedia (Stanford), Cosmopedia (OpenStax), Cosmopedia (WikiHow), Cosmopedia (Khan Academy), FineMath 4+, SmolTalk, github-code-clean).

Part of the LittleKedi family. Not affiliated with DeepSeek.

This is the base model, before any instruction tuning. It continues text; it doesn't follow instructions. It's an educational and research artifact: it writes fluent, on-topic English but knows few facts and confidently makes things up. Don't use it for anything that matters.

The LittleKedi family

LittleKedi is a family of small DeepSeek-V3-style Mixture-of-Experts language models trained from scratch. Every size comes in three variants that differ in how sparse they are: how many routed experts they have, and what share of their parameters a token actually passes through.

Variant Routed experts
standard (<size>) a few wide experts; the densest variant
sparse (<size>-sparse) many narrower, fine-grained experts, more of them active per token
wide-sparse (<size>-wide-sparse) more experts again; the largest total and smallest active share

More experts means more total capacity to store knowledge, but each expert sees fewer tokens (tokens × top-k ÷ experts per layer), so sparser variants need more training data before their experts are well trained. The variants of a size share a tokenizer and dataset, so they can be compared directly. Shapes change as the family scales; the table below is generated from the current presets.

This card is for tiny-wide-sparse, the wide-sparse tiny model: 1.6M of its 3.8M parameters (42%) are active per token. Its siblings are tiny and tiny-sparse.

Model Variant Vocab Total Active Active non-emb. Active % Routed experts
tiny standard 8K 2.0M 1.5M 0.5M 78% 8 × 64, top-2
tiny-sparse sparse 8K 2.6M 1.6M 0.5M 60% 32 × 32, top-4
tiny-wide-sparse ← this model wide-sparse 8K 3.8M 1.6M 0.5M 42% 64 × 32, top-4
mini standard 8K 9.4M 4.8M 3.2M 51% 12 × 128, top-3
mini-sparse sparse 8K 19.8M 4.8M 3.2M 24% 64 × 64, top-6
mini-wide-sparse wide-sparse 8K 36.4M 4.9M 3.3M 13% 128 × 64, top-6
small standard 8K 12.2M 6.3M 4.2M 52% 16 × 128, top-4
small-sparse sparse 8K 24.1M 6.4M 4.3M 27% 80 × 64, top-8
small-wide-sparse wide-sparse 8K 43.9M 6.5M 4.4M 15% 160 × 64, top-8
moderate standard 16K 49.8M 18.8M 12.5M 38% 24 × 192, top-4
moderate-sparse sparse 16K 105.8M 19.1M 12.8M 18% 120 × 96, top-8
moderate-wide-sparse wide-sparse 16K 199.0M 19.4M 13.1M 10% 240 × 96, top-8
medium standard 16K 85.9M 30.8M 22.4M 36% 24 × 256, top-4
medium-sparse sparse 16K 185.3M 31.2M 22.8M 17% 120 × 128, top-8
medium-wide-sparse wide-sparse 16K 350.9M 31.6M 23.2M 9% 240 × 128, top-8
base standard 32K 258.7M 105.4M 80.2M 41% 32 × 256, top-6
base-sparse sparse 32K 542.8M 106.3M 81.2M 20% 160 × 128, top-12
base-wide-sparse wide-sparse 32K 1015.9M 107.6M 82.4M 11% 320 × 128, top-12

"Active" counts everything a token passes through: embeddings, attention, the dense layer, shared experts and its top-k routed experts. "Active non-embedding" leaves out the embedding table (tied with the output head), which costs memory but almost no compute; it's the fair number for comparing compute between models.

Architecture

Component This model
Parameters 3.78M total, 1.57M active per token (42%), 0.52M active non-embedding
Layers / hidden size 4 / 128 (layer 0 has a dense SwiGLU FFN of 352; layers 1–3 are MoE)
Attention Multi-head Latent Attention (MLA): 4 heads, KV latent 32, decoupled RoPE dim 16, head dims 16+16 (q/k) and 16 (v). The KV cache stores only the 32-dim latent plus a 16-dim RoPE key per token
MoE DeepSeekMoE: 64 routed experts (SwiGLU, 32 hidden), top-4, plus 2 shared experts; sigmoid gating; group-limited routing (4 groups, top-2)
Load balancing Auxiliary-loss-free bias balancing (update speed 0.001) plus a small sequence-wise loss (α = 0.0001). Over the second half of training the median batch put 1.23× the mean load on the busiest expert (1.28× in the worst layer), and 0.00% of expert assignments were dropped for capacity.
Expert dispatch Training: fixed capacity (capacity factor 2.0; overflow assignments are dropped). Inference is always dropless
Embeddings Input embedding and output head tied
Context 1,024 tokens

Training

Data 1,000M tokens, 801,792 documents (~264 tokens per parameter, ~637 per active parameter), one pass, sources interleaved throughout (mix below)
Tokenizer Byte-level BPE, 8,192 tokens (~3.75 bytes/token on this mix); digits not split (v1 tokenizer); no chat tokens, so chat data is plain User: / Assistant: text
Steps 15,258 × 65,536 tokens (batch 16 × 1024 × 4 accumulation)
Optimizer Muon (momentum 0.95, Nesterov, 5 Newton-Schulz steps; each expert's matrix orthogonalized separately) for hidden matrices, AdamW (β = 0.9, 0.95) for embeddings, norms and the router; weight decay 0.1, gradient clipping 1.0
Learning rate 2e-3 peak; WSD: 200 warm-up steps, constant, then linear decay to 0 over the last 20% of steps (schedule set for 15,258 steps)
Precision bf16 autocast, torch.compile
Compute ~3.1e15 FLOPs (6 × active non-embedding parameters × tokens)
Hardware One AMD Radeon 890M integrated GPU (Ryzen AI 300 laptop), ROCm on Windows: about 10.4 hours at a median ~26.7K tokens/s
Source What it is Tokens Share
FineWeb-Edu sample/10BT educational web pages 700M 70%
Cosmopedia (Stanford) synthetic textbooks 70M 7%
Cosmopedia (OpenStax) synthetic textbooks 30M 3%
Cosmopedia (WikiHow) synthetic how-to articles 30M 3%
Cosmopedia (Khan Academy) synthetic lessons 20M 2%
FineMath 4+ maths web pages 80M 8%
SmolTalk chat conversations 50M 5%
github-code-clean source code 20M 2%

Evaluation

Metric Value
Validation loss (nats/token, held-out mix) 3.736
Perplexity 41.9
Bits per byte 1.389
Experts never chosen on the validation sample 0

Validation loss over training: 5.59 (16M tok) → 4.07 (164M tok) → 3.98 (295M tok) → 3.93 (442M tok) → 3.91 (590M tok) → 3.89 (737M tok) → 3.84 (868M tok) → 3.74 (1,000M tok)

Zero-shot benchmarks (evals.py)

Task acc_norm acc Items Chance
SciQ 53.9% ± 1.6% 60.4% 1,000 25%
ARC-Easy 32.4% ± 1.0% 32.4% 2,376 25%
HellaSwag 26.3% ± 0.4% 26.3% 10,042 25%

Each answer choice is scored by the log-probability of its text after the prompt, using lm-evaluation-harness prompts (SciQ includes its supporting passage). acc_norm divides by the answer's length in characters; acc doesn't. Splits: SciQ test, ARC-Easy test, HellaSwag validation.

MMLU (mmlu_eval.py)

Format Accuracy Questions scored Average few-shot examples
letter 25.2% 14,037 (skipped 5) 4.5
cloze 25.7% 14,037 (skipped 5) 0.0
Category Subjects Questions Letter Cloze
STEM 19 3,153 27.3% 24.2%
Humanities 13 4,705 23.5% 26.0%
Social Sciences 12 3,077 24.0% 26.2%
Other 13 3,102 26.6% 26.2%

Biology & medicine (cloze): 27.8% on 2,089 questions, +2.8 points against 25% chance (about 2.9 standard errors, significant). This pools the 10 life-science and health subjects; it isn't an official MMLU category.

In the letter format the model compares " A"–" D", with same-subject examples added while they fit the context. In the cloze format each answer's text is scored by log-probability per byte; at this size the cloze results are the meaningful ones. Chance is 25%.

Compared with its siblings and other small models

This model tiny Pythia-70M GPT-2 Supra-50M
Parameters (total) 3.8M 2.0M 70.4M 124.4M 51.8M
Active non-embedding parameters 0.5M 0.5M 18.9M 85.1M 35.4M
Training tokens 1,000M 1,000M 300B (the Pile) undisclosed (WebText) 20B (FineWeb-Edu)
SciQ (acc_norm) 53.9% 56.3% 55.2% 53.2% 77.2%
ARC-Easy (acc_norm) 32.4% 32.7% 35.0% 42.0% 52.2%
HellaSwag (acc_norm) 26.3% 26.7% – 29.5% 31.8%
MMLU (cloze) 25.7% 25.8% – – –
MMLU biology & medicine (cloze) 27.8% 28.4% – – –
Validation loss (held-out mix) 3.736 3.869 – – –

LittleKedi columns are measured with this repo's evals.py and mmlu_eval.py on the checkpoints named. Validation losses are only comparable between models that share a tokenizer and validation set. The other columns are published figures: Pythia-70M from EleutherAI's evaluations as quoted on the Wisp-15M card, GPT-2 and Supra-50M from the Supra-50M-Base card. Their prompts may differ from ours, so treat gaps of a few points as noise.

What it can and can't do

This is the sparsest tiny model: the same 0.5M active non-embedding parameters per token as tiny, but with 64 narrower experts and ~3× tiny's expert capacity. It models text measurably better than tiny (validation loss 3.736 against 3.869 on the same data, matching tiny's loss with roughly a third of the tokens), but in use it's a model of the same kind: it has learned the form of English much better than its content. It's a research artifact for studying sparsity in small MoE models, not a source of information.

What it does

  • ✅ Writes grammatical, fluent sentences for a paragraph or two, in the register of educational web text and textbooks.
  • ✅ Stays roughly on topic, and sometimes gets the broad frame right where tiny didn't: "The French Revolution began in" → "…the late 18th century" (with sampling; tiny said 1906). Treat this as an occasional improvement, not a reliable one: the same prompt decoded greedily gives "the late 19th century".
  • ✅ Picks up document structure from its data mix: markdown headings and # math: blocks after maths prompts, bullet lists, dialogue turns after User: / Assistant: text.
  • ✅ Benchmark signal above chance, on a par with tiny (SciQ 53.9%, ARC-Easy 32.4%, MMLU biology & medicine 28% in the cloze format). Its lower loss hasn't turned into higher multiple-choice accuracy at this size.

What it can't do

  • ❌ Facts. Names and dates are plausible in kind but usually wrong: "Albert Einstein was" → "a professor of medicine at the University of California"; "DNA is made of" → "a coin, a piece of paper". Treat every name, date and number it writes as made up.
  • ❌ Arithmetic. "2 + 2 =" → "6", "7 + 6 =" → "6".
  • ❌ Code. Given a Python function signature it produces docstring fragments and whitespace.
  • ❌ Instructions or questions. It's a base model: it continues text rather than answering it.
  • ❌ Long-range coherence. Topics drift after a sentence or two ("Photosynthesis is the process by which a person's body is able to move in a controlled way").

Decoding

Greedy decoding falls into loops even faster than tiny's ("the heart, the heart, the heart…"). Use sampling with temperature ≈ 0.7, top-p ≈ 0.9 and a repetition penalty of about 1.15.

Usage

PyTorch (needs the littlekedi package from the training code):

import json, torch
from safetensors.torch import load_file
from littlekedi import LittleKediModel, ModelConfig
from littlekedi.tokenizer import BPETokenizer

model = LittleKediModel(ModelConfig.from_dict(json.load(open("config.json")))).eval()
model.load_state_dict(load_file("model.safetensors"))
tok = BPETokenizer.from_file("tokenizer.json")
ids = torch.tensor([[tok.eot_id] + tok.encode("Photosynthesis is")])
print(tok.decode(model.generate(ids, 60, temperature=0.7, top_p=0.9, repetition_penalty=1.15)[0].tolist()))

Or interactively: python sample.py --ckpt <checkpoint>.pt. Base models are released as PyTorch weights only; GGUF builds come with the instruct models.

Files

File Contents
model.safetensors PyTorch weights (embedding tied with the output head), littlekedi parameter names
config.json ModelConfig for littlekedi
tokenizer.json Byte-level BPE tokenizer (Hugging Face tokenizers format)
little_kedi.jpg The LittleKedi logo

Acknowledgements

  • The architecture follows DeepSeek-V3 (DeepSeek-AI, 2024): MLA, DeepSeekMoE and auxiliary-loss-free balancing.
  • Muon (Keller Jordan et al., 2024) with Moonlight's update scaling (Liu et al., 2025).
  • Pretraining data: FineWeb-Edu (ODC-BY 1.0).
  • Pretraining data: Cosmopedia (Apache 2.0).
  • Pretraining data: FineMath 4+ (ODC-BY 1.0).
  • Pretraining data: SmolTalk (Apache 2.0 for its new subsets, with the datasets it includes under their own licenses).
  • Pretraining data: github-code-clean (Apache 2.0, with each file under its own open-source license).

Card generated 2026-10-10 by prepare_release.py from tiny-wide-sparse.pt (step 15257).

Downloads last month
21
Safetensors
Model size
3.78M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train edededdy/LittleKedi-tiny-wide-sparse-base

Collection including edededdy/LittleKedi-tiny-wide-sparse-base