Download README.md from edededdy/LittleKedi-tiny-base: direct link, hf CLI and curl.
- Browser
- Download file 14.7 kB
-
https://huggingface.co/edededdy/LittleKedi-tiny-base/resolve/main/README.md
- Command line
-
hf download hf://edededdy/LittleKedi-tiny-base/README.md
-
curl -L -o README.md https://huggingface.co/edededdy/LittleKedi-tiny-base/resolve/main/README.md
license: apache-2.0
language:
- en
pipeline_tag: text-generation
datasets:
- HuggingFaceFW/fineweb-edu
- HuggingFaceTB/cosmopedia
- HuggingFaceTB/finemath
- HuggingFaceTB/smoltalk
- codeparrot/github-code-clean
tags:
- mixture-of-experts
- moe
- deepseek
- multi-head-latent-attention
- mla
- from-scratch
- tiny
- littlekedi
- base

LittleKedi-tiny-base
LittleKedi-tiny-base is a small DeepSeek-V3-style Mixture-of-Experts language model trained from scratch. It has
2.0M parameters in total, of which 1.5M are active per token
(0.50M excluding the embedding table). It was pretrained on a 1,000M-token mix (FineWeb-Edu, Cosmopedia (Stanford), Cosmopedia (OpenStax), Cosmopedia (WikiHow), Cosmopedia (Khan Academy), FineMath 4+, SmolTalk, github-code-clean).
Part of the LittleKedi family. Not affiliated with DeepSeek.
This is the base model, before any instruction tuning. It continues text; it doesn't follow instructions. It's an educational and research artifact: it writes fluent, on-topic English but knows few facts and confidently makes things up. Don't use it for anything that matters.
The LittleKedi family
LittleKedi is a family of small DeepSeek-V3-style Mixture-of-Experts language models trained from scratch. Every size comes in three variants that differ in how sparse they are: how many routed experts they have, and what share of their parameters a token actually passes through.
| Variant | Routed experts |
|---|---|
standard (<size>) |
a few wide experts; the densest variant |
sparse (<size>-sparse) |
many narrower, fine-grained experts, more of them active per token |
wide-sparse (<size>-wide-sparse) |
more experts again; the largest total and smallest active share |
More experts means more total capacity to store knowledge, but each expert sees fewer tokens (tokens Γ top-k Γ· experts per layer), so sparser variants need more training data before their experts are well trained. The variants of a size share a tokenizer and dataset, so they can be compared directly. Shapes change as the family scales; the table below is generated from the current presets.
This card is for tiny, the standard tiny model: 1.5M of its 2.0M parameters (78%) are active per token. Its siblings are tiny-sparse and tiny-wide-sparse.
| Model | Variant | Vocab | Total | Active | Active non-emb. | Active % | Routed experts |
|---|---|---|---|---|---|---|---|
tiny β this model |
standard | 8K | 2.0M | 1.5M | 0.5M | 78% | 8 Γ 64, top-2 |
tiny-sparse |
sparse | 8K | 2.6M | 1.6M | 0.5M | 60% | 32 Γ 32, top-4 |
tiny-wide-sparse |
wide-sparse | 8K | 3.8M | 1.6M | 0.5M | 42% | 64 Γ 32, top-4 |
mini |
standard | 8K | 9.4M | 4.8M | 3.2M | 51% | 12 Γ 128, top-3 |
mini-sparse |
sparse | 8K | 19.8M | 4.8M | 3.2M | 24% | 64 Γ 64, top-6 |
mini-wide-sparse |
wide-sparse | 8K | 36.4M | 4.9M | 3.3M | 13% | 128 Γ 64, top-6 |
small |
standard | 8K | 12.2M | 6.3M | 4.2M | 52% | 16 Γ 128, top-4 |
small-sparse |
sparse | 8K | 24.1M | 6.4M | 4.3M | 27% | 80 Γ 64, top-8 |
small-wide-sparse |
wide-sparse | 8K | 43.9M | 6.5M | 4.4M | 15% | 160 Γ 64, top-8 |
moderate |
standard | 16K | 49.8M | 18.8M | 12.5M | 38% | 24 Γ 192, top-4 |
moderate-sparse |
sparse | 16K | 105.8M | 19.1M | 12.8M | 18% | 120 Γ 96, top-8 |
moderate-wide-sparse |
wide-sparse | 16K | 199.0M | 19.4M | 13.1M | 10% | 240 Γ 96, top-8 |
medium |
standard | 16K | 85.9M | 30.8M | 22.4M | 36% | 24 Γ 256, top-4 |
medium-sparse |
sparse | 16K | 185.3M | 31.2M | 22.8M | 17% | 120 Γ 128, top-8 |
medium-wide-sparse |
wide-sparse | 16K | 350.9M | 31.6M | 23.2M | 9% | 240 Γ 128, top-8 |
base |
standard | 32K | 258.7M | 105.4M | 80.2M | 41% | 32 Γ 256, top-6 |
base-sparse |
sparse | 32K | 542.8M | 106.3M | 81.2M | 20% | 160 Γ 128, top-12 |
base-wide-sparse |
wide-sparse | 32K | 1015.9M | 107.6M | 82.4M | 11% | 320 Γ 128, top-12 |
"Active" counts everything a token passes through: embeddings, attention, the dense layer, shared experts and its top-k routed experts. "Active non-embedding" leaves out the embedding table (tied with the output head), which costs memory but almost no compute; it's the fair number for comparing compute between models.
Architecture
| Component | This model |
|---|---|
| Parameters | 1.99M total, 1.55M active per token (78%), 0.50M active non-embedding |
| Layers / hidden size | 4 / 128 (layer 0 has a dense SwiGLU FFN of 352; layers 1β3 are MoE) |
| Attention | Multi-head Latent Attention (MLA): 4 heads, KV latent 32, decoupled RoPE dim 16, head dims 16+16 (q/k) and 16 (v). The KV cache stores only the 32-dim latent plus a 16-dim RoPE key per token |
| MoE | DeepSeekMoE: 8 routed experts (SwiGLU, 64 hidden), top-2, plus 1 shared expert; sigmoid gating |
| Load balancing | Auxiliary-loss-free bias balancing (update speed 0.001) plus a small sequence-wise loss (Ξ± = 0.0001). Over the second half of training the median batch put 1.07Γ the mean load on the busiest expert (1.08Γ in the worst layer), and 0.00% of expert assignments were dropped for capacity. |
| Expert dispatch | Training: fixed capacity (capacity factor 2.0; overflow assignments are dropped). Inference is always dropless |
| Embeddings | Input embedding and output head tied |
| Context | 1,024 tokens |
Training
| Data | 1,000M tokens, 801,792 documents (~502 tokens per parameter, ~646 per active parameter), one pass, sources interleaved throughout (mix below) |
| Tokenizer | Byte-level BPE, 8,192 tokens (~3.75 bytes/token on this mix); digits not split (v1 tokenizer); no chat tokens, so chat data is plain User: / Assistant: text |
| Steps | 15,258 Γ 65,536 tokens (batch 16 Γ 1024 Γ 4 accumulation) |
| Optimizer | Muon (momentum 0.95, Nesterov, 5 Newton-Schulz steps; each expert's matrix orthogonalized separately) for hidden matrices, AdamW (Ξ² = 0.9, 0.95) for embeddings, norms and the router; weight decay 0.1, gradient clipping 1.0 |
| Learning rate | 2e-3 peak; WSD: 200 warm-up steps, constant, then linear decay to 0 over the last 20% of steps (schedule set for 15,258 steps) |
| Precision | bf16 autocast, torch.compile |
| Compute | ~3e15 FLOPs (6 Γ active non-embedding parameters Γ tokens) |
| Hardware | One AMD Radeon 890M integrated GPU (Ryzen AI 300 laptop), ROCm on Windows: about 8.9 hours at a median ~31.1K tokens/s |
| Source | What it is | Tokens | Share |
|---|---|---|---|
FineWeb-Edu sample/10BT |
educational web pages | 700M | 70% |
| Cosmopedia (Stanford) | synthetic textbooks | 70M | 7% |
| Cosmopedia (OpenStax) | synthetic textbooks | 30M | 3% |
| Cosmopedia (WikiHow) | synthetic how-to articles | 30M | 3% |
| Cosmopedia (Khan Academy) | synthetic lessons | 20M | 2% |
| FineMath 4+ | maths web pages | 80M | 8% |
| SmolTalk | chat conversations | 50M | 5% |
| github-code-clean | source code | 20M | 2% |
Evaluation
| Metric | Value |
|---|---|
| Validation loss (nats/token, held-out mix) | 3.869 |
| Perplexity | 47.9 |
| Bits per byte | 1.438 |
| Experts never chosen on the validation sample | 0 |
Validation loss over training: 5.60 (16M tok) β 4.17 (164M tok) β 4.08 (295M tok) β 4.04 (442M tok) β 4.02 (590M tok) β 4.01 (737M tok) β 3.96 (868M tok) β 3.87 (1,000M tok)
Zero-shot benchmarks (evals.py)
| Task | acc_norm | acc | Items | Chance |
|---|---|---|---|---|
| SciQ | 56.3% Β± 1.6% | 59.8% | 1,000 | 25% |
| ARC-Easy | 32.7% Β± 1.0% | 32.3% | 2,376 | 25% |
| HellaSwag | 26.7% Β± 0.4% | 26.4% | 10,042 | 25% |
Each answer choice is scored by the log-probability of its text after the prompt, using lm-evaluation-harness prompts (SciQ includes its supporting passage). acc_norm divides by the answer's length in characters; acc doesn't. Splits: SciQ test, ARC-Easy test, HellaSwag validation.
MMLU (mmlu_eval.py)
| Format | Accuracy | Questions scored | Average few-shot examples |
|---|---|---|---|
| letter | 25.4% | 14,037 (skipped 5) | 4.5 |
| cloze | 25.8% | 14,037 (skipped 5) | 0.0 |
| Category | Subjects | Questions | Letter | Cloze |
|---|---|---|---|---|
| STEM | 19 | 3,153 | 27.6% | 23.5% |
| Humanities | 13 | 4,705 | 24.4% | 25.5% |
| Social Sciences | 12 | 3,077 | 23.6% | 26.6% |
| Other | 13 | 3,102 | 26.4% | 27.9% |
Biology & medicine (cloze): 28.4% on 2,089 questions, +3.4 points against 25% chance (about 3.6 standard errors, significant). This pools the 10 life-science and health subjects; it isn't an official MMLU category.
In the letter format the model compares " A"β" D", with same-subject examples added while they fit the context. In the cloze format each answer's text is scored by log-probability per byte; at this size the cloze results are the meaningful ones. Chance is 25%.
Compared with its siblings and other small models
| This model | Pythia-70M | GPT-2 | Supra-50M | |
|---|---|---|---|---|
| Parameters (total) | 2.0M | 70.4M | 124.4M | 51.8M |
| Active non-embedding parameters | 0.5M | 18.9M | 85.1M | 35.4M |
| Training tokens | 1,000M | 300B (the Pile) | undisclosed (WebText) | 20B (FineWeb-Edu) |
| SciQ (acc_norm) | 56.3% | 55.2% | 53.2% | 77.2% |
| ARC-Easy (acc_norm) | 32.7% | 35.0% | 42.0% | 52.2% |
| HellaSwag (acc_norm) | 26.7% | β | 29.5% | 31.8% |
| MMLU (cloze) | 25.8% | β | β | β |
| MMLU biology & medicine (cloze) | 28.4% | β | β | β |
| Validation loss (held-out mix) | 3.869 | β | β | β |
LittleKedi columns are measured with this repo's evals.py and mmlu_eval.py on the checkpoints named. Validation
losses are only comparable between models that share a tokenizer and validation set. The other columns are
published figures: Pythia-70M from EleutherAI's evaluations as quoted on the
Wisp-15M card, GPT-2 and Supra-50M from the
Supra-50M-Base card. Their prompts may differ from ours,
so treat gaps of a few points as noise.
What it can and can't do
At 0.5M active non-embedding parameters this model has learned the form of English much better than its content. It's a research artifact for studying small MoE models, not a source of information.
What it does
- β Writes grammatical, fluent sentences for a paragraph or two, in the register of its training data (educational web text and textbooks): "Photosynthesis is the process by which" β "β¦it is used to monitor and measure the concentration of nutrients."
- β Stays roughly on topic: a prompt about the heart leads to blood vessels, the body and disease; a story opening leads to characters and a plot.
- β
Picks up document structure from its data mix: markdown headings and worked "Example 1:" sections
after maths prompts, dialogue turns after
User:/Assistant:text. - β Shows some benchmark signal above chance (SciQ 56.3%, ARC-Easy 32.7%, MMLU biology & medicine 28.4% in the cloze format), mostly from matching answers to a passage or recognising which option sounds right, not from recalling facts.
What it can't do
- β Facts. It reliably names the right kind of thing but the wrong thing: "The French Revolution began in" β "1906β¦"; "Albert Einstein was" β "a co-director of the French Marxi Palm." Treat every name, date and number it writes as made up.
- β Arithmetic. "2 + 2 =" β "3", "7 + 6 =" β "6".
- β Code. Code was 2% of its training data; given a Python function signature it produces empty docstrings and whitespace.
- β Instructions or questions. It's a base model: it continues text rather than answering it, and with only 0.5M non-embedding parameters an instruction-tuned version won't be a reliable assistant either.
- β Long-range coherence. Topics drift after a few sentences, and invented names appear ("the Grovi Palace").
Decoding
Greedy or low-temperature decoding falls into loops almost immediately ("The process of the plant is formed by the plant. The process of the plant is formed by the plantβ¦"). Use sampling with temperature β 0.7, top-p β 0.9 and a repetition penalty of about 1.15, which is what produced the examples above.
Usage
PyTorch (needs the littlekedi package from the training code):
import json, torch
from safetensors.torch import load_file
from littlekedi import LittleKediModel, ModelConfig
from littlekedi.tokenizer import BPETokenizer
model = LittleKediModel(ModelConfig.from_dict(json.load(open("config.json")))).eval()
model.load_state_dict(load_file("model.safetensors"))
tok = BPETokenizer.from_file("tokenizer.json")
ids = torch.tensor([[tok.eot_id] + tok.encode("Photosynthesis is")])
print(tok.decode(model.generate(ids, 60, temperature=0.7, top_p=0.9, repetition_penalty=1.15)[0].tolist()))
Or interactively: python sample.py --ckpt <checkpoint>.pt. Base models are released as PyTorch weights only;
GGUF builds come with the instruct models.
Files
| File | Contents |
|---|---|
model.safetensors |
PyTorch weights (embedding tied with the output head), littlekedi parameter names |
config.json |
ModelConfig for littlekedi |
tokenizer.json |
Byte-level BPE tokenizer (Hugging Face tokenizers format) |
little_kedi.jpg |
The LittleKedi logo |
Acknowledgements
- The architecture follows DeepSeek-V3 (DeepSeek-AI, 2024): MLA, DeepSeekMoE and auxiliary-loss-free balancing.
- Muon (Keller Jordan et al., 2024) with Moonlight's update scaling (Liu et al., 2025).
- Pretraining data: FineWeb-Edu (ODC-BY 1.0).
- Pretraining data: Cosmopedia (Apache 2.0).
- Pretraining data: FineMath 4+ (ODC-BY 1.0).
- Pretraining data: SmolTalk (Apache 2.0 for its new subsets, with the datasets it includes under their own licenses).
- Pretraining data: github-code-clean (Apache 2.0, with each file under its own open-source license).
Card generated 2026-10-09 by prepare_release.py from tiny.pt (step 15257).