--- license: apache-2.0 language: - en pipeline_tag: text-generation datasets: - HuggingFaceFW/fineweb-edu - HuggingFaceTB/cosmopedia - HuggingFaceTB/finemath - HuggingFaceTB/smoltalk - codeparrot/github-code-clean tags: - mixture-of-experts - moe - deepseek - multi-head-latent-attention - mla - from-scratch - tiny - littlekedi - base ---

LittleKedi logo

# LittleKedi-tiny-base `LittleKedi-tiny-base` is a **small DeepSeek-V3-style Mixture-of-Experts language model trained from scratch**. It has **2.0M parameters** in total, of which **1.5M are active per token** (0.50M excluding the embedding table). It was pretrained on a 1,000M-token mix (FineWeb-Edu, Cosmopedia (Stanford), Cosmopedia (OpenStax), Cosmopedia (WikiHow), Cosmopedia (Khan Academy), FineMath 4+, SmolTalk, github-code-clean). *Part of the LittleKedi family. Not affiliated with DeepSeek.* This is the **base model**, before any instruction tuning. It continues text; it doesn't follow instructions. It's an educational and research artifact: it writes fluent, on-topic English but **knows few facts and confidently makes things up**. Don't use it for anything that matters. ## The LittleKedi family LittleKedi is a family of small DeepSeek-V3-style Mixture-of-Experts language models trained from scratch. Every size comes in three variants that differ in **how sparse** they are: how many routed experts they have, and what share of their parameters a token actually passes through. | Variant | Routed experts | |---|---| | **standard** (``) | a few wide experts; the densest variant | | **sparse** (`-sparse`) | many narrower, fine-grained experts, more of them active per token | | **wide-sparse** (`-wide-sparse`) | more experts again; the largest total and smallest active share | More experts means more total capacity to store knowledge, but each expert sees fewer tokens (tokens × top-k ÷ experts per layer), so sparser variants need more training data before their experts are well trained. The variants of a size share a tokenizer and dataset, so they can be compared directly. Shapes change as the family scales; the table below is generated from the current presets. This card is for **`tiny`**, the **standard** tiny model: 1.5M of its 2.0M parameters (78%) are active per token. Its siblings are `tiny-sparse` and `tiny-wide-sparse`. | Model | Variant | Vocab | Total | Active | Active non-emb. | Active % | Routed experts | |---|---|---|---|---|---|---|---| | **`tiny` ← this model** | **standard** | **8K** | **2.0M** | **1.5M** | **0.5M** | **78%** | **8 × 64, top-2** | | `tiny-sparse` | sparse | 8K | 2.6M | 1.6M | 0.5M | 60% | 32 × 32, top-4 | | `tiny-wide-sparse` | wide-sparse | 8K | 3.8M | 1.6M | 0.5M | 42% | 64 × 32, top-4 | | `mini` | standard | 8K | 9.4M | 4.8M | 3.2M | 51% | 12 × 128, top-3 | | `mini-sparse` | sparse | 8K | 19.8M | 4.8M | 3.2M | 24% | 64 × 64, top-6 | | `mini-wide-sparse` | wide-sparse | 8K | 36.4M | 4.9M | 3.3M | 13% | 128 × 64, top-6 | | `small` | standard | 8K | 12.2M | 6.3M | 4.2M | 52% | 16 × 128, top-4 | | `small-sparse` | sparse | 8K | 24.1M | 6.4M | 4.3M | 27% | 80 × 64, top-8 | | `small-wide-sparse` | wide-sparse | 8K | 43.9M | 6.5M | 4.4M | 15% | 160 × 64, top-8 | | `moderate` | standard | 16K | 49.8M | 18.8M | 12.5M | 38% | 24 × 192, top-4 | | `moderate-sparse` | sparse | 16K | 105.8M | 19.1M | 12.8M | 18% | 120 × 96, top-8 | | `moderate-wide-sparse` | wide-sparse | 16K | 199.0M | 19.4M | 13.1M | 10% | 240 × 96, top-8 | | `medium` | standard | 16K | 85.9M | 30.8M | 22.4M | 36% | 24 × 256, top-4 | | `medium-sparse` | sparse | 16K | 185.3M | 31.2M | 22.8M | 17% | 120 × 128, top-8 | | `medium-wide-sparse` | wide-sparse | 16K | 350.9M | 31.6M | 23.2M | 9% | 240 × 128, top-8 | | `base` | standard | 32K | 258.7M | 105.4M | 80.2M | 41% | 32 × 256, top-6 | | `base-sparse` | sparse | 32K | 542.8M | 106.3M | 81.2M | 20% | 160 × 128, top-12 | | `base-wide-sparse` | wide-sparse | 32K | 1015.9M | 107.6M | 82.4M | 11% | 320 × 128, top-12 | "Active" counts everything a token passes through: embeddings, attention, the dense layer, shared experts and its top-k routed experts. "Active non-embedding" leaves out the embedding table (tied with the output head), which costs memory but almost no compute; it's the fair number for comparing compute between models. ## Architecture | Component | This model | |---|---| | Parameters | 1.99M total, 1.55M active per token (78%), 0.50M active non-embedding | | Layers / hidden size | 4 / 128 (layer 0 has a dense SwiGLU FFN of 352; layers 1–3 are MoE) | | Attention | **Multi-head Latent Attention (MLA)**: 4 heads, KV latent 32, decoupled RoPE dim 16, head dims 16+16 (q/k) and 16 (v). The KV cache stores only the 32-dim latent plus a 16-dim RoPE key per token | | MoE | **DeepSeekMoE**: 8 routed experts (SwiGLU, 64 hidden), top-2, plus 1 shared expert; sigmoid gating | | Load balancing | **Auxiliary-loss-free** bias balancing (update speed 0.001) plus a small sequence-wise loss (α = 0.0001). Over the second half of training the median batch put 1.07× the mean load on the busiest expert (1.08× in the worst layer), and 0.00% of expert assignments were dropped for capacity. | | Expert dispatch | Training: fixed capacity (capacity factor 2.0; overflow assignments are dropped). Inference is always dropless | | Embeddings | Input embedding and output head **tied** | | Context | 1,024 tokens | ## Training | | | |---|---| | Data | 1,000M tokens, 801,792 documents (~502 tokens per parameter, ~646 per active parameter), one pass, sources interleaved throughout (mix below) | | Tokenizer | Byte-level BPE, 8,192 tokens (~3.75 bytes/token on this mix); digits not split (v1 tokenizer); no chat tokens, so chat data is plain `User:` / `Assistant:` text | | Steps | 15,258 × 65,536 tokens (batch 16 × 1024 × 4 accumulation) | | Optimizer | **Muon** (momentum 0.95, Nesterov, 5 Newton-Schulz steps; each expert's matrix orthogonalized separately) for hidden matrices, AdamW (β = 0.9, 0.95) for embeddings, norms and the router; weight decay 0.1, gradient clipping 1.0 | | Learning rate | 2e-3 peak; **WSD**: 200 warm-up steps, constant, then linear decay to 0 over the last 20% of steps (schedule set for 15,258 steps) | | Precision | bf16 autocast, `torch.compile` | | Compute | ~3e15 FLOPs (6 × active non-embedding parameters × tokens) | | Hardware | One AMD Radeon 890M integrated GPU (Ryzen AI 300 laptop), ROCm on Windows: about 8.9 hours at a median ~31.1K tokens/s | | Source | What it is | Tokens | Share | |---|---|---|---| | FineWeb-Edu `sample/10BT` | educational web pages | 700M | 70% | | Cosmopedia (Stanford) | synthetic textbooks | 70M | 7% | | Cosmopedia (OpenStax) | synthetic textbooks | 30M | 3% | | Cosmopedia (WikiHow) | synthetic how-to articles | 30M | 3% | | Cosmopedia (Khan Academy) | synthetic lessons | 20M | 2% | | FineMath 4+ | maths web pages | 80M | 8% | | SmolTalk | chat conversations | 50M | 5% | | github-code-clean | source code | 20M | 2% | ## Evaluation | Metric | Value | |---|---| | Validation loss (nats/token, held-out mix) | **3.869** | | Perplexity | 47.9 | | Bits per byte | 1.438 | | Experts never chosen on the validation sample | 0 | Validation loss over training: 5.60 (16M tok) → 4.17 (164M tok) → 4.08 (295M tok) → 4.04 (442M tok) → 4.02 (590M tok) → 4.01 (737M tok) → 3.96 (868M tok) → 3.87 (1,000M tok) ### Zero-shot benchmarks (`evals.py`) | Task | acc_norm | acc | Items | Chance | |---|---|---|---|---| | SciQ | 56.3% ± 1.6% | 59.8% | 1,000 | 25% | | ARC-Easy | 32.7% ± 1.0% | 32.3% | 2,376 | 25% | | HellaSwag | 26.7% ± 0.4% | 26.4% | 10,042 | 25% | Each answer choice is scored by the log-probability of its text after the prompt, using lm-evaluation-harness prompts (SciQ includes its supporting passage). **acc_norm** divides by the answer's length in characters; **acc** doesn't. Splits: SciQ test, ARC-Easy test, HellaSwag validation. ### MMLU (`mmlu_eval.py`) | Format | Accuracy | Questions scored | Average few-shot examples | |---|---|---|---| | letter | 25.4% | 14,037 (skipped 5) | 4.5 | | cloze | 25.8% | 14,037 (skipped 5) | 0.0 | | Category | Subjects | Questions | Letter | Cloze | |---|---|---|---|---| | STEM | 19 | 3,153 | 27.6% | 23.5% | | Humanities | 13 | 4,705 | 24.4% | 25.5% | | Social Sciences | 12 | 3,077 | 23.6% | 26.6% | | Other | 13 | 3,102 | 26.4% | 27.9% | **Biology & medicine (cloze): 28.4% on 2,089 questions**, +3.4 points against 25% chance (about 3.6 standard errors, significant). This pools the 10 life-science and health subjects; it isn't an official MMLU category. In the **letter** format the model compares " A"–" D", with same-subject examples added while they fit the context. In the **cloze** format each answer's text is scored by log-probability per byte; at this size the cloze results are the meaningful ones. Chance is 25%. ### Compared with its siblings and other small models | | **This model** | [Pythia-70M](https://huggingface.co/EleutherAI/pythia-70m) | [GPT-2](https://huggingface.co/openai-community/gpt2) | [Supra-50M](https://huggingface.co/SupraLabs/Supra-50M-Base) | |---|---|---|---|---| | Parameters (total) | 2.0M | 70.4M | 124.4M | 51.8M | | Active non-embedding parameters | 0.5M | 18.9M | 85.1M | 35.4M | | Training tokens | 1,000M | 300B (the Pile) | undisclosed (WebText) | 20B (FineWeb-Edu) | | SciQ (acc_norm) | 56.3% | 55.2% | 53.2% | 77.2% | | ARC-Easy (acc_norm) | 32.7% | 35.0% | 42.0% | 52.2% | | HellaSwag (acc_norm) | 26.7% | – | 29.5% | 31.8% | | MMLU (cloze) | 25.8% | – | – | – | | MMLU biology & medicine (cloze) | 28.4% | – | – | – | | Validation loss (held-out mix) | 3.869 | – | – | – | LittleKedi columns are measured with this repo's `evals.py` and `mmlu_eval.py` on the checkpoints named. Validation losses are only comparable between models that share a tokenizer and validation set. The other columns are **published figures**: Pythia-70M from EleutherAI's evaluations as quoted on the [Wisp-15M card](https://huggingface.co/DedeProGames/Wisp-15M), GPT-2 and Supra-50M from the [Supra-50M-Base card](https://huggingface.co/SupraLabs/Supra-50M-Base). Their prompts may differ from ours, so treat gaps of a few points as noise. ## What it can and can't do At 0.5M active non-embedding parameters this model has learned **the form of English much better than its content**. It's a research artifact for studying small MoE models, not a source of information. **What it does** - ✅ Writes grammatical, fluent sentences for a paragraph or two, in the register of its training data (educational web text and textbooks): *"Photosynthesis is the process by which"* → *"…it is used to monitor and measure the concentration of nutrients."* - ✅ Stays roughly on topic: a prompt about the heart leads to blood vessels, the body and disease; a story opening leads to characters and a plot. - ✅ Picks up document structure from its data mix: markdown headings and worked "Example 1:" sections after maths prompts, dialogue turns after `User:` / `Assistant:` text. - ✅ Shows some benchmark signal above chance (SciQ 56.3%, ARC-Easy 32.7%, MMLU biology & medicine 28.4% in the cloze format), mostly from matching answers to a passage or recognising which option sounds right, not from recalling facts. **What it can't do** - ❌ **Facts.** It reliably names the right kind of thing but the wrong thing: *"The French Revolution began in"* → *"1906…"*; *"Albert Einstein was"* → *"a co-director of the French Marxi Palm."* Treat every name, date and number it writes as made up. - ❌ **Arithmetic.** *"2 + 2 ="* → *"3"*, *"7 + 6 ="* → *"6"*. - ❌ **Code.** Code was 2% of its training data; given a Python function signature it produces empty docstrings and whitespace. - ❌ **Instructions or questions.** It's a base model: it continues text rather than answering it, and with only 0.5M non-embedding parameters an instruction-tuned version won't be a reliable assistant either. - ❌ **Long-range coherence.** Topics drift after a few sentences, and invented names appear (*"the Grovi Palace"*). **Decoding** Greedy or low-temperature decoding falls into loops almost immediately (*"The process of the plant is formed by the plant. The process of the plant is formed by the plant…"*). Use sampling with **temperature ≈ 0.7, top-p ≈ 0.9 and a repetition penalty of about 1.15**, which is what produced the examples above. ## Usage **PyTorch** (needs the `littlekedi` package from [the training code](https://github.com/EdmundMartin/LittleKedi)): ```python import json, torch from safetensors.torch import load_file from littlekedi import LittleKediModel, ModelConfig from littlekedi.tokenizer import BPETokenizer model = LittleKediModel(ModelConfig.from_dict(json.load(open("config.json")))).eval() model.load_state_dict(load_file("model.safetensors")) tok = BPETokenizer.from_file("tokenizer.json") ids = torch.tensor([[tok.eot_id] + tok.encode("Photosynthesis is")]) print(tok.decode(model.generate(ids, 60, temperature=0.7, top_p=0.9, repetition_penalty=1.15)[0].tolist())) ``` Or interactively: `python sample.py --ckpt .pt`. Base models are released as PyTorch weights only; GGUF builds come with the instruct models. ## Files | File | Contents | |---|---| | `model.safetensors` | PyTorch weights (embedding tied with the output head), `littlekedi` parameter names | | `config.json` | `ModelConfig` for `littlekedi` | | `tokenizer.json` | Byte-level BPE tokenizer (Hugging Face `tokenizers` format) | | `little_kedi.jpg` | The LittleKedi logo | ## Acknowledgements - The architecture follows **DeepSeek-V3** (DeepSeek-AI, 2024): MLA, DeepSeekMoE and auxiliary-loss-free balancing. - **Muon** (Keller Jordan et al., 2024) with Moonlight's update scaling (Liu et al., 2025). - Pretraining data: **FineWeb-Edu** (ODC-BY 1.0). - Pretraining data: **Cosmopedia** (Apache 2.0). - Pretraining data: **FineMath 4+** (ODC-BY 1.0). - Pretraining data: **SmolTalk** (Apache 2.0 for its new subsets, with the datasets it includes under their own licenses). - Pretraining data: **github-code-clean** (Apache 2.0, with each file under its own open-source license). *Card generated 2026-10-09 by `prepare_release.py` from `tiny.pt` (step 15257).*