|
Download README.md from edededdy/LittleKedi-tiny-base: direct link, hf CLI and curl.
- Browser
- Download file 14.7 kB
-
https://huggingface.co/edededdy/LittleKedi-tiny-base/resolve/main/README.md
- Command line
-
hf download hf://edededdy/LittleKedi-tiny-base/README.md
-
curl -L -o README.md https://huggingface.co/edededdy/LittleKedi-tiny-base/resolve/main/README.md
14.7 kB
| license: apache-2.0 | |
| language: | |
| - en | |
| pipeline_tag: text-generation | |
| datasets: | |
| - HuggingFaceFW/fineweb-edu | |
| - HuggingFaceTB/cosmopedia | |
| - HuggingFaceTB/finemath | |
| - HuggingFaceTB/smoltalk | |
| - codeparrot/github-code-clean | |
| tags: | |
| - mixture-of-experts | |
| - moe | |
| - deepseek | |
| - multi-head-latent-attention | |
| - mla | |
| - from-scratch | |
| - tiny | |
| - littlekedi | |
| - base | |
| <p align="center"><img src="little_kedi.jpg" alt="LittleKedi logo" width="640"></p> | |
| # LittleKedi-tiny-base | |
| `LittleKedi-tiny-base` is a **small DeepSeek-V3-style Mixture-of-Experts language model trained from scratch**. It has | |
| **2.0M parameters** in total, of which **1.5M are active per token** | |
| (0.50M excluding the embedding table). It was pretrained on a 1,000M-token mix (FineWeb-Edu, Cosmopedia (Stanford), Cosmopedia (OpenStax), Cosmopedia (WikiHow), Cosmopedia (Khan Academy), FineMath 4+, SmolTalk, github-code-clean). | |
| *Part of the LittleKedi family. Not affiliated with DeepSeek.* | |
| This is the **base model**, before any instruction tuning. It continues text; it doesn't follow instructions. It's an educational and research artifact: it writes fluent, on-topic English but **knows few facts and confidently makes things up**. Don't use it for anything that matters. | |
| ## The LittleKedi family | |
| LittleKedi is a family of small DeepSeek-V3-style Mixture-of-Experts language models trained from | |
| scratch. Every size comes in three variants that differ in **how sparse** they are: how many routed experts | |
| they have, and what share of their parameters a token actually passes through. | |
| | Variant | Routed experts | | |
| |---|---| | |
| | **standard** (`<size>`) | a few wide experts; the densest variant | | |
| | **sparse** (`<size>-sparse`) | many narrower, fine-grained experts, more of them active per token | | |
| | **wide-sparse** (`<size>-wide-sparse`) | more experts again; the largest total and smallest active share | | |
| More experts means more total capacity to store knowledge, but each expert | |
| sees fewer tokens (tokens Γ top-k Γ· experts per layer), so sparser variants need more training data before | |
| their experts are well trained. The variants of a size share a tokenizer and dataset, so they can be | |
| compared directly. Shapes change as the family scales; the table below is generated from the current | |
| presets. | |
| This card is for **`tiny`**, the **standard** tiny model: 1.5M of its 2.0M parameters (78%) are active per token. Its siblings are `tiny-sparse` and `tiny-wide-sparse`. | |
| | Model | Variant | Vocab | Total | Active | Active non-emb. | Active % | Routed experts | | |
| |---|---|---|---|---|---|---|---| | |
| | **`tiny` β this model** | **standard** | **8K** | **2.0M** | **1.5M** | **0.5M** | **78%** | **8 Γ 64, top-2** | | |
| | `tiny-sparse` | sparse | 8K | 2.6M | 1.6M | 0.5M | 60% | 32 Γ 32, top-4 | | |
| | `tiny-wide-sparse` | wide-sparse | 8K | 3.8M | 1.6M | 0.5M | 42% | 64 Γ 32, top-4 | | |
| | `mini` | standard | 8K | 9.4M | 4.8M | 3.2M | 51% | 12 Γ 128, top-3 | | |
| | `mini-sparse` | sparse | 8K | 19.8M | 4.8M | 3.2M | 24% | 64 Γ 64, top-6 | | |
| | `mini-wide-sparse` | wide-sparse | 8K | 36.4M | 4.9M | 3.3M | 13% | 128 Γ 64, top-6 | | |
| | `small` | standard | 8K | 12.2M | 6.3M | 4.2M | 52% | 16 Γ 128, top-4 | | |
| | `small-sparse` | sparse | 8K | 24.1M | 6.4M | 4.3M | 27% | 80 Γ 64, top-8 | | |
| | `small-wide-sparse` | wide-sparse | 8K | 43.9M | 6.5M | 4.4M | 15% | 160 Γ 64, top-8 | | |
| | `moderate` | standard | 16K | 49.8M | 18.8M | 12.5M | 38% | 24 Γ 192, top-4 | | |
| | `moderate-sparse` | sparse | 16K | 105.8M | 19.1M | 12.8M | 18% | 120 Γ 96, top-8 | | |
| | `moderate-wide-sparse` | wide-sparse | 16K | 199.0M | 19.4M | 13.1M | 10% | 240 Γ 96, top-8 | | |
| | `medium` | standard | 16K | 85.9M | 30.8M | 22.4M | 36% | 24 Γ 256, top-4 | | |
| | `medium-sparse` | sparse | 16K | 185.3M | 31.2M | 22.8M | 17% | 120 Γ 128, top-8 | | |
| | `medium-wide-sparse` | wide-sparse | 16K | 350.9M | 31.6M | 23.2M | 9% | 240 Γ 128, top-8 | | |
| | `base` | standard | 32K | 258.7M | 105.4M | 80.2M | 41% | 32 Γ 256, top-6 | | |
| | `base-sparse` | sparse | 32K | 542.8M | 106.3M | 81.2M | 20% | 160 Γ 128, top-12 | | |
| | `base-wide-sparse` | wide-sparse | 32K | 1015.9M | 107.6M | 82.4M | 11% | 320 Γ 128, top-12 | | |
| "Active" counts everything a token passes through: embeddings, attention, the dense layer, shared | |
| experts and its top-k routed experts. "Active non-embedding" leaves out the embedding table (tied with the | |
| output head), which costs memory but almost no compute; it's the fair number for comparing compute | |
| between models. | |
| ## Architecture | |
| | Component | This model | | |
| |---|---| | |
| | Parameters | 1.99M total, 1.55M active per token (78%), 0.50M active non-embedding | | |
| | Layers / hidden size | 4 / 128 (layer 0 has a dense SwiGLU FFN of 352; layers 1β3 are MoE) | | |
| | Attention | **Multi-head Latent Attention (MLA)**: 4 heads, KV latent 32, decoupled RoPE dim 16, head dims 16+16 (q/k) and 16 (v). The KV cache stores only the 32-dim latent plus a 16-dim RoPE key per token | | |
| | MoE | **DeepSeekMoE**: 8 routed experts (SwiGLU, 64 hidden), top-2, plus 1 shared expert; sigmoid gating | | |
| | Load balancing | **Auxiliary-loss-free** bias balancing (update speed 0.001) plus a small sequence-wise loss (Ξ± = 0.0001). Over the second half of training the median batch put 1.07Γ the mean load on the busiest expert (1.08Γ in the worst layer), and 0.00% of expert assignments were dropped for capacity. | | |
| | Expert dispatch | Training: fixed capacity (capacity factor 2.0; overflow assignments are dropped). Inference is always dropless | | |
| | Embeddings | Input embedding and output head **tied** | | |
| | Context | 1,024 tokens | | |
| ## Training | |
| | | | | |
| |---|---| | |
| | Data | 1,000M tokens, 801,792 documents (~502 tokens per parameter, ~646 per active parameter), one pass, sources interleaved throughout (mix below) | | |
| | Tokenizer | Byte-level BPE, 8,192 tokens (~3.75 bytes/token on this mix); digits not split (v1 tokenizer); no chat tokens, so chat data is plain `User:` / `Assistant:` text | | |
| | Steps | 15,258 Γ 65,536 tokens (batch 16 Γ 1024 Γ 4 accumulation) | | |
| | Optimizer | **Muon** (momentum 0.95, Nesterov, 5 Newton-Schulz steps; each expert's matrix orthogonalized separately) for hidden matrices, AdamW (Ξ² = 0.9, 0.95) for embeddings, norms and the router; weight decay 0.1, gradient clipping 1.0 | | |
| | Learning rate | 2e-3 peak; **WSD**: 200 warm-up steps, constant, then linear decay to 0 over the last 20% of steps (schedule set for 15,258 steps) | | |
| | Precision | bf16 autocast, `torch.compile` | | |
| | Compute | ~3e15 FLOPs (6 Γ active non-embedding parameters Γ tokens) | | |
| | Hardware | One AMD Radeon 890M integrated GPU (Ryzen AI 300 laptop), ROCm on Windows: about 8.9 hours at a median ~31.1K tokens/s | | |
| | Source | What it is | Tokens | Share | | |
| |---|---|---|---| | |
| | FineWeb-Edu `sample/10BT` | educational web pages | 700M | 70% | | |
| | Cosmopedia (Stanford) | synthetic textbooks | 70M | 7% | | |
| | Cosmopedia (OpenStax) | synthetic textbooks | 30M | 3% | | |
| | Cosmopedia (WikiHow) | synthetic how-to articles | 30M | 3% | | |
| | Cosmopedia (Khan Academy) | synthetic lessons | 20M | 2% | | |
| | FineMath 4+ | maths web pages | 80M | 8% | | |
| | SmolTalk | chat conversations | 50M | 5% | | |
| | github-code-clean | source code | 20M | 2% | | |
| ## Evaluation | |
| | Metric | Value | | |
| |---|---| | |
| | Validation loss (nats/token, held-out mix) | **3.869** | | |
| | Perplexity | 47.9 | | |
| | Bits per byte | 1.438 | | |
| | Experts never chosen on the validation sample | 0 | | |
| Validation loss over training: 5.60 (16M tok) β 4.17 (164M tok) β 4.08 (295M tok) β 4.04 (442M tok) β 4.02 (590M tok) β 4.01 (737M tok) β 3.96 (868M tok) β 3.87 (1,000M tok) | |
| ### Zero-shot benchmarks (`evals.py`) | |
| | Task | acc_norm | acc | Items | Chance | | |
| |---|---|---|---|---| | |
| | SciQ | 56.3% Β± 1.6% | 59.8% | 1,000 | 25% | | |
| | ARC-Easy | 32.7% Β± 1.0% | 32.3% | 2,376 | 25% | | |
| | HellaSwag | 26.7% Β± 0.4% | 26.4% | 10,042 | 25% | | |
| Each answer choice is scored by the log-probability of its text after the prompt, using lm-evaluation-harness | |
| prompts (SciQ includes its supporting passage). **acc_norm** divides by the answer's length in characters; | |
| **acc** doesn't. Splits: SciQ test, ARC-Easy test, HellaSwag validation. | |
| ### MMLU (`mmlu_eval.py`) | |
| | Format | Accuracy | Questions scored | Average few-shot examples | | |
| |---|---|---|---| | |
| | letter | 25.4% | 14,037 (skipped 5) | 4.5 | | |
| | cloze | 25.8% | 14,037 (skipped 5) | 0.0 | | |
| | Category | Subjects | Questions | Letter | Cloze | | |
| |---|---|---|---|---| | |
| | STEM | 19 | 3,153 | 27.6% | 23.5% | | |
| | Humanities | 13 | 4,705 | 24.4% | 25.5% | | |
| | Social Sciences | 12 | 3,077 | 23.6% | 26.6% | | |
| | Other | 13 | 3,102 | 26.4% | 27.9% | | |
| **Biology & medicine (cloze): 28.4% on 2,089 questions**, +3.4 points against 25% chance (about 3.6 standard errors, significant). This pools the 10 life-science and health subjects; it isn't an official MMLU category. | |
| In the **letter** format the model compares " A"β" D", with same-subject examples added while they fit the | |
| context. In the **cloze** format each answer's text is scored by log-probability per byte; at this size the | |
| cloze results are the meaningful ones. Chance is 25%. | |
| ### Compared with its siblings and other small models | |
| | | **This model** | [Pythia-70M](https://huggingface.co/EleutherAI/pythia-70m) | [GPT-2](https://huggingface.co/openai-community/gpt2) | [Supra-50M](https://huggingface.co/SupraLabs/Supra-50M-Base) | | |
| |---|---|---|---|---| | |
| | Parameters (total) | 2.0M | 70.4M | 124.4M | 51.8M | | |
| | Active non-embedding parameters | 0.5M | 18.9M | 85.1M | 35.4M | | |
| | Training tokens | 1,000M | 300B (the Pile) | undisclosed (WebText) | 20B (FineWeb-Edu) | | |
| | SciQ (acc_norm) | 56.3% | 55.2% | 53.2% | 77.2% | | |
| | ARC-Easy (acc_norm) | 32.7% | 35.0% | 42.0% | 52.2% | | |
| | HellaSwag (acc_norm) | 26.7% | β | 29.5% | 31.8% | | |
| | MMLU (cloze) | 25.8% | β | β | β | | |
| | MMLU biology & medicine (cloze) | 28.4% | β | β | β | | |
| | Validation loss (held-out mix) | 3.869 | β | β | β | | |
| LittleKedi columns are measured with this repo's `evals.py` and `mmlu_eval.py` on the checkpoints named. Validation | |
| losses are only comparable between models that share a tokenizer and validation set. The other columns are | |
| **published figures**: Pythia-70M from EleutherAI's evaluations as quoted on the | |
| [Wisp-15M card](https://huggingface.co/DedeProGames/Wisp-15M), GPT-2 and Supra-50M from the | |
| [Supra-50M-Base card](https://huggingface.co/SupraLabs/Supra-50M-Base). Their prompts may differ from ours, | |
| so treat gaps of a few points as noise. | |
| ## What it can and can't do | |
| At 0.5M active non-embedding parameters this model has learned **the form of English much better than its | |
| content**. It's a research artifact for studying small MoE models, not a source of information. | |
| **What it does** | |
| - β Writes grammatical, fluent sentences for a paragraph or two, in the register of its training data | |
| (educational web text and textbooks): *"Photosynthesis is the process by which"* β *"β¦it is used to monitor | |
| and measure the concentration of nutrients."* | |
| - β Stays roughly on topic: a prompt about the heart leads to blood vessels, the body and disease; a story | |
| opening leads to characters and a plot. | |
| - β Picks up document structure from its data mix: markdown headings and worked "Example 1:" sections | |
| after maths prompts, dialogue turns after `User:` / `Assistant:` text. | |
| - β Shows some benchmark signal above chance (SciQ 56.3%, ARC-Easy 32.7%, MMLU biology & medicine 28.4% in | |
| the cloze format), mostly from matching answers to a passage or recognising which option sounds right, | |
| not from recalling facts. | |
| **What it can't do** | |
| - β **Facts.** It reliably names the right kind of thing but the wrong thing: *"The French Revolution began | |
| in"* β *"1906β¦"*; *"Albert Einstein was"* β *"a co-director of the French Marxi Palm."* Treat every name, | |
| date and number it writes as made up. | |
| - β **Arithmetic.** *"2 + 2 ="* β *"3"*, *"7 + 6 ="* β *"6"*. | |
| - β **Code.** Code was 2% of its training data; given a Python function signature it produces empty | |
| docstrings and whitespace. | |
| - β **Instructions or questions.** It's a base model: it continues text rather than answering it, and with | |
| only 0.5M non-embedding parameters an instruction-tuned version won't be a reliable assistant either. | |
| - β **Long-range coherence.** Topics drift after a few sentences, and invented names appear | |
| (*"the Grovi Palace"*). | |
| **Decoding** | |
| Greedy or low-temperature decoding falls into loops almost immediately (*"The process of the plant is formed | |
| by the plant. The process of the plant is formed by the plantβ¦"*). Use sampling with **temperature β 0.7, | |
| top-p β 0.9 and a repetition penalty of about 1.15**, which is what produced the examples above. | |
| ## Usage | |
| **PyTorch** (needs the `littlekedi` package from [the training code](https://github.com/EdmundMartin/LittleKedi)): | |
| ```python | |
| import json, torch | |
| from safetensors.torch import load_file | |
| from littlekedi import LittleKediModel, ModelConfig | |
| from littlekedi.tokenizer import BPETokenizer | |
| model = LittleKediModel(ModelConfig.from_dict(json.load(open("config.json")))).eval() | |
| model.load_state_dict(load_file("model.safetensors")) | |
| tok = BPETokenizer.from_file("tokenizer.json") | |
| ids = torch.tensor([[tok.eot_id] + tok.encode("Photosynthesis is")]) | |
| print(tok.decode(model.generate(ids, 60, temperature=0.7, top_p=0.9, repetition_penalty=1.15)[0].tolist())) | |
| ``` | |
| Or interactively: `python sample.py --ckpt <checkpoint>.pt`. Base models are released as PyTorch weights only; | |
| GGUF builds come with the instruct models. | |
| ## Files | |
| | File | Contents | | |
| |---|---| | |
| | `model.safetensors` | PyTorch weights (embedding tied with the output head), `littlekedi` parameter names | | |
| | `config.json` | `ModelConfig` for `littlekedi` | | |
| | `tokenizer.json` | Byte-level BPE tokenizer (Hugging Face `tokenizers` format) | | |
| | `little_kedi.jpg` | The LittleKedi logo | | |
| ## Acknowledgements | |
| - The architecture follows **DeepSeek-V3** (DeepSeek-AI, 2024): MLA, DeepSeekMoE and auxiliary-loss-free balancing. | |
| - **Muon** (Keller Jordan et al., 2024) with Moonlight's update scaling (Liu et al., 2025). | |
| - Pretraining data: **FineWeb-Edu** (ODC-BY 1.0). | |
| - Pretraining data: **Cosmopedia** (Apache 2.0). | |
| - Pretraining data: **FineMath 4+** (ODC-BY 1.0). | |
| - Pretraining data: **SmolTalk** (Apache 2.0 for its new subsets, with the datasets it includes under their own licenses). | |
| - Pretraining data: **github-code-clean** (Apache 2.0, with each file under its own open-source license). | |
| *Card generated 2026-10-09 by `prepare_release.py` from `tiny.pt` (step 15257).* | |