edededdy commited on
Commit
277acf1
Β·
verified Β·
1 Parent(s): 9bf30e3

Upload folder using huggingface_hub

Browse files
Files changed (6) hide show
  1. .gitattributes +1 -0
  2. README.md +260 -0
  3. config.json +29 -0
  4. little_kedi.jpg +3 -0
  5. model.safetensors +3 -0
  6. tokenizer.json +0 -0
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ little_kedi.jpg filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,260 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ pipeline_tag: text-generation
6
+ datasets:
7
+ - HuggingFaceFW/fineweb-edu
8
+ - HuggingFaceTB/cosmopedia
9
+ - HuggingFaceTB/finemath
10
+ - HuggingFaceTB/smoltalk
11
+ - codeparrot/github-code-clean
12
+ tags:
13
+ - mixture-of-experts
14
+ - moe
15
+ - deepseek
16
+ - multi-head-latent-attention
17
+ - mla
18
+ - from-scratch
19
+ - tiny
20
+ - littlekedi
21
+ - base
22
+ ---
23
+
24
+ <p align="center"><img src="little_kedi.jpg" alt="LittleKedi logo" width="640"></p>
25
+
26
+ # LittleKedi-tiny-base
27
+
28
+ `LittleKedi-tiny-base` is a **small DeepSeek-V3-style Mixture-of-Experts language model trained from scratch**. It has
29
+ **2.0M parameters** in total, of which **1.5M are active per token**
30
+ (0.50M excluding the embedding table). It was pretrained on a 1,000M-token mix (FineWeb-Edu, Cosmopedia (Stanford), Cosmopedia (OpenStax), Cosmopedia (WikiHow), Cosmopedia (Khan Academy), FineMath 4+, SmolTalk, github-code-clean).
31
+
32
+ *Part of the LittleKedi family. Not affiliated with DeepSeek.*
33
+
34
+ This is the **base model**, before any instruction tuning. It continues text; it doesn't follow instructions. It's an educational and research artifact: it writes fluent, on-topic English but **knows few facts and confidently makes things up**. Don't use it for anything that matters.
35
+
36
+ ## The LittleKedi family
37
+
38
+ LittleKedi is a family of small DeepSeek-V3-style Mixture-of-Experts language models trained from
39
+ scratch. Every size comes in three variants that differ in **how sparse** they are: how many routed experts
40
+ they have, and what share of their parameters a token actually passes through.
41
+
42
+ | Variant | Routed experts |
43
+ |---|---|
44
+ | **standard** (`<size>`) | a few wide experts; the densest variant |
45
+ | **sparse** (`<size>-sparse`) | many narrower, fine-grained experts, more of them active per token |
46
+ | **wide-sparse** (`<size>-wide-sparse`) | more experts again; the largest total and smallest active share |
47
+
48
+ More experts means more total capacity to store knowledge, but each expert
49
+ sees fewer tokens (tokens Γ— top-k Γ· experts per layer), so sparser variants need more training data before
50
+ their experts are well trained. The variants of a size share a tokenizer and dataset, so they can be
51
+ compared directly. Shapes change as the family scales; the table below is generated from the current
52
+ presets.
53
+
54
+ This card is for **`tiny`**, the **standard** tiny model: 1.5M of its 2.0M parameters (78%) are active per token. Its siblings are `tiny-sparse` and `tiny-wide-sparse`.
55
+
56
+ | Model | Variant | Vocab | Total | Active | Active non-emb. | Active % | Routed experts | Tokens/expert (1B) |
57
+ |---|---|---|---|---|---|---|---|---|
58
+ | **`tiny` ← this model** | **standard** | **8K** | **2.0M** | **1.5M** | **0.5M** | **78%** | **8 Γ— 64, top-2** | **250M** |
59
+ | `tiny-sparse` | sparse | 8K | 2.6M | 1.6M | 0.5M | 60% | 32 Γ— 32, top-4 | 125M |
60
+ | `tiny-wide-sparse` | wide-sparse | 8K | 3.8M | 1.6M | 0.5M | 42% | 64 Γ— 32, top-4 | 62M |
61
+ | `mini` | standard | 8K | 9.4M | 4.8M | 3.2M | 51% | 12 Γ— 128, top-3 | 250M |
62
+ | `mini-sparse` | sparse | 8K | 19.8M | 4.8M | 3.2M | 24% | 64 Γ— 64, top-6 | 94M |
63
+ | `mini-wide-sparse` | wide-sparse | 8K | 36.4M | 4.9M | 3.3M | 13% | 128 Γ— 64, top-6 | 47M |
64
+ | `small` | standard | 8K | 12.2M | 6.3M | 4.2M | 52% | 16 Γ— 128, top-4 | 250M |
65
+ | `small-sparse` | sparse | 8K | 24.1M | 6.4M | 4.3M | 27% | 80 Γ— 64, top-8 | 100M |
66
+ | `small-wide-sparse` | wide-sparse | 8K | 43.9M | 6.5M | 4.4M | 15% | 160 Γ— 64, top-8 | 50M |
67
+ | `moderate` | standard | 16K | 49.8M | 18.8M | 12.5M | 38% | 24 Γ— 192, top-4 | 167M |
68
+ | `moderate-sparse` | sparse | 16K | 105.8M | 19.1M | 12.8M | 18% | 120 Γ— 96, top-8 | 67M |
69
+ | `moderate-wide-sparse` | wide-sparse | 16K | 199.0M | 19.4M | 13.1M | 10% | 240 Γ— 96, top-8 | 33M |
70
+ | `medium` | standard | 16K | 85.9M | 30.8M | 22.4M | 36% | 24 Γ— 256, top-4 | 167M |
71
+ | `medium-sparse` | sparse | 16K | 185.3M | 31.2M | 22.8M | 17% | 120 Γ— 128, top-8 | 67M |
72
+ | `medium-wide-sparse` | wide-sparse | 16K | 350.9M | 31.6M | 23.2M | 9% | 240 Γ— 128, top-8 | 33M |
73
+ | `base` | standard | 32K | 258.7M | 105.4M | 80.2M | 41% | 32 Γ— 256, top-6 | 188M |
74
+ | `base-sparse` | sparse | 32K | 542.8M | 106.3M | 81.2M | 20% | 160 Γ— 128, top-12 | 75M |
75
+ | `base-wide-sparse` | wide-sparse | 32K | 1015.9M | 107.6M | 82.4M | 11% | 320 Γ— 128, top-12 | 38M |
76
+
77
+ "Active" counts everything a token passes through: embeddings, attention, the dense layer, shared
78
+ experts and its top-k routed experts. "Active non-embedding" leaves out the embedding table (tied with the
79
+ output head), which costs memory but almost no compute; it's the fair number for comparing compute
80
+ between models. Tokens per expert is per MoE layer for a 1B-token run with balanced routing; below
81
+ ~25–50M, experts are probably undertrained.
82
+
83
+ ## Architecture
84
+
85
+ | Component | This model |
86
+ |---|---|
87
+ | Parameters | 1.99M total, 1.55M active per token (78%), 0.50M active non-embedding |
88
+ | Layers / hidden size | 4 / 128 (layer 0 has a dense SwiGLU FFN of 352; layers 1–3 are MoE) |
89
+ | Attention | **Multi-head Latent Attention (MLA)**: 4 heads, KV latent 32, decoupled RoPE dim 16, head dims 16+16 (q/k) and 16 (v). The KV cache stores only the 32-dim latent plus a 16-dim RoPE key per token |
90
+ | MoE | **DeepSeekMoE**: 8 routed experts (SwiGLU, 64 hidden), top-2, plus 1 shared expert; sigmoid gating |
91
+ | Load balancing | **Auxiliary-loss-free** bias balancing (update speed 0.001) plus a small sequence-wise loss (Ξ± = 0.0001). Over the second half of training the median batch put 1.07Γ— the mean load on the busiest expert (1.08Γ— in the worst layer), and 0.00% of expert assignments were dropped for capacity. |
92
+ | Expert dispatch | Training: fixed capacity (capacity factor 2.0; overflow assignments are dropped). Inference is always dropless |
93
+ | Embeddings | Input embedding and output head **tied** |
94
+ | Context | 1,024 tokens |
95
+
96
+ ## Training
97
+
98
+ | | |
99
+ |---|---|
100
+ | Data | 1,000M tokens, 801,792 documents (~502 tokens per parameter, ~646 per active parameter), one pass, sources interleaved throughout (mix below) |
101
+ | Tokenizer | Byte-level BPE, 8,192 tokens (~3.75 bytes/token on this mix); digits not split (v1 tokenizer); no chat tokens, so chat data is plain `User:` / `Assistant:` text |
102
+ | Steps | 15,258 Γ— 65,536 tokens (batch 16 Γ— 1024 Γ— 4 accumulation) |
103
+ | Optimizer | **Muon** (momentum 0.95, Nesterov, 5 Newton-Schulz steps; each expert's matrix orthogonalized separately) for hidden matrices, AdamW (Ξ² = 0.9, 0.95) for embeddings, norms and the router; weight decay 0.1, gradient clipping 1.0 |
104
+ | Learning rate | 2e-3 peak; **WSD**: 200 warm-up steps, constant, then linear decay to 0 over the last 20% of steps (schedule set for 15,258 steps) |
105
+ | Precision | bf16 autocast, `torch.compile` |
106
+ | Compute | ~3e15 FLOPs (6 Γ— active non-embedding parameters Γ— tokens) |
107
+ | Hardware | One AMD Radeon 890M integrated GPU (Ryzen AI 300 laptop), ROCm on Windows: about 8.9 hours at a median ~31.1K tokens/s |
108
+
109
+ | Source | What it is | Tokens | Share |
110
+ |---|---|---|---|
111
+ | FineWeb-Edu `sample/10BT` | educational web pages | 700M | 70% |
112
+ | Cosmopedia (Stanford) | synthetic textbooks | 70M | 7% |
113
+ | Cosmopedia (OpenStax) | synthetic textbooks | 30M | 3% |
114
+ | Cosmopedia (WikiHow) | synthetic how-to articles | 30M | 3% |
115
+ | Cosmopedia (Khan Academy) | synthetic lessons | 20M | 2% |
116
+ | FineMath 4+ | maths web pages | 80M | 8% |
117
+ | SmolTalk | chat conversations | 50M | 5% |
118
+ | github-code-clean | source code | 20M | 2% |
119
+
120
+ ## Evaluation
121
+
122
+ | Metric | Value |
123
+ |---|---|
124
+ | Validation loss (nats/token, held-out mix) | **3.869** |
125
+ | Perplexity | 47.9 |
126
+ | Bits per byte | 1.438 |
127
+ | Experts never chosen on the validation sample | 0 |
128
+
129
+ Validation loss over training: 5.60 (16M tok) β†’ 4.17 (164M tok) β†’ 4.08 (295M tok) β†’ 4.04 (442M tok) β†’ 4.02 (590M tok) β†’ 4.01 (737M tok) β†’ 3.96 (868M tok) β†’ 3.87 (1,000M tok)
130
+
131
+ ### Zero-shot benchmarks (`evals.py`)
132
+
133
+ | Task | acc_norm | acc | Items | Chance |
134
+ |---|---|---|---|---|
135
+ | SciQ | 56.3% Β± 1.6% | 59.8% | 1,000 | 25% |
136
+ | ARC-Easy | 32.7% Β± 1.0% | 32.3% | 2,376 | 25% |
137
+ | HellaSwag | 26.7% Β± 0.4% | 26.4% | 10,042 | 25% |
138
+
139
+ Each answer choice is scored by the log-probability of its text after the prompt, using lm-evaluation-harness
140
+ prompts (SciQ includes its supporting passage). **acc_norm** divides by the answer's length in characters;
141
+ **acc** doesn't. Splits: SciQ test, ARC-Easy test, HellaSwag validation.
142
+
143
+ ### MMLU (`mmlu_eval.py`)
144
+
145
+ | Format | Accuracy | Questions scored | Average few-shot examples |
146
+ |---|---|---|---|
147
+ | letter | 25.4% | 14,037 (skipped 5) | 4.5 |
148
+ | cloze | 25.8% | 14,037 (skipped 5) | 0.0 |
149
+
150
+ | Category | Subjects | Questions | Letter | Cloze |
151
+ |---|---|---|---|---|
152
+ | STEM | 19 | 3,153 | 27.6% | 23.5% |
153
+ | Humanities | 13 | 4,705 | 24.4% | 25.5% |
154
+ | Social Sciences | 12 | 3,077 | 23.6% | 26.6% |
155
+ | Other | 13 | 3,102 | 26.4% | 27.9% |
156
+
157
+ **Biology & medicine (cloze): 28.4% on 2,089 questions**, +3.4 points against 25% chance (about 3.6 standard errors, significant). This pools the 10 life-science and health subjects; it isn't an official MMLU category.
158
+
159
+ In the **letter** format the model compares " A"–" D", with same-subject examples added while they fit the
160
+ context. In the **cloze** format each answer's text is scored by log-probability per byte; at this size the
161
+ cloze results are the meaningful ones. Chance is 25%.
162
+
163
+ ### Compared with its siblings and other small models
164
+
165
+ | | **This model** | [Pythia-70M](https://huggingface.co/EleutherAI/pythia-70m) | [GPT-2](https://huggingface.co/openai-community/gpt2) | [Supra-50M](https://huggingface.co/SupraLabs/Supra-50M-Base) |
166
+ |---|---|---|---|---|
167
+ | Parameters (total) | 2.0M | 70.4M | 124.4M | 51.8M |
168
+ | Active non-embedding parameters | 0.5M | 18.9M | 85.1M | 35.4M |
169
+ | Training tokens | 1,000M | 300B (the Pile) | undisclosed (WebText) | 20B (FineWeb-Edu) |
170
+ | SciQ (acc_norm) | 56.3% | 55.2% | 53.2% | 77.2% |
171
+ | ARC-Easy (acc_norm) | 32.7% | 35.0% | 42.0% | 52.2% |
172
+ | HellaSwag (acc_norm) | 26.7% | – | 29.5% | 31.8% |
173
+ | MMLU (cloze) | 25.8% | – | – | – |
174
+ | MMLU biology & medicine (cloze) | 28.4% | – | – | – |
175
+ | Validation loss (held-out mix) | 3.869 | – | – | – |
176
+
177
+ LittleKedi columns are measured with this repo's `evals.py` and `mmlu_eval.py` on the checkpoints named. Validation
178
+ losses are only comparable between models that share a tokenizer and validation set. The other columns are
179
+ **published figures**: Pythia-70M from EleutherAI's evaluations as quoted on the
180
+ [Wisp-15M card](https://huggingface.co/DedeProGames/Wisp-15M), GPT-2 and Supra-50M from the
181
+ [Supra-50M-Base card](https://huggingface.co/SupraLabs/Supra-50M-Base). Their prompts may differ from ours,
182
+ so treat gaps of a few points as noise.
183
+
184
+ ## What it can and can't do
185
+
186
+ At 0.5M active non-embedding parameters this model has learned **the form of English much better than its
187
+ content**. It's a research artifact for studying small MoE models, not a source of information.
188
+
189
+ **What it does**
190
+
191
+ - βœ… Writes grammatical, fluent sentences for a paragraph or two, in the register of its training data
192
+ (educational web text and textbooks): *"Photosynthesis is the process by which"* β†’ *"…it is used to monitor
193
+ and measure the concentration of nutrients."*
194
+ - βœ… Stays roughly on topic: a prompt about the heart leads to blood vessels, the body and disease; a story
195
+ opening leads to characters and a plot.
196
+ - βœ… Picks up document structure from its data mix: markdown headings and worked "Example 1:" sections
197
+ after maths prompts, dialogue turns after `User:` / `Assistant:` text.
198
+ - βœ… Shows some benchmark signal above chance (SciQ 56.3%, ARC-Easy 32.7%, MMLU biology & medicine 28.4% in
199
+ the cloze format), mostly from matching answers to a passage or recognising which option sounds right,
200
+ not from recalling facts.
201
+
202
+ **What it can't do**
203
+
204
+ - ❌ **Facts.** It reliably names the right kind of thing but the wrong thing: *"The French Revolution began
205
+ in"* β†’ *"1906…"*; *"Albert Einstein was"* β†’ *"a co-director of the French Marxi Palm."* Treat every name,
206
+ date and number it writes as made up.
207
+ - ❌ **Arithmetic.** *"2 + 2 ="* β†’ *"3"*, *"7 + 6 ="* β†’ *"6"*.
208
+ - ❌ **Code.** Code was 2% of its training data; given a Python function signature it produces empty
209
+ docstrings and whitespace.
210
+ - ❌ **Instructions or questions.** It's a base model: it continues text rather than answering it, and with
211
+ only 0.5M non-embedding parameters an instruction-tuned version won't be a reliable assistant either.
212
+ - ❌ **Long-range coherence.** Topics drift after a few sentences, and invented names appear
213
+ (*"the Grovi Palace"*).
214
+
215
+ **Decoding**
216
+
217
+ Greedy or low-temperature decoding falls into loops almost immediately (*"The process of the plant is formed
218
+ by the plant. The process of the plant is formed by the plant…"*). Use sampling with **temperature β‰ˆ 0.7,
219
+ top-p β‰ˆ 0.9 and a repetition penalty of about 1.15**, which is what produced the examples above.
220
+
221
+ ## Usage
222
+
223
+ **PyTorch** (needs the `littlekedi` package from [the training code](https://github.com/EdmundMartin/LittleKedi)):
224
+
225
+ ```python
226
+ import json, torch
227
+ from safetensors.torch import load_file
228
+ from littlekedi import LittleKediModel, ModelConfig
229
+ from littlekedi.tokenizer import BPETokenizer
230
+
231
+ model = LittleKediModel(ModelConfig.from_dict(json.load(open("config.json")))).eval()
232
+ model.load_state_dict(load_file("model.safetensors"))
233
+ tok = BPETokenizer.from_file("tokenizer.json")
234
+ ids = torch.tensor([[tok.eot_id] + tok.encode("Photosynthesis is")])
235
+ print(tok.decode(model.generate(ids, 60, temperature=0.7, top_p=0.9, repetition_penalty=1.15)[0].tolist()))
236
+ ```
237
+
238
+ Or interactively: `python sample.py --ckpt <checkpoint>.pt`. Base models are released as PyTorch weights only;
239
+ GGUF builds come with the instruct models.
240
+
241
+ ## Files
242
+
243
+ | File | Contents |
244
+ |---|---|
245
+ | `model.safetensors` | PyTorch weights (embedding tied with the output head), `littlekedi` parameter names |
246
+ | `config.json` | `ModelConfig` for `littlekedi` |
247
+ | `tokenizer.json` | Byte-level BPE tokenizer (Hugging Face `tokenizers` format) |
248
+ | `little_kedi.jpg` | The LittleKedi logo |
249
+
250
+ ## Acknowledgements
251
+
252
+ - The architecture follows **DeepSeek-V3** (DeepSeek-AI, 2024): MLA, DeepSeekMoE and auxiliary-loss-free balancing.
253
+ - **Muon** (Keller Jordan et al., 2024) with Moonlight's update scaling (Liu et al., 2025).
254
+ - Pretraining data: **FineWeb-Edu** (ODC-BY 1.0).
255
+ - Pretraining data: **Cosmopedia** (Apache 2.0).
256
+ - Pretraining data: **FineMath 4+** (ODC-BY 1.0).
257
+ - Pretraining data: **SmolTalk** (Apache 2.0 for its new subsets, with the datasets it includes under their own licenses).
258
+ - Pretraining data: **github-code-clean** (Apache 2.0, with each file under its own open-source license).
259
+
260
+ *Card generated 2026-10-09 by `prepare_release.py` from `tiny.pt` (step 15257).*
config.json ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "vocab_size": 8192,
3
+ "max_seq_len": 1024,
4
+ "dim": 128,
5
+ "n_layers": 4,
6
+ "n_dense_layers": 1,
7
+ "dense_hidden_dim": 352,
8
+ "norm_eps": 1e-06,
9
+ "init_std": 0.02,
10
+ "tie_embeddings": true,
11
+ "n_heads": 4,
12
+ "q_lora_rank": 0,
13
+ "kv_lora_rank": 32,
14
+ "qk_nope_head_dim": 16,
15
+ "qk_rope_head_dim": 16,
16
+ "v_head_dim": 16,
17
+ "rope_theta": 10000.0,
18
+ "n_routed_experts": 8,
19
+ "n_shared_experts": 1,
20
+ "n_activated_experts": 2,
21
+ "moe_hidden_dim": 64,
22
+ "n_expert_groups": 1,
23
+ "n_limited_groups": 1,
24
+ "route_scale": 1.0,
25
+ "bias_update_speed": 0.001,
26
+ "seq_aux_loss_alpha": 0.0001,
27
+ "moe_backend": "capacity",
28
+ "capacity_factor": 2.0
29
+ }
little_kedi.jpg ADDED

Git LFS Details

  • SHA256: 2eaa8b24fd3885aad7a44660c49fb4625027d375bbb00fb448e6579b23985310
  • Pointer size: 132 Bytes
  • Size of remote file: 1.87 MB
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:369319cb33abff7f4a0c178586a42aab0c6b77437fdcd2ba8392fe33e3264343
3
+ size 7969152
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff