Upload folder using huggingface_hub
Browse files- .gitattributes +1 -0
- README.md +260 -0
- config.json +29 -0
- little_kedi.jpg +3 -0
- model.safetensors +3 -0
- tokenizer.json +0 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
little_kedi.jpg filter=lfs diff=lfs merge=lfs -text
|
README.md
ADDED
|
@@ -0,0 +1,260 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
language:
|
| 4 |
+
- en
|
| 5 |
+
pipeline_tag: text-generation
|
| 6 |
+
datasets:
|
| 7 |
+
- HuggingFaceFW/fineweb-edu
|
| 8 |
+
- HuggingFaceTB/cosmopedia
|
| 9 |
+
- HuggingFaceTB/finemath
|
| 10 |
+
- HuggingFaceTB/smoltalk
|
| 11 |
+
- codeparrot/github-code-clean
|
| 12 |
+
tags:
|
| 13 |
+
- mixture-of-experts
|
| 14 |
+
- moe
|
| 15 |
+
- deepseek
|
| 16 |
+
- multi-head-latent-attention
|
| 17 |
+
- mla
|
| 18 |
+
- from-scratch
|
| 19 |
+
- tiny
|
| 20 |
+
- littlekedi
|
| 21 |
+
- base
|
| 22 |
+
---
|
| 23 |
+
|
| 24 |
+
<p align="center"><img src="little_kedi.jpg" alt="LittleKedi logo" width="640"></p>
|
| 25 |
+
|
| 26 |
+
# LittleKedi-tiny-base
|
| 27 |
+
|
| 28 |
+
`LittleKedi-tiny-base` is a **small DeepSeek-V3-style Mixture-of-Experts language model trained from scratch**. It has
|
| 29 |
+
**2.0M parameters** in total, of which **1.5M are active per token**
|
| 30 |
+
(0.50M excluding the embedding table). It was pretrained on a 1,000M-token mix (FineWeb-Edu, Cosmopedia (Stanford), Cosmopedia (OpenStax), Cosmopedia (WikiHow), Cosmopedia (Khan Academy), FineMath 4+, SmolTalk, github-code-clean).
|
| 31 |
+
|
| 32 |
+
*Part of the LittleKedi family. Not affiliated with DeepSeek.*
|
| 33 |
+
|
| 34 |
+
This is the **base model**, before any instruction tuning. It continues text; it doesn't follow instructions. It's an educational and research artifact: it writes fluent, on-topic English but **knows few facts and confidently makes things up**. Don't use it for anything that matters.
|
| 35 |
+
|
| 36 |
+
## The LittleKedi family
|
| 37 |
+
|
| 38 |
+
LittleKedi is a family of small DeepSeek-V3-style Mixture-of-Experts language models trained from
|
| 39 |
+
scratch. Every size comes in three variants that differ in **how sparse** they are: how many routed experts
|
| 40 |
+
they have, and what share of their parameters a token actually passes through.
|
| 41 |
+
|
| 42 |
+
| Variant | Routed experts |
|
| 43 |
+
|---|---|
|
| 44 |
+
| **standard** (`<size>`) | a few wide experts; the densest variant |
|
| 45 |
+
| **sparse** (`<size>-sparse`) | many narrower, fine-grained experts, more of them active per token |
|
| 46 |
+
| **wide-sparse** (`<size>-wide-sparse`) | more experts again; the largest total and smallest active share |
|
| 47 |
+
|
| 48 |
+
More experts means more total capacity to store knowledge, but each expert
|
| 49 |
+
sees fewer tokens (tokens Γ top-k Γ· experts per layer), so sparser variants need more training data before
|
| 50 |
+
their experts are well trained. The variants of a size share a tokenizer and dataset, so they can be
|
| 51 |
+
compared directly. Shapes change as the family scales; the table below is generated from the current
|
| 52 |
+
presets.
|
| 53 |
+
|
| 54 |
+
This card is for **`tiny`**, the **standard** tiny model: 1.5M of its 2.0M parameters (78%) are active per token. Its siblings are `tiny-sparse` and `tiny-wide-sparse`.
|
| 55 |
+
|
| 56 |
+
| Model | Variant | Vocab | Total | Active | Active non-emb. | Active % | Routed experts | Tokens/expert (1B) |
|
| 57 |
+
|---|---|---|---|---|---|---|---|---|
|
| 58 |
+
| **`tiny` β this model** | **standard** | **8K** | **2.0M** | **1.5M** | **0.5M** | **78%** | **8 Γ 64, top-2** | **250M** |
|
| 59 |
+
| `tiny-sparse` | sparse | 8K | 2.6M | 1.6M | 0.5M | 60% | 32 Γ 32, top-4 | 125M |
|
| 60 |
+
| `tiny-wide-sparse` | wide-sparse | 8K | 3.8M | 1.6M | 0.5M | 42% | 64 Γ 32, top-4 | 62M |
|
| 61 |
+
| `mini` | standard | 8K | 9.4M | 4.8M | 3.2M | 51% | 12 Γ 128, top-3 | 250M |
|
| 62 |
+
| `mini-sparse` | sparse | 8K | 19.8M | 4.8M | 3.2M | 24% | 64 Γ 64, top-6 | 94M |
|
| 63 |
+
| `mini-wide-sparse` | wide-sparse | 8K | 36.4M | 4.9M | 3.3M | 13% | 128 Γ 64, top-6 | 47M |
|
| 64 |
+
| `small` | standard | 8K | 12.2M | 6.3M | 4.2M | 52% | 16 Γ 128, top-4 | 250M |
|
| 65 |
+
| `small-sparse` | sparse | 8K | 24.1M | 6.4M | 4.3M | 27% | 80 Γ 64, top-8 | 100M |
|
| 66 |
+
| `small-wide-sparse` | wide-sparse | 8K | 43.9M | 6.5M | 4.4M | 15% | 160 Γ 64, top-8 | 50M |
|
| 67 |
+
| `moderate` | standard | 16K | 49.8M | 18.8M | 12.5M | 38% | 24 Γ 192, top-4 | 167M |
|
| 68 |
+
| `moderate-sparse` | sparse | 16K | 105.8M | 19.1M | 12.8M | 18% | 120 Γ 96, top-8 | 67M |
|
| 69 |
+
| `moderate-wide-sparse` | wide-sparse | 16K | 199.0M | 19.4M | 13.1M | 10% | 240 Γ 96, top-8 | 33M |
|
| 70 |
+
| `medium` | standard | 16K | 85.9M | 30.8M | 22.4M | 36% | 24 Γ 256, top-4 | 167M |
|
| 71 |
+
| `medium-sparse` | sparse | 16K | 185.3M | 31.2M | 22.8M | 17% | 120 Γ 128, top-8 | 67M |
|
| 72 |
+
| `medium-wide-sparse` | wide-sparse | 16K | 350.9M | 31.6M | 23.2M | 9% | 240 Γ 128, top-8 | 33M |
|
| 73 |
+
| `base` | standard | 32K | 258.7M | 105.4M | 80.2M | 41% | 32 Γ 256, top-6 | 188M |
|
| 74 |
+
| `base-sparse` | sparse | 32K | 542.8M | 106.3M | 81.2M | 20% | 160 Γ 128, top-12 | 75M |
|
| 75 |
+
| `base-wide-sparse` | wide-sparse | 32K | 1015.9M | 107.6M | 82.4M | 11% | 320 Γ 128, top-12 | 38M |
|
| 76 |
+
|
| 77 |
+
"Active" counts everything a token passes through: embeddings, attention, the dense layer, shared
|
| 78 |
+
experts and its top-k routed experts. "Active non-embedding" leaves out the embedding table (tied with the
|
| 79 |
+
output head), which costs memory but almost no compute; it's the fair number for comparing compute
|
| 80 |
+
between models. Tokens per expert is per MoE layer for a 1B-token run with balanced routing; below
|
| 81 |
+
~25β50M, experts are probably undertrained.
|
| 82 |
+
|
| 83 |
+
## Architecture
|
| 84 |
+
|
| 85 |
+
| Component | This model |
|
| 86 |
+
|---|---|
|
| 87 |
+
| Parameters | 1.99M total, 1.55M active per token (78%), 0.50M active non-embedding |
|
| 88 |
+
| Layers / hidden size | 4 / 128 (layer 0 has a dense SwiGLU FFN of 352; layers 1β3 are MoE) |
|
| 89 |
+
| Attention | **Multi-head Latent Attention (MLA)**: 4 heads, KV latent 32, decoupled RoPE dim 16, head dims 16+16 (q/k) and 16 (v). The KV cache stores only the 32-dim latent plus a 16-dim RoPE key per token |
|
| 90 |
+
| MoE | **DeepSeekMoE**: 8 routed experts (SwiGLU, 64 hidden), top-2, plus 1 shared expert; sigmoid gating |
|
| 91 |
+
| Load balancing | **Auxiliary-loss-free** bias balancing (update speed 0.001) plus a small sequence-wise loss (Ξ± = 0.0001). Over the second half of training the median batch put 1.07Γ the mean load on the busiest expert (1.08Γ in the worst layer), and 0.00% of expert assignments were dropped for capacity. |
|
| 92 |
+
| Expert dispatch | Training: fixed capacity (capacity factor 2.0; overflow assignments are dropped). Inference is always dropless |
|
| 93 |
+
| Embeddings | Input embedding and output head **tied** |
|
| 94 |
+
| Context | 1,024 tokens |
|
| 95 |
+
|
| 96 |
+
## Training
|
| 97 |
+
|
| 98 |
+
| | |
|
| 99 |
+
|---|---|
|
| 100 |
+
| Data | 1,000M tokens, 801,792 documents (~502 tokens per parameter, ~646 per active parameter), one pass, sources interleaved throughout (mix below) |
|
| 101 |
+
| Tokenizer | Byte-level BPE, 8,192 tokens (~3.75 bytes/token on this mix); digits not split (v1 tokenizer); no chat tokens, so chat data is plain `User:` / `Assistant:` text |
|
| 102 |
+
| Steps | 15,258 Γ 65,536 tokens (batch 16 Γ 1024 Γ 4 accumulation) |
|
| 103 |
+
| Optimizer | **Muon** (momentum 0.95, Nesterov, 5 Newton-Schulz steps; each expert's matrix orthogonalized separately) for hidden matrices, AdamW (Ξ² = 0.9, 0.95) for embeddings, norms and the router; weight decay 0.1, gradient clipping 1.0 |
|
| 104 |
+
| Learning rate | 2e-3 peak; **WSD**: 200 warm-up steps, constant, then linear decay to 0 over the last 20% of steps (schedule set for 15,258 steps) |
|
| 105 |
+
| Precision | bf16 autocast, `torch.compile` |
|
| 106 |
+
| Compute | ~3e15 FLOPs (6 Γ active non-embedding parameters Γ tokens) |
|
| 107 |
+
| Hardware | One AMD Radeon 890M integrated GPU (Ryzen AI 300 laptop), ROCm on Windows: about 8.9 hours at a median ~31.1K tokens/s |
|
| 108 |
+
|
| 109 |
+
| Source | What it is | Tokens | Share |
|
| 110 |
+
|---|---|---|---|
|
| 111 |
+
| FineWeb-Edu `sample/10BT` | educational web pages | 700M | 70% |
|
| 112 |
+
| Cosmopedia (Stanford) | synthetic textbooks | 70M | 7% |
|
| 113 |
+
| Cosmopedia (OpenStax) | synthetic textbooks | 30M | 3% |
|
| 114 |
+
| Cosmopedia (WikiHow) | synthetic how-to articles | 30M | 3% |
|
| 115 |
+
| Cosmopedia (Khan Academy) | synthetic lessons | 20M | 2% |
|
| 116 |
+
| FineMath 4+ | maths web pages | 80M | 8% |
|
| 117 |
+
| SmolTalk | chat conversations | 50M | 5% |
|
| 118 |
+
| github-code-clean | source code | 20M | 2% |
|
| 119 |
+
|
| 120 |
+
## Evaluation
|
| 121 |
+
|
| 122 |
+
| Metric | Value |
|
| 123 |
+
|---|---|
|
| 124 |
+
| Validation loss (nats/token, held-out mix) | **3.869** |
|
| 125 |
+
| Perplexity | 47.9 |
|
| 126 |
+
| Bits per byte | 1.438 |
|
| 127 |
+
| Experts never chosen on the validation sample | 0 |
|
| 128 |
+
|
| 129 |
+
Validation loss over training: 5.60 (16M tok) β 4.17 (164M tok) β 4.08 (295M tok) β 4.04 (442M tok) β 4.02 (590M tok) β 4.01 (737M tok) β 3.96 (868M tok) β 3.87 (1,000M tok)
|
| 130 |
+
|
| 131 |
+
### Zero-shot benchmarks (`evals.py`)
|
| 132 |
+
|
| 133 |
+
| Task | acc_norm | acc | Items | Chance |
|
| 134 |
+
|---|---|---|---|---|
|
| 135 |
+
| SciQ | 56.3% Β± 1.6% | 59.8% | 1,000 | 25% |
|
| 136 |
+
| ARC-Easy | 32.7% Β± 1.0% | 32.3% | 2,376 | 25% |
|
| 137 |
+
| HellaSwag | 26.7% Β± 0.4% | 26.4% | 10,042 | 25% |
|
| 138 |
+
|
| 139 |
+
Each answer choice is scored by the log-probability of its text after the prompt, using lm-evaluation-harness
|
| 140 |
+
prompts (SciQ includes its supporting passage). **acc_norm** divides by the answer's length in characters;
|
| 141 |
+
**acc** doesn't. Splits: SciQ test, ARC-Easy test, HellaSwag validation.
|
| 142 |
+
|
| 143 |
+
### MMLU (`mmlu_eval.py`)
|
| 144 |
+
|
| 145 |
+
| Format | Accuracy | Questions scored | Average few-shot examples |
|
| 146 |
+
|---|---|---|---|
|
| 147 |
+
| letter | 25.4% | 14,037 (skipped 5) | 4.5 |
|
| 148 |
+
| cloze | 25.8% | 14,037 (skipped 5) | 0.0 |
|
| 149 |
+
|
| 150 |
+
| Category | Subjects | Questions | Letter | Cloze |
|
| 151 |
+
|---|---|---|---|---|
|
| 152 |
+
| STEM | 19 | 3,153 | 27.6% | 23.5% |
|
| 153 |
+
| Humanities | 13 | 4,705 | 24.4% | 25.5% |
|
| 154 |
+
| Social Sciences | 12 | 3,077 | 23.6% | 26.6% |
|
| 155 |
+
| Other | 13 | 3,102 | 26.4% | 27.9% |
|
| 156 |
+
|
| 157 |
+
**Biology & medicine (cloze): 28.4% on 2,089 questions**, +3.4 points against 25% chance (about 3.6 standard errors, significant). This pools the 10 life-science and health subjects; it isn't an official MMLU category.
|
| 158 |
+
|
| 159 |
+
In the **letter** format the model compares " A"β" D", with same-subject examples added while they fit the
|
| 160 |
+
context. In the **cloze** format each answer's text is scored by log-probability per byte; at this size the
|
| 161 |
+
cloze results are the meaningful ones. Chance is 25%.
|
| 162 |
+
|
| 163 |
+
### Compared with its siblings and other small models
|
| 164 |
+
|
| 165 |
+
| | **This model** | [Pythia-70M](https://huggingface.co/EleutherAI/pythia-70m) | [GPT-2](https://huggingface.co/openai-community/gpt2) | [Supra-50M](https://huggingface.co/SupraLabs/Supra-50M-Base) |
|
| 166 |
+
|---|---|---|---|---|
|
| 167 |
+
| Parameters (total) | 2.0M | 70.4M | 124.4M | 51.8M |
|
| 168 |
+
| Active non-embedding parameters | 0.5M | 18.9M | 85.1M | 35.4M |
|
| 169 |
+
| Training tokens | 1,000M | 300B (the Pile) | undisclosed (WebText) | 20B (FineWeb-Edu) |
|
| 170 |
+
| SciQ (acc_norm) | 56.3% | 55.2% | 53.2% | 77.2% |
|
| 171 |
+
| ARC-Easy (acc_norm) | 32.7% | 35.0% | 42.0% | 52.2% |
|
| 172 |
+
| HellaSwag (acc_norm) | 26.7% | β | 29.5% | 31.8% |
|
| 173 |
+
| MMLU (cloze) | 25.8% | β | β | β |
|
| 174 |
+
| MMLU biology & medicine (cloze) | 28.4% | β | β | β |
|
| 175 |
+
| Validation loss (held-out mix) | 3.869 | β | β | β |
|
| 176 |
+
|
| 177 |
+
LittleKedi columns are measured with this repo's `evals.py` and `mmlu_eval.py` on the checkpoints named. Validation
|
| 178 |
+
losses are only comparable between models that share a tokenizer and validation set. The other columns are
|
| 179 |
+
**published figures**: Pythia-70M from EleutherAI's evaluations as quoted on the
|
| 180 |
+
[Wisp-15M card](https://huggingface.co/DedeProGames/Wisp-15M), GPT-2 and Supra-50M from the
|
| 181 |
+
[Supra-50M-Base card](https://huggingface.co/SupraLabs/Supra-50M-Base). Their prompts may differ from ours,
|
| 182 |
+
so treat gaps of a few points as noise.
|
| 183 |
+
|
| 184 |
+
## What it can and can't do
|
| 185 |
+
|
| 186 |
+
At 0.5M active non-embedding parameters this model has learned **the form of English much better than its
|
| 187 |
+
content**. It's a research artifact for studying small MoE models, not a source of information.
|
| 188 |
+
|
| 189 |
+
**What it does**
|
| 190 |
+
|
| 191 |
+
- β
Writes grammatical, fluent sentences for a paragraph or two, in the register of its training data
|
| 192 |
+
(educational web text and textbooks): *"Photosynthesis is the process by which"* β *"β¦it is used to monitor
|
| 193 |
+
and measure the concentration of nutrients."*
|
| 194 |
+
- β
Stays roughly on topic: a prompt about the heart leads to blood vessels, the body and disease; a story
|
| 195 |
+
opening leads to characters and a plot.
|
| 196 |
+
- β
Picks up document structure from its data mix: markdown headings and worked "Example 1:" sections
|
| 197 |
+
after maths prompts, dialogue turns after `User:` / `Assistant:` text.
|
| 198 |
+
- β
Shows some benchmark signal above chance (SciQ 56.3%, ARC-Easy 32.7%, MMLU biology & medicine 28.4% in
|
| 199 |
+
the cloze format), mostly from matching answers to a passage or recognising which option sounds right,
|
| 200 |
+
not from recalling facts.
|
| 201 |
+
|
| 202 |
+
**What it can't do**
|
| 203 |
+
|
| 204 |
+
- β **Facts.** It reliably names the right kind of thing but the wrong thing: *"The French Revolution began
|
| 205 |
+
in"* β *"1906β¦"*; *"Albert Einstein was"* β *"a co-director of the French Marxi Palm."* Treat every name,
|
| 206 |
+
date and number it writes as made up.
|
| 207 |
+
- β **Arithmetic.** *"2 + 2 ="* β *"3"*, *"7 + 6 ="* β *"6"*.
|
| 208 |
+
- β **Code.** Code was 2% of its training data; given a Python function signature it produces empty
|
| 209 |
+
docstrings and whitespace.
|
| 210 |
+
- β **Instructions or questions.** It's a base model: it continues text rather than answering it, and with
|
| 211 |
+
only 0.5M non-embedding parameters an instruction-tuned version won't be a reliable assistant either.
|
| 212 |
+
- β **Long-range coherence.** Topics drift after a few sentences, and invented names appear
|
| 213 |
+
(*"the Grovi Palace"*).
|
| 214 |
+
|
| 215 |
+
**Decoding**
|
| 216 |
+
|
| 217 |
+
Greedy or low-temperature decoding falls into loops almost immediately (*"The process of the plant is formed
|
| 218 |
+
by the plant. The process of the plant is formed by the plantβ¦"*). Use sampling with **temperature β 0.7,
|
| 219 |
+
top-p β 0.9 and a repetition penalty of about 1.15**, which is what produced the examples above.
|
| 220 |
+
|
| 221 |
+
## Usage
|
| 222 |
+
|
| 223 |
+
**PyTorch** (needs the `littlekedi` package from [the training code](https://github.com/EdmundMartin/LittleKedi)):
|
| 224 |
+
|
| 225 |
+
```python
|
| 226 |
+
import json, torch
|
| 227 |
+
from safetensors.torch import load_file
|
| 228 |
+
from littlekedi import LittleKediModel, ModelConfig
|
| 229 |
+
from littlekedi.tokenizer import BPETokenizer
|
| 230 |
+
|
| 231 |
+
model = LittleKediModel(ModelConfig.from_dict(json.load(open("config.json")))).eval()
|
| 232 |
+
model.load_state_dict(load_file("model.safetensors"))
|
| 233 |
+
tok = BPETokenizer.from_file("tokenizer.json")
|
| 234 |
+
ids = torch.tensor([[tok.eot_id] + tok.encode("Photosynthesis is")])
|
| 235 |
+
print(tok.decode(model.generate(ids, 60, temperature=0.7, top_p=0.9, repetition_penalty=1.15)[0].tolist()))
|
| 236 |
+
```
|
| 237 |
+
|
| 238 |
+
Or interactively: `python sample.py --ckpt <checkpoint>.pt`. Base models are released as PyTorch weights only;
|
| 239 |
+
GGUF builds come with the instruct models.
|
| 240 |
+
|
| 241 |
+
## Files
|
| 242 |
+
|
| 243 |
+
| File | Contents |
|
| 244 |
+
|---|---|
|
| 245 |
+
| `model.safetensors` | PyTorch weights (embedding tied with the output head), `littlekedi` parameter names |
|
| 246 |
+
| `config.json` | `ModelConfig` for `littlekedi` |
|
| 247 |
+
| `tokenizer.json` | Byte-level BPE tokenizer (Hugging Face `tokenizers` format) |
|
| 248 |
+
| `little_kedi.jpg` | The LittleKedi logo |
|
| 249 |
+
|
| 250 |
+
## Acknowledgements
|
| 251 |
+
|
| 252 |
+
- The architecture follows **DeepSeek-V3** (DeepSeek-AI, 2024): MLA, DeepSeekMoE and auxiliary-loss-free balancing.
|
| 253 |
+
- **Muon** (Keller Jordan et al., 2024) with Moonlight's update scaling (Liu et al., 2025).
|
| 254 |
+
- Pretraining data: **FineWeb-Edu** (ODC-BY 1.0).
|
| 255 |
+
- Pretraining data: **Cosmopedia** (Apache 2.0).
|
| 256 |
+
- Pretraining data: **FineMath 4+** (ODC-BY 1.0).
|
| 257 |
+
- Pretraining data: **SmolTalk** (Apache 2.0 for its new subsets, with the datasets it includes under their own licenses).
|
| 258 |
+
- Pretraining data: **github-code-clean** (Apache 2.0, with each file under its own open-source license).
|
| 259 |
+
|
| 260 |
+
*Card generated 2026-10-09 by `prepare_release.py` from `tiny.pt` (step 15257).*
|
config.json
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"vocab_size": 8192,
|
| 3 |
+
"max_seq_len": 1024,
|
| 4 |
+
"dim": 128,
|
| 5 |
+
"n_layers": 4,
|
| 6 |
+
"n_dense_layers": 1,
|
| 7 |
+
"dense_hidden_dim": 352,
|
| 8 |
+
"norm_eps": 1e-06,
|
| 9 |
+
"init_std": 0.02,
|
| 10 |
+
"tie_embeddings": true,
|
| 11 |
+
"n_heads": 4,
|
| 12 |
+
"q_lora_rank": 0,
|
| 13 |
+
"kv_lora_rank": 32,
|
| 14 |
+
"qk_nope_head_dim": 16,
|
| 15 |
+
"qk_rope_head_dim": 16,
|
| 16 |
+
"v_head_dim": 16,
|
| 17 |
+
"rope_theta": 10000.0,
|
| 18 |
+
"n_routed_experts": 8,
|
| 19 |
+
"n_shared_experts": 1,
|
| 20 |
+
"n_activated_experts": 2,
|
| 21 |
+
"moe_hidden_dim": 64,
|
| 22 |
+
"n_expert_groups": 1,
|
| 23 |
+
"n_limited_groups": 1,
|
| 24 |
+
"route_scale": 1.0,
|
| 25 |
+
"bias_update_speed": 0.001,
|
| 26 |
+
"seq_aux_loss_alpha": 0.0001,
|
| 27 |
+
"moe_backend": "capacity",
|
| 28 |
+
"capacity_factor": 2.0
|
| 29 |
+
}
|
little_kedi.jpg
ADDED
|
Git LFS Details
|
model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:369319cb33abff7f4a0c178586a42aab0c6b77437fdcd2ba8392fe33e3264343
|
| 3 |
+
size 7969152
|
tokenizer.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|