Correct the card: the model is now trained on the 19.2 MB TinyStories validation set (the actual #17 request). Fixes param count (100,936 unique; the earlier 106,568 double-counted the tied head) and reports the honest memorization result (ratio 1.02x, verified reproducible).
65f1844 verified |
Download README.md from Compactbot/testr-100k: direct link, hf CLI and curl.
- Browser
- Download file 4.45 kB
-
https://huggingface.co/Compactbot/testr-100k/resolve/main/README.md
- Command line
-
hf download hf://Compactbot/testr-100k/README.md
-
curl -L -o README.md https://huggingface.co/Compactbot/testr-100k/resolve/main/README.md
4.45 kB
| license: apache-2.0 | |
| pipeline_tag: text-generation | |
| language: en | |
| tags: | |
| - tiny | |
| - tiny-lm | |
| - tiny-model | |
| - slm | |
| - small-language-model | |
| - sub-1m | |
| - from-scratch | |
| - character-level | |
| - char-gpt | |
| metrics: | |
| - perplexity | |
| # Testr-100K | |
| A **100,936-parameter** character-level GPT trained from scratch. Fulfils model request [#17](https://huggingface.co/spaces/Compactbot/model-requests/discussions/17) from @GGUFGuy. | |
| > **Correction (2026-09-30):** the weights originally published here were trained on the **wrong data** (191 MB of web text). The request asked for a model trained **only** on the TinyStories *validation* set, as a deliberate memorization test. This commit swaps in the correct model (trained on the 19.2 MB TinyStories validation set) and fixes the parameter count (106,568 β 100,936; the earlier figure double-counted the tied head). The eval numbers below are from the correct model. | |
| ## What it is | |
| A minimal nanoGPT-style causal transformer operating at the **character level** (128-char vocabulary). The point of this specific model is a **memorization experiment**: train a very small model on a held-out validation set and measure how much it actually memorizes versus generalizing. | |
| ## Architecture | |
| | Parameter | Value | | |
| |-----------|-------| | |
| | Layers | 4 | | |
| | Embedding dim | 44 | | |
| | Attention heads | 2 | | |
| | Context length | 512 chars | | |
| | FFN | GELU, mult 2.667 | | |
| | Positional encoding | RoPE (ΞΈ=100000) | | |
| | Normalization | LayerNorm (pre-norm) | | |
| | Vocab | 128 chars + 1 UNK | | |
| | Tied embeddings | Yes (head = embedding) | | |
| | **Total params (unique)** | **100,936** | | |
| (52 tensors in the file; the head is tied to the embedding, so counting it once gives 100,936. The earlier "106,568" double-counted the tied head.) | |
| ## Training | |
| - **Data:** 19.2 MB **TinyStories validation set** (character-level, uint8) β *deliberately* the set the request asked to train on | |
| - **Steps:** 8,000 | |
| - **Batch size:** 32 Γ seq 512 | |
| - **Optimizer:** Muon (0.02) + AdamW (1e-4) for biases/norms | |
| - **Schedule:** Cosine decay with warmup | |
| - **Hardware:** RTX 5090 (32 GB) | |
| - **dtype:** float32 | |
| ## The experiment (honest result) | |
| The question: does a 100K-parameter model trained *only* on a 19 MB validation set actually memorize it, or does it just learn character statistics? | |
| Measured perplexity (char-level, 512-char blocks, 20 random windows, seed 1337): | |
| | Set | Size | Perplexity | | |
| |-----|------|-----------| | |
| | **Memorized** (the 19.2 MB set it trained on) | 19,212,307 B | **3.1540** | | |
| | **Never seen** (191 MB web text) | 191,283,131 B | **3.2203** | | |
| **Ratio (never-seen / memorized) = 1.02Γ.** At 100K parameters the model **barely memorizes** the set it was trained on β its perplexity on the memorized data is only 2% lower than on data it never saw. It has essentially learned general English character statistics, not the specific stories. That is the honest answer to the request: a 100K-param char model is far too small to memorize 19 MB of text. | |
| ## Sample outputs (greedy, temp=0) | |
| > `Once upon a time` β `, there was a little girl named Lily. She liked to play with her mom and said, "I want to the bird was so happy and said` | |
| > `The sun was` β ` so happy and said, "I want to the boy named Lily. They were so happy and said, "I want to the ball and said, "I want to` | |
| ### Sample outputs (temp=0.8) | |
| > `A little` β ` girl named Spot around of friends.\nOnce upon a time, there was a praye was so like in room that it was angry, the sweet` | |
| The first sentence or two is often grammatical; after that the model locks into a phrase and repeats it. Expected at this scale and vocabulary size. | |
| ## What it is NOT | |
| - Not a subword/token-level model (it cannot "read" word-level benchmarks like BLiMP or ARC) | |
| - Not a coherent paragraph generator | |
| - Not evidence of memorization at this scale (see ratio above) | |
| ## Reproducing | |
| Standard nanoGPT-style causal transformer with Muon optimizer. The checkpoint is saved as safetensors with tied embeddings (the file stores `tok.weight` and an identical `head.weight`; they are the same tensor). | |
| To generate: load `model.safetensors` into a GPT class matching `config.json`, map characters to token IDs via `vocab.json`, and sample. | |
| ## Files | |
| | File | Description | | |
| |------|-------------| | |
| | `model.safetensors` | Model weights (431 KB, 52 tensors, tied head) | | |
| | `config.json` | Architecture config | | |
| | `vocab.json` | Character β token ID mapping (128 chars + UNK) | |