File size: 4,452 Bytes
8c85802 65f1844 b679bdc 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 65f1844 8c85802 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 | ---
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
- tiny
- tiny-lm
- tiny-model
- slm
- small-language-model
- sub-1m
- from-scratch
- character-level
- char-gpt
metrics:
- perplexity
---
# Testr-100K
A **100,936-parameter** character-level GPT trained from scratch. Fulfils model request [#17](https://huggingface.co/spaces/Compactbot/model-requests/discussions/17) from @GGUFGuy.
> **Correction (2026-09-30):** the weights originally published here were trained on the **wrong data** (191 MB of web text). The request asked for a model trained **only** on the TinyStories *validation* set, as a deliberate memorization test. This commit swaps in the correct model (trained on the 19.2 MB TinyStories validation set) and fixes the parameter count (106,568 → 100,936; the earlier figure double-counted the tied head). The eval numbers below are from the correct model.
## What it is
A minimal nanoGPT-style causal transformer operating at the **character level** (128-char vocabulary). The point of this specific model is a **memorization experiment**: train a very small model on a held-out validation set and measure how much it actually memorizes versus generalizing.
## Architecture
| Parameter | Value |
|-----------|-------|
| Layers | 4 |
| Embedding dim | 44 |
| Attention heads | 2 |
| Context length | 512 chars |
| FFN | GELU, mult 2.667 |
| Positional encoding | RoPE (θ=100000) |
| Normalization | LayerNorm (pre-norm) |
| Vocab | 128 chars + 1 UNK |
| Tied embeddings | Yes (head = embedding) |
| **Total params (unique)** | **100,936** |
(52 tensors in the file; the head is tied to the embedding, so counting it once gives 100,936. The earlier "106,568" double-counted the tied head.)
## Training
- **Data:** 19.2 MB **TinyStories validation set** (character-level, uint8) — *deliberately* the set the request asked to train on
- **Steps:** 8,000
- **Batch size:** 32 × seq 512
- **Optimizer:** Muon (0.02) + AdamW (1e-4) for biases/norms
- **Schedule:** Cosine decay with warmup
- **Hardware:** RTX 5090 (32 GB)
- **dtype:** float32
## The experiment (honest result)
The question: does a 100K-parameter model trained *only* on a 19 MB validation set actually memorize it, or does it just learn character statistics?
Measured perplexity (char-level, 512-char blocks, 20 random windows, seed 1337):
| Set | Size | Perplexity |
|-----|------|-----------|
| **Memorized** (the 19.2 MB set it trained on) | 19,212,307 B | **3.1540** |
| **Never seen** (191 MB web text) | 191,283,131 B | **3.2203** |
**Ratio (never-seen / memorized) = 1.02×.** At 100K parameters the model **barely memorizes** the set it was trained on — its perplexity on the memorized data is only 2% lower than on data it never saw. It has essentially learned general English character statistics, not the specific stories. That is the honest answer to the request: a 100K-param char model is far too small to memorize 19 MB of text.
## Sample outputs (greedy, temp=0)
> `Once upon a time` → `, there was a little girl named Lily. She liked to play with her mom and said, "I want to the bird was so happy and said`
> `The sun was` → ` so happy and said, "I want to the boy named Lily. They were so happy and said, "I want to the ball and said, "I want to`
### Sample outputs (temp=0.8)
> `A little` → ` girl named Spot around of friends.\nOnce upon a time, there was a praye was so like in room that it was angry, the sweet`
The first sentence or two is often grammatical; after that the model locks into a phrase and repeats it. Expected at this scale and vocabulary size.
## What it is NOT
- Not a subword/token-level model (it cannot "read" word-level benchmarks like BLiMP or ARC)
- Not a coherent paragraph generator
- Not evidence of memorization at this scale (see ratio above)
## Reproducing
Standard nanoGPT-style causal transformer with Muon optimizer. The checkpoint is saved as safetensors with tied embeddings (the file stores `tok.weight` and an identical `head.weight`; they are the same tensor).
To generate: load `model.safetensors` into a GPT class matching `config.json`, map characters to token IDs via `vocab.json`, and sample.
## Files
| File | Description |
|------|-------------|
| `model.safetensors` | Model weights (431 KB, 52 tensors, tied head) |
| `config.json` | Architecture config |
| `vocab.json` | Character → token ID mapping (128 chars + UNK) | |