Download README.md from Compactbot/testr-100k: direct link, hf CLI and curl.
- Browser
- Download file 4.45 kB
-
https://huggingface.co/Compactbot/testr-100k/resolve/main/README.md
- Command line
-
hf download hf://Compactbot/testr-100k/README.md
-
curl -L -o README.md https://huggingface.co/Compactbot/testr-100k/resolve/main/README.md
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
- tiny
- tiny-lm
- tiny-model
- slm
- small-language-model
- sub-1m
- from-scratch
- character-level
- char-gpt
metrics:
- perplexity
Testr-100K
A 100,936-parameter character-level GPT trained from scratch. Fulfils model request #17 from @GGUFGuy.
Correction (2026-09-30): the weights originally published here were trained on the wrong data (191 MB of web text). The request asked for a model trained only on the TinyStories validation set, as a deliberate memorization test. This commit swaps in the correct model (trained on the 19.2 MB TinyStories validation set) and fixes the parameter count (106,568 → 100,936; the earlier figure double-counted the tied head). The eval numbers below are from the correct model.
What it is
A minimal nanoGPT-style causal transformer operating at the character level (128-char vocabulary). The point of this specific model is a memorization experiment: train a very small model on a held-out validation set and measure how much it actually memorizes versus generalizing.
Architecture
| Parameter | Value |
|---|---|
| Layers | 4 |
| Embedding dim | 44 |
| Attention heads | 2 |
| Context length | 512 chars |
| FFN | GELU, mult 2.667 |
| Positional encoding | RoPE (θ=100000) |
| Normalization | LayerNorm (pre-norm) |
| Vocab | 128 chars + 1 UNK |
| Tied embeddings | Yes (head = embedding) |
| Total params (unique) | 100,936 |
(52 tensors in the file; the head is tied to the embedding, so counting it once gives 100,936. The earlier "106,568" double-counted the tied head.)
Training
- Data: 19.2 MB TinyStories validation set (character-level, uint8) — deliberately the set the request asked to train on
- Steps: 8,000
- Batch size: 32 × seq 512
- Optimizer: Muon (0.02) + AdamW (1e-4) for biases/norms
- Schedule: Cosine decay with warmup
- Hardware: RTX 5090 (32 GB)
- dtype: float32
The experiment (honest result)
The question: does a 100K-parameter model trained only on a 19 MB validation set actually memorize it, or does it just learn character statistics?
Measured perplexity (char-level, 512-char blocks, 20 random windows, seed 1337):
| Set | Size | Perplexity |
|---|---|---|
| Memorized (the 19.2 MB set it trained on) | 19,212,307 B | 3.1540 |
| Never seen (191 MB web text) | 191,283,131 B | 3.2203 |
Ratio (never-seen / memorized) = 1.02×. At 100K parameters the model barely memorizes the set it was trained on — its perplexity on the memorized data is only 2% lower than on data it never saw. It has essentially learned general English character statistics, not the specific stories. That is the honest answer to the request: a 100K-param char model is far too small to memorize 19 MB of text.
Sample outputs (greedy, temp=0)
Once upon a time→, there was a little girl named Lily. She liked to play with her mom and said, "I want to the bird was so happy and said
The sun was→so happy and said, "I want to the boy named Lily. They were so happy and said, "I want to the ball and said, "I want to
Sample outputs (temp=0.8)
A little→girl named Spot around of friends.\nOnce upon a time, there was a praye was so like in room that it was angry, the sweet
The first sentence or two is often grammatical; after that the model locks into a phrase and repeats it. Expected at this scale and vocabulary size.
What it is NOT
- Not a subword/token-level model (it cannot "read" word-level benchmarks like BLiMP or ARC)
- Not a coherent paragraph generator
- Not evidence of memorization at this scale (see ratio above)
Reproducing
Standard nanoGPT-style causal transformer with Muon optimizer. The checkpoint is saved as safetensors with tied embeddings (the file stores tok.weight and an identical head.weight; they are the same tensor).
To generate: load model.safetensors into a GPT class matching config.json, map characters to token IDs via vocab.json, and sample.
Files
| File | Description |
|---|---|
model.safetensors |
Model weights (431 KB, 52 tensors, tied head) |
config.json |
Architecture config |
vocab.json |
Character → token ID mapping (128 chars + UNK) |