testr-100k / README.md
Compactbot's picture
Correct the card: the model is now trained on the 19.2 MB TinyStories validation set (the actual #17 request). Fixes param count (100,936 unique; the earlier 106,568 double-counted the tied head) and reports the honest memorization result (ratio 1.02x, verified reproducible).
65f1844 verified
|
Raw History Blame Contribute Delete
4.45 kB
metadata
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
  - tiny
  - tiny-lm
  - tiny-model
  - slm
  - small-language-model
  - sub-1m
  - from-scratch
  - character-level
  - char-gpt
metrics:
  - perplexity

Testr-100K

A 100,936-parameter character-level GPT trained from scratch. Fulfils model request #17 from @GGUFGuy.

Correction (2026-09-30): the weights originally published here were trained on the wrong data (191 MB of web text). The request asked for a model trained only on the TinyStories validation set, as a deliberate memorization test. This commit swaps in the correct model (trained on the 19.2 MB TinyStories validation set) and fixes the parameter count (106,568 → 100,936; the earlier figure double-counted the tied head). The eval numbers below are from the correct model.

What it is

A minimal nanoGPT-style causal transformer operating at the character level (128-char vocabulary). The point of this specific model is a memorization experiment: train a very small model on a held-out validation set and measure how much it actually memorizes versus generalizing.

Architecture

Parameter Value
Layers 4
Embedding dim 44
Attention heads 2
Context length 512 chars
FFN GELU, mult 2.667
Positional encoding RoPE (θ=100000)
Normalization LayerNorm (pre-norm)
Vocab 128 chars + 1 UNK
Tied embeddings Yes (head = embedding)
Total params (unique) 100,936

(52 tensors in the file; the head is tied to the embedding, so counting it once gives 100,936. The earlier "106,568" double-counted the tied head.)

Training

  • Data: 19.2 MB TinyStories validation set (character-level, uint8) — deliberately the set the request asked to train on
  • Steps: 8,000
  • Batch size: 32 × seq 512
  • Optimizer: Muon (0.02) + AdamW (1e-4) for biases/norms
  • Schedule: Cosine decay with warmup
  • Hardware: RTX 5090 (32 GB)
  • dtype: float32

The experiment (honest result)

The question: does a 100K-parameter model trained only on a 19 MB validation set actually memorize it, or does it just learn character statistics?

Measured perplexity (char-level, 512-char blocks, 20 random windows, seed 1337):

Set Size Perplexity
Memorized (the 19.2 MB set it trained on) 19,212,307 B 3.1540
Never seen (191 MB web text) 191,283,131 B 3.2203

Ratio (never-seen / memorized) = 1.02×. At 100K parameters the model barely memorizes the set it was trained on — its perplexity on the memorized data is only 2% lower than on data it never saw. It has essentially learned general English character statistics, not the specific stories. That is the honest answer to the request: a 100K-param char model is far too small to memorize 19 MB of text.

Sample outputs (greedy, temp=0)

Once upon a time → , there was a little girl named Lily. She liked to play with her mom and said, "I want to the bird was so happy and said

The sun was → so happy and said, "I want to the boy named Lily. They were so happy and said, "I want to the ball and said, "I want to

Sample outputs (temp=0.8)

A little → girl named Spot around of friends.\nOnce upon a time, there was a praye was so like in room that it was angry, the sweet

The first sentence or two is often grammatical; after that the model locks into a phrase and repeats it. Expected at this scale and vocabulary size.

What it is NOT

  • Not a subword/token-level model (it cannot "read" word-level benchmarks like BLiMP or ARC)
  • Not a coherent paragraph generator
  • Not evidence of memorization at this scale (see ratio above)

Reproducing

Standard nanoGPT-style causal transformer with Muon optimizer. The checkpoint is saved as safetensors with tied embeddings (the file stores tok.weight and an identical head.weight; they are the same tensor).

To generate: load model.safetensors into a GPT class matching config.json, map characters to token IDs via vocab.json, and sample.

Files

File Description
model.safetensors Model weights (431 KB, 52 tensors, tied head)
config.json Architecture config
vocab.json Character → token ID mapping (128 chars + UNK)