micro500
A language model with exactly 500 parameters. Not 500 million, not 500 thousand. 500.
It was trained from scratch on a slice of Tiny Shakespeare. It will not write good Shakespeare. It will write Shakespeare-shaped noise: the right letters in roughly the right neighbourhoods, with the occasional real word ("the", "and", "to") surfacing by accident.
Model details
| Parameters | 500 (see note on padding below) |
| Architecture | Character-level Elman RNN with tied input/output embeddings |
| Vocabulary | 27 characters: space + a-z |
| Training data | First 200,000 characters of Tiny Shakespeare, lowercased, with all other characters mapped to space and whitespace collapsed |
| Training | 3,000 steps, AdamW, lr 3e-2, batch 64, sequence length 32 |
| File format | safetensors |
Architecture: embedding (27 x d) -> tanh RNN cell (hidden size h) -> linear projection
back to the embedding dimension -> logits computed against the transposed embedding matrix
(weight tying, so the output layer adds no extra weights) plus a per-character output bias.
The values of d and h are stored in micro500_config.json.
Note on padding: the architecture search picks the largest configuration that fits
within 500 parameters, then adds a small unused pad tensor to make the total exactly 500.
That tensor never touches the forward pass, so the effective parameter count is
500 - pad, where pad is recorded in the config file.
Files
micro500.safetensors: the weightsmicro500_config.json: vocabulary and layer sizesmodel.py: the model classchat.py: interactive prompt loop that reports characters (tokens) per second
Usage
pip install torch safetensors
Interactive
python chat.py
The model loads once and stays in memory. Type a prompt, get a continuation and a tokens/sec figure. Exit with Ctrl+C.
In Python
import json
import torch
import torch.nn.functional as F
from safetensors.torch import load_file
from model import MicroLM
cfg = json.load(open("micro500_config.json"))
chars = cfg["chars"]
stoi = {c: i for i, c in enumerate(chars)}
itos = {i: c for c, i in stoi.items()}
net = MicroLM(len(chars), cfg["d"], cfg["h"], cfg["pad"])
net.load_state_dict(load_file("micro500.safetensors"))
net.eval()
@torch.no_grad()
def generate(prompt="the ", n=200, temp=0.8):
ids = [stoi.get(c, 0) for c in prompt.lower()]
for _ in range(n):
logits = net(torch.tensor([ids[-64:]]))[0, -1] / temp
ids.append(torch.multinomial(F.softmax(logits, -1), 1).item())
return "".join(itos[i] for i in ids)
print(generate("to be or "))
Lower temperature (0.5 to 0.8) gives more recognisable letter patterns; higher temperature (1.0+) gives funnier chaos.
Limitations
- Everything is lowercase a-z and spaces. No punctuation, digits, or capitals.
- It has no understanding of grammar, meaning, or anything else. It learned roughly which letters tend to follow which.
- Prompts containing characters outside
a-zare lowercased, and unknown characters are mapped to the space token. - It is a joke model. Please do not use it for anything that matters.
Training data
Trained on Tiny Shakespeare,
originally distributed with Andrej Karpathy's char-rnn. The Hugging Face dataset
karpathy/tiny_shakespeare is a loading script that points at this file.
Citation
@misc{karpathy2015charrnn,
author = {Karpathy, Andrej},
title = {char-rnn},
year = {2015},
howpublished = {\url{https://github.com/karpathy/char-rnn}}
}