micro500

A language model with exactly 500 parameters. Not 500 million, not 500 thousand. 500.

It was trained from scratch on a slice of Tiny Shakespeare. It will not write good Shakespeare. It will write Shakespeare-shaped noise: the right letters in roughly the right neighbourhoods, with the occasional real word ("the", "and", "to") surfacing by accident.

Model details

Parameters 500 (see note on padding below)
Architecture Character-level Elman RNN with tied input/output embeddings
Vocabulary 27 characters: space + a-z
Training data First 200,000 characters of Tiny Shakespeare, lowercased, with all other characters mapped to space and whitespace collapsed
Training 3,000 steps, AdamW, lr 3e-2, batch 64, sequence length 32
File format safetensors

Architecture: embedding (27 x d) -> tanh RNN cell (hidden size h) -> linear projection back to the embedding dimension -> logits computed against the transposed embedding matrix (weight tying, so the output layer adds no extra weights) plus a per-character output bias. The values of d and h are stored in micro500_config.json.

Note on padding: the architecture search picks the largest configuration that fits within 500 parameters, then adds a small unused pad tensor to make the total exactly 500. That tensor never touches the forward pass, so the effective parameter count is 500 - pad, where pad is recorded in the config file.

Files

  • micro500.safetensors: the weights
  • micro500_config.json: vocabulary and layer sizes
  • model.py: the model class
  • chat.py: interactive prompt loop that reports characters (tokens) per second

Usage

pip install torch safetensors

Interactive

python chat.py

The model loads once and stays in memory. Type a prompt, get a continuation and a tokens/sec figure. Exit with Ctrl+C.

In Python

import json
import torch
import torch.nn.functional as F
from safetensors.torch import load_file
from model import MicroLM

cfg = json.load(open("micro500_config.json"))
chars = cfg["chars"]
stoi = {c: i for i, c in enumerate(chars)}
itos = {i: c for c, i in stoi.items()}

net = MicroLM(len(chars), cfg["d"], cfg["h"], cfg["pad"])
net.load_state_dict(load_file("micro500.safetensors"))
net.eval()

@torch.no_grad()
def generate(prompt="the ", n=200, temp=0.8):
    ids = [stoi.get(c, 0) for c in prompt.lower()]
    for _ in range(n):
        logits = net(torch.tensor([ids[-64:]]))[0, -1] / temp
        ids.append(torch.multinomial(F.softmax(logits, -1), 1).item())
    return "".join(itos[i] for i in ids)

print(generate("to be or "))

Lower temperature (0.5 to 0.8) gives more recognisable letter patterns; higher temperature (1.0+) gives funnier chaos.

Limitations

  • Everything is lowercase a-z and spaces. No punctuation, digits, or capitals.
  • It has no understanding of grammar, meaning, or anything else. It learned roughly which letters tend to follow which.
  • Prompts containing characters outside a-z are lowercased, and unknown characters are mapped to the space token.
  • It is a joke model. Please do not use it for anything that matters.

Training data

Trained on Tiny Shakespeare, originally distributed with Andrej Karpathy's char-rnn. The Hugging Face dataset karpathy/tiny_shakespeare is a loading script that points at this file.

Citation

@misc{karpathy2015charrnn,
  author = {Karpathy, Andrej},
  title = {char-rnn},
  year = {2015},
  howpublished = {\url{https://github.com/karpathy/char-rnn}}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train ohnah/micro500