--- language: - en library_name: pytorch pipeline_tag: text-generation datasets: - karpathy/tiny_shakespeare tags: - tiny - character-level - shakespeare - safetensors - joke --- # micro500 A language model with **exactly 500 parameters**. Not 500 million, not 500 thousand. 500. It was trained from scratch on a slice of Tiny Shakespeare. It will not write good Shakespeare. It will write Shakespeare-shaped noise: the right letters in roughly the right neighbourhoods, with the occasional real word ("the", "and", "to") surfacing by accident. ## Model details | | | |---|---| | Parameters | 500 (see note on padding below) | | Architecture | Character-level Elman RNN with tied input/output embeddings | | Vocabulary | 27 characters: space + `a-z` | | Training data | First 200,000 characters of Tiny Shakespeare, lowercased, with all other characters mapped to space and whitespace collapsed | | Training | 3,000 steps, AdamW, lr 3e-2, batch 64, sequence length 32 | | File format | `safetensors` | **Architecture:** embedding (27 x d) -> tanh RNN cell (hidden size h) -> linear projection back to the embedding dimension -> logits computed against the transposed embedding matrix (weight tying, so the output layer adds no extra weights) plus a per-character output bias. The values of `d` and `h` are stored in `micro500_config.json`. **Note on padding:** the architecture search picks the largest configuration that fits within 500 parameters, then adds a small unused `pad` tensor to make the total exactly 500. That tensor never touches the forward pass, so the *effective* parameter count is `500 - pad`, where `pad` is recorded in the config file. ## Files - `micro500.safetensors`: the weights - `micro500_config.json`: vocabulary and layer sizes - `model.py`: the model class - `chat.py`: interactive prompt loop that reports characters (tokens) per second ## Usage ```bash pip install torch safetensors ``` ### Interactive ```bash python chat.py ``` The model loads once and stays in memory. Type a prompt, get a continuation and a tokens/sec figure. Exit with Ctrl+C. ### In Python ```python import json import torch import torch.nn.functional as F from safetensors.torch import load_file from model import MicroLM cfg = json.load(open("micro500_config.json")) chars = cfg["chars"] stoi = {c: i for i, c in enumerate(chars)} itos = {i: c for c, i in stoi.items()} net = MicroLM(len(chars), cfg["d"], cfg["h"], cfg["pad"]) net.load_state_dict(load_file("micro500.safetensors")) net.eval() @torch.no_grad() def generate(prompt="the ", n=200, temp=0.8): ids = [stoi.get(c, 0) for c in prompt.lower()] for _ in range(n): logits = net(torch.tensor([ids[-64:]]))[0, -1] / temp ids.append(torch.multinomial(F.softmax(logits, -1), 1).item()) return "".join(itos[i] for i in ids) print(generate("to be or ")) ``` Lower temperature (0.5 to 0.8) gives more recognisable letter patterns; higher temperature (1.0+) gives funnier chaos. ## Limitations - Everything is lowercase a-z and spaces. No punctuation, digits, or capitals. - It has no understanding of grammar, meaning, or anything else. It learned roughly which letters tend to follow which. - Prompts containing characters outside `a-z` are lowercased, and unknown characters are mapped to the space token. - It is a joke model. Please do not use it for anything that matters. ## Training data Trained on [Tiny Shakespeare](https://github.com/karpathy/char-rnn/blob/master/data/tinyshakespeare/input.txt), originally distributed with Andrej Karpathy's char-rnn. The Hugging Face dataset `karpathy/tiny_shakespeare` is a loading script that points at this file. ## Citation ```bibtex @misc{karpathy2015charrnn, author = {Karpathy, Andrej}, title = {char-rnn}, year = {2015}, howpublished = {\url{https://github.com/karpathy/char-rnn}} } ```