File size: 3,866 Bytes
ccb8705 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 | ---
language:
- en
library_name: pytorch
pipeline_tag: text-generation
datasets:
- karpathy/tiny_shakespeare
tags:
- tiny
- character-level
- shakespeare
- safetensors
- joke
---
# micro500
A language model with **exactly 500 parameters**. Not 500 million, not 500 thousand. 500.
It was trained from scratch on a slice of Tiny Shakespeare. It will not write good
Shakespeare. It will write Shakespeare-shaped noise: the right letters in roughly the
right neighbourhoods, with the occasional real word ("the", "and", "to") surfacing by
accident.
## Model details
| | |
|---|---|
| Parameters | 500 (see note on padding below) |
| Architecture | Character-level Elman RNN with tied input/output embeddings |
| Vocabulary | 27 characters: space + `a-z` |
| Training data | First 200,000 characters of Tiny Shakespeare, lowercased, with all other characters mapped to space and whitespace collapsed |
| Training | 3,000 steps, AdamW, lr 3e-2, batch 64, sequence length 32 |
| File format | `safetensors` |
**Architecture:** embedding (27 x d) -> tanh RNN cell (hidden size h) -> linear projection
back to the embedding dimension -> logits computed against the transposed embedding matrix
(weight tying, so the output layer adds no extra weights) plus a per-character output bias.
The values of `d` and `h` are stored in `micro500_config.json`.
**Note on padding:** the architecture search picks the largest configuration that fits
within 500 parameters, then adds a small unused `pad` tensor to make the total exactly 500.
That tensor never touches the forward pass, so the *effective* parameter count is
`500 - pad`, where `pad` is recorded in the config file.
## Files
- `micro500.safetensors`: the weights
- `micro500_config.json`: vocabulary and layer sizes
- `model.py`: the model class
- `chat.py`: interactive prompt loop that reports characters (tokens) per second
## Usage
```bash
pip install torch safetensors
```
### Interactive
```bash
python chat.py
```
The model loads once and stays in memory. Type a prompt, get a continuation and a
tokens/sec figure. Exit with Ctrl+C.
### In Python
```python
import json
import torch
import torch.nn.functional as F
from safetensors.torch import load_file
from model import MicroLM
cfg = json.load(open("micro500_config.json"))
chars = cfg["chars"]
stoi = {c: i for i, c in enumerate(chars)}
itos = {i: c for c, i in stoi.items()}
net = MicroLM(len(chars), cfg["d"], cfg["h"], cfg["pad"])
net.load_state_dict(load_file("micro500.safetensors"))
net.eval()
@torch.no_grad()
def generate(prompt="the ", n=200, temp=0.8):
ids = [stoi.get(c, 0) for c in prompt.lower()]
for _ in range(n):
logits = net(torch.tensor([ids[-64:]]))[0, -1] / temp
ids.append(torch.multinomial(F.softmax(logits, -1), 1).item())
return "".join(itos[i] for i in ids)
print(generate("to be or "))
```
Lower temperature (0.5 to 0.8) gives more recognisable letter patterns; higher
temperature (1.0+) gives funnier chaos.
## Limitations
- Everything is lowercase a-z and spaces. No punctuation, digits, or capitals.
- It has no understanding of grammar, meaning, or anything else. It learned roughly which
letters tend to follow which.
- Prompts containing characters outside `a-z` are lowercased, and unknown characters are
mapped to the space token.
- It is a joke model. Please do not use it for anything that matters.
## Training data
Trained on [Tiny Shakespeare](https://github.com/karpathy/char-rnn/blob/master/data/tinyshakespeare/input.txt),
originally distributed with Andrej Karpathy's char-rnn. The Hugging Face dataset
`karpathy/tiny_shakespeare` is a loading script that points at this file.
## Citation
```bibtex
@misc{karpathy2015charrnn,
author = {Karpathy, Andrej},
title = {char-rnn},
year = {2015},
howpublished = {\url{https://github.com/karpathy/char-rnn}}
}
``` |