File size: 3,866 Bytes
ccb8705
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
---
language:
- en
library_name: pytorch
pipeline_tag: text-generation
datasets:
- karpathy/tiny_shakespeare
tags:
- tiny
- character-level
- shakespeare
- safetensors
- joke
---

# micro500

A language model with **exactly 500 parameters**. Not 500 million, not 500 thousand. 500.

It was trained from scratch on a slice of Tiny Shakespeare. It will not write good
Shakespeare. It will write Shakespeare-shaped noise: the right letters in roughly the
right neighbourhoods, with the occasional real word ("the", "and", "to") surfacing by
accident.

## Model details

| | |
|---|---|
| Parameters | 500 (see note on padding below) |
| Architecture | Character-level Elman RNN with tied input/output embeddings |
| Vocabulary | 27 characters: space + `a-z` |
| Training data | First 200,000 characters of Tiny Shakespeare, lowercased, with all other characters mapped to space and whitespace collapsed |
| Training | 3,000 steps, AdamW, lr 3e-2, batch 64, sequence length 32 |
| File format | `safetensors` |

**Architecture:** embedding (27 x d) -> tanh RNN cell (hidden size h) -> linear projection
back to the embedding dimension -> logits computed against the transposed embedding matrix
(weight tying, so the output layer adds no extra weights) plus a per-character output bias.
The values of `d` and `h` are stored in `micro500_config.json`.

**Note on padding:** the architecture search picks the largest configuration that fits
within 500 parameters, then adds a small unused `pad` tensor to make the total exactly 500.
That tensor never touches the forward pass, so the *effective* parameter count is
`500 - pad`, where `pad` is recorded in the config file.

## Files

- `micro500.safetensors`: the weights
- `micro500_config.json`: vocabulary and layer sizes
- `model.py`: the model class
- `chat.py`: interactive prompt loop that reports characters (tokens) per second

## Usage

```bash
pip install torch safetensors
```

### Interactive

```bash
python chat.py
```

The model loads once and stays in memory. Type a prompt, get a continuation and a
tokens/sec figure. Exit with Ctrl+C.

### In Python

```python
import json
import torch
import torch.nn.functional as F
from safetensors.torch import load_file
from model import MicroLM

cfg = json.load(open("micro500_config.json"))
chars = cfg["chars"]
stoi = {c: i for i, c in enumerate(chars)}
itos = {i: c for c, i in stoi.items()}

net = MicroLM(len(chars), cfg["d"], cfg["h"], cfg["pad"])
net.load_state_dict(load_file("micro500.safetensors"))
net.eval()

@torch.no_grad()
def generate(prompt="the ", n=200, temp=0.8):
    ids = [stoi.get(c, 0) for c in prompt.lower()]
    for _ in range(n):
        logits = net(torch.tensor([ids[-64:]]))[0, -1] / temp
        ids.append(torch.multinomial(F.softmax(logits, -1), 1).item())
    return "".join(itos[i] for i in ids)

print(generate("to be or "))
```

Lower temperature (0.5 to 0.8) gives more recognisable letter patterns; higher
temperature (1.0+) gives funnier chaos.

## Limitations

- Everything is lowercase a-z and spaces. No punctuation, digits, or capitals.
- It has no understanding of grammar, meaning, or anything else. It learned roughly which
  letters tend to follow which.
- Prompts containing characters outside `a-z` are lowercased, and unknown characters are
  mapped to the space token.
- It is a joke model. Please do not use it for anything that matters.

## Training data

Trained on [Tiny Shakespeare](https://github.com/karpathy/char-rnn/blob/master/data/tinyshakespeare/input.txt),
originally distributed with Andrej Karpathy's char-rnn. The Hugging Face dataset
`karpathy/tiny_shakespeare` is a loading script that points at this file.

## Citation

```bibtex
@misc{karpathy2015charrnn,
  author = {Karpathy, Andrej},
  title = {char-rnn},
  year = {2015},
  howpublished = {\url{https://github.com/karpathy/char-rnn}}
}
```