micro500 / README.md
ohnah's picture
make le readme
ccb8705 verified
|
Raw History Blame Contribute Delete
3.87 kB
---
language:
- en
library_name: pytorch
pipeline_tag: text-generation
datasets:
- karpathy/tiny_shakespeare
tags:
- tiny
- character-level
- shakespeare
- safetensors
- joke
---
# micro500
A language model with **exactly 500 parameters**. Not 500 million, not 500 thousand. 500.
It was trained from scratch on a slice of Tiny Shakespeare. It will not write good
Shakespeare. It will write Shakespeare-shaped noise: the right letters in roughly the
right neighbourhoods, with the occasional real word ("the", "and", "to") surfacing by
accident.
## Model details
| | |
|---|---|
| Parameters | 500 (see note on padding below) |
| Architecture | Character-level Elman RNN with tied input/output embeddings |
| Vocabulary | 27 characters: space + `a-z` |
| Training data | First 200,000 characters of Tiny Shakespeare, lowercased, with all other characters mapped to space and whitespace collapsed |
| Training | 3,000 steps, AdamW, lr 3e-2, batch 64, sequence length 32 |
| File format | `safetensors` |
**Architecture:** embedding (27 x d) -> tanh RNN cell (hidden size h) -> linear projection
back to the embedding dimension -> logits computed against the transposed embedding matrix
(weight tying, so the output layer adds no extra weights) plus a per-character output bias.
The values of `d` and `h` are stored in `micro500_config.json`.
**Note on padding:** the architecture search picks the largest configuration that fits
within 500 parameters, then adds a small unused `pad` tensor to make the total exactly 500.
That tensor never touches the forward pass, so the *effective* parameter count is
`500 - pad`, where `pad` is recorded in the config file.
## Files
- `micro500.safetensors`: the weights
- `micro500_config.json`: vocabulary and layer sizes
- `model.py`: the model class
- `chat.py`: interactive prompt loop that reports characters (tokens) per second
## Usage
```bash
pip install torch safetensors
```
### Interactive
```bash
python chat.py
```
The model loads once and stays in memory. Type a prompt, get a continuation and a
tokens/sec figure. Exit with Ctrl+C.
### In Python
```python
import json
import torch
import torch.nn.functional as F
from safetensors.torch import load_file
from model import MicroLM
cfg = json.load(open("micro500_config.json"))
chars = cfg["chars"]
stoi = {c: i for i, c in enumerate(chars)}
itos = {i: c for c, i in stoi.items()}
net = MicroLM(len(chars), cfg["d"], cfg["h"], cfg["pad"])
net.load_state_dict(load_file("micro500.safetensors"))
net.eval()
@torch.no_grad()
def generate(prompt="the ", n=200, temp=0.8):
ids = [stoi.get(c, 0) for c in prompt.lower()]
for _ in range(n):
logits = net(torch.tensor([ids[-64:]]))[0, -1] / temp
ids.append(torch.multinomial(F.softmax(logits, -1), 1).item())
return "".join(itos[i] for i in ids)
print(generate("to be or "))
```
Lower temperature (0.5 to 0.8) gives more recognisable letter patterns; higher
temperature (1.0+) gives funnier chaos.
## Limitations
- Everything is lowercase a-z and spaces. No punctuation, digits, or capitals.
- It has no understanding of grammar, meaning, or anything else. It learned roughly which
letters tend to follow which.
- Prompts containing characters outside `a-z` are lowercased, and unknown characters are
mapped to the space token.
- It is a joke model. Please do not use it for anything that matters.
## Training data
Trained on [Tiny Shakespeare](https://github.com/karpathy/char-rnn/blob/master/data/tinyshakespeare/input.txt),
originally distributed with Andrej Karpathy's char-rnn. The Hugging Face dataset
`karpathy/tiny_shakespeare` is a loading script that points at this file.
## Citation
```bibtex
@misc{karpathy2015charrnn,
author = {Karpathy, Andrej},
title = {char-rnn},
year = {2015},
howpublished = {\url{https://github.com/karpathy/char-rnn}}
}
```