|
Download README.md from ohnah/micro500: direct link, hf CLI and curl.
- Browser
- Download file 3.87 kB
-
https://huggingface.co/ohnah/micro500/resolve/main/README.md
- Command line
-
hf download hf://ohnah/micro500/README.md
-
curl -L -o README.md https://huggingface.co/ohnah/micro500/resolve/main/README.md
3.87 kB
| language: | |
| - en | |
| library_name: pytorch | |
| pipeline_tag: text-generation | |
| datasets: | |
| - karpathy/tiny_shakespeare | |
| tags: | |
| - tiny | |
| - character-level | |
| - shakespeare | |
| - safetensors | |
| - joke | |
| # micro500 | |
| A language model with **exactly 500 parameters**. Not 500 million, not 500 thousand. 500. | |
| It was trained from scratch on a slice of Tiny Shakespeare. It will not write good | |
| Shakespeare. It will write Shakespeare-shaped noise: the right letters in roughly the | |
| right neighbourhoods, with the occasional real word ("the", "and", "to") surfacing by | |
| accident. | |
| ## Model details | |
| | | | | |
| |---|---| | |
| | Parameters | 500 (see note on padding below) | | |
| | Architecture | Character-level Elman RNN with tied input/output embeddings | | |
| | Vocabulary | 27 characters: space + `a-z` | | |
| | Training data | First 200,000 characters of Tiny Shakespeare, lowercased, with all other characters mapped to space and whitespace collapsed | | |
| | Training | 3,000 steps, AdamW, lr 3e-2, batch 64, sequence length 32 | | |
| | File format | `safetensors` | | |
| **Architecture:** embedding (27 x d) -> tanh RNN cell (hidden size h) -> linear projection | |
| back to the embedding dimension -> logits computed against the transposed embedding matrix | |
| (weight tying, so the output layer adds no extra weights) plus a per-character output bias. | |
| The values of `d` and `h` are stored in `micro500_config.json`. | |
| **Note on padding:** the architecture search picks the largest configuration that fits | |
| within 500 parameters, then adds a small unused `pad` tensor to make the total exactly 500. | |
| That tensor never touches the forward pass, so the *effective* parameter count is | |
| `500 - pad`, where `pad` is recorded in the config file. | |
| ## Files | |
| - `micro500.safetensors`: the weights | |
| - `micro500_config.json`: vocabulary and layer sizes | |
| - `model.py`: the model class | |
| - `chat.py`: interactive prompt loop that reports characters (tokens) per second | |
| ## Usage | |
| ```bash | |
| pip install torch safetensors | |
| ``` | |
| ### Interactive | |
| ```bash | |
| python chat.py | |
| ``` | |
| The model loads once and stays in memory. Type a prompt, get a continuation and a | |
| tokens/sec figure. Exit with Ctrl+C. | |
| ### In Python | |
| ```python | |
| import json | |
| import torch | |
| import torch.nn.functional as F | |
| from safetensors.torch import load_file | |
| from model import MicroLM | |
| cfg = json.load(open("micro500_config.json")) | |
| chars = cfg["chars"] | |
| stoi = {c: i for i, c in enumerate(chars)} | |
| itos = {i: c for c, i in stoi.items()} | |
| net = MicroLM(len(chars), cfg["d"], cfg["h"], cfg["pad"]) | |
| net.load_state_dict(load_file("micro500.safetensors")) | |
| net.eval() | |
| @torch.no_grad() | |
| def generate(prompt="the ", n=200, temp=0.8): | |
| ids = [stoi.get(c, 0) for c in prompt.lower()] | |
| for _ in range(n): | |
| logits = net(torch.tensor([ids[-64:]]))[0, -1] / temp | |
| ids.append(torch.multinomial(F.softmax(logits, -1), 1).item()) | |
| return "".join(itos[i] for i in ids) | |
| print(generate("to be or ")) | |
| ``` | |
| Lower temperature (0.5 to 0.8) gives more recognisable letter patterns; higher | |
| temperature (1.0+) gives funnier chaos. | |
| ## Limitations | |
| - Everything is lowercase a-z and spaces. No punctuation, digits, or capitals. | |
| - It has no understanding of grammar, meaning, or anything else. It learned roughly which | |
| letters tend to follow which. | |
| - Prompts containing characters outside `a-z` are lowercased, and unknown characters are | |
| mapped to the space token. | |
| - It is a joke model. Please do not use it for anything that matters. | |
| ## Training data | |
| Trained on [Tiny Shakespeare](https://github.com/karpathy/char-rnn/blob/master/data/tinyshakespeare/input.txt), | |
| originally distributed with Andrej Karpathy's char-rnn. The Hugging Face dataset | |
| `karpathy/tiny_shakespeare` is a loading script that points at this file. | |
| ## Citation | |
| ```bibtex | |
| @misc{karpathy2015charrnn, | |
| author = {Karpathy, Andrej}, | |
| title = {char-rnn}, | |
| year = {2015}, | |
| howpublished = {\url{https://github.com/karpathy/char-rnn}} | |
| } | |
| ``` |