list the temperature-trained runs
Browse files
README.md
CHANGED
|
@@ -1,54 +1,55 @@
|
|
| 1 |
-
---
|
| 2 |
-
license: mit
|
| 3 |
-
tags:
|
| 4 |
-
- arithmetic
|
| 5 |
-
- length-generalization
|
| 6 |
-
- grokking
|
| 7 |
-
- interpretability
|
| 8 |
-
---
|
| 9 |
-
|
| 10 |
-
# carrybit checkpoints
|
| 11 |
-
|
| 12 |
-
Trained weights from [carrybit](https://github.com/Supergoatscriptguy/carrybit),
|
| 13 |
-
a small research project on tiny transformers learning exact integer
|
| 14 |
-
arithmetic. The code, configs, figures and the full write-up live in that
|
| 15 |
-
repo. This repo holds the final checkpoint of every run in the write-up so the
|
| 16 |
-
analysis experiments can be run without retraining.
|
| 17 |
-
|
| 18 |
-
Every checkpoint is a plain PyTorch `state_dict` for `carrybit.model.Transformer`
|
| 19 |
-
(or `TwoHotMLP` for `modular_mlp`). Each folder has the exact `config.json` it
|
| 20 |
-
was trained with and the `metrics.csv` logged during training. Folder names
|
| 21 |
-
match `runs/` in the GitHub repo, so every experiment script there works on
|
| 22 |
-
these files as they are.
|
| 23 |
-
|
| 24 |
-
```python
|
| 25 |
-
import json, torch
|
| 26 |
-
from carrybit.config import load_config
|
| 27 |
-
from carrybit.model import Transformer
|
| 28 |
-
|
| 29 |
-
run = "addition_big/position_coupling_s1"
|
| 30 |
-
raw = json.load(open(f"{run}/config.json"))
|
| 31 |
-
cfg = load_config("configs/addition_big.yaml",
|
| 32 |
-
[f"{s}.{k}={json.dumps(v)}" for s in ("task", "model")
|
| 33 |
-
for k, v in raw[s].items() if k != "kind"])
|
| 34 |
-
model = Transformer(16, cfg.model)
|
| 35 |
-
model.load_state_dict(torch.load(f"{run}/step_60000.pt"))
|
| 36 |
-
```
|
| 37 |
-
|
| 38 |
-
## Contents
|
| 39 |
-
|
| 40 |
-
| folder | what | runs |
|
| 41 |
-
|---|---|---|
|
| 42 |
-
| `modular_add`, `modular_add_wd0.1`, `modular_add_wd0` | one-layer transformer, a + b mod 113, weight decay 1 / 0.1 / 0 | 1 each |
|
| 43 |
-
| `modular_mlp` | ReLU MLP on two-hot inputs, p = 97 (Swaroop 2026 setup) | 1 |
|
| 44 |
-
| `addition` | the length ladder: 3.4M params, trained on 1 to 20 digits, nine formats | 3 seeds each, 6 for position coupling |
|
| 45 |
-
| `addition_big` | 11M params, abacus and position coupling, trained on 1 to 30 digits | 2 seeds
|
| 46 |
-
| `
|
| 47 |
-
| `
|
| 48 |
-
| `
|
| 49 |
-
| `
|
| 50 |
-
|
| 51 |
-
|
| 52 |
-
|
| 53 |
-
|
| 54 |
-
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
tags:
|
| 4 |
+
- arithmetic
|
| 5 |
+
- length-generalization
|
| 6 |
+
- grokking
|
| 7 |
+
- interpretability
|
| 8 |
+
---
|
| 9 |
+
|
| 10 |
+
# carrybit checkpoints
|
| 11 |
+
|
| 12 |
+
Trained weights from [carrybit](https://github.com/Supergoatscriptguy/carrybit),
|
| 13 |
+
a small research project on tiny transformers learning exact integer
|
| 14 |
+
arithmetic. The code, configs, figures and the full write-up live in that
|
| 15 |
+
repo. This repo holds the final checkpoint of every run in the write-up so the
|
| 16 |
+
analysis experiments can be run without retraining.
|
| 17 |
+
|
| 18 |
+
Every checkpoint is a plain PyTorch `state_dict` for `carrybit.model.Transformer`
|
| 19 |
+
(or `TwoHotMLP` for `modular_mlp`). Each folder has the exact `config.json` it
|
| 20 |
+
was trained with and the `metrics.csv` logged during training. Folder names
|
| 21 |
+
match `runs/` in the GitHub repo, so every experiment script there works on
|
| 22 |
+
these files as they are.
|
| 23 |
+
|
| 24 |
+
```python
|
| 25 |
+
import json, torch
|
| 26 |
+
from carrybit.config import load_config
|
| 27 |
+
from carrybit.model import Transformer
|
| 28 |
+
|
| 29 |
+
run = "addition_big/position_coupling_s1"
|
| 30 |
+
raw = json.load(open(f"{run}/config.json"))
|
| 31 |
+
cfg = load_config("configs/addition_big.yaml",
|
| 32 |
+
[f"{s}.{k}={json.dumps(v)}" for s in ("task", "model")
|
| 33 |
+
for k, v in raw[s].items() if k != "kind"])
|
| 34 |
+
model = Transformer(16, cfg.model)
|
| 35 |
+
model.load_state_dict(torch.load(f"{run}/step_60000.pt"))
|
| 36 |
+
```
|
| 37 |
+
|
| 38 |
+
## Contents
|
| 39 |
+
|
| 40 |
+
| folder | what | runs |
|
| 41 |
+
|---|---|---|
|
| 42 |
+
| `modular_add`, `modular_add_wd0.1`, `modular_add_wd0` | one-layer transformer, a + b mod 113, weight decay 1 / 0.1 / 0 | 1 each |
|
| 43 |
+
| `modular_mlp` | ReLU MLP on two-hot inputs, p = 97 (Swaroop 2026 setup) | 1 |
|
| 44 |
+
| `addition` | the length ladder: 3.4M params, trained on 1 to 20 digits, nine formats | 3 seeds each, 6 for position coupling |
|
| 45 |
+
| `addition_big` | 11M params, abacus and position coupling, trained on 1 to 30 digits | 2 abacus seeds, 4 coupling seeds |
|
| 46 |
+
| `addition_sharp`, `addition_big_sharp` | position coupling trained with attention logits scaled by 2 (`model.attn_scale`) | 3 and 2 seeds |
|
| 47 |
+
| `blankspace_big` | 11M params, fixed-width aligned blankspace, trained on 1 to 20 digits | 2 seeds |
|
| 48 |
+
| `addition_constant_lr`, `addition_no_wd` | position coupling ablations | 3 and 6 seeds |
|
| 49 |
+
| `addition_carry_heavy`, `addition_fixed_length` | position coupling with carry-heavy or single-length training data | 3 seeds each |
|
| 50 |
+
| `subtraction` | position coupling and fixed blankspace on a - b | 3 seeds each |
|
| 51 |
+
|
| 52 |
+
The one to try first is `addition_big/position_coupling_s1`. Trained on up to
|
| 53 |
+
30 digits, it scores 0% exact match at 100 digits as is. Set
|
| 54 |
+
`block.attn.scale = 2.0` on every block before decoding and it scores 100% at
|
| 55 |
+
200. The GitHub README explains why.
|