SuperGoatScriptGuy commited on
Commit
561bdf2
·
verified ·
1 Parent(s): 834675c

list the temperature-trained runs

Browse files
Files changed (1) hide show
  1. README.md +55 -54
README.md CHANGED
@@ -1,54 +1,55 @@
1
- ---
2
- license: mit
3
- tags:
4
- - arithmetic
5
- - length-generalization
6
- - grokking
7
- - interpretability
8
- ---
9
-
10
- # carrybit checkpoints
11
-
12
- Trained weights from [carrybit](https://github.com/Supergoatscriptguy/carrybit),
13
- a small research project on tiny transformers learning exact integer
14
- arithmetic. The code, configs, figures and the full write-up live in that
15
- repo. This repo holds the final checkpoint of every run in the write-up so the
16
- analysis experiments can be run without retraining.
17
-
18
- Every checkpoint is a plain PyTorch `state_dict` for `carrybit.model.Transformer`
19
- (or `TwoHotMLP` for `modular_mlp`). Each folder has the exact `config.json` it
20
- was trained with and the `metrics.csv` logged during training. Folder names
21
- match `runs/` in the GitHub repo, so every experiment script there works on
22
- these files as they are.
23
-
24
- ```python
25
- import json, torch
26
- from carrybit.config import load_config
27
- from carrybit.model import Transformer
28
-
29
- run = "addition_big/position_coupling_s1"
30
- raw = json.load(open(f"{run}/config.json"))
31
- cfg = load_config("configs/addition_big.yaml",
32
- [f"{s}.{k}={json.dumps(v)}" for s in ("task", "model")
33
- for k, v in raw[s].items() if k != "kind"])
34
- model = Transformer(16, cfg.model)
35
- model.load_state_dict(torch.load(f"{run}/step_60000.pt"))
36
- ```
37
-
38
- ## Contents
39
-
40
- | folder | what | runs |
41
- |---|---|---|
42
- | `modular_add`, `modular_add_wd0.1`, `modular_add_wd0` | one-layer transformer, a + b mod 113, weight decay 1 / 0.1 / 0 | 1 each |
43
- | `modular_mlp` | ReLU MLP on two-hot inputs, p = 97 (Swaroop 2026 setup) | 1 |
44
- | `addition` | the length ladder: 3.4M params, trained on 1 to 20 digits, nine formats | 3 seeds each, 6 for position coupling |
45
- | `addition_big` | 11M params, abacus and position coupling, trained on 1 to 30 digits | 2 seeds each |
46
- | `blankspace_big` | 11M params, fixed-width aligned blankspace, trained on 1 to 20 digits | 2 seeds |
47
- | `addition_constant_lr`, `addition_no_wd` | position coupling ablations | 3 and 6 seeds |
48
- | `addition_carry_heavy`, `addition_fixed_length` | position coupling with carry-heavy or single-length training data | 3 seeds each |
49
- | `subtraction` | position coupling and fixed blankspace on a - b | 3 seeds each |
50
-
51
- The one to try first is `addition_big/position_coupling_s1`. Trained on up to
52
- 30 digits, it scores 0% exact match at 100 digits as is. Set
53
- `block.attn.scale = 2.0` on every block before decoding and it scores 100% at
54
- 200. The GitHub README explains why.
 
 
1
+ ---
2
+ license: mit
3
+ tags:
4
+ - arithmetic
5
+ - length-generalization
6
+ - grokking
7
+ - interpretability
8
+ ---
9
+
10
+ # carrybit checkpoints
11
+
12
+ Trained weights from [carrybit](https://github.com/Supergoatscriptguy/carrybit),
13
+ a small research project on tiny transformers learning exact integer
14
+ arithmetic. The code, configs, figures and the full write-up live in that
15
+ repo. This repo holds the final checkpoint of every run in the write-up so the
16
+ analysis experiments can be run without retraining.
17
+
18
+ Every checkpoint is a plain PyTorch `state_dict` for `carrybit.model.Transformer`
19
+ (or `TwoHotMLP` for `modular_mlp`). Each folder has the exact `config.json` it
20
+ was trained with and the `metrics.csv` logged during training. Folder names
21
+ match `runs/` in the GitHub repo, so every experiment script there works on
22
+ these files as they are.
23
+
24
+ ```python
25
+ import json, torch
26
+ from carrybit.config import load_config
27
+ from carrybit.model import Transformer
28
+
29
+ run = "addition_big/position_coupling_s1"
30
+ raw = json.load(open(f"{run}/config.json"))
31
+ cfg = load_config("configs/addition_big.yaml",
32
+ [f"{s}.{k}={json.dumps(v)}" for s in ("task", "model")
33
+ for k, v in raw[s].items() if k != "kind"])
34
+ model = Transformer(16, cfg.model)
35
+ model.load_state_dict(torch.load(f"{run}/step_60000.pt"))
36
+ ```
37
+
38
+ ## Contents
39
+
40
+ | folder | what | runs |
41
+ |---|---|---|
42
+ | `modular_add`, `modular_add_wd0.1`, `modular_add_wd0` | one-layer transformer, a + b mod 113, weight decay 1 / 0.1 / 0 | 1 each |
43
+ | `modular_mlp` | ReLU MLP on two-hot inputs, p = 97 (Swaroop 2026 setup) | 1 |
44
+ | `addition` | the length ladder: 3.4M params, trained on 1 to 20 digits, nine formats | 3 seeds each, 6 for position coupling |
45
+ | `addition_big` | 11M params, abacus and position coupling, trained on 1 to 30 digits | 2 abacus seeds, 4 coupling seeds |
46
+ | `addition_sharp`, `addition_big_sharp` | position coupling trained with attention logits scaled by 2 (`model.attn_scale`) | 3 and 2 seeds |
47
+ | `blankspace_big` | 11M params, fixed-width aligned blankspace, trained on 1 to 20 digits | 2 seeds |
48
+ | `addition_constant_lr`, `addition_no_wd` | position coupling ablations | 3 and 6 seeds |
49
+ | `addition_carry_heavy`, `addition_fixed_length` | position coupling with carry-heavy or single-length training data | 3 seeds each |
50
+ | `subtraction` | position coupling and fixed blankspace on a - b | 3 seeds each |
51
+
52
+ The one to try first is `addition_big/position_coupling_s1`. Trained on up to
53
+ 30 digits, it scores 0% exact match at 100 digits as is. Set
54
+ `block.attn.scale = 2.0` on every block before decoding and it scores 100% at
55
+ 200. The GitHub README explains why.