ANLP Assignment 2, Task 2: hand-written optimizers
The dense 26.45M-parameter decoder of Task 1 (variant v1), pretrained from scratch on browndw/human-ai-parallel-corpus for 1x the training data (44.43M tokens) with AdamW, MARS, Lion and Muon, one per category of 'Fantastic Pretraining Optimizers'.
| optimizer | val loss | val ppl | human | LLM | BLEU | BLEU (sm.) | ROUGE-L | state MB | peak MB | tok/s | min | reaches AdamW-final at | speedup vs AdamW |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AdamW | 3.5921 | 36.31 | 4.288 | 3.371 | 0.77 | 0.79 | 13.06 | 201.8 | 5,187 | 72,168 | 10.3 | 1.00x data | 1.00x |
| MARS | 3.4849 | 32.62 | 4.203 | 3.256 | 0.68 | 0.71 | 13.14 | 302.7 | 5,287 | 71,163 | 10.4 | 0.71x data | 1.41x |
| Lion | 3.5988 | 36.55 | 4.290 | 3.379 | 0.56 | 0.59 | 13.06 | 100.9 | 5,086 | 76,248 | 9.7 | never | 0.90x |
| Muon | 3.2929 | 26.92 | 4.035 | 3.057 | 0.80 | 0.82 | 13.53 | 147.8 | 5,133 | 70,142 | 10.6 | 0.60x data | 1.67x |
<optimizer>/final.pt holds {'model': state_dict, 'config': ...}; the tokenizer is
Task 1's (tokenizer.json), and the code is in the submission repository.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support