Optimizer study — tuned base models
AdamW vs Muon pretrained bases: one model per (size, token budget, optimizer) at that cell's tuned LR. WSD, wd 0.1, 1M batch, DCLM.
This collection has no items.
AdamW vs Muon pretrained bases: one model per (size, token budget, optimizer) at that cell's tuned LR. WSD, wd 0.1, 1M batch, DCLM.