--- datasets: - mlfoundations/dclm-baseline-1.0 library_name: megatron-lm license: mit tags: - mixture-of-experts - moe - distributionally-robust-optimization - megatron pipeline_tag: text-generation --- # DRMoET checkpoints **Paper:** [Distributionally Robust Mixture-of-Experts Training](https://huggingface.co/papers/2610.07207) **Project page:** **Code:** The weights in this repository are the checkpoints from the DRMoET paper: two DRMoET models (activation credit), each paired with the FLAME-MoE baseline trained with the same data, seed, and schedule. The folders below describe each checkpoint. ## Checkpoints | Folder | Method | Experts (top-k) | Layers / hidden | Train iters | DRO (β, η, α, credit) | |---|---|---|---|---|---| | `DRMoET-1.7B` | DRMoET, activation credit | 64 (top-6) + shared | 18 / 2048 | 32,000 | 0.999, 0.1, 1.0, L2 activation norm | | `FLAME-MoE-1.7B` | FLAME-MoE baseline | 64 (top-6) + shared | 18 / 2048 | 32,000 | — | | `DRMoET-290M-32E` | DRMoET, activation credit | 32 (top-6) + shared | 9 / 1024 | 16,000 | 0.999, 0.001, 1.0, L2 activation norm | | `FLAME-MoE-290M-32E` | FLAME-MoE baseline | 32 (top-6) + shared | 9 / 1024 | 16,000 | — | All models use sequence length 2048, global batch size 1024, learning rate 3e-4, auxiliary load-balancing loss coefficient 0.01, and the `EleutherAI/pythia-12b` tokenizer, trained on DCLM. Each DRMoET checkpoint and its baseline share the same seed (1,234 for 1.7B and 3,407 for 290M-32E).