|
Download README.md from TonyTeng/DRMoET: direct link, hf CLI and curl.
- Browser
- Download file 1.55 kB
-
https://huggingface.co/TonyTeng/DRMoET/resolve/main/README.md
- Command line
-
hf download hf://TonyTeng/DRMoET/README.md
-
curl -L -o README.md https://huggingface.co/TonyTeng/DRMoET/resolve/main/README.md
1.55 kB
metadata
datasets:
- mlfoundations/dclm-baseline-1.0
library_name: megatron-lm
license: mit
tags:
- mixture-of-experts
- moe
- distributionally-robust-optimization
- megatron
pipeline_tag: text-generation
DRMoET checkpoints
Paper: Distributionally Robust Mixture-of-Experts Training
Project page: https://drmoet.github.io Code: https://github.com/MAPS-research/DRMoET
The weights in this repository are the checkpoints from the DRMoET paper: two DRMoET models (activation credit), each paired with the FLAME-MoE baseline trained with the same data, seed, and schedule. The folders below describe each checkpoint.
Checkpoints
| Folder | Method | Experts (top-k) | Layers / hidden | Train iters | DRO (β, η, α, credit) |
|---|---|---|---|---|---|
DRMoET-1.7B |
DRMoET, activation credit | 64 (top-6) + shared | 18 / 2048 | 32,000 | 0.999, 0.1, 1.0, L2 activation norm |
FLAME-MoE-1.7B |
FLAME-MoE baseline | 64 (top-6) + shared | 18 / 2048 | 32,000 | — |
DRMoET-290M-32E |
DRMoET, activation credit | 32 (top-6) + shared | 9 / 1024 | 16,000 | 0.999, 0.001, 1.0, L2 activation norm |
FLAME-MoE-290M-32E |
FLAME-MoE baseline | 32 (top-6) + shared | 9 / 1024 | 16,000 | — |
All models use sequence length 2048, global batch size 1024, learning rate 3e-4, auxiliary load-balancing loss coefficient 0.01, and the EleutherAI/pythia-12b tokenizer, trained on DCLM. Each DRMoET checkpoint and its baseline share the same seed (1,234 for 1.7B and 3,407 for 290M-32E).