DRMoET / README.md
TonyTeng's picture nielsr's picture
nielsr HF Staff
Add pipeline tag and paper link (#1)
7cd3987
|
Raw History Blame Contribute Delete
1.55 kB
metadata
datasets:
  - mlfoundations/dclm-baseline-1.0
library_name: megatron-lm
license: mit
tags:
  - mixture-of-experts
  - moe
  - distributionally-robust-optimization
  - megatron
pipeline_tag: text-generation

DRMoET checkpoints

Paper: Distributionally Robust Mixture-of-Experts Training

Project page: https://drmoet.github.io Code: https://github.com/MAPS-research/DRMoET

The weights in this repository are the checkpoints from the DRMoET paper: two DRMoET models (activation credit), each paired with the FLAME-MoE baseline trained with the same data, seed, and schedule. The folders below describe each checkpoint.

Checkpoints

Folder Method Experts (top-k) Layers / hidden Train iters DRO (β, η, α, credit)
DRMoET-1.7B DRMoET, activation credit 64 (top-6) + shared 18 / 2048 32,000 0.999, 0.1, 1.0, L2 activation norm
FLAME-MoE-1.7B FLAME-MoE baseline 64 (top-6) + shared 18 / 2048 32,000 —
DRMoET-290M-32E DRMoET, activation credit 32 (top-6) + shared 9 / 1024 16,000 0.999, 0.001, 1.0, L2 activation norm
FLAME-MoE-290M-32E FLAME-MoE baseline 32 (top-6) + shared 9 / 1024 16,000 —

All models use sequence length 2048, global batch size 1024, learning rate 3e-4, auxiliary load-balancing loss coefficient 0.01, and the EleutherAI/pythia-12b tokenizer, trained on DCLM. Each DRMoET checkpoint and its baseline share the same seed (1,234 for 1.7B and 3,407 for 290M-32E).