MA-FPPO: Multi-Agent Flow-Pretrained Policy Optimization
Guowei Zou, Haonan Chen, Haitao Wang, Beiwen Zhang, Na Yan, Hejun Wu.
Sun Yat-sen University · National University of Singapore.
Paper · Project page · Code
Download
Download the ZIP archives for the benchmarks you need and extract them into the same directory:
| Benchmark | Archive | Checkpoints |
|---|---|---|
| MPE | main_mpe.zip | 20 |
| MA-MuJoCo | main_mamujoco.zip | 20 |
| OMIGA | main_omiga.zip | 24 |
| SMAC | main_smac.zip | 24 |
| SMACv2 | main_smacv2.zip | 10 |
Verify downloaded archives with shasum -a 256 -c ARCHIVE_SHA256SUMS. After extraction, use SHA256SUMS to verify individual checkpoints. If downloading only selected archives, check only their corresponding checksum entries.
Contents
This repository packages the historical training-seed-0 checkpoints for 48 main-experiment settings. Each setting includes pretraining and online policy weights. Two additional 100M checkpoints reproduce the 2halfcheetah Good/Medium main-table endpoints, for 98 files in total. The implementation includes one additional SMACv2 Terran Random recipe, but its completed trained model was not located and is not included.
The code's default seeds 0/1/2 support new training runs. The archived models are not three independently trained seeds. See MANIFEST.json for task, data quality, recipe, stage, actual training step, source checkpoint SHA-256, and export SHA-256.
Export format
All model tensors are preserved. Optimizer state and training RNG state are omitted, and machine-specific path strings are replaced by portable placeholders. Use these checkpoints for evaluation and initialization, not exact training resumption. The original files remain unchanged on the training server.
The named code release provides a discrete-policy evaluation loader that restores weights without depending on the original server paths or optimizer state. Consult its MODELS.md for continuous and discrete evaluation commands. Models require the matching observation layout and simulator versions.
Data and provenance
The shared CoFlow dataset repository currently contains MPE Tag and World. Other benchmark datasets and conversion procedures are documented in the code's DATA.md. The main release includes MPE, local-observation MA-MuJoCo, full-observation OMIGA tasks, SMAC, and SMACv2.
These are research policies for their specified simulation benchmarks. The checkpoints do not establish transfer to physical systems or unseen simulator configurations. Third-party software and datasets retain their original licenses and access terms. No repository-wide license is assigned here to components whose release license has not been specified.
Integrity
After extraction, run shasum -a 256 -c SHA256SUMS. Archive-level checksums are listed separately in ARCHIVE_SHA256SUMS.
For 2halfcheetah Good and Medium, online/checkpoint_final.pt is the 50M endpoint used in matched-budget comparisons, while online/checkpoint_100000000.pt is the later endpoint used in the main table.