MA-FPPO: Multi-Agent Flow-Pretrained Policy Optimization

Guowei Zou, Haonan Chen, Haitao Wang, Beiwen Zhang, Na Yan, Hejun Wu.

Sun Yat-sen University · National University of Singapore.

Paper · Project page · Code

Download

Download the ZIP archives for the benchmarks you need and extract them into the same directory:

Benchmark Archive Checkpoints
MPE main_mpe.zip 20
MA-MuJoCo main_mamujoco.zip 20
OMIGA main_omiga.zip 24
SMAC main_smac.zip 24
SMACv2 main_smacv2.zip 10

Verify downloaded archives with shasum -a 256 -c ARCHIVE_SHA256SUMS. After extraction, use SHA256SUMS to verify individual checkpoints. If downloading only selected archives, check only their corresponding checksum entries.

Contents

This repository packages the historical training-seed-0 checkpoints for 48 main-experiment settings. Each setting includes pretraining and online policy weights. Two additional 100M checkpoints reproduce the 2halfcheetah Good/Medium main-table endpoints, for 98 files in total. The implementation includes one additional SMACv2 Terran Random recipe, but its completed trained model was not located and is not included.

The code's default seeds 0/1/2 support new training runs. The archived models are not three independently trained seeds. See MANIFEST.json for task, data quality, recipe, stage, actual training step, source checkpoint SHA-256, and export SHA-256.

Export format

All model tensors are preserved. Optimizer state and training RNG state are omitted, and machine-specific path strings are replaced by portable placeholders. Use these checkpoints for evaluation and initialization, not exact training resumption. The original files remain unchanged on the training server.

The named code release provides a discrete-policy evaluation loader that restores weights without depending on the original server paths or optimizer state. Consult its MODELS.md for continuous and discrete evaluation commands. Models require the matching observation layout and simulator versions.

Data and provenance

The shared CoFlow dataset repository currently contains MPE Tag and World. Other benchmark datasets and conversion procedures are documented in the code's DATA.md. The main release includes MPE, local-observation MA-MuJoCo, full-observation OMIGA tasks, SMAC, and SMACv2.

These are research policies for their specified simulation benchmarks. The checkpoints do not establish transfer to physical systems or unseen simulator configurations. Third-party software and datasets retain their original licenses and access terms. No repository-wide license is assigned here to components whose release license has not been specified.

Integrity

After extraction, run shasum -a 256 -c SHA256SUMS. Archive-level checksums are listed separately in ARCHIVE_SHA256SUMS.

For 2halfcheetah Good and Medium, online/checkpoint_final.pt is the 50M endpoint used in matched-budget comparisons, while online/checkpoint_100000000.pt is the later endpoint used in the main table.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Paper for ma-fppo/MA-FPPO