Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding
Abstract
Decentralized multi-agent path finding (MAPF) with communication requires agents to reach individual goals without collisions under partial observability. Learnable policies trained on expert data provide an effective approach to this problem. However, when several coordinated joint actions are valid in the same context, independently sampling from per-agent distributions can recombine locally valid choices into incompatible joint actions. This failure can arise from the final sampling mechanism even when the per-agent action distributions are learned correctly. DMM (Decentralized Master-Mind) addresses this by replacing one-shot action sampling with discrete, iterative refinement of action intents across communication rounds, inspired by denoising in diffusion models. Agents initialize random action intents and refine them through local communication, coupling their choices before commitment. DMM is pretrained with imitation learning on expert MAPF solutions and further optimized with MICPO, a critic-free group-relative reinforcement-learning method designed for multi-agent, multi-round action refinement. DMM generally achieves higher success rates and lower solution costs than the evaluated learnable baselines. On 1,600 MovingAI tasks, DMM fine-tuned with MICPO solves 1,598, the highest coverage among the evaluated methods, while achieving solution costs close to those of the strongest baselines. DMM also scales to over one million simultaneously acting agents in obstacle-rich environments. These results show that round-level intent refinement can improve joint-action coordination while preserving decentralized execution.
Community
We found a serious problem in learned multi-agent pathfinding (MAPF). Even when agents learn to communicate, each one still picks its final action independently, and we call this the decentralized factorization gap. When several joint solutions are valid, agents can combine pieces of different ones. Each choice makes sense on its own, but together they lead to collisions or deadlocks. A bigger model or a bigger dataset doesn't fix this, in theory or in practice.
We propose DMM (Decentralized Master-Mind), where agents agree first and act second. Inspired by diffusion, each agent starts from a random "intent" and refines it over several rounds, exchanging intents with its neighbors. DMM is trained on expert data and then fine-tuned with our multi-agent variant of GRPO.
Contributions:
- Agents actually agree. In a symmetric corridor, DMM picks a valid joint action 95–99.9% of the time, while the previous SOTA (LC-MAPF, MAGAT+, HMAGAT) manage about 50%.
- New SOTA on the POGEMA benchmark.
- Rivals centralized search: DMM solves 1,598 of 1,600 MovingAI tasks, more than LG-LaCAM and MAPF-LNS2.
- Scales to one million agents in huge mazes.
Get this paper in your agent:
hf papers read 2609.32019 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper