MobiAgent: Dual-Loop Recursive Policy Self-Improvement for Long-Horizon Mobile Manipulation
Abstract
Long-horizon mobile manipulation presents significant challenges due to compounding execution errors and capacity interference between locomotion and arm control. While recent Vision-Language-Action models excel at short-horizon tasks, they lack the hierarchical reasoning required for multi-stage objectives. Furthermore, existing hierarchical agents suffer from rigid sub-task mapping, inflexible replanning, and a lack of continuous learning. To address these limitations, we introduce MobiAgent, a dual-loop agentic framework that bridges robust deployment execution and recursive policy self-improvement. During deployment, the Inner Loop decouples high-level reasoning from low-level control through highly composable atomic skills. It employs Vision-Language models for receding-horizon planning and visual reflection, dynamically composing skills to ensure robust error recovery. These skills are executed by specialized flow-matching experts that share a unified VLM backbone, maximizing reusability while mitigating capacity interference. Concurrently, the Outer Loop drives automated lifelong learning by autonomously segmenting and verifying deployment rollouts, clustering them to discover atomic skills, and continuously fine-tuning the skill library without human annotations. Evaluations on RoboCasa, BEHAVIOR-1K, and real-world tasks demonstrate the effectiveness of MobiAgent. It outperforms π_{0.5}-TA by 22.5 percentage points on BEHAVIOR-1K and enables robust recovery from execution failures. Through autonomous data recycling, success improves from 7.50% to 27.50% on RoboCasa and from 32.5% to 57.5% on Astribot S1.
Community
Introducing MobiAgent, our CoRL 2026 work on long-horizon mobile manipulation. It combines robust task execution with autonomous skill discovery and recursive policy self-improvement.
Project page | Paper · Model weights
Real-robot demo
The robot navigates to a bottle, picks it up, carries it to a bin, and disposes of it. The clip is shown at 3x speed.
Dual-loop framework
- Inner loop: the Planner composes atomic skills, the Executor runs skill-specific action experts, and the Critic checks progress to advance, retry, or replan.
- Outer loop: deployment rollouts are segmented and verified, grouped into reusable skills, and used to update the skill policies without human annotations.
- Shared backbone: skill-specific flow-matching experts share a unified VLM backbone.
Get this paper in your agent:
hf papers read 2610.03476 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper
