Papers
arxiv:2610.03476

MobiAgent: Dual-Loop Recursive Policy Self-Improvement for Long-Horizon Mobile Manipulation

Published on Oct 2
Authors:
,
,
,
,
,
,

Abstract

Long-horizon mobile manipulation presents significant challenges due to compounding execution errors and capacity interference between locomotion and arm control. While recent Vision-Language-Action models excel at short-horizon tasks, they lack the hierarchical reasoning required for multi-stage objectives. Furthermore, existing hierarchical agents suffer from rigid sub-task mapping, inflexible replanning, and a lack of continuous learning. To address these limitations, we introduce MobiAgent, a dual-loop agentic framework that bridges robust deployment execution and recursive policy self-improvement. During deployment, the Inner Loop decouples high-level reasoning from low-level control through highly composable atomic skills. It employs Vision-Language models for receding-horizon planning and visual reflection, dynamically composing skills to ensure robust error recovery. These skills are executed by specialized flow-matching experts that share a unified VLM backbone, maximizing reusability while mitigating capacity interference. Concurrently, the Outer Loop drives automated lifelong learning by autonomously segmenting and verifying deployment rollouts, clustering them to discover atomic skills, and continuously fine-tuning the skill library without human annotations. Evaluations on RoboCasa, BEHAVIOR-1K, and real-world tasks demonstrate the effectiveness of MobiAgent. It outperforms π_{0.5}-TA by 22.5 percentage points on BEHAVIOR-1K and enables robust recovery from execution failures. Through autonomous data recycling, success improves from 7.50% to 27.50% on RoboCasa and from 32.5% to 57.5% on Astribot S1.

Community

Introducing MobiAgent, our CoRL 2026 work on long-horizon mobile manipulation. It combines robust task execution with autonomous skill discovery and recursive policy self-improvement.

Project page | Paper · Model weights

Real-robot demo

The robot navigates to a bottle, picks it up, carries it to a bin, and disposes of it. The clip is shown at 3x speed.

Dual-loop framework

  • Inner loop: the Planner composes atomic skills, the Executor runs skill-specific action experts, and the Critic checks progress to advance, retry, or replan.
  • Outer loop: deployment rollouts are segmented and verified, grouped into reusable skills, and used to update the skill policies without human annotations.
  • Shared backbone: skill-specific flow-matching experts share a unified VLM backbone.

MobiAgent dual-loop framework

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.03476
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.03476 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.03476 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.