NavHarness: Towards Lifelong Embodied Navigation
Abstract
Frontier models can now perform well on individual embodied navigation tasks through multi-round multimodal reasoning with simple tools. Across successive tasks, however, an agent must also rely on an evolving map and earlier search records, both of which may be incomplete or conflict with new observations. We present NavHarness, a training-free embodied harness towards lifelong navigation that makes memory processing part of the navigation loop. During navigation, its multi-round agentic session draws on maps, task records, and house knowledge, checking them against observations and recording corrections to guide its actions. NavHarness preserves this experience across fresh conversations for new tasks or recovery attempts, while outcome verification and run-end summaries support its later reuse. On GOAT-Bench, NavHarness improves s-SR over context-only independent sessions by 18.6 points with Astra and 22.6 with Opus 5. Using SLAM-estimated poses, NavHarness with GPT-6 Astra achieves state-of-the-art task success of 83.7 s-SR with 36.9 e-SR on GOAT-Bench and 85.9 s-SR on IR2R-CE. To understand these gains, we examine how experience is carried between sessions and find that structured recovery handovers outperform length-matched summaries. In extended deployments across houses, consolidation improves navigation beyond retaining maps and task records, with case studies showing how agents use earlier experience to interpret new goals, investigate unresolved questions, and resume failed searches. We suggest that progress towards lifelong navigation depends on how successive reasoning sessions build on prior experience, alongside improvements in single-task capability.
Community
Frontier models can now perform well on individual embodied navigation tasks through multi-round multimodal reasoning with simple tools. Across successive tasks, however, a robot must also rely on an evolving map and earlier search records, both of which may be incomplete or conflict with new observations.
NavHarness is a training-free embodied harness towards lifelong navigation that makes memory processing part of the navigation loop. During navigation, a multi-round agentic session draws on maps, task records and house knowledge, checks them against its observations, and records corrections that guide later decisions. NavHarness preserves this experience across fresh conversations for new tasks and recovery attempts, while outcome verification and run-end consolidation decide what later sessions inherit.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- EvoNav-Bench: Benchmarking Lifelong Navigation in Evolving Environments (2026)
- GPT-6-Astra Lights Up Embodied Navigation: Evaluation in Zero-Shot Vision-and-Language Navigation in Continuous Environments (2026)
- Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents (2026)
- HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness (2026)
- CGFM-Nav: Cognitive Graph-Field Memory for Semantic-Guided Lifelong Multimodal Embodied Navigation (2026)
- AdaHVLA: Adaptive Harnesses for Long-Horizon Vision-Language-Action Execution (2026)
- NavProbe: Evidence-Grounded Reasoning with Active Memory Retrieval for Zero-Shot Navigation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper