Papers
arxiv:2610.04929

RobotUse: Allocating Computation, Context, and Decisions

Published on Oct 4
· Submitted by
Junhoo
on Oct 6
Authors:
,
,
,
,
,
,

Abstract

Robot agents must connect their intended actions to observed outcomes while retaining the context needed to revise their choices over repeated attempts. Existing interfaces often leave these choices inside predefined tools or require agents to manage detailed execution code and its growing history. We introduce RobotUse, a robot agent harness that organizes computation, context, and decisions around specifying and revising physical actions. Agents visually select targets and poses, while the backend handles geometry, motion planning, and control. Subagents retain detailed interactions within each subgoal and return the information needed for subsequent decisions. Continual harnessing lets agents learn from execution by updating a persistent playbook. On RoboLab, RobotUse achieves 45% task success, outperforming CaP-X by 6.7 percentage points while maintaining compact decision contexts and reducing reliance on predefined action abstractions. Furthermore, we show that RobotUse learns from real-world execution despite imperfect feedback and transfers what it learns to subsequent tasks. Project page is available at https://robotuse-team.github.io/.

Community

Paper submitter

RobotUse is a robot agent harness that connects language-level goals to visual decisions and physical execution. Robot tasks require agents to plan toward an overall goal while selecting targets, choosing gripper poses, and revising actions based on observed outcomes. RobotUse organizes computation, context, and decisions across a Main agent, a Subagent, and a Backend.

The Main agent plans in language, delegating subgoals with relevant constraints. The Subagent makes visual action decisions: it selects targets and locations in images, inspects candidate gripper poses, and adjusts them using execution feedback. The Backend performs geometric computation, motion planning, and control, returning fresh observations and diagnostics. The Subagent then reports the outcome, its causes, and the current state to the Main agent, informing the next decision.

Detailed observations and local recovery attempts remain within each Subagent’s context, while the Main agent receives the information needed to continue planning. RobotUse also refines a persistent playbook for continual harness , allowing accumulated guidance to improve subsequent decisions without changing model parameters or robot control code.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.04929
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.04929 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.04929 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.04929 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.