Robot agents must connect their intended actions to observed outcomes while retaining the context needed to revise their choices over repeated attempts. Existing interfaces often leave these choices inside predefined tools or require agents to manage detailed execution code and its growing history. We introduce RobotUse, a robot agent harness that organizes computation, context, and decisions around specifying and revising physical actions. Agents visually select targets and poses, while the backend handles geometry, motion planning, and control. Subagents retain detailed interactions within each subgoal and return the information needed for subsequent decisions. Continual harnessing lets agents learn from execution by updating a persistent playbook. On RoboLab, RobotUse achieves 45% task success, outperforming CaP-X by 6.7 percentage points while maintaining compact decision contexts and reducing reliance on predefined action abstractions. Furthermore, we show that RobotUse learns from real-world execution despite imperfect feedback and transfers what it learns to subsequent tasks. Project page is available at https://robotuse-team.github.io/.
Figures & tables
Figure 1: RobotUse. The main agent decides in language: from the current state, the playbook, and returned reports, it issues an instruction. The subagent decides on the image, selecting points and poses that the robot backend plans and executes, and returns a report of the outcome and its cause. Between episodes (dashed), a refiner reads only the main-agent turns and edits the playbook.
Figure 2: From code-based execution to agent-specified actions. (a) Conventional code-as-policy composes high-level calls that determine grasp poses internally. (b) A naive expansion exposes perception, grasp generation, and coordinate transformations, but the illustrated routine still selects the highest-scoring grasp; the agent’s intended grasp has no explicit selection interface. (c) RobotUse exposes grasp candidates for visual selection and pose previews for adjustment. The agent specifies the intended action, and the backend computes and executes the corresponding motion.
Figure 3: Qualitative comparison with CaP-X. (a–g) RobotUse checks imperfect tool results in the images and succeeds (3/3); CaP-X’s code execution fails in 10 of 11 attempts, and its name-queried pointing tool marks the wrong box (0/3). (h–n) Without the grasp tool, RobotUse corrects its grasp visually; CaP-X uses one orientation for every cube in programs 3–6.
Figure 4: Simulation and real-robot evaluation environments. (a) RoboLab scenes: food pickup, shelf placement, and bottle collection (left to right). (b) Franka Panda setups: Pick and Place, Stack Cube, and Press Button (left to right).
Task success (%) ↑
Method
Overall
Simple
Moderate
Complex
(40 tasks)
(21 tasks)
(13 tasks)
(6 tasks)
Direct-action policies
π0.5
30.75
30.48
33.08
26.67
Cosmos 3
42.25
44.76
45.38
26.67
Code- and skill-based agents
Table 1: Task success on RoboLab. Language-agent methods use 120 episodes each; direct-action policies use 400. Success rates are reported overall and by task difficulty.
Figure 5: Robustness of RobotUse. Success with and without the grasp tool; 40 tasks per condition.
Figure 6: Context handoff. Token usage for RobotUse and single-agent execution across 40 tasks and three seeds (120 episodes per condition). Mean and peak input are averaged across episodes; total tokens sum input and output within each episode.
Figure 7: Learning from experience. 40 tasks per revision.
Figure 8: Qualitative playbook changes. Left: robot-frame guidance grounds “in front of the bowl.” Right: checking the original stack enables complete unstacking.
Figure 9: Real-robot task success. Franka Panda; 25 evaluation trials per task. RobotUse freezes its playbook after its first learning success.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Run
Main turns/ep.
Mean input (k)
Peak input (k)
Coverage
Baseline seed 0 / round 0
10.62
9.12
12.58
424/425
Baseline seed 2
12.40
9.75
13.79
496/496
Refinement round 1
11.08
10.01
13.64
443/443
Refinement round 2
10.70
13.38
18.73
427/428
Refinement round 3
11.25
14.37
20.13
450/450
Appendix
Table 2: Additional RobotUse main-agent accounting. Each row contains 40 episodes. Definitions and coverage follow Table 9 . Baseline seed 0 also serves as refinement round 0.
Configuration
Success
Mean input
Peak input
Total tokens
(%) ↑
(k/call)
(k)
(k/ep.)
RobotUse
45.0
13.07
27.45
596.8
Single-agent
40.8
55.01
95.74
2909
Appendix
Table 3: Context handoff ablation on RoboLab. Forty tasks across seeds 0–2, one episode per task and seed (120 episodes per condition). Input/call is the mean input tokens per model call and Peak is the largest input to a single call, both computed per episode and averaged across episodes. Total tokens sum input and output over all calls per episode.
Success (%) ↑
Usage per episode
Round
Overall
Simple
Moderate
Complex
Cost ($)
Main turns
0
47.50
61.90
30.77
33.33
0.4787
10.62
1
50.00
66.67
38.46
16.67
0.5409
11.08
2
55.00
71.43
46.15
16.67
0.5277
10.70
3
55.00
76.19
30.77
33.33
0.5401
11.25
Appendix
Table 4: Continual harnessing on RoboLab. Each round contains 40 seed-0 episodes. Cost includes all agent roles; turns count main-agent decisions. Protocol details appear in Appendix A .
Tokens per episode (k)
Time (s/ep.)
Cost ($/success)
Round
Total
Effective
0
555.69
502.59
1774.55
1.0077
1
628.00
562.84
2082.74
1.0817
2
632.88
552.59
1924.91
0.9594
3
660.72
566.46
1785.98
0.9820
Appendix
Table 5: Evaluation usage across harness revisions. Per-episode token counts are in thousands. Cost/success divides evaluation costs by confirmed successes.
Version
Cause (observed issue)
Change
v0
—
Baseline, brief skeletal guide.
v1
Targets were lost or miscounted mid-task.
(+) Build a checklist of targets, counts, order, and relations.
Rotation restrictions blocked a needed pose.
( − ) Remove the yaw-only constraint and allow tilt.
v2
Screen and robot directions were confused.
(+) Define front/back/left/right in the robot frame.
Rejected grasps were retried unchanged.
(+) Diagnose the rejection before retrying.
v3
Residual stacked objects were overlooked.
(+) Re-check the original location before declaring unstacking complete.
Appendix
Table 6: Playbook changes across harness revisions. Representative additions (+), removals ( − ), and modifications (M).
Method
$/ep.
Total tokens (k)
Effective (k)
Time (s)
$/success
CaP-X
0.1672
119.94
119.20
642.80
0.4417
Open Robot Skill
1.3800
5779.41
1230.81
1185.77
7.1346
RobotUse
0.5110
596.80
555.33
1880.51
1.1984
Appendix
Table 7: Additional inference usage. Means over three runs; mini-SWE costs $0.1706 per episode; token counts are in thousands per episode and time is seconds per episode. Cost/success is the mean of the three run-level ratios.
Figure 10: Task completion against reported API cost. Each point summarizes the three reported runs of RobotUse, CaP-X, or Open Robot Skill. API cost averages all 120 episodes per method and excludes local perception and GPU costs. RobotUse attains higher success with higher reported cost than CaP-X and lower cost than Open Robot Skill.
Visual
Relational
Procedural
Method
Color
Sem.
Size
Conj.
Count.
Spatial
Afford.
Reor.
Sort.
Stack.
(8)
(16)
(4)
(4)
(3)
(8)
(4)
(2)
(4)
(3)
Direct-action policies
π0.5
21.25
26.88
55.00
50.00
83.33
13.75
17.50
25.00
22.50
16.67
Cosmos 3
28.75
31.88
42.50
72.50
90.00
48.75
30.00
20.00
22.50
6.67
Language agents
Appendix
Table 8: Task success by RoboLab task attribute. Success rates (%) over the episodes of Table 1 . Attributes and categories follow RoboLab’s task metadata; a task can carry several attributes, and parentheses give the number of tasks.
Method
Main turns/ep.
Mean input (k)
Peak input (k)
Coverage
CaP-X
6.48
7.37
11.90
259/259
RobotUse
10.62
9.12
12.58
424/425
Appendix
Table 9: Main-agent requests and input context on the 40 seed-0 tasks. RobotUse counts role prime ; CaP-X counts role code , excluding VDM and Molmo calls. Mean input averages per-episode mean input lengths; peak input averages observed per-episode maxima. Input lengths include cached tokens and use thousands of tokens. Coverage counts main requests with recorded input usage.
Method
Grasp tool
Overall (%)
Simple
Moderate
Complex
$/ep.
CaP-X
With
45.00%
52.38%
38.46%
33.33%
0.1704
CaP-X
Without
17.50%
23.81%
7.69%
16.67%
0.2024
RobotUse
With
50.00%
47.62%
53.85%
50.00%
0.5704
RobotUse
Without
37.50%
47.62%
23.08%
33.33%
0.4100
Appendix
Table 10: Performance with and without the grasp tool. Each condition contains 40 seed-1 tasks. CaP-X success decreases by 27.50 percentage points and RobotUse by 12.50 points. RobotUse specifies no-grasp motions through a clicked location, height, and pose adjustments.
Method
Grasp tool
Turns
Total (k)
Effective (k)
Time (s)
$/success
CaP-X
With
13.70
126.43
125.52
622.63
0.3788
CaP-X
Without
15.75
145.66
143.09
963.83
1.1568
RobotUse
With
—
641.37
584.98
1760.33
1.1407
RobotUse
Without
—
420.13
370.87
1359.01
1.0933
Appendix
Table 11: Resource use in the grasp-tool comparison. Tokens include all agent roles and are thousands per episode.
Frontier models can control robots, but reasoning through every reach, grasp, and retreat makes manipulation slow and token-intensive. We revisit code as policy with a different division of labor: models build executable tools, code handles multi-phase motions, and models decide what to do next. We introduce URAI (Universal Robot-Agent Interface), which couples a programming agent that constructs robot tools with an execution agent that uses them in a feedback loop. The programming agent writes reusable and task-specific tools from task intent and refines them through execution feedback and human guidance. The execution agent selects and parameterizes these tools from current observations; each call runs a complete motion locally before returning control to the agent. Unlike delegating subsequent decisions to a generated program, this design retains model-level decision-making between tool executions. Validated tool revisions persist across episodes without updating foundation-model weights, and a shared GUI and API make the same tools available to humans and agents. Across five RoboDojo tasks and four frozen execution agents, URAI raises aggregate success from 18.0% to 53.0% relative to direct fingertip control, with the largest gain on Swap Blocks; with the same tools, a program written in advance reaches only 24% against 56% for two agents deciding after each call. Three of the four agents also finish episodes 1.3-1.5 times faster with 1.5-1.7 times fewer execution-agent output tokens; DeepSeek-V4-Flash's cost barely changes. We further evaluate URAI on seven real-world AgileX dual-arm tasks, spanning object manipulation, cloth folding, and human-interactive tic-tac-toe. URAI connects the coding and decision-making capabilities of frontier agents, organizing robot control around reusable tools that agents can both invoke and revise.
Shijia Ge, Alex Zhou, Jianshu Zeng +12
Hexafuture Inc. · Peking University · Beijing Institute of Technology +1
Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, they often incur high execution costs by reasoning and exploring the physical world from scratch. To reduce these costs, we introduce RoboSkill, a framework that connects skill acquisition and reuse through an Explore, Execute, Evolve loop. Within this loop, the agent explores to gather task-relevant information, executes tasks while adapting to feedback, and evolves its skill library based on execution records. It then reuses these skills to guide exploration and execution in the next cycle, closing the loop. To improve loop efficiency, we complement vision with tactile feedback to reduce uncertainty during physical interaction. We further augment textual guidance with reusable code to reduce reasoning overhead during skill reuse. On LIBERO-10, RoboSkill improves first-episode success rates by 12.5--25.0 percentage points and reduces average runtime by 7.6--72.4% across four agents. On real robots, it improves success rates by 8.3 percentage points and reduces average runtime for successful trials by at least 14.4%.
Sicheng Xie, Yitong Chen, Haidong Cao +3
Institute of Trustworthy Embodied AI, Fudan University · Shanghai Innovation Institute · NeoteAI.
A central goal in robotics is to enable manipulation across changing tasks and environments. Vision-language-action (VLA) models provide broad manipulation capabilities but can struggle when execution requires diagnosing failures and adapting behavior. Strong agents can discover effective interventions through interaction with these policies. We propose Recursive Harness Distillation to accumulate this experience as reusable guidance across agents. A strong agent distills its experience into a playbook for a light agent, then recursively refines the playbook using the light agent's execution feedback. The resulting playbook enables agents to reuse accumulated intervention knowledge in new task instances without updating model parameters. In real-world manipulation, the harness improves success from 37.3% to 64.0%. On SimplerEnv Bridge, the light agent with the playbook achieves 66.7% success, compared with 41.7% for the GR00T-only baseline, and outperforms the strong agent without a playbook. The same playbook also benefits the strong agent, which reaches 79.2% success. These results demonstrate the feasibility of harness distillation for robotics: intervention experience can be accumulated, refined through execution, and reused across agents to improve manipulation.
Seungyeon Kim, Junhoo Lee, Minkyu Kim +2
Seoul National University · Korea Advanced Institute of Science and Technology