RobotUse: Allocating Computation, Context, and Decisions
Organizations: KAIST
Abstract
Robot agents must connect their intended actions to observed outcomes while retaining the context needed to revise their choices over repeated attempts. Existing interfaces often leave these choices inside predefined tools or require agents to manage detailed execution code and its growing history. We introduce RobotUse, a robot agent harness that organizes computation, context, and decisions around specifying and revising physical actions. Agents visually select targets and poses, while the backend handles geometry, motion planning, and control. Subagents retain detailed interactions within each subgoal and return the information needed for subsequent decisions. Continual harnessing lets agents learn from execution by updating a persistent playbook. On RoboLab, RobotUse achieves 45% task success, outperforming CaP-X by 6.7 percentage points while maintaining compact decision contexts and reducing reliance on predefined action abstractions. Furthermore, we show that RobotUse learns from real-world execution despite imperfect feedback and transfers what it learns to subsequent tasks. Project page is available at https://robotuse-team.github.io/.
Figures & tables
| Task success (%) | ||||
| Method | Overall | Simple | Moderate | Complex |
| (40 tasks) | (21 tasks) | (13 tasks) | (6 tasks) | |
| Direct-action policies | ||||
| 30.75 | 30.48 | 33.08 | 26.67 | |
| Cosmos 3 | 42.25 | 44.76 | 45.38 | 26.67 |
| Code- and skill-based agents | ||||
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Run | Main turns/ep. | Mean input (k) | Peak input (k) | Coverage |
|---|---|---|---|---|
| Baseline seed 0 / round 0 | 10.62 | 9.12 | 12.58 | 424/425 |
| Baseline seed 2 | 12.40 | 9.75 | 13.79 | 496/496 |
| Refinement round 1 | 11.08 | 10.01 | 13.64 | 443/443 |
| Refinement round 2 | 10.70 | 13.38 | 18.73 | 427/428 |
| Refinement round 3 | 11.25 | 14.37 | 20.13 | 450/450 |
| Configuration | Success | Mean input | Peak input | Total tokens |
|---|---|---|---|---|
| (%) | (k/call) | (k) | (k/ep.) | |
| RobotUse | 45.0 | 13.07 | 27.45 | 596.8 |
| Single-agent | 40.8 | 55.01 | 95.74 | 2909 |
| Success (%) | Usage per episode | |||||
|---|---|---|---|---|---|---|
| Round | Overall | Simple | Moderate | Complex | Cost ($) | Main turns |
| 0 | 47.50 | 61.90 | 30.77 | 33.33 | 0.4787 | 10.62 |
| 1 | 50.00 | 66.67 | 38.46 | 16.67 | 0.5409 | 11.08 |
| 2 | 55.00 | 71.43 | 46.15 | 16.67 | 0.5277 | 10.70 |
| 3 | 55.00 | 76.19 | 30.77 | 33.33 | 0.5401 | 11.25 |
| Tokens per episode (k) | Time (s/ep.) | Cost ($/success) | ||
|---|---|---|---|---|
| Round | Total | Effective | ||
| 0 | 555.69 | 502.59 | 1774.55 | 1.0077 |
| 1 | 628.00 | 562.84 | 2082.74 | 1.0817 |
| 2 | 632.88 | 552.59 | 1924.91 | 0.9594 |
| 3 | 660.72 | 566.46 | 1785.98 | 0.9820 |
| Version | Cause (observed issue) | Change |
|---|---|---|
| v0 | — | Baseline, brief skeletal guide. |
| v1 | Targets were lost or miscounted mid-task. | (+) Build a checklist of targets, counts, order, and relations. |
| Rotation restrictions blocked a needed pose. | ( ) Remove the yaw-only constraint and allow tilt. | |
| v2 | Screen and robot directions were confused. | (+) Define front/back/left/right in the robot frame. |
| Rejected grasps were retried unchanged. | (+) Diagnose the rejection before retrying. | |
| v3 | Residual stacked objects were overlooked. | (+) Re-check the original location before declaring unstacking complete. |
| Method | $/ep. | Total tokens (k) | Effective (k) | Time (s) | $/success |
|---|---|---|---|---|---|
| CaP-X | 0.1672 | 119.94 | 119.20 | 642.80 | 0.4417 |
| Open Robot Skill | 1.3800 | 5779.41 | 1230.81 | 1185.77 | 7.1346 |
| RobotUse | 0.5110 | 596.80 | 555.33 | 1880.51 | 1.1984 |
| Visual | Relational | Procedural | ||||||||
| Method | Color | Sem. | Size | Conj. | Count. | Spatial | Afford. | Reor. | Sort. | Stack. |
| (8) | (16) | (4) | (4) | (3) | (8) | (4) | (2) | (4) | (3) | |
| Direct-action policies | ||||||||||
| 21.25 | 26.88 | 55.00 | 50.00 | 83.33 | 13.75 | 17.50 | 25.00 | 22.50 | 16.67 | |
| Cosmos 3 | 28.75 | 31.88 | 42.50 | 72.50 | 90.00 | 48.75 | 30.00 | 20.00 | 22.50 | 6.67 |
| Language agents | ||||||||||
| Method | Main turns/ep. | Mean input (k) | Peak input (k) | Coverage |
|---|---|---|---|---|
| CaP-X | 6.48 | 7.37 | 11.90 | 259/259 |
| RobotUse | 10.62 | 9.12 | 12.58 | 424/425 |
| Method | Grasp tool | Overall (%) | Simple | Moderate | Complex | $/ep. |
|---|---|---|---|---|---|---|
| CaP-X | With | 45.00% | 52.38% | 38.46% | 33.33% | 0.1704 |
| CaP-X | Without | 17.50% | 23.81% | 7.69% | 16.67% | 0.2024 |
| RobotUse | With | 50.00% | 47.62% | 53.85% | 50.00% | 0.5704 |
| RobotUse | Without | 37.50% | 47.62% | 23.08% | 33.33% | 0.4100 |
| Method | Grasp tool | Turns | Total (k) | Effective (k) | Time (s) | $/success |
|---|---|---|---|---|---|---|
| CaP-X | With | 13.70 | 126.43 | 125.52 | 622.63 | 0.3788 |
| CaP-X | Without | 15.75 | 145.66 | 143.09 | 963.83 | 1.1568 |
| RobotUse | With | — | 641.37 | 584.98 | 1760.33 | 1.1407 |
| RobotUse | Without | — | 420.13 | 370.87 | 1359.01 | 1.0933 |