Make Code as Policy Great Again: Frontier Agents Write, Call, and Evolve Robot Tools
Organizations: Hexafuture Inc. · Peking University · Beijing Institute of Technology · Tsinghua University
Abstract
Frontier models can control robots, but reasoning through every reach, grasp, and retreat makes manipulation slow and token-intensive. We revisit code as policy with a different division of labor: models build executable tools, code handles multi-phase motions, and models decide what to do next. We introduce URAI (Universal Robot-Agent Interface), which couples a programming agent that constructs robot tools with an execution agent that uses them in a feedback loop. The programming agent writes reusable and task-specific tools from task intent and refines them through execution feedback and human guidance. The execution agent selects and parameterizes these tools from current observations; each call runs a complete motion locally before returning control to the agent. Unlike delegating subsequent decisions to a generated program, this design retains model-level decision-making between tool executions. Validated tool revisions persist across episodes without updating foundation-model weights, and a shared GUI and API make the same tools available to humans and agents. Across five RoboDojo tasks and four frozen execution agents, URAI raises aggregate success from 18.0% to 53.0% relative to direct fingertip control, with the largest gain on Swap Blocks; with the same tools, a program written in advance reaches only 24% against 56% for two agents deciding after each call. Three of the four agents also finish episodes 1.3-1.5 times faster with 1.5-1.7 times fewer execution-agent output tokens; DeepSeek-V4-Flash's cost barely changes. We further evaluate URAI on seven real-world AgileX dual-arm tasks, spanning object manipulation, cloth folding, and human-interactive tic-tac-toe. URAI connects the coding and decision-making capabilities of frontier agents, organizing robot control around reusable tools that agents can both invoke and revise.
Figures & tables
| Method | Pour liquid into cup | Pour balls into vase | Fold clothes | Swap blocks a | Stack blocks | Avg. SR | Time (min) | Tokens (k) |
|---|---|---|---|---|---|---|---|---|
| 28.0 | 13.3 | 38.7 | 0.0 | 13.3 | 18.7 | – | – | |
| GPT-6 Astra, RoboProbe | 8.0 | 4.0 | 72.0 | 0.0 | 88.0 | 34.4 | – | – |
| Xiaomi-Robotics-1 | 64.0 | 46.0 | 73.3 | 0.0 | 21.3 | 40.9 | – | – |
| Liber-0 Preview | 57.0 | 31.0 | 79.0 | 1.0 | 72.0 | 48.0 | – | – |
| Simate-beta | 57.3 | 45.3 | 80.0 | 38.7 | 69.3 | 58.1 | – | – |
| DeepSeek-V4-Flash, native | 0.0 | 0.0 | 0.0 | 0.0 | 40.0 | 8.0 | 21.4 | 44.3 |
| Condition | Motion calls | s/call | Motion (%) | Obs./probe (%) | Other (%) | Other (s) |
|---|---|---|---|---|---|---|
| Claude Fable 5.1 + URAI | 4.3 | 63 | 31 | 31 | 38 | 426 |
| Claude Fable 5.1, native | 22.8 | 20 | 34 | 16 | 51 | 772 |
| Claude Opus 5.5 + URAI | 3.5 | 72 | 37 | 35 | 28 | 250 |
| Claude Opus 5.5, native | 21.8 | 20 | 44 | 18 | 38 | 405 |
| GPT-6 Astra + URAI | 5.6 | 44 | 47 | 21 | 33 | 188 |
| GPT-6 Astra, native | 21.0 | 19 | 50 | 16 | 34 | 276 |
| Task | Method | Success | Decisions / model calls | Tokens | Time (min) |
|---|---|---|---|---|---|
| Unscrew bottle cap | GPT-Policy, published (R3) | 3/3 | 54.7 | – | 17.9 |
| GPT-6 Astra + URAI | 3/3 | 24.7 | 1.52k | 5.1 | |
| Tic-tac-toe | GPT-Policy, published (R8) | 3/3 | 69.7 | – | 13.6 |
| GPT-6 Astra + URAI | 3/3 | 37.0 | 3.26k | 5.0 | |
| Block into bowl | Robocurve, published | 19/20 | 20 (budget) | 2.1k | 2.5 |
| GPT-6 Astra + URAI | 3/3 | 6.3 | 443 | 1.0 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| System | Who writes the code, and when | Unit of one model decision | Model at run time | Persistence across episodes | Human interface |
|---|---|---|---|---|---|
| Code as Policies [ 26 ] | LM, one program per instruction | the whole program | program decides; perception calls inside | none | – |
| CaP-X / CaP-Agent0 [ 11 ] | coding agent, one program per turn | one program per turn | re-plans between turns from execution feedback | synthesized skills (CaP-Agent0) | – |
| Agent as Policy [ 21 ] | preparation agent per task type; execution agent per trial | program and motion commands | revises from physical feedback | saved procedures reused | – |
| ASPIRE [ 30 ] | agent writes and repairs programs between rollouts | program | program runs | validated skill library | – |
| ENPIRE [ 42 ] | coding agent revises the policy between rollouts | policy | policy runs | retained policies | harness |
| LATM / CRAFT [ 5 , 45 ] | tool-maker LLM, offline | one tool call by the tool user | tool user decides each call | toolset | – |