Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens
Organizations: Peking University · National University of Singapore · NVIDIA · Impossible Research
Abstract
Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant observations incur substantial token overhead. We introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with selective observation: the agent composes classical robot primitives and learned vision-language-action (VLA) policies into Python cells that perform conditional checks and local retries, returning only explicitly requested images and state feedback for replanning. Across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, we compare PyRUA-Lean with a tool-calling baseline using the same GPT-6 Astra planner and underlying robot primitives. Under equal LLM-call budgets, PyRUA-Lean increases overall success from 63.1% to 71.7%. On instances solved by both agents, it uses 49% fewer LLM calls and 65% fewer input tokens.
Figures & tables
| Tool (RPent) | Code (PyRUA-Lean) | |
| Primitives, the same in both | ||
| Arm motion | move_to , move_pose , move_delta , rotate_wrist , rotate_pitch | |
| Gripper | set_gripper , release , scripted_grasp | |
| VLA policy | pi0_pick , pi0_doubled , lingbot_act , rldx_skill , rldx_arm | |
| Mobile base | navigate_to , move_base | |
| Perception | segment , back_project , back_project_batch , sample_world_xyz , query_world_map | |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | Tasks | Seeds per task | Task instances | LLM calls per episode |
|---|---|---|---|---|
| LIBERO-PRO | 40 | 5 | 200 | 40 |
| RoboTwin 2.0 | 50 | 5 | 250 | 40 |
| RC365 atomic | 18 | 5 | 90 | 40 |
| RC365 composite | 32 | 5 | 160 | 100 |
| All | 140 | 700 |
| Tool , images after every move | Tool , images on demand | Code | |
|---|---|---|---|
| Success rate (%), all task instances | 83.0 | 68.5 | 94.0 |
| Failures: gave up / call budget | 9 / 25 | 17 / 46 | 5 / 7 |
| LLM calls per solved episode ∗ | 19.9 | 22.8 | 7.4 |
| Prompt tokens per solved episode ∗ | 1.07M | 1.07M | 222k |