RACaP: Agentic Reasoning, Acting, and Coding as Policies for Evolvable Robot Learning
Organizations: The Chinese University of Hong Kong · Knowin AI · The Hong Kong University of Science and Technology (Guangzhou)
Abstract
General-purpose robot agents must learn from experience, transfer to new tasks, and act efficiently. Code as Policies (CaP) methods generate and repair programs at runtime, incurring latency and entangling reusable mechanisms with task-specific decisions. We introduce RACaP, an agentic framework that moves coding to evolution and uses a Reasoning-and-Acting (ReAct) loop to call frozen, typed Policy APIs at deployment. A two-phase strategy combines capability curriculum learning with autonomous self-evolution to improve the APIs, the ReAct harness, and experience memory. The APIs encode reusable physical mechanisms while exposing arguments for runtime adaptation. ReAct combines task-specific working memory, long-term experience memory, and visual feedback to select actions, verify outcomes, and recover from failures without modifying source code. RACaP achieves 54.4% success on LIBERO-90, 45.0% on zero-shot LIBERO-PRO, and 46.0% on LIBERO-Long, compared with at most 4.0% for CaP baselines on long-horizon tasks. On LIBERO-PRO, it achieves 2.5 times the success rate of CaP baselines and a 1.9-fold speedup in median policy time. For efficient on-robot deployment, rejection-sampled fine-tuning distills GPT-5.6 ReAct decisions into Qwen3-VL-8B-Instruct, yielding a 13.2-fold per-decision inference speedup and reducing repeated physical calls from 16 to 4. These results show that separating reusable code from runtime decisions supports continued evolution, effective transfer, and efficient long-horizon control.
Figures & tables
| Object | Goal | Spatial | ||||||||
| Method | Pos. | Task | Pos. | Task | Pos. | Task | Avg. | Time (s) | Calls | Cost (\downarrow$ |
| CaP-X | 20.0 | 30.0 | 13.3 | 6.7 | 3.3 | 6.7 | 13.3 | 845 | 13.0 | 0.68 |
| RATS-base | 36.7 | 26.7 | 16.7 | 10.0 | 13.3 | 3.3 | 17.8 | 726 | 28.5 | 0.91 |
| RATS-90 | 10.0 | 10.0 | 16.7 | 6.7 | 0.0 | 6.7 | 8.3 | 954 | 45.2 | 1.25 |
| RACaP -Phase 1 | 26.7 | 50.0 | 56.7 | 6.7 | 23.3 | 33.3 | 32.8 | 454 | 16.6 | 0.28 |
| RACaP -Phase 2 | 73.3 | 60.0 | 46.7 | 13.3 | 30.0 | 46.7 | 45.0 | 380 | 15.9 | 0.24 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Layer | Owns | Excludes |
| Policy API | Local contact geometry, motion execution, physical measurements, failure reporting | Whole-task instruction, causal goal ordering, semantic state checks |
| Runtime ReAct | Task decomposition, tool and argument selection, failure recovery, stopping logic | Source code modifications, privileged evaluator state, hidden task branches |
| Working Memory | Current instruction, remaining subgoals, recent API calls, visual outcomes | Cross-episode learning, source code edits, simulator state |
| Experience Memory | Reusable action prerequisites, recovery lessons, API capability boundaries | Task answers, absolute scene coordinates, evaluator state |
| Evolution Harness | Rollout analysis, code candidate generation, paired regression, version control | Deployment action choice, simulator state in policy input |
| Scope | Role | Trigger Condition | Public Observation Input | Action-Bearing JSON Output |
| Full ReAct | task_plan | Instruction start or replanning trigger | Task instruction, current RGB-D image, skill schemas, Full memory | Ordered goals with specified skill and typed arguments |
| Full ReAct | task_plan_review | Syntax or ordering warning from checker | Instruction, candidate plan, checker warnings, RGB-D image, Full memory | Revised ordered goals following identical schema |
| Full ReAct | state_preflight | Prior to articulation or control motion | Typed state goal, RGB-D image, advisory state report, Full memory | goal_satisfied boolean flag |
| Full ReAct | state_decide | Following articulation or control motion | State goal, attempted call, tool report, before/after/home images, visual crop | goal_satisfied , retry flag, and updated mechanism arguments |
| Transport ReAct | transport_decide | Local loop initialization or recovery step | Single transport goal, current image, local call history, gripper state, schemas | Target action and typed args dictionary |
| Transport ReAct | transport_reflect | Following transport execution | Instruction, attempted call, tool reports, gripper state, before/after images | Visual outcome next action and optional retry_args |
| Stage | Capability | Benchmark Tasks and Visual Evidence | Primary System Artifact Modified |
| S0 | Direct Transport | Moving visible objects to open targets or containers | pickplace geometry routines |
| S1 | Transport ReAct | Recovery from empty grasps and offset placements | Local ReAct prompts and retry parameters |
| S2 | Bounded Placement | Distinguishing target identity and valid interior regions | Destination grounding and check reports |
| S3a | Insertion | Insertion across books, mugs, caddies, and shelves | insert footprint and orientation control |
| S3b | Stacking | Aligning support surfaces and hollow nesting geometries | stack support stability verifications |
| S3c | Articulation | Door, drawer, and cabinet opening along line/arc paths | articulate trajectory contact paths |
| Turn | API Action | Observation Feedback and Reactive Decision |
| 1–2 | pickplace , check | Measured offset is 19.5 cm from target center. Initiate retry with measured nudge. |
| 3–4 | pickplace , check | Error decreases to 6.3 cm. Adjust nudge direction rather than repeating past trajectory. |
| 5–6 | pickplace , check | Error decreases to 3.2 cm. Switch to local push instead of full regrasp. |
| 7–8 | push , check | Grounding verification yields 2.0 cm alignment, satisfying placement threshold. |
| 9 | done | ReAct terminates execution based on visual validation. Native evaluator confirms success. |
| Lineage | Iteration Units | Simulator Episodes | Hosted API Calls | Total Cost ($) |
| RACaP-Phase 2 | 32 proposals | 596 | n/a | n/a |
| RATS (Committed) | 15 rounds | 206 | 2,474 | 62.55 |
| RATS (Valid Append-Only) | 15 rounds | 252 | 3,056 | 74.43 |
| Method | Native Success (%) | In (%) | On (%) | Turnon (%) | Close (%) | ReAct Turns |
| RACaP-Phase 1 | 32.0 | 26.7 | 40.0 | 20.0 | 0.0 | 8.22 |
| RACaP-Phase 2 | 46.0 | 33.3 | 65.0 | 50.0 | 0.0 | 8.72 |
| Method | Zero-Shot Success (%) | One-Trial Success (%) | Raw -value |
| CaP-X | 18.3 | 15.0 | 0.791 |
| RATS-base | 16.7 | 21.7 | 0.629 |
| RATS-90 | 11.7 | 3.3 | 0.180 |
| RACaP-Phase 1 | 40.0 | 48.3 | 0.302 |
| RACaP-Phase 2 | 38.3 | 48.3 | 0.146 |
| Method | Native Success (%) | Median Policy Time (s) | API Calls | Cost ($) |
| CaP-X | 25.7 | 82.4 | 2.1 | 0.111 |
| RATS-90 | 17.1 | 104.3 | 3.0 | 0.169 |
| RACaP-Phase 2 frozen | 17.1 | 44.8 | 7.8 | 0.106 |
| RATS-RS | 31.4 | 90.7 | 3.0 | 0.159 |
| RACaP-RS | 31.4 | 32.9 | 6.7 | 0.089 |