RoboAware: Learning to Coordinate Embodied Skills from Counterfactual Outcomes
Organizations: The Chinese University of Hong Kong · The Hong Kong University of Science and Technology · Knowin AI · The University of Hong Kong · Peking University · University of the Chinese Academy of Sciences
Abstract
Embodied coding agents can combine modular robot skills with frozen end-to-end policies, yet effective composition requires anticipating which policy family will succeed in the current physical state. We present RoboAware, which builds on coding agents' skill orchestration by learning only a state-conditioned responsibility coordinator from counterfactual outcomes. Inspired by the success of REPL, we propose the schema and formulate a hierarchical MDP based on it. organizes skills uniformly into five semantic stages, defining where responsibility can be compared. To address the lack of counterfactual branch outcomes in existing work, we introduce State-Locked Counterfactual Branching (SCB), which restores the same training state to generate and execute a code block from each admissible family, exposing outcomes that selected-branch experience leaves unobserved. Building on this, we propose Execution-Aware Learning (EAL), which combines Monte Carlo tree search with Q-learning to distill these outcomes into family-conditioned values. At deployment, the coordinator selects the policy family according to observable context, and the frozen coding agent generates the next local code block. Comprehensive single-episode evaluations on 100 tasks show that RoboAware reaches a 77.0% overall success rate, with SOTA averages of 90.0% on RoboSuite, 73.8% on diverse LIBERO-Pro task clusters, and 90.0% on challenging RoboTwin bimanual tasks, outperforming existing code-as-policy and VLA-harness baselines.
Figures & tables
| Method | LIBERO-Pro | RoboSuite | RoboTwin | Overall | |||||||
| Spat-T | Spat-S | Obj-T | Obj-S | Goal-T | Goal-S | L10-T | L10-S | ||||
| End-to-end policies | |||||||||||
| [ 7 ] | 0.0 | 0.0 | 0.0 | 2.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | / | 0.2 |
| [ 20 ] | 1.0 | 20.0 | 1.0 | 17.0 | 2.0 | 38.0 | 1.0 | 8.0 | 0.0 | / | 9.8 |
| FastWAM [ 21 ] | 30.0 | 0.0 | 10.0 | 0.0 | 10.0 | 0.0 | 10.0 | 0.0 | 0.0 | 81.9 | 14.2 |
| FasterWAM [ 25 ] | 50.0 | 0.0 | 10.0 | 10.0 | 10.0 | 10.0 | 10.0 | 0.0 | 0.0 | 85.6 | 18.6 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Fail | ||||
| perceive | target ID/pose | typed target info | retry / | |
| propose | pose/mask eef target | valid eef pose | retry / perceive / | |
| pre-manipulate | approach target available | pre-contact reached | replan / / | |
| perform | pre-manip. | contact/gripper postcond. | retry / post-manip. / | |
| post-manipulate | perform | retreat/release/stabilize | retry / / |
| Setting | Value |
| Fine-tuning method | full-FT (backbone + MLP value head) |
| Thinking | Disabled |
| Optimizer | AdamW |
| Minibatch size | |
| Learning rate / schedule | , cosine, warmup |
| Gradient updates |
| Setting | Value |
| Backbone | Qwen3.5-0.8B |
| Image viewpoint | agent-view RGB + depth visualization |
| Image resolution | |
| Proprioception encoding | structured text |
| Depth encoding | colormap image |
| History window (full transitions) |
| API | Stage | Input | Output |
| Perceive_object_pose | perceive | task instruction + agent-view RGB-D + calibration | target pose + mask/geometry provenance |
| Propose_eef_pose | propose | RGB-D geometry + target mask + robot state + task context | ranked end-effector target + screening provenance |
| Plan_to_eef_pose | pre / post-manipulate | scene observation, robot state, target pose, optional attachment | continuous, smooth, collision-aware qpos sequence |
| Perform_gripper | perform | gripper open/close command | updated gripper state |
| Perform_E2E | perceive , pre-manipulate , perform , post-manipulate | route ID + task instruction + RGB-D + proprioception | closed-loop control until postcondition, failure, or budget |
| Policy family | Code-block space | Permitted API pattern | Backend choice |
| stage-admissible modular API composition | fixed selected modular backends | ||
| Perform_E2E | retained frozen backend selected inside |
| Method | Success@1 | Success@3 | Success@5 | Proposal time (s) | Screening time (s) | Time@1 (s) |
| AnyGrasp [ 61 ] | 13/18 | 16/18 | 17/18 | 7.98 | 3.46 | 36.39 |
| GraspGen [ 62 ] | 11/18 | 12/18 | 12/18 | 4.54 | 3.16 | 24.84 |
| NeuGraspNet [ 18 ] | 12/18 | 13/18 | 13/18 | 12.90 | 1.56 | 39.05 |
| EconomicGrasp [ 63 ] | 10/18 | 12/18 | 12/18 | 5.03 | 1.11 | 20.68 |
| RNGNet [ 64 ] | 7/18 | 8/18 | 8/18 | 18.89 | 1.34 | 46.33 |
| NeuGraspNet+ (ours) | 18/18 | 18/18 | 18/18 | 14.36 | 2.23 | 49.04 |
| Planner-native fresh (main) | Fixed grasp | Conditional diagnostic | |||||
| Planner | Plannable | Task success | Pipeline / Plan (s) | Task success | Plan (s) | Grasp success | Plan (s) |
| cuRoboV2 [ 19 ] | 18/18 | 18/18 | 3.82 / 0.254 | 15/18 | 0.192 | 18/18 | 0.935 |
| DRP/IMPACT [ 65 ] | 16/18 | 14/18 | 6.23 / 0.418 | 11/18 | 0.576 | 16/18 | 1.526 |
| VAMP [ 66 ] | 12/18 | 11/18 | 73.30 / 9.320 | 6/18 | 0.931 | 14/18 | 1.598 |
| Avoid Everything [ 67 ] | – | – | – | – | – | 2/18 | 3.677 |
| LIBERO-Pro | RoboSuite | RoboTwin 2.0 | ||||
| Method | Full | G–L | Full | G–L | Full | G–L |
| Vision-Language-Action Policies | ||||||
| OpenVLA-OFT [ 68 ] | 4/14 | 5/14 | 0/4 | 0/4 | 0/10 | 1/10 |
| PRTS [ 69 ] | 6/14 | 7/14 | 1/4 | 1/4 | 0/10 | 4/10 |
| RDT2 [ 70 ] | 0/14 | 0/14 | 0/4 | 0/4 | 0/10 | 0/10 |
| SpatialVLA [ 71 ] | 0/14 | 0/14 | 0/4 | 0/4 | 0/10 | 1/10 |
| Full | G–L | |||||||||
| All trials | Successful only | All trials | Successful only | |||||||
| Method | Steps | Time (s) | Hz | Mean (s) | Median (s) | Steps | Time (s) | Hz | Mean (s) | Median (s) |
| Vision–Language–Action policies | ||||||||||
| OpenVLA-OFT | ||||||||||
| PRTS | ||||||||||
| RDT2 | — | — | — | — | ||||||
| Full | G–L | |||||||||
| All trials | Successful only | All trials | Successful only | |||||||
| Method | Steps | Time (s) | Hz | Mean (s) | Median (s) | Steps | Time (s) | Hz | Mean (s) | Median (s) |
| Vision–Language–Action policies | ||||||||||
| OpenVLA-OFT | — | — | — | — | ||||||
| PRTS | ||||||||||
| RDT2 | — | — | — | — | ||||||
| Full | G–L | |||||||||
| All trials | Successful trials | All trials | Successful trials | |||||||
| Method | Steps | Time (s) | Hz | Mean (s) | Median (s) | Steps | Time (s) | Hz | Mean (s) | Median (s) |
| Vision–Language–Action policies | ||||||||||
| OpenVLA-OFT | — | — | ||||||||
| PRTS | — | — | ||||||||
| RDT2 | — | — | — | — | ||||||
| Parameter | Symbol | Setting |
| Discount factor | ||
| Perception reward weight | ||
| Manipulation reward weight | ||
| Task-success reward weight | ||
| Smooth- threshold | — | |
| Neighboring-root draws |
| Method | Planner | Memory | Seed-0 reference | SAM3 | Scripted hybrid | VLA |
| A: Full Harness VLA | SFT | |||||
| B: w/o Memory | SFT | |||||
| C: w/o Reference | SFT | |||||
| D: w/o SAM3 | SFT | |||||
| E: w/o Memory & SAM3 | SFT | |||||
| F: Bare SFT | SFT |
| Method | Success / valid | Success rate (%) | Drop (pp) | 95% CI (pp) | S F | F S |
| A: Full Harness VLA | 15/24 | 62.5 | – | – | – | – |
| B: w/o Memory | 11/24 | 45.8 | 16.7 | 7 | 3 | |
| C: w/o Reference | 12/24 | 50.0 | 12.5 | 6 | 3 | |
| D: w/o SAM3 | 14/24 | 58.3 | 4.2 | 2 | 1 | |
| E: w/o Memory & SAM3 | 9/24 | 37.5 | 25.0 | 7 | 1 | |
| F: Bare SFT | 10/24 | 41.7 | 20.8 | 6 | 1 |
| Method | Success | Failure | Error | Valid / scheduled | Valid SR (%) | Coverage (%) | Paired (pp) | 95% CI (pp) | |
| A: Full Harness VLA | 64 | 49 | 127 | 113/240 | 56.6 | 47.1 | – | – | – |
| B: w/o Memory | 20 | 98 | 122 | 118/240 | 16.9 | 49.2 | 113 | ||
| E: w/o Memory & SAM3 | 24 | 91 | 125 | 115/240 | 20.9 | 47.9 | 109 |