cs.ROOct 8, 2026

RoboAware: Learning to Coordinate Embodied Skills from Counterfactual Outcomes

Authors: Bohan Zhou, Xingbei Chen, Emily Huang, Weilin Ruan, Haojian Huang, Yehang Zhang, Zexi Li, Wenqian Li, +12 more

Organizations: The Chinese University of Hong Kong · The Hong Kong University of Science and Technology · Knowin AI · The University of Hong Kong · Peking University · University of the Chinese Academy of Sciences

Abstract

Embodied coding agents can combine modular robot skills with frozen end-to-end policies, yet effective composition requires anticipating which policy family will succeed in the current physical state. We present RoboAware, which builds on coding agents' skill orchestration by learning only a state-conditioned responsibility coordinator from counterfactual outcomes. Inspired by the success of REPL, we propose the P5P^5 schema and formulate a hierarchical MDP based on it. P5P^5 organizes skills uniformly into five semantic stages, defining where responsibility can be compared. To address the lack of counterfactual branch outcomes in existing work, we introduce State-Locked Counterfactual Branching (SCB), which restores the same training state to generate and execute a code block from each admissible family, exposing outcomes that selected-branch experience leaves unobserved. Building on this, we propose Execution-Aware Learning (EAL), which combines Monte Carlo tree search with Q-learning to distill these outcomes into family-conditioned values. At deployment, the coordinator selects the policy family according to observable context, and the frozen coding agent generates the next local code block. Comprehensive single-episode evaluations on 100 tasks show that RoboAware reaches a 77.0% overall success rate, with SOTA averages of 90.0% on RoboSuite, 73.8% on diverse LIBERO-Pro task clusters, and 90.0% on challenging RoboTwin bimanual tasks, outperforming existing code-as-policy and VLA-harness baselines.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Playful Agentic Robot Learning

    Jun 17, 2026Junyi Zhang, Jiaxin Ge, Hanjun Yoo +17Robot Policy LearningRobot Skill Learning

  2. Encore: Few-Shot Agentic Discovery of Manipulation Strategies

    Sep 29, 2026Yifan Kang, Zihan Wang, Zhiwen Fan +1Robot Policy LearningLanguage-Conditioned Robot Manipulation

  3. RHO: Your Coding Agent is Secretly a Roboticist

    Jun 15, 2026Karim Elmaaroufi, Justin Svegliato, Sarunas Kalade +3Robot Policy LearningAI Coding Agents