cs.ROSep 30, 2026

PRICE the Action Chunks: Physical Relational Credit Assignment for Embodied Reinforcement Learning

Authors: Yangang Zou, Jiajun Lu, Weitao Zhou, Haibao Yu, Bozhou Zhang, Jiawei Wang, Honglong Tian, Minglei Li, +1 more

Organizations: School of Data Science, Fudan University · Shanghai Innovation Institute · Beihang University · Simple AI · Tsinghua University · Tuojing Intelligence · The University of Hong Kong

Abstract

Outcome-based reinforcement learning (RL) post-trains vision--language--action policies using terminal success signals, but assigns the same trajectory-level advantage to every action chunk. A failed episode can thus penalize useful early actions as if they caused the failure. Existing approaches seek finer-grained feedback through learned evaluators, adding task-specific supervision or additional model training. We explore, for the first time to our knowledge, whether physical relations across trajectories can provide action-chunk credit in embodied RL from terminal outcomes alone, without an auxiliary evaluator. The key insight is that rollouts reaching corresponding physical situations can serve as references for one another: their terminal outcomes provide evidence for assessing local progress. We introduce Physical Relations for Inferring Credit from Episodes(PRICE), with two components: (i) a physical relational graph that pools current and historical outcomes at corresponding chunk boundaries to estimate success potentials; and (ii) confidence-gated credit assignment that uses changes in these potentials to refine trajectory-level supervision. Our analysis connects oracle potential changes to the terminal-success objective and provides a finite-sample directional bound for outcome-independent evidence pools. Independent continuation tests show that PRICE's retained credits align with local progress, while experiments on LIBERO, RoboTwin 2.0, and real robots demonstrate improved task success over outcome-based baselines and faster learning.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

    Jun 30, 2026Yuanda Xu, Zhengze Zhou, Hejian Sang +6Credit AssignmentAgentic Reinforcement Learning

  2. ProCredit: From Outcome Rewards to Progress Credit in Agentic Reinforcement Learning

    Sep 23, 2026Ming Ma, Yi Zhu, Yiran Zhong +7Agentic Reinforcement LearningReward Signal

  3. Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization

    Aug 31, 2026Jingxiao Yang, Wangjie Gan, Yingxuan Zhuang +3Single Trajectory