cs.ROSep 27, 2026

Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies

Authors: Duo Wu, Haifeng Wang, Rongwei Lu, Jinghe Wang, Tianyi Xiong, Zhimin Wang, Chao Yu, Shuai Ma, +1 more

Organizations: Tsinghua University · Yuanxing Robotics · Pengcheng Laboratory

Abstract

Reinforcement learning (RL) enables generalist robot policies to improve through trial-and-error interaction, yet its effectiveness is fundamentally constrained by sparse task rewards. Existing general-purpose reward models typically alleviate this issue by learning task progress from expert demonstrations, but introduce a distribution mismatch with the mixed-quality rollouts encountered during policy optimization, making their estimates unreliable on suboptimal and failed behaviors from which the policy must learn. In this work, we introduce a demonstration-free reward learning paradigm where dense reward feedback can be learned directly from sparse task outcomes and policy experience. We theoretically show that terminal task outcomes implicitly define dense success-probability feedback at intermediate timesteps, which can be recursively learned through bootstrapping. Based on this insight, we introduce eVTA0_0, which learns success probabilities from mixed-quality policy rollouts through temporal-difference-style bootstrapping, without expert demonstrations or intermediate annotations. We further introduce RL with Evolving Rewards (RLER), a closed-loop framework that adapts eVTA0_0 using newly collected rollouts as the policy evolves. Experiments show that eVTA0_0 provides more informative rewards than state-of-the-art reward models and achieves the best average policy performance across all LIBERO task suites under the same RL training budget, improving success rates by 5.4%-13.8% over the initial policy. In real-world manipulation, RLER further improves overall success rates by 20%-26%, with 35%-36% gains under out-of-distribution conditions. These results demonstrate the effectiveness of demonstration-free reward learning and adapting rewards as the policy evolves. Project webpage: https://duowuyms.github.io/evta0.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models

    May 22, 2026Christian Gumbsch, Leonardo Barcellona, Lennard Schünemann +7Scalable Robot Learning

  2. TaRL: Learning General and Physical Rewards from Tactile Demonstrations

    Sep 29, 2026Po-Yi Wu, Dao-Jan Chang, Shang-Ya Hsiao +3TactileRobotic Grasping

  3. Generalizable Dense Reward for Long-Horizon Robotic Tasks

    Mar 31, 2026Silong Yong, Stephen Sheng, Carl Qi +6