Reinforcement learning (RL) enables generalist robot policies to improve through trial-and-error interaction, yet its effectiveness is fundamentally constrained by sparse task rewards. Existing general-purpose reward models typically alleviate this issue by learning task progress from expert demonstrations, but introduce a distribution mismatch with the mixed-quality rollouts encountered during policy optimization, making their estimates unreliable on suboptimal and failed behaviors from which the policy must learn. In this work, we introduce a demonstration-free reward learning paradigm where dense reward feedback can be learned directly from sparse task outcomes and policy experience. We theoretically show that terminal task outcomes implicitly define dense success-probability feedback at intermediate timesteps, which can be recursively learned through bootstrapping. Based on this insight, we introduce eVTA0, which learns success probabilities from mixed-quality policy rollouts through temporal-difference-style bootstrapping, without expert demonstrations or intermediate annotations. We further introduce RL with Evolving Rewards (RLER), a closed-loop framework that adapts eVTA0 using newly collected rollouts as the policy evolves. Experiments show that eVTA0 provides more informative rewards than state-of-the-art reward models and achieves the best average policy performance across all LIBERO task suites under the same RL training budget, improving success rates by 5.4%-13.8% over the initial policy. In real-world manipulation, RLER further improves overall success rates by 20%-26%, with 35%-36% gains under out-of-distribution conditions. These results demonstrate the effectiveness of demonstration-free reward learning and adapting rewards as the policy evolves. Project webpage: https://duowuyms.github.io/evta0.
Figures & tables
Figure 1: Overview of eVTA 0 . The left panel illustrates the model architecture, which estimates the probability of eventual task success conditioned on the task description and recent observation history. The right panel illustrates the training procedure, where eVTA 0 is trained on mixed-quality policy rollouts through TD-style bootstrapping.
Figure 2: Illustration of RL training with a fixed eVTA 0 (left) and the proposed RL with Evolving Rewards (RLER) framework (right). RLER adapts eVTA 0 using newly collected policy rollouts, enabling reward learning and policy optimization to evolve together.
LIBERO
VOC ( ↑ Better)
VROC ( ↑ Better)
Method
Spatial
Object
Goal
Long
Avg.
Spatial
Object
Goal
Long
Avg.
GVL (Qwen3-VL-4B)
0.14
0.59
0.12
0.28
0.28
0.00
0.16
0.04
0.03
0.06
GVL (Qwen3-VL-8B)
0.25
0.46
0.33
0.37
0.35
-0.01
-0.04
-0.03
0.10
0.00
GVL (GPT-5)
0.79
0.79
0.72
0.66
0.74
0.30
0.53
0.52
0.57
0.48
TOPReward (Qwen3-VL-4B)
0.70
0.73
0.74
0.68
0.71
-0.73
-0.33
-0.71
-0.55
-0.58
Table 1: VOC and VROC results on LIBERO and MetaWorld. The best and second-best average results of each benchmark are shown in bold and underlined, respectively.
LIBERO
MSE ( ↓ Better)
Kendall’s τa ( ↑ Better)
Method
Spatial
Object
Goal
Long
Avg.
Spatial
Object
Goal
Long
Avg.
GVL (Qwen3-VL-4B)
0.66
0.54
0.39
0.44
0.51
0.00
0.73
0.13
0.00
0.21
GVL (Qwen3-VL-8B)
0.72
0.76
0.50
0.43
0.60
-0.17
0.00
0.18
0.02
0.01
GVL (GPT-5)
0.64
0.49
0.40
0.40
0.48
0.37
0.60
0.30
-0.05
0.31
TOPReward (Qwen3-VL-4B)
0.73
0.76
0.53
0.56
0.64
-0.27
0.62
0.21
0.25
0.20
Table 2: MSE and Kendall’s τa results on LIBERO and MetaWorld. The best and second-best average results of each benchmark are shown in bold and underlined, respectively.
Figure 3: Qualitative comparison on successful and failed trajectories under the same task.
Figure 4: Policy performance with different rewards under the same RL training budget.
Figure 5: Real-world success rates across RLER rounds compared with SFT and fixed eVTA 0 .
Algorithm 1 Training eVTA 0 via TD-Style Bootstrapping
Parameter
Configuration
RL configuration
RL algorithm
GRPO
Advantage estimator
Group-relative advantages
Training budget
50 epochs with 6,400 interaction episodes
Parallel environments
16
Rollouts per epoch
8
Appendix
Table 3: RL and π0.5 training configurations for the policy learning experiments on LIBERO-Spatial / LIBERO-Object / LIBERO-Goal / LIBERO-Long.
Parameter
Configuration
Number of tasks
10 / 10 / 10 / 10
Evaluation rollouts per task
50
Total rollouts
500 / 500 / 500 / 500
Maximum episode length
240 / 240 / 320 / 480 steps
Action chunk length
10
Denoising steps
5
Appendix
Table 4: π0.5 evaluation configurations for the policy learning experiments on LIBERO-Spatial / LIBERO-Object / LIBERO-Goal / LIBERO-Long.
Parameter
Configuration
Training epochs
5
Expert demonstrations
20 / 60
Batch size
4
Gradient accumulation steps
2
Learning rate
1×10−5
Gradient clipping
1.0
Appendix
Table 5: Configurations of SFT policies in real-world experiments ( Pick Carrot / Put Shuttlecock ).
Figure 6: The physical workspace used in our real-world experiments, which is equipped with a Franka Research 3 robot arm, two Intel RealSense D435i cameras and optionally a moving conveyor.
Figure 7: The ID and OOD configurations for the real-world experiments.
Parameter
Configuration
IQL configuration
Number of trajectories
50 / 50 / 100
Training epochs
30
Batch size
4
Gradient accumulation steps
4
Discount factor γ
1.0
Appendix
Table 6: Training configurations for IQL and eVTA 0 in real-world experiments ( RLER Round 1 / RLER Round 2 / Fixed eVTA 0 ) .
Figure 8: Ablation of hyperparameters in eVTA 0 . The top panel reports the effect of the context window size w , while the bottom part shows the effect of the target-network update interval K .
Figure 9: Additional qualitative examples of reward predictions on successful trajectories.
Figure 10: Additional qualitative examples of reward predictions on failed trajectories.
Policy
ID
OOD Distance
OOD Rotation 1
OOD Rotation 2
OOD Rotation 3
OOD Rotation 4
OOD Rotation 5
Average
SFT
24/25
1/5
3/4
1/4
0/4
2/4
1/4
32/50
eVTA 0 -RLER (Round 1)
25/25
0/5
4/4
0/4
3/4
3/4
0/4
35/50
eVTA 0 -RLER (Round 2)
25/25
5/5
4/4
2/4
2/4
3/4
1/4
42/50
Fixed eVTA 0
24/25
1/5
4/4
1/4
2/4
2/4
2/4
36/50
Appendix
Table 7: Detailed real-world results for “pick up the carrot on the moving conveyor and place it on the plate” . Each entry reports the number of successful trials over the total number of trials.
Policy
ID
OOD-Color
OOD-Position
Average
SFT
21/30
9/10
0/10
30/50
eVTA 0 -RLER (Round 1)
26/30
9/10
3/10
38/50
eVTA 0 -RLER (Round 2)
27/30
9/10
7/10
43/50
Fixed eVTA 0
26/30
9/10
4/10
39/50
Appendix
Table 8: Detailed real-world results for “put the shuttlecock from the table onto the shuttlecock in the plate” . Each entry reports the number of successful trials over the total number of trials.
Figure 11: Real-world rollouts under OOD conditions for “pick up the carrot on the moving conveyor and place it on the plate”.
Figure 12: Real-world rollouts under OOD conditions for “put the shuttlecock from the table onto the shuttlecock in the plate” .
Figure 13: Comparison of implicit value estimation and explicit success-probability reward learning. The left panel illustrates policy learning curves of PPO with sparse binary rewards and eVTA 0 rewards. The right panel depicts the estimation quality of the learned signals from the PPO critic and eVTA 0 .
Parameter
Configuration
Training budget
10 epochs with 1,280 interaction episodes
Parallel environments
16
Rollouts per epoch
8
Episodes per epoch
128
Maximum episode length
240 steps
Batch size
4
Appendix
Table 9: PPO training configurations for the critic comparison experiment on LIBERO-Spatial in Appendix E.1 . The binary reward and eVTA 0 reward use the same settings.
Task suite
Method
VOC ↑
VROC ↑
MSE ↓
Kendall’s τa↑
LIBERO-Spatial
TD
0.87
0.73
0.06
0.98
MC
0.24
0.17
0.05
0.96
LIBERO-Object
TD
0.90
0.88
0.03
1.00
MC
0.35
0.51
0.03
0.98
LIBERO-Goal
TD
0.84
0.74
0.07
1.00
MC
0.35
0.32
0.04
0.97
Appendix
Table 10: Reward-quality comparison on LIBERO between temporal-difference (TD) bootstrapping and Monte Carlo (MC) regression. The best result is shown in bold.
Figure 14: Qualitative comparison of eVTA 0 trained by TD bootstrapping and MC regression on a successful trajectory (top) and a failed trajectory (bottom).