Reinforcement learning (RL) enables generalist robot policies to improve through trial-and-error interaction, yet its effectiveness is fundamentally constrained by sparse task rewards. Existing general-purpose reward models typically alleviate this issue by learning task progress from expert demonstrations, but introduce a distribution mismatch with the mixed-quality rollouts encountered during policy optimization, making their estimates unreliable on suboptimal and failed behaviors from which the policy must learn. In this work, we introduce a demonstration-free reward learning paradigm where dense reward feedback can be learned directly from sparse task outcomes and policy experience. We theoretically show that terminal task outcomes implicitly define dense success-probability feedback at intermediate timesteps, which can be recursively learned through bootstrapping. Based on this insight, we introduce eVTA0, which learns success probabilities from mixed-quality policy rollouts through temporal-difference-style bootstrapping, without expert demonstrations or intermediate annotations. We further introduce RL with Evolving Rewards (RLER), a closed-loop framework that adapts eVTA0 using newly collected rollouts as the policy evolves. Experiments show that eVTA0 provides more informative rewards than state-of-the-art reward models and achieves the best average policy performance across all LIBERO task suites under the same RL training budget, improving success rates by 5.4%-13.8% over the initial policy. In real-world manipulation, RLER further improves overall success rates by 20%-26%, with 35%-36% gains under out-of-distribution conditions. These results demonstrate the effectiveness of demonstration-free reward learning and adapting rewards as the policy evolves. Project webpage: https://duowuyms.github.io/evta0.
Figures & tables
Figure 1: Overview of eVTA 0 . The left panel illustrates the model architecture, which estimates the probability of eventual task success conditioned on the task description and recent observation history. The right panel illustrates the training procedure, where eVTA 0 is trained on mixed-quality policy rollouts through TD-style bootstrapping.
Figure 2: Illustration of RL training with a fixed eVTA 0 (left) and the proposed RL with Evolving Rewards (RLER) framework (right). RLER adapts eVTA 0 using newly collected policy rollouts, enabling reward learning and policy optimization to evolve together.
LIBERO
VOC ( ↑ Better)
VROC ( ↑ Better)
Method
Spatial
Object
Goal
Long
Avg.
Spatial
Object
Goal
Long
Avg.
GVL (Qwen3-VL-4B)
0.14
0.59
0.12
0.28
0.28
0.00
0.16
0.04
0.03
0.06
GVL (Qwen3-VL-8B)
0.25
0.46
0.33
0.37
0.35
-0.01
-0.04
-0.03
0.10
0.00
GVL (GPT-5)
0.79
0.79
0.72
0.66
0.74
0.30
0.53
0.52
0.57
0.48
TOPReward (Qwen3-VL-4B)
0.70
0.73
0.74
0.68
0.71
-0.73
-0.33
-0.71
-0.55
-0.58
Table 1: VOC and VROC results on LIBERO and MetaWorld. The best and second-best average results of each benchmark are shown in bold and underlined, respectively.
LIBERO
MSE ( ↓ Better)
Kendall’s τa ( ↑ Better)
Method
Spatial
Object
Goal
Long
Avg.
Spatial
Object
Goal
Long
Avg.
GVL (Qwen3-VL-4B)
0.66
0.54
0.39
0.44
0.51
0.00
0.73
0.13
0.00
0.21
GVL (Qwen3-VL-8B)
0.72
0.76
0.50
0.43
0.60
-0.17
0.00
0.18
0.02
0.01
GVL (GPT-5)
0.64
0.49
0.40
0.40
0.48
0.37
0.60
0.30
-0.05
0.31
TOPReward (Qwen3-VL-4B)
0.73
0.76
0.53
0.56
0.64
-0.27
0.62
0.21
0.25
0.20
Table 2: MSE and Kendall’s τa results on LIBERO and MetaWorld. The best and second-best average results of each benchmark are shown in bold and underlined, respectively.
Figure 3: Qualitative comparison on successful and failed trajectories under the same task.
Figure 4: Policy performance with different rewards under the same RL training budget.
Figure 5: Real-world success rates across RLER rounds compared with SFT and fixed eVTA 0 .
Algorithm 1 Training eVTA 0 via TD-Style Bootstrapping
Parameter
Configuration
RL configuration
RL algorithm
GRPO
Advantage estimator
Group-relative advantages
Training budget
50 epochs with 6,400 interaction episodes
Parallel environments
16
Rollouts per epoch
8
Appendix
Table 3: RL and π0.5 training configurations for the policy learning experiments on LIBERO-Spatial / LIBERO-Object / LIBERO-Goal / LIBERO-Long.
Parameter
Configuration
Number of tasks
10 / 10 / 10 / 10
Evaluation rollouts per task
50
Total rollouts
500 / 500 / 500 / 500
Maximum episode length
240 / 240 / 320 / 480 steps
Action chunk length
10
Denoising steps
5
Appendix
Table 4: π0.5 evaluation configurations for the policy learning experiments on LIBERO-Spatial / LIBERO-Object / LIBERO-Goal / LIBERO-Long.
Parameter
Configuration
Training epochs
5
Expert demonstrations
20 / 60
Batch size
4
Gradient accumulation steps
2
Learning rate
1×10−5
Gradient clipping
1.0
Appendix
Table 5: Configurations of SFT policies in real-world experiments ( Pick Carrot / Put Shuttlecock ).
Figure 6: The physical workspace used in our real-world experiments, which is equipped with a Franka Research 3 robot arm, two Intel RealSense D435i cameras and optionally a moving conveyor.
Figure 7: The ID and OOD configurations for the real-world experiments.
Parameter
Configuration
IQL configuration
Number of trajectories
50 / 50 / 100
Training epochs
30
Batch size
4
Gradient accumulation steps
4
Discount factor γ
1.0
Appendix
Table 6: Training configurations for IQL and eVTA 0 in real-world experiments ( RLER Round 1 / RLER Round 2 / Fixed eVTA 0 ) .
Figure 8: Ablation of hyperparameters in eVTA 0 . The top panel reports the effect of the context window size w , while the bottom part shows the effect of the target-network update interval K .
Figure 9: Additional qualitative examples of reward predictions on successful trajectories.
Figure 10: Additional qualitative examples of reward predictions on failed trajectories.
Policy
ID
OOD Distance
OOD Rotation 1
OOD Rotation 2
OOD Rotation 3
OOD Rotation 4
OOD Rotation 5
Average
SFT
24/25
1/5
3/4
1/4
0/4
2/4
1/4
32/50
eVTA 0 -RLER (Round 1)
25/25
0/5
4/4
0/4
3/4
3/4
0/4
35/50
eVTA 0 -RLER (Round 2)
25/25
5/5
4/4
2/4
2/4
3/4
1/4
42/50
Fixed eVTA 0
24/25
1/5
4/4
1/4
2/4
2/4
2/4
36/50
Appendix
Table 7: Detailed real-world results for “pick up the carrot on the moving conveyor and place it on the plate” . Each entry reports the number of successful trials over the total number of trials.
Policy
ID
OOD-Color
OOD-Position
Average
SFT
21/30
9/10
0/10
30/50
eVTA 0 -RLER (Round 1)
26/30
9/10
3/10
38/50
eVTA 0 -RLER (Round 2)
27/30
9/10
7/10
43/50
Fixed eVTA 0
26/30
9/10
4/10
39/50
Appendix
Table 8: Detailed real-world results for “put the shuttlecock from the table onto the shuttlecock in the plate” . Each entry reports the number of successful trials over the total number of trials.
Figure 11: Real-world rollouts under OOD conditions for “pick up the carrot on the moving conveyor and place it on the plate”.
Figure 12: Real-world rollouts under OOD conditions for “put the shuttlecock from the table onto the shuttlecock in the plate” .
Figure 13: Comparison of implicit value estimation and explicit success-probability reward learning. The left panel illustrates policy learning curves of PPO with sparse binary rewards and eVTA 0 rewards. The right panel depicts the estimation quality of the learned signals from the PPO critic and eVTA 0 .
Parameter
Configuration
Training budget
10 epochs with 1,280 interaction episodes
Parallel environments
16
Rollouts per epoch
8
Episodes per epoch
128
Maximum episode length
240 steps
Batch size
4
Appendix
Table 9: PPO training configurations for the critic comparison experiment on LIBERO-Spatial in Appendix E.1 . The binary reward and eVTA 0 reward use the same settings.
Task suite
Method
VOC ↑
VROC ↑
MSE ↓
Kendall’s τa↑
LIBERO-Spatial
TD
0.87
0.73
0.06
0.98
MC
0.24
0.17
0.05
0.96
LIBERO-Object
TD
0.90
0.88
0.03
1.00
MC
0.35
0.51
0.03
0.98
LIBERO-Goal
TD
0.84
0.74
0.07
1.00
MC
0.35
0.32
0.04
0.97
Appendix
Table 10: Reward-quality comparison on LIBERO between temporal-difference (TD) bootstrapping and Monte Carlo (MC) regression. The best result is shown in bold.
Figure 14: Qualitative comparison of eVTA 0 trained by TD bootstrapping and MC regression on a successful trajectory (top) and a failed trajectory (bottom).
Reinforcement learning relies on accurate reward functions, which are often hand-crafted or even unavailable in real-world applications, such as robotics. Recent work has explored the zero-shot reasoning capabilities of pre-trained Vision-Language Models (VLMs) as reward models. However, without careful prompt engineering, these approaches tend to produce suboptimal rewards, where false positive predictions can severely degrade downstream policy learning. In robotics, limited datasets comprising expert demonstrations are often collected to bootstrap policy learning. This scenario provides an opportunity to optimize a reward model prior policy training. We propose Demo2Reward a test-time adaptation technique to optimize the language instruction of a reward model based on a few demonstrations (3-10 trajectories) to reduce false positives while preserving true positives. Crucially, this requires no additional model training or computation resources during policy learning. We show that Demo2Reward consistently outperforms existing zero- and few-shot VLM reward models across a range of simulated robotic tasks and policy backbones. Finally, we demonstrate that Demo2Reward effectively transfers to a real-world robotic learning scenario, enabling policy learning without manually engineering a reward function.
Christian Gumbsch, Leonardo Barcellona, Lennard Schünemann +7
University of Amsterdam · 2Catholic University of Leuven · 3Toyota Research Institute +1
Contact-rich manipulation requires robots to sequence precise contacts, maintain stable grasps, and apply directed forces. Reinforcement learning (RL) can acquire such behaviors automatically, but its performance hinges on reward design: sparse rewards reduce the learning efficiency, while dense rewards are hard to specify. Visual reward learning addresses this by inferring rewards from action-free demonstrations. Because it conditions only on visual observations, it fails to capture rewards beyond visual goals. We propose Tactile Reward Learning (TaRL), a framework that learns rewards from tactile demonstrations. TaRL takes a sequence of tactile deformation maps as input, and regresses task-completion progress from both successful and failed demonstrations. Because TaRL captures local robot-object interaction, it provides informative feedback to learn firm grasps and correctly directed forces; meanwhile, it is robust to changes in scene layout such as object position. We evaluate TaRL on four manipulation tasks in simulation and two in the real world. Used as a shaping reward, it substantially improves both sample efficiency and final success rate, raising success on Nut threading from 34% to 56% in simulation and on cube pickup from 37% to 97% in the real world. Combining tactile with visual rewards improves performance further. TaRL also generalizes across object instances: trained on box placement and directly deployed to can placement, it significantly improves policy learning on the new task. Project page is available at https://embodiedai-ntu.github.io/tarl.
Existing robotic foundation policies are trained primarily via large-scale imitation learning. While such models demonstrate strong capabilities, they often struggle with long-horizon tasks due to distribution shift and error accumulation. While reinforcement learning (RL) can finetune these models, it cannot work well across diverse tasks without manual reward engineering. We propose VLLR, a dense reward framework combining (1) an extrinsic reward from Large Language Models (LLMs) and Vision-Language Models (VLMs) for task progress recognition, and (2) an intrinsic reward based on policy self-certainty. VLLR uses LLMs to decompose tasks into verifiable subtasks and then VLMs to estimate progress to initialize the value function for a brief warm-up phase, avoiding prohibitive inference cost during full training; and self-certainty provides per-step intrinsic guidance throughout PPO finetuning. Ablation studies reveal complementary benefits: VLM-based value initialization primarily improves task completion efficiency, while self-certainty primarily enhances success rates, particularly on out-of-distribution tasks. On the CHORES benchmark covering mobile manipulation and navigation, VLLR achieves up to 56% absolute success rate gains over the pretrained policy, up to 5% gains over state-of-the-art RL finetuning methods on in-distribution tasks, and up to 10% gains on out-of-distribution tasks, all without manual reward engineering. Additional visualizations can be found in https://silongyong.github.io/vllr_project_page/
Silong Yong, Stephen Sheng, Carl Qi +6
Carnegie Mellon University · Amazon Robotics · UT Austin