Existing robotic foundation policies are trained primarily via large-scale imitation learning. While such models demonstrate strong capabilities, they often struggle with long-horizon tasks due to distribution shift and error accumulation. While reinforcement learning (RL) can finetune these models, it cannot work well across diverse tasks without manual reward engineering. We propose VLLR, a dense reward framework combining (1) an extrinsic reward from Large Language Models (LLMs) and Vision-Language Models (VLMs) for task progress recognition, and (2) an intrinsic reward based on policy self-certainty. VLLR uses LLMs to decompose tasks into verifiable subtasks and then VLMs to estimate progress to initialize the value function for a brief warm-up phase, avoiding prohibitive inference cost during full training; and self-certainty provides per-step intrinsic guidance throughout PPO finetuning. Ablation studies reveal complementary benefits: VLM-based value initialization primarily improves task completion efficiency, while self-certainty primarily enhances success rates, particularly on out-of-distribution tasks. On the CHORES benchmark covering mobile manipulation and navigation, VLLR achieves up to 56% absolute success rate gains over the pretrained policy, up to 5% gains over state-of-the-art RL finetuning methods on in-distribution tasks, and up to 10% gains on out-of-distribution tasks, all without manual reward engineering. Additional visualizations can be found in https://silongyong.github.io/vllr_project_page/
Figures & tables
Fig. 1: Method overview. VLLR is a reward model with three components, i.e. sparse task success reward, intrinsic reward and extrinsic reward (Sec. III-D ). The environment provides task instruction and a scene graph context to the LLM for task decomposition (Sec. III-B1 ). Then the decomposed subgoals are fed into a VLM together with the current and last observations (Sec. III-B2 ). The VLM provides a noisy progress estimate which is then smoothed out (Sec. III-B3 ) and used as a reward signal for initializing the value function (Sec. III-B4 ) used for PPO. The intrinsic reward is calculated using the policy’s output action distribution (Sec. III-C ). It is fed together with the task success signal into the PPO algorithm and then produces updates for finetuning the pretrained policy (Sec. III-D ).
Fig. 2: We showcase an example to compare different VLM progress estimation including Qwen, Nova Pro and CLIP. The VLMs are tasked to estimate progress for finding a laptop that is not in plain sight. We found that Qwen tends to over-saturate the estimation, mostly caused by falsely recognizing the laptop, CLIP is unstable, wrong and unable to identify overall task completion, and Nova Pro is the best at progress estimation, which means it provides correct signal when seeing the laptop and approaching it. Here the rollout is collected using A* and the ground truth progress should be linearly increasing.
Fig. 3: We showcase an example where raw progress estimation from Nova Pro is noisy and our method is able to identify the actual progress made by the agent. The task is to find a clock and grasp it. In the first column, we can see that throughout the rollout, the progress estimation is noisy and problematic, unable to provide meaningful reward signals. In the second column, our method is able to provide three clear jumps in terms of progress estimation. The first jump rewards the fact that the robot sees the clock, the second jump rewards the robot positioning itself in front of the clock and the third jump rewards the robot actually picking up the clock. In the third column, we showcase the actual observation Nova Pro is provided. We highlight the clock in red box in frame 1 and 20. Frame 1 showcase the fact that the robot sees the clock and frame 20 demonstrates the robot positioning itself in front of the clock. We connected these two observations to the first two jumps in our progress estimation.
Success (SEL)
IL+RL: Dense Reward
IL+RL: Sparse Reward
IL Only
VLLR (Ours)
FLaRe
PIRLNav
JSRL
SPOC
ObjectNav
87.6 (77.7)
85.0 (67.6)
20.0 (7.0)
21.0 (15.6)
55.0 (42.2)
Fetch
70.7 (63.2)
65.23 (54.7)
0.0 (0.0)
2.9 (2.8)
14.0 (10.5)
PickUp
97.0 (96.4)
91.8 (90.4)
0.0 (0.0)
50.9 (47.7)
90.1 (86.9)
RoomVisit
68.3 (64.8)
60.87 ( 65.7 )
12.5 (11.0)
19.0 (18.6)
40.5 (35.7)
TABLE I: Success and Episode-length weighted Success (SEL) on the CHORES [ 1 ] benchmark. Our method outperforms the previous SoTA methods by providing generalizable reward signal across tasks. Baselines are chosen from FLaRe [ 2 ] .
Success (SEL)
VLLR (Ours)
FLaRe
Poliformer (Sp)
SPOC++
Poliformer (De)
ObjNavRel
67.4 (61.5)
66.3 (60.6)
6.7 (6.7)
54.5 (44.6)
36.1 (32.4)
ObjNavAff
90.1 (78.2)
79.7 (70.6)
35.5 (29.4)
62.4 (50.6)
53.8 (43.1)
TABLE II: Our method can serve as the reward signal for out-of-distribution tasks that are never seen by the base model, and make the policy achieve state-of-the-art performance. Baselines are chosen from FLaRe [ 2 ]
Success (SEL)
Base
SCR
VLM
Full
Fetch
65.23 (54.7)
69.18 (60.4)
68.00 (63.8)
70.7 (63.2)
ObjectNav
85.0 (67.6)
86.7 (76.6)
87.1 (77.5)
87.6 (77.7)
ObjNavAff
79.7 (70.6)
89.7 (76.5)
89.0 (77.1)
90.1 (78.2)
TABLE III: Ablation study on three representative tasks. We exclude PickUp (near-ceiling at 97%) and RoomVisit (SEL characteristics discussed in Sec. IV-C ) as they offer limited signal for component-level comparison. Base: sparse reward only; SCR: self-certainty reward; VLM: VLM progress reward; Full: complete VLLR.
Reinforcement Learning (RL) has shown strong potential for improving robotic manipulation policies, yet its practical use remains bottlenecked by the difficulty of specifying reward functions that are both semantically meaningful and reusable across tasks. In this paper, we propose Large Reward Models (LRMs), a framework that adapts foundation VLMs into frame-level reward generators for robot policy refinement. We specialize a state-of-the-art VLM on a multi-source dataset spanning real-world robot trajectories, human-object interactions, and simulated manipulation environments. Unlike prior approaches that mainly evaluate trajectories post-hoc, LRMs expose multiple reward interfaces from visual observations: progress estimation, task completion, and temporal contrastive comparison. Starting from an imitation-learned policy, we use these VLM-derived rewards to guide PPO refinement on held-out long-horizon manipulation tasks. Our experiments show that LRM progress rewards provide the strongest non-privileged online refinement signal, improving the IL baseline and narrowing the gap to privileged environment rewards. We further deploy progress rewards for progress-weighted behavioral cloning on four real-world manipulation tasks spanning two robot platforms, improving over SFT on all four tasks. These results suggest that modality-specific specialization of foundation VLMs can provide practical visual reward signals for both simulated policy refinement and physical robot self-improvement without hand-coded task rewards.
Fine-tuning vision-language-action (VLA) policies for long-horizon manipulation still relies heavily on behavior cloning, which requires costly high-quality demonstrations and keeps policies near the demonstration distribution. Reward models can reduce this dependence by reweighting demonstrations and providing dense supervision for on-robot reinforcement learning (RL), but they must be dense, accurate, and general. Existing methods fall short: task-specific stage-aware models are accurate but require per-task annotations, while general vision-language-model (VLM) reward models are broadly applicable but too coarse for fine-grained long-horizon progress. We introduce RM, a multi-task stage-aware reward model that combines an action-primitive-based stage estimator with a multi-gate Mixture-of-Experts (MMoE) value head to produce dense per-step rewards across manipulation tasks. Building on RM, we further propose SPIRAL (Self-Policy Improvement via Reward-Aligned Learning), an on-policy reward-guided framework that improves VLA policies from cheap autonomous rollouts. On a 10-task benchmark, RM reduces value-estimation MSE by 80% over the strongest baselines; when used in SPIRAL, it improves task success from around 50% to near-perfect performance on Folding Shorts (58% to 100%) and Cleaning Whiteboard (50% to 90%), showing that high-quality dense rewards are key to a stable robot data flywheel. Project website: https://qianzhong-chen.github.io/sarm2.github.io/.
Qianzhong Chen, Hau Zheng, Justin Yu +8
Stanford University · UC Berkeley · Shanghai Jiao Tong University
Reinforcement learning holds great promise for improving robot policies beyond the limits of imitation learning. However, its practical adoption remains bottlenecked by the lack of reliable vision-language reward models that provide dense and informative feedback. Two key challenges remain: acquiring diverse failure data at scale and obtaining fine-grained reward signals beyond sparse trajectory-level success labels. Collecting failure trajectories typically requires laborious human effort, while pseudo-failures constructed by relabeling successful demonstrations fail to capture the diverse physical failure modes that arise during robot execution. Meanwhile, existing reward models often predict sparse binary or trajectory-level rewards, which provide limited guidance for efficient policy optimization. We introduce DenseReward, a dense robotic reward model that addresses both challenges. To train DenseReward, we develop an automated failure data generation pipeline that synthesizes physically realistic failure trajectories in simulation without human labeling, covering diverse failure modes such as collisions, missed grasps, object drops, and recovery behaviors. DenseReward predicts dense frame-level reward scores from visual observations and language instructions, enabling fine-grained estimation of task progress throughout an episode. Experiments show that DenseReward outperforms general-purpose VLMs and existing robotic reward models in dense reward prediction across both simulated and real-world manipulation. We further demonstrate that DenseReward provides effective reward guidance for downstream model predictive control and reinforcement learning. We release the dataset, trained reward models, and evaluation suite to support the development of failure-aware dense reward modeling for robot learning.
Yu Fang, Wanxi Dong, Jiaqi Liu +7
University of North Carolina at Chapel Hill · Carnegie Mellon University · Shanghai Jiao Tong University +1