Reinforcement Learning (RL) has shown strong potential for improving robotic manipulation policies, yet its practical use remains bottlenecked by the difficulty of specifying reward functions that are both semantically meaningful and reusable across tasks. In this paper, we propose Large Reward Models (LRMs), a framework that adapts foundation VLMs into frame-level reward generators for robot policy refinement. We specialize a state-of-the-art VLM on a multi-source dataset spanning real-world robot trajectories, human-object interactions, and simulated manipulation environments. Unlike prior approaches that mainly evaluate trajectories post-hoc, LRMs expose multiple reward interfaces from visual observations: progress estimation, task completion, and temporal contrastive comparison. Starting from an imitation-learned policy, we use these VLM-derived rewards to guide PPO refinement on held-out long-horizon manipulation tasks. Our experiments show that LRM progress rewards provide the strongest non-privileged online refinement signal, improving the IL baseline and narrowing the gap to privileged environment rewards. We further deploy progress rewards for progress-weighted behavioral cloning on four real-world manipulation tasks spanning two robot platforms, improving over SFT on all four tasks. These results suggest that modality-specific specialization of foundation VLMs can provide practical visual reward signals for both simulated policy refinement and physical robot self-improvement without hand-coded task rewards.
Figures & tables
Fig. 1: Overview of Large Reward Models. LRMs instantiate three modality-specific reward models from the same Qwen3-VL-8B-Instruct initialization: Temporal Contrastive Reward ( rcont ) for relative temporal ranking, Absolute Progress Reward ( rprog ) for dense progress estimation, and Task Completion Reward ( rcomp ) for terminal success verification. Completion and contrastive use separate LoRA adapters on a frozen backbone, while progress additionally updates the vision encoder and multimodal merger. The outputs represent separately trained reward interfaces rather than simultaneous predictions from a multi-task head. During single-interface deployment, one reward model is selected for the downstream setting, and its structured text response is parsed into scalar feedback.
Model
Kendall’s τ
Spearman’s ρ
Qwen3-VL
0.257
0.257
Our LRM
0.296
0.296
Improvement
+15.3%
+15.3%
TABLE I: Contrastive Discrimination performance. LRM specialization improves ranking correlation as measured by Kendall’s τ and Spearman’s ρ , indicating more consistent temporal ordering.
Metric
Qwen3-VL
Our LRM
Δ
Exact Acc
13.10%
13.87%
+0.77%
Acc@ ± 0.2
41.95%
50.58%
+8.63%
MAE
0.378
0.302
− 20.0%
RMSE
0.490
0.395
− 19.3%
TABLE II: Progress estimation against temporal proxy labels. LRM has lower MAE and RMSE and higher prediction accuracy than zero-shot Qwen3-VL.
Fig. 2: Cumulative accuracy at varying tolerance thresholds ( ±Δ ). LRM improves over zero-shot Qwen3-VL by 8.63 percentage points at ±0.2 .
Group
Method
Success Rate
π0.5 SFT Baseline
55.62±3.30
Privileged
Env Reward
60.69±1.90
Zero-shot VLM
Qwen3-VL- rcomp
57.63±2.53
Zero-shot VLM
Qwen3-VL- rcont
53.31±1.20
Zero-shot VLM
Qwen3-VL- rprog
54.37±2.53
LRM
LRM- rcomp
57.00±2.22
TABLE III: Closed-loop success rate (%) on ManiSkill3. Mean ± standard deviation (SD) over 5 evaluation seeds for 1 trained checkpoint per condition. LRM- rprog has the highest non-privileged mean success rate.
Progress Reward Model
Success Rate
RoboReward-8B [ 19 ]
53.37±1.85
Robometer-4B [ 20 ]
57.37±1.18
TOPReward [ 18 ]
51.06±1.63
LRM- rprog
58.94±2.18
TABLE IV: Progress reward comparison on ManiSkill3 (%). All methods use the same PPO setting and K=10 . Mean ± SD over 5 evaluation seeds for 1 trained checkpoint per method; LRM- rprog has the highest mean success rate.
rcomp
rprog
Metric
SFT
RL
SFT
RL
ROC-AUC
0.660
0.795
0.874
0.950
Pairwise Acc (%)
45.4
63.9
80.1
93.4
Global Pearson
0.398
0.757
0.739
0.773
Per-traj Pearson
0.257
0.331
0.577
0.671
TABLE V: Open-loop reward alignment. Fixed LRMs show higher agreement with simulator signals on RL-policy rollouts than on SFT-policy rollouts. Columns identify the rollout policy for each reward interface.
Metric
SFT
RL- rcont
Direction Acc (%)
85.2
85.5
Progress Recall (%)
85.6
86.0
Monotonicity (success)
0.535
0.559
TABLE VI: Open-loop step-level reliability ( rcont ). The fixed contrastive LRM shows similar directional accuracy on SFT and RL rollouts, with slightly higher metrics on RL rollouts.
Reward
K=5
K=10
K=16
rcomp
55.50±2.83
57.00±2.22
56.00±0.81
rcont
57.12±2.05
56.56±1.15
52.62±1.42
rprog
57.81±3.68
58.94±2.18
56.81±2.65
TABLE VII: Query interval sweep (%). Mean ± SD over 5 evaluation seeds for 1 trained checkpoint per configuration. Progress and completion remain stable across the tested intervals, while contrastive reward drops at K=16 , supporting K=10 as a practical trade-off.
Reward Query
Qwen3-VL
LRM
Progress
409±9
360±19
Completion
342±6
340±5
Contrastive
6838±2184
7093±2307
TABLE VIII: Single-query reward inference latency. HTTP round-trip time in ms (mean ± SD). Progress and completion are sub-second; contrastive queries include reasoning generation and are substantially slower.
Fig. 3: Qualitative comparison of real-world rollouts. On the giraffe-in-bowl task, the SFT baseline places the toy beside the target bowl, whereas the policy after LRM-guided progress-weighted fine-tuning places it inside the bowl. On the tissue-out-of-box task, the SFT baseline fails to grasp the tissue, whereas the policy after LRM-guided progress-weighted fine-tuning successfully grasps and pulls it out of the box. These examples illustrate that LRM-guided adaptation corrects distinct policy failure modes and improves real-world task execution.
Task
SFT Baseline
Qwen3-VL- rprog
LRM- rprog
Giraffe-in-bowl
38.3 (23/60)
45.0 (27/60)
56.7 (34/60)
Can-on-cube
43.3 (26/60)
31.7 (19/60)
55.0 (33/60)
Open-drawer
46.7 (28/60)
18.3 (11/60)
61.7 (37/60)
Tissue-out-of-box
8.3 (5/60)
95.0 (57/60)
88.3 (53/60)
TABLE IX: Real-world progress-weighted BC. Success rates (%) with successful trials out of 60. Across two robot platforms, LRM scoring consistently improves over SFT on all four tasks, whereas zero-shot Qwen3-VL scoring produces task-dependent improvements and degradations.
Reinforcement learning (RL) for robotic manipulation often requires manually designing a dense reward function, which is difficult to tune and often fragile, or learning a reward from human demonstrations or preferences, which can be expensive. A recent line of work uses pretrained vision-language models (VLMs) as zero-shot reward models, replacing these costs with a single text prompt. However, we argue that a single global prompt is too coarse for long-horizon manipulation tasks with randomized initial conditions. The single-prompt VLM reward is near-flat for much of the trajectory, making early progress hard for the agent to detect. We propose Reinforced Micro-Task Learning (RMTL), an approach that decomposes a manipulation task into a small set of language-described micro-tasks and trains the agent to switch between them. At each step, the agent receives a multi-view VLM reward computed using the prompt of the currently active micro-task and averaged across multiple camera views to reduce the effect of view-specific occlusions. A reverse curriculum gradually exposes the agent to harder initial conditions, while a PPO worker is first trained with a fixed distance-based rule that selects the active micro-task. We then replace this rule with a learned hierarchical manager, turning rule-based phase selection into a fully learned hierarchical policy. We instantiate RMTL on the Fetch manipulation environment using three short stage-specific prompts and without additional prompt tuning. Experiments show that RMTL provides more informative reward signals than single-prompt VLM rewards, enabling faster learning. These results suggest that decomposing VLM rewards into micro-task-specific language prompts can substantially improve the scalability of language-guided reinforcement learning for robotic manipulation.
Fine-tuning vision-language-action (VLA) policies for long-horizon manipulation still relies heavily on behavior cloning, which requires costly high-quality demonstrations and keeps policies near the demonstration distribution. Reward models can reduce this dependence by reweighting demonstrations and providing dense supervision for on-robot reinforcement learning (RL), but they must be dense, accurate, and general. Existing methods fall short: task-specific stage-aware models are accurate but require per-task annotations, while general vision-language-model (VLM) reward models are broadly applicable but too coarse for fine-grained long-horizon progress. We introduce RM, a multi-task stage-aware reward model that combines an action-primitive-based stage estimator with a multi-gate Mixture-of-Experts (MMoE) value head to produce dense per-step rewards across manipulation tasks. Building on RM, we further propose SPIRAL (Self-Policy Improvement via Reward-Aligned Learning), an on-policy reward-guided framework that improves VLA policies from cheap autonomous rollouts. On a 10-task benchmark, RM reduces value-estimation MSE by 80% over the strongest baselines; when used in SPIRAL, it improves task success from around 50% to near-perfect performance on Folding Shorts (58% to 100%) and Cleaning Whiteboard (50% to 90%), showing that high-quality dense rewards are key to a stable robot data flywheel. Project website: https://qianzhong-chen.github.io/sarm2.github.io/.
Qianzhong Chen, Hau Zheng, Justin Yu +8
Stanford University · UC Berkeley · Shanghai Jiao Tong University
Reinforcement learning for robot manipulation is often bottlenecked by reward design, especially in long-horizon tasks: sparse success rewards provide weak supervision, while hand-crafted dense rewards are tedious to design and generalize poorly across tasks. Progress-based reward models offer a promising alternative by estimating how far an observation has advanced toward task completion, but existing approaches often require task-specific demonstrations or progress labels, and can assign high rewards to visually plausible but physically incorrect states. We introduce the Reference-Anchored Reward Model (RARM), a lightweight visual comparator that converts a single successful demonstration into a dense, progress-aware reward. RARM is trained once on general-purpose videos with a contrastive temporal objective, requiring no robot-specific data, task-specific reward labels, or per-task reward engineering. At deployment, RARM matches rollout clips to reference clips and rewards only confident forward progress, suppressing uncertain matches that may otherwise produce false-positive rewards. Across 9 simulated manipulation tasks from LIBERO and MetaWorld and 4 real-world tasks, RARM achieves the best overall success rates in subsequent RL training, with particularly large gains on long-horizon tasks such as cloth folding, where unreliable progress estimates are especially harmful.
Pengzhi Yang, Xinyu Wang, Pengyu Jing +7
† NUS Human-Centered Robotic Lab · ‡ Booking.com · § University of Cambridge +1