Reinforcement Learning (RL) has shown strong potential for improving robotic manipulation policies, yet its practical use remains bottlenecked by the difficulty of specifying reward functions that are both semantically meaningful and reusable across tasks. In this paper, we propose Large Reward Models (LRMs), a framework that adapts foundation VLMs into frame-level reward generators for robot policy refinement. We specialize a state-of-the-art VLM on a multi-source dataset spanning real-world robot trajectories, human-object interactions, and simulated manipulation environments. Unlike prior approaches that mainly evaluate trajectories post-hoc, LRMs expose multiple reward interfaces from visual observations: progress estimation, task completion, and temporal contrastive comparison. Starting from an imitation-learned policy, we use these VLM-derived rewards to guide PPO refinement on held-out long-horizon manipulation tasks. Our experiments show that LRM progress rewards provide the strongest non-privileged online refinement signal, improving the IL baseline and narrowing the gap to privileged environment rewards. We further deploy progress rewards for progress-weighted behavioral cloning on four real-world manipulation tasks spanning two robot platforms, improving over SFT on all four tasks. These results suggest that modality-specific specialization of foundation VLMs can provide practical visual reward signals for both simulated policy refinement and physical robot self-improvement without hand-coded task rewards.
Figures & tables
Fig. 1: Overview of Large Reward Models. LRMs instantiate three modality-specific reward models from the same Qwen3-VL-8B-Instruct initialization: Temporal Contrastive Reward ( rcont ) for relative temporal ranking, Absolute Progress Reward ( rprog ) for dense progress estimation, and Task Completion Reward ( rcomp ) for terminal success verification. Completion and contrastive use separate LoRA adapters on a frozen backbone, while progress additionally updates the vision encoder and multimodal merger. The outputs represent separately trained reward interfaces rather than simultaneous predictions from a multi-task head. During single-interface deployment, one reward model is selected for the downstream setting, and its structured text response is parsed into scalar feedback.
Model
Kendall’s τ
Spearman’s ρ
Qwen3-VL
0.257
0.257
Our LRM
0.296
0.296
Improvement
+15.3%
+15.3%
TABLE I: Contrastive Discrimination performance. LRM specialization improves ranking correlation as measured by Kendall’s τ and Spearman’s ρ , indicating more consistent temporal ordering.
Metric
Qwen3-VL
Our LRM
Δ
Exact Acc
13.10%
13.87%
+0.77%
Acc@ ± 0.2
41.95%
50.58%
+8.63%
MAE
0.378
0.302
− 20.0%
RMSE
0.490
0.395
− 19.3%
TABLE II: Progress estimation against temporal proxy labels. LRM has lower MAE and RMSE and higher prediction accuracy than zero-shot Qwen3-VL.
Fig. 2: Cumulative accuracy at varying tolerance thresholds ( ±Δ ). LRM improves over zero-shot Qwen3-VL by 8.63 percentage points at ±0.2 .
Group
Method
Success Rate
π0.5 SFT Baseline
55.62±3.30
Privileged
Env Reward
60.69±1.90
Zero-shot VLM
Qwen3-VL- rcomp
57.63±2.53
Zero-shot VLM
Qwen3-VL- rcont
53.31±1.20
Zero-shot VLM
Qwen3-VL- rprog
54.37±2.53
LRM
LRM- rcomp
57.00±2.22
TABLE III: Closed-loop success rate (%) on ManiSkill3. Mean ± standard deviation (SD) over 5 evaluation seeds for 1 trained checkpoint per condition. LRM- rprog has the highest non-privileged mean success rate.
Progress Reward Model
Success Rate
RoboReward-8B [ 19 ]
53.37±1.85
Robometer-4B [ 20 ]
57.37±1.18
TOPReward [ 18 ]
51.06±1.63
LRM- rprog
58.94±2.18
TABLE IV: Progress reward comparison on ManiSkill3 (%). All methods use the same PPO setting and K=10 . Mean ± SD over 5 evaluation seeds for 1 trained checkpoint per method; LRM- rprog has the highest mean success rate.
rcomp
rprog
Metric
SFT
RL
SFT
RL
ROC-AUC
0.660
0.795
0.874
0.950
Pairwise Acc (%)
45.4
63.9
80.1
93.4
Global Pearson
0.398
0.757
0.739
0.773
Per-traj Pearson
0.257
0.331
0.577
0.671
TABLE V: Open-loop reward alignment. Fixed LRMs show higher agreement with simulator signals on RL-policy rollouts than on SFT-policy rollouts. Columns identify the rollout policy for each reward interface.
Metric
SFT
RL- rcont
Direction Acc (%)
85.2
85.5
Progress Recall (%)
85.6
86.0
Monotonicity (success)
0.535
0.559
TABLE VI: Open-loop step-level reliability ( rcont ). The fixed contrastive LRM shows similar directional accuracy on SFT and RL rollouts, with slightly higher metrics on RL rollouts.
Reward
K=5
K=10
K=16
rcomp
55.50±2.83
57.00±2.22
56.00±0.81
rcont
57.12±2.05
56.56±1.15
52.62±1.42
rprog
57.81±3.68
58.94±2.18
56.81±2.65
TABLE VII: Query interval sweep (%). Mean ± SD over 5 evaluation seeds for 1 trained checkpoint per configuration. Progress and completion remain stable across the tested intervals, while contrastive reward drops at K=16 , supporting K=10 as a practical trade-off.
Reward Query
Qwen3-VL
LRM
Progress
409±9
360±19
Completion
342±6
340±5
Contrastive
6838±2184
7093±2307
TABLE VIII: Single-query reward inference latency. HTTP round-trip time in ms (mean ± SD). Progress and completion are sub-second; contrastive queries include reasoning generation and are substantially slower.
Fig. 3: Qualitative comparison of real-world rollouts. On the giraffe-in-bowl task, the SFT baseline places the toy beside the target bowl, whereas the policy after LRM-guided progress-weighted fine-tuning places it inside the bowl. On the tissue-out-of-box task, the SFT baseline fails to grasp the tissue, whereas the policy after LRM-guided progress-weighted fine-tuning successfully grasps and pulls it out of the box. These examples illustrate that LRM-guided adaptation corrects distinct policy failure modes and improves real-world task execution.
Task
SFT Baseline
Qwen3-VL- rprog
LRM- rprog
Giraffe-in-bowl
38.3 (23/60)
45.0 (27/60)
56.7 (34/60)
Can-on-cube
43.3 (26/60)
31.7 (19/60)
55.0 (33/60)
Open-drawer
46.7 (28/60)
18.3 (11/60)
61.7 (37/60)
Tissue-out-of-box
8.3 (5/60)
95.0 (57/60)
88.3 (53/60)
TABLE IX: Real-world progress-weighted BC. Success rates (%) with successful trials out of 60. Across two robot platforms, LRM scoring consistently improves over SFT on all four tasks, whereas zero-shot Qwen3-VL scoring produces task-dependent improvements and degradations.