cs.ROMar 17, 2026

Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models

Authors: Yanru Wu, Weiduo Yuan, Esteban Martinez Licon, Ang Qi, Vitor Guizilini, Jiageng Mao, Yue Wang

Organizations: USC Physical Superintelligence Lab

Abstract

Reinforcement Learning (RL) has shown strong potential for improving robotic manipulation policies, yet its practical use remains bottlenecked by the difficulty of specifying reward functions that are both semantically meaningful and reusable across tasks. In this paper, we propose Large Reward Models (LRMs), a framework that adapts foundation VLMs into frame-level reward generators for robot policy refinement. We specialize a state-of-the-art VLM on a multi-source dataset spanning real-world robot trajectories, human-object interactions, and simulated manipulation environments. Unlike prior approaches that mainly evaluate trajectories post-hoc, LRMs expose multiple reward interfaces from visual observations: progress estimation, task completion, and temporal contrastive comparison. Starting from an imitation-learned policy, we use these VLM-derived rewards to guide PPO refinement on held-out long-horizon manipulation tasks. Our experiments show that LRM progress rewards provide the strongest non-privileged online refinement signal, improving the IL baseline and narrowing the gap to privileged environment rewards. We further deploy progress rewards for progress-weighted behavioral cloning on four real-world manipulation tasks spanning two robot platforms, improving over SFT on all four tasks. These results suggest that modality-specific specialization of foundation VLMs can provide practical visual reward signals for both simulated policy refinement and physical robot self-improvement without hand-coded task rewards.

Figures & tables

Explore similar work

CardsList
  1. RMTL: Reinforced Micro-task Learning for Long-Horizon Manipulation with VLM Rewards

    Jun 24, 2026Anıl Can Ateş, Orhan Kahraman, Cihan TopalRobotic ManipulationLong-Horizon Manipulation

  2. SARM2: Multi-Task Stage Aware Reward Modeling for Self Improving Robotic Manipulation

    Jun 9, 2026Qianzhong Chen, Hau Zheng, Justin Yu +8Robotic ManipulationScalable Robot Learning

  3. RARM: Confidence-Gated Progress Reward Modeling for RL in Manipulation

    Jun 20, 2026Pengzhi Yang, Xinyu Wang, Pengyu Jing +7Robotic ManipulationProgress Reward Modeling