cs.ROSep 28, 2026

ARS: Agentic Reward System for Robot Learning

Authors: Sheng Hu, Weiyi Lu, Lingbing Zeng, Gan Weng, Weiwei Zhang, Kai Xie, Xiaofeng Mou, Yi Xu

Organizations: AIRC, Midea Group

Abstract

Progress reward modeling is the problem of estimating how a robot's behavior changes task progress over time. Reliable estimation requires distinguishing meaningful state changes from failed attempts and task-irrelevant actions. We introduce the Agentic Reward System (ARS), an inference framework for progress reward modeling with general-purpose vision-language models (VLMs), without additional reward-model training. Given an offline trajectory and a task instruction, ARS uses adaptive visual inspection for both event proposal and verification. A subagent proposes a task-relevant event timeline, which a primary agent verifies and revises before estimating per-frame progress. ARS can incorporate optional terminal outcome labels and visual references to inform its judgments. It can also audit progress estimates from external reward models. We evaluate ARS with a 27B VLM on a controlled semantic-mismatch benchmark and downstream policy learning in simulation and on a real robot. The benchmark reveals that several evaluated reward baselines assign spurious progress to wrong-object manipulation even in simple pick-and-place scenes. ARS better suppresses these errors and outperforms these baselines in simulation policy learning. We further demonstrate that ARS supports long-horizon policy learning from mixed-quality offline experience on real-robot multi-screw fastening in a full-scale laboratory replica of an industrial washing-machine assembly line. These results suggest that structured inference and verification can improve the usefulness of general-purpose VLMs for robot reward modeling. Code is at https://github.com/midea-ai/ars

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RARM: Confidence-Gated Progress Reward Modeling for RL in Manipulation

    Jun 20, 2026Pengzhi Yang, Xinyu Wang, Pengyu Jing +7Robotic ManipulationProgress Reward Modeling

  2. Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models

    Mar 17, 2026Yanru Wu, Weiduo Yuan, Esteban Martinez Licon +4Robotic Manipulation PoliciesRobot Systems

  3. Progress Reward Modeling for Robotic Learning: A Comprehensive Survey

    Jul 22, 2026Jianshu Zhang, Keliang Wu, Haoran Lu +8Progress Reward ModelingScalable Robot Learning