ARS: Agentic Reward System for Robot Learning
Organizations: AIRC, Midea Group
Abstract
Progress reward modeling is the problem of estimating how a robot's behavior changes task progress over time. Reliable estimation requires distinguishing meaningful state changes from failed attempts and task-irrelevant actions. We introduce the Agentic Reward System (ARS), an inference framework for progress reward modeling with general-purpose vision-language models (VLMs), without additional reward-model training. Given an offline trajectory and a task instruction, ARS uses adaptive visual inspection for both event proposal and verification. A subagent proposes a task-relevant event timeline, which a primary agent verifies and revises before estimating per-frame progress. ARS can incorporate optional terminal outcome labels and visual references to inform its judgments. It can also audit progress estimates from external reward models. We evaluate ARS with a 27B VLM on a controlled semantic-mismatch benchmark and downstream policy learning in simulation and on a real robot. The benchmark reveals that several evaluated reward baselines assign spurious progress to wrong-object manipulation even in simple pick-and-place scenes. ARS better suppresses these errors and outperforms these baselines in simulation policy learning. We further demonstrate that ARS supports long-horizon policy learning from mixed-quality offline experience on real-robot multi-screw fastening in a full-scale laboratory replica of an industrial washing-machine assembly line. These results suggest that structured inference and verification can improve the usefulness of general-purpose VLMs for robot reward modeling. Code is at https://github.com/midea-ai/ars
Figures & tables
| Group | Method | VOC | FPSR | CMPG |
|---|---|---|---|---|
| Reward-trained | LRM | 0.828 | 0.893 | 0.373 |
| R 2 VLM | 0.916 | 0.184 | -0.050 | |
| SOLE-R1 | 0.902 | 0.361 | 0.266 | |
| Robometer | 0.983 | 0.240 | 0.366 | |
| Audit | ARS w/ Robometer proposal | 0.988 | 0.915 | 0.891 |
| Reward-training-free | TOPReward + Qwen3-VL-8B | 0.930 | 0.368 | 0.685 |
| Group | Method | Success (%) | vs. Vanilla | Retained (%) |
|---|---|---|---|---|
| Baseline | Vanilla BC | 44.2 0.3 | — | 100.0 |
| Reward-trained | LRM | 38.6 0.5 | -5.6 | 13.7 |
| R 2 VLM | 41.5 0.4 | -2.7 | 13.6 | |
| SOLE-R1 | 44.8 0.4 | +0.6 | 52.6 | |
| Robometer | 46.8 0.9 | +2.6 | 89.5 | |
| Reward-training-free | TOPReward + Qwen3-VL-8B | 43.1 0.5 | -1.1 | 77.0 |
| Configuration | VOC | FPSR | CMPG |
|---|---|---|---|
| Robometer -4B | 0.983 | 0.240 | 0.366 |
| ARS + Qwen3.5-4B | 0.986 | 0.625 | 0.633 |
| GVL-style + Qwen3.6-27B | 0.637 | 0.900 | 0.354 |
| ARS + Qwen3.6-27B | 0.995 | 0.918 | 0.904 |
| Configuration | Success (%) | Retained (%) | Queries | Input (K) | Output (K) |
|---|---|---|---|---|---|
| TOPReward + Qwen3-VL-32B | 47.6 0.7 | 67.9 | 390.3 | 2,866.3 | 0.0 |
| ARS w/o verification & inspection | 45.4 0.3 | 78.1 | 7.3 | 242.7 | 10.1 |
| ARS w/o verification | 52.3 0.8 | 70.3 | 10.2 | 413.2 | 9.9 |
| ARS w/o context separation | 51.8 0.5 | 68.7 | 19.0 | 1,097.2 | 15.2 |
| Full ARS | 57.9 0.4 | 66.9 | 24.9 | 980.6 | 21.1 |
| Configuration | Success (%) |
|---|---|
| Demo + all rollouts BC | 50.5 |
| Demo-only BC | 58.0 |
| ARS | 72.0 |
| ARS + visual reference | 81.5 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Main instructions |
|---|---|
| Shared instructions | Ground visual claims in inspected trajectory media. Use the task instruction to identify relevant entities and goal relations. Describe visible state changes and localize events using original frame indices. Request focused temporal windows and spatial crops when additional evidence could change the judgment. Express uncertainty through confidence scores and concise explanations. |
| Event proposal | Construct an event timeline covering the entire trajectory. Organize intervals around meaningful changes in task state, including preparation, control acquisition, sustained manipulation, completion, recovery, and periods without progress. Associate each event with its frame interval, supporting visual evidence, and confidence. |
| Independent verification | Treat proposed events as provisional claims. Acquire visual evidence during verification to examine each event’s description, temporal boundaries, and confidence. Inspect neighboring frames to assess boundary placement. Revise, split, or merge events when necessary, and confirm each resulting event before proceeding. |
| Progress estimation | For sequential tasks, infer subgoals and estimate total progress from their currently achieved contributions, including partial completion. Assign limited credit to useful preparation, localized increases to verified achievements, and gradual increases to continued advancement. Use plateaus for preserved states and decreases for deterioration. Later observations may verify an earlier event; assign credit to the supported event boundary. Preserve meaningful partial achievement in incomplete trajectories. |
| Candidate-curve auditing | When an external progress curve is supplied, inspect its increases, decreases, and terminal values against the visual evidence. Retain supported estimates and revise unsupported intervals. Without a candidate curve, generate progress values for the complete trajectory. |
| Optional auxiliary context | Use terminal outcome labels to constrain completion judgments while determining event timing from visual evidence. Use reference examples to interpret task states and goal relations. Apply reference progress values only when the corresponding state is visibly supported in the evaluated trajectory. |
| Tool | Stage | Function |
|---|---|---|
| read_robot_instruction | E, V, P | Retrieve the task instruction and any explicitly supplied task clarification. |
| read_metadata | E, V, P | Retrieve technical video metadata, including frame count, frame dimensions, presentation frame rate, and frame-indexing conventions. This tool provides no additional task annotations or simulator state. |
| inspect_visual | E, V, P | Retrieve a selected frame interval as a video clip or image tile, with configurable temporal sampling and an optional spatial crop. Return the media with its evidence identifier and frame indices. |
| read_observations | E, V, P | Retrieve the current event timeline and coverage diagnostics. |
| add_observations revise_observations | E, V | Create or revise event intervals, descriptions, evidence links, and confidence scores. |
| merge_observations | V | Combine events that describe one continuous phase. |
| Proposed event | Frame interval |
|---|---|
| Initial state and approach to the stove knob | |
| Knob grasping and turning | |
| Burner switching off and withdrawal | |
| Repositioning and approach to the moka pot | |
| Moka pot grasp acquisition | |
| Lifting and transporting the moka pot |
| Method | Exposure | Evidence |
|---|---|---|
| LRM | N | ManiSkill is described as zero-shot to the reward models ( Wu et al., 2026 ) . |
| R 2 VLM | N | No ManiSkill training is reported for the evaluated checkpoint ( Zhang et al., 2026d ) . |
| SOLE-R1 | N | No ManiSkill training source is identified in the published training mixture ( Schroeder et al., 2026 ) . |
| Robometer | RBM-1M includes FAILSafe trajectories collected in ManiSkill ( Liang et al., 2026 ) . | |
| TOPReward | N | No ManiSkill training is reported in the Qwen model cards; no additional reward-model training is performed ( Chen et al., 2026b ) . |
| ARS | N | No ManiSkill training is reported in the Qwen model cards; no additional reward-model training is performed. |
| Method | |||||
|---|---|---|---|---|---|
| LRM | |||||
| R 2 VLM | |||||
| SOLE-R1 , | |||||
| Robometer | |||||
| TOPReward -8B, linear | |||||
| TOPReward -32B, linear |
| Method | Exposure | Evidence |
|---|---|---|
| LRM | LIBERO is an explicit source in the training mixture ( Wu et al., 2026 ) . | |
| R 2 VLM | N | No LIBERO training is reported for the evaluated checkpoint ( Zhang et al., 2026d ) . |
| SOLE-R1 | Supervised fine-tuning includes 504,544 ecot_libero_all examples ( Schroeder et al., 2026 ) . | |
| Robometer | RBM-1M includes LIBERO-{Long, Object, Spatial, Goal} and generated failures ( Liang et al., 2026 ) . | |
| TOPReward | N | No LIBERO training is reported in the Qwen model cards; no additional reward-model training is performed ( Chen et al., 2026b ) . |
| ARS | N | No LIBERO training is reported in the Qwen model cards; no additional reward-model training is performed. |
| Original instruction | Appended clarification |
|---|---|
| put both the alphabet soup and the tomato sauce in the basket | alphabet soup: blue-and-orange can with alphabet-letter soup graphics. tomato sauce: red/green/white can with tomato pictures. |
| put both the alphabet soup and the cream cheese box in the basket | alphabet soup: blue-and-orange can with alphabet-letter soup graphics. cream cheese: white rectangular box with a large purple-and-blue label. |
| put both the cream cheese box and the butter in the basket | cream cheese: white rectangular box with a large purple-and-blue label. butter: red-orange rectangular box with a blue-and-white panel. |
| put the white mug on the left plate and put the yellow and white mug on the right plate | Plate sides are robot-relative: white mug goes to robot-left (image right); yellow-and-white mug goes to robot-right (image left). |
| Method | Success (%) |
|---|---|
| Vanilla BC | 40.1 1.8 |
| LRM | 36.8 1.5 |
| R 2 VLM | 34.1 1.1 |
| SOLE-R1 | 40.4 2.8 |
| Robometer | 42.1 2.3 |
| TOPReward + Qwen3-VL-8B | 37.9 2.5 |
| Configuration | Queries | Input (K) | Output (K) | Input / query (K) |
|---|---|---|---|---|
| Full ARS : subagent | 5.8 | 187.0 | 3.9 | 32.5 |
| Full ARS : primary agent | 19.2 | 793.6 | 17.1 | 41.4 |
| Full ARS : total | 24.9 | 980.6 | 21.1 | 39.3 |
| ARS w/o context separation | 19.0 | 1,097.2 | 15.2 | 57.6 |
| Configuration | VOC | FPSR | CMPG |
|---|---|---|---|
| ARS w/o verification & inspection | 0.991 | 0.915 | 0.895 |
| ARS w/o verification | 0.995 | 0.916 | 0.910 |
| ARS w/o context separation | 0.993 | 0.913 | 0.910 |
| Full ARS | 0.995 | 0.918 | 0.904 |
| Frame order | Success (%) | Retained (%) |
|---|---|---|
| Chronological | 0.0 | 3.58 |
| Shuffled | 14.80 |
| Thinking mode | Semantic-mismatch | LIBERO-Long | |||
|---|---|---|---|---|---|
| VOC | FPSR | CMPG | Success (%) | Retained (%) | |
| Disabled | 0.990 | 0.915 | 0.877 | 55.4 0.7 | 62.0 |
| Enabled | 0.995 | 0.918 | 0.904 | 57.9 0.4 | 66.9 |