cs.CVSep 29, 2026

HaPRL: Human-Anchored Process Reinforcement Learning for Visual Search Agent

Authors: Zhangquan Chen, Yaoxin Niu, Xiang An, Mingze Sun, Zhumei Wang, Chih-Ting Liao, Hongkun Cao, Ruqi Huang

Organizations: Tsinghua University · Peng Cheng Laboratory · LMMs-Lab · Beijing Institute of Technology · University of New South Wales

Abstract

Multi-turn visual search agents answer questions about high-resolution images by iteratively deciding where to look. Reinforcement learning for these agents rewards only the final answer, leaving the search process unsupervised. Consequently, faulty routes in which the reasoning process is erroneous yet the final result is correct arise frequently, which in turn leads to ineffective training, i.e., scaling along the wrong paths. In this paper, we introduce HaPRL, the first framework to reinforce the search process with human search behavior. We first build an annotation platform and collect 1K+ human-annotated data with fine-grained behavioral signals. During training, a carefully designed judge scores each rollout with task-adaptive weights, anchored on the distilled trace of how a human annotator actually searched the same image. Extensive experiments show that HaPRL consistently outperforms outcome-based RL, and early-stage process supervision yields 6.7x more improvement in subsequent outcome-based scaling. Our results also demonstrate the importance of aligning model behavior with human process annotation signals, which offer new insight into the training of foundation models.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Dense Process Supervision for Search Agents via Fact Utility Estimation

    Sep 1, 2026Rongzhi Zhu, Xiangyu Liu, Yi Liu +7Process Reward ModelCredit Assignment

  2. ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards

    Apr 22, 2026Wentao Yan, Shengqin Wang, Huichi Zhou +4Multimodal Search AgentsVisual Reasoning