Multi-turn visual search agents answer questions about high-resolution images by iteratively deciding where to look. Reinforcement learning for these agents rewards only the final answer, leaving the search process unsupervised. Consequently, faulty routes in which the reasoning process is erroneous yet the final result is correct arise frequently, which in turn leads to ineffective training, i.e., scaling along the wrong paths. In this paper, we introduce HaPRL, the first framework to reinforce the search process with human search behavior. We first build an annotation platform and collect 1K+ human-annotated data with fine-grained behavioral signals. During training, a carefully designed judge scores each rollout with task-adaptive weights, anchored on the distilled trace of how a human annotator actually searched the same image. Extensive experiments show that HaPRL consistently outperforms outcome-based RL, and early-stage process supervision yields 6.7x more improvement in subsequent outcome-based scaling. Our results also demonstrate the importance of aligning model behavior with human process annotation signals, which offer new insight into the training of foundation models.
Figures & tables
VisualProbe
High-resolution QA
Model
Easy
Medium
Hard
V ∗ QA
HR-4K
HR-8K
Avg.
Qwen3-VL-4B-Instruct
Base
36.88
16.42
13.21
67.16
55.50
51.25
40.07
+SFT
27.66
14.55
6.60
46.07
48.88
44.38
31.36
+SFT+Outcome-RL
29.08
17.91
8.49
58.64
47.13
42.63
33.98
+SFT+HaPRL
41.13
26.49
20.75
68.06
56.50
52.63
44.26
(+10.28)
Table 1: Accuracy comparison of generalist VLMs, SFT, +outcome-based RL, and +our method (HaPRL) on VisualProbe Easy/Medium/Hard ( Lai et al., 2025 ) , V ∗ QA ( Wu and Xie, 2024 ) , HR-4K/8K ( Wang et al., 2025c ) . All RL post-training methods are trained for three epochs. Avg. is the unweighted mean over the six sets, and the parenthesized value is the HaPRL margin over the matched Outcome-RL control. Best per column of the same backbone in bold .
VisualProbe
High-resolution QA
Model
Easy
Medium
Hard
V ∗ QA
HR-4K
HR-8K
Avg.
+SFT
84.90
86.05
71.54
80.34
86.49
82.64
81.99
+SFT+Outcome-RL
83.55
86.92
65.56
83.22
87.24
85.96
82.08
Δ
− 1.35
+0.87
− 5.98
+2.88
+0.75
+3.32
+0.08
+SFT+HaPRL
86.71
87.27
75.28
83.24
88.66
87.47
84.77
Δ
+1.81
+1.22
+3.74
+2.90
+2.17
+4.83
+2.78
Table 2: The process score (%) on answer-correct rollouts of different methods on Qwen3-VL-4B; Δ against the cold start. The best results are in bold . Outcome-RL raises answer accuracy while leaving some search quality flat, whereas HaPRL maintains a positive gain on every split.
VisualProbe
High-resolution QA
Annotated prompts
Easy
Medium
Hard
V ∗ QA
HR-4K
HR-8K
Avg.
200
31.21
14.18
16.98
57.07
51.38
48.88
36.62
400
32.70
15.19
17.92
56.97
52.88
48.12
37.30
600
35.46
17.16
18.87
60.40
54.25
49.13
39.21
800
37.99
20.19
19.21
63.40
55.88
50.94
41.27
1000
40.15
24.34
20.55
65.69
56.12
51.68
43.09
Table 3: Accuracy of three-epoch training at different data sizes on Qwen3-VL-4B. Accuracy rises monotonically with the number of process-annotated prompts. The best is bold .
Figure 3: HaPRL scales with the amount of process-annotated data , and injecting human-aligned process data in the early stage enables stronger scaling even when later training uses only the outcome signal. (a) Accuracy grows monotonically with the process-annotation budget. 1/5 annotated data already beat Outcome-RL trained on all of the them. (b) A second, outcome-only stage returns +0.91 points to an Outcome-RL-trained policy and +6.13 points to a HaPRL-trained one ( 6.7× ).
VisualProbe
High-resolution QA
Training schedule
Easy
Medium
Hard
V ∗ QA
HR-4K
HR-8K
Avg.
Stage 1 only ( 400 prompts)
Outcome-RL
28.35
14.94
11.21
54.84
46.12
40.30
32.63
HaPRL
32.70
15.19
17.92
56.97
52.88
44.12
36.63
Δ
+4.35
+0.25
+6.71
+2.13
+6.76
+3.82
+4.00
Stage 1 → outcome-only stage 2 ( 600 prompts)
Table 4: Accuracy of one-epoch training on Qwen3-VL-4B. Stage 1 trains on 400 data under either reward (outcome-based vs. HaPRL); stage 2 continues under the same outcome reward on another 600 data. The best is bold .
Variant
Human trace
Correct pairing
Avg.
Δ
Shuffled process labels
✓
28.63
− 15.63
Judge-only process
36.41
− 7.85
+10% trace noise
✓
✓
41.74
− 2.52
HaPRL (full)
✓
✓
44.26
–
Table 5: Ablation study of different training variants. Six-benchmark average accuracy of three-epoch training on Qwen3-VL-4B; Δ against full HaPRL. Shuffled re-pairs each reference trace with another question’s trajectory, keeping the label distribution intact. Judge-only removes the trace from the judge prompt and leaves the rubric and task weighting intact. Noise perturbs 10% of the recorded dwell and hover events.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Profile
Tar.
Evd.
Prog.
Tool
Com.
General search
.20
.30
.25
.15
.10
OCR / text
.15
.40
.20
.15
.10
Count / relation
.20
.35
.25
.10
.10
Attribute
.25
.30
.20
.15
.10
Appendix
Table 6: Task-adaptive weights over the five process dimensions. Columns denote target semantics (Tar.), evidence acquisition (Evd.), search progress (Prog.), tool discipline (Tool), and communication discipline (Com.).
Setting
Value
Algorithm
GRPO ( Shao et al., 2024b )
Framework
verl ( Sheng et al., 2025 ) , vLLM rollout ( Kwon et al., 2023 )
Rollouts per prompt
8
Prompt batch size
48
PPO mini / micro batch
8 / 1 per GPU
Actor learning rate
5×10−7
Appendix
Table 7: RL configuration. The last row is the only difference between the two conditions.
Reward reference
Specifies
Judge
Avg.
Δ vs. HaPRL
None (Outcome-RL)
—
—
33.98
−10.28
Hand-designed tool-use bonus
crop novelty
—
35.72
−8.54
Rubric only, no reference
rubric dimensions
✓
36.41
−7.85
Final evidence box, coverage reward
location
—
38.61
−5.65
Teacher-generated evidence path
location, ordering
✓
39.85
−4.41
Human ordered trace (HaPRL)
location, ordering, dead ends
✓
44.26
—
Appendix
Table 8: Reward reference and what it specifies about the search. All arms train Qwen3-VL-4B for three epochs on the same 1,104 prompts and differ only in what the process reward is anchored to. Avg. is the macro average over the six evaluation splits. The best is bold .
Cold-start recipe
SFT
+ Outcome-RL
+ HaPRL
Δ (HaPRL − Outcome-RL)
A: filtered corpus, five epochs (ours)
31.36
33.98
44.26
+10.28
B: full corpus, over-turn masking, two epochs
39.84
43.15
49.02
+5.87
Appendix
Table 9: Cold-start recipe crossed with reward condition on Qwen3-VL-4B. Recipe A is the cold start used throughout the main text; Recipe B keeps the unfiltered corpus, masks the loss beyond the six-round cap, and stops at two epochs. The backbone scores 40.07 before any agent training. Entries are macro averages over the six evaluation splits.
Training schedule
Seed 1
Seed 2
Seed 3
Mean ± SD
Stage-2 gain
Stage 1 only ( 400 prompts)
Outcome-RL
32.63
32.05
32.55
32.41±0.31
—
HaPRL
36.63
35.94
36.48
36.35±0.36
—
Stage 1 → outcome-only stage 2 ( 600 prompts)
Outcome-RL → Outcome-RL
33.54
32.86
33.47
33.29±0.37
+0.88±0.06
HaPRL → Outcome-RL
42.76
41.62
42.16
42.18±0.57
+5.83±0.26
Appendix
Table 10: Three seeds of the two-stage schedule of Table 4 on Qwen3-VL-4B. Seed 1 is the run reported in the main text. Entries are macro averages over the six evaluation splits; the last column is the stage-2 gain of each row over its own stage-1 checkpoint. Mean ± SD over three seeds.
Rounds / crops and verdict
Fig.
Task profile
Source
+SFT
+Outcome-RL
+HaPRL
4
OCR / text
HR-Bench 4K-500
3 / 2 ×
6 / 5 ✓
4 / 3 ✓
5
OCR / text
HR-Bench 8K-39
4 / 3 ×
7 / 6 ×
3 / 2 ✓
6
OCR / text
VisualProbe Easy-128
4 / 3 ×
7 / 6 ×
2 / 1 ✓
7
OCR / relation
VisualProbe Hard-54
7 / 6 ×
6 / 5 ✓
2 / 1 ✓
8
Attribute
VisualProbe Medium-151
7 / 6 ×
7 / 6 ×
3 / 2 ✓
Appendix
Table 11: The six qualitative cases. Rounds and crops are counted from the rendered trajectories in Figures 4 – 9 . In every case HaPRL reaches the correct answer with the shortest evidence path of the three conditions.
Figure 4: Fine-grained OCR through visual search (HR-Bench 4K-500). The cold start zooms twice, misreads the word on the wing as “SUNNY”, and answers the wrong option. Outcome-RL needs five crops, two of which land on water and rock, before reading the word correctly. HaPRL re-centers once on the upper-middle region, targets the wing text, and answers in four rounds.
Figure 5: Fine-grained device-brand OCR (HR-Bench 8K-39). The cold start over-zooms until the text is illegible and guesses “Sony”. Outcome-RL repeats the same bottom-region crop three times, abandons it, jumps to an unrelated device, and exhausts its round budget without a valid answer. HaPRL covers the left-side device, recognizes the vertical insignia text, and answers in three rounds.
Figure 6: Historical-date OCR (VisualProbe Easy-128). The cold start zooms into irrelevant publication text and reports a date that is not in the image. Outcome-RL chases several date-like lines across six crops and returns the truncated “March 3”. HaPRL crops the article’s displayed header directly and reads “March, 1935” with one crop.
Figure 7: Relational OCR in a crowded scene (VisualProbe Hard-54). The question asks for the text to the right of beer , so the two words must be legible in one view. The cold start re-crops the same unreadable region four times and returns no valid answer. Outcome-RL reaches the right answer but repeats a nearly identical crop and spends a further round confirming what it had already read. HaPRL crops the joint sign region once, which makes the relation readable immediately.
Figure 8: Text-color recognition on a dense shelf (VisualProbe Medium-151). Both controls lose the target among repeated shelf crops and end without a usable answer, the cold start concluding the color cannot be determined. HaPRL first locates the shelf holding the queried word, then tightens while keeping the word visible, and reads its color in three rounds.
Figure 9: Small-object color disambiguation (V ∗ direct_attributes-58). Both controls fix on a salient flagpole rather than the queried flag and report the colors of the wrong object, arriving at the same incorrect option by different routes. HaPRL crops the roadside flag directly and answers in two rounds. A wrong target reached fluently is the failure the rubric’s target-semantics dimension is meant to price.
State Key Laboratory for Novel Software Technology, Nanjing University, China · Ant Group, China · National Institute of Healthcare Data Science, Nanjing University, China