High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover detail, but it asks the main model to choose a region before obtaining a reliable overview. We introduce VPS, a visual parallel-search framework in which a main agent first invokes grid_search to inspect image tiles in parallel with question-conditioned sub-agents, and then adaptively invokes zoom_in on a precise or merged region. The same interface supports both training-free inference and post-training of the main and sub-agents. Across five benchmark splits and three model sizes, VPS improves mean accuracy over dedicated zoom-only search in 14 of 15 same-model comparisons, with gains up to 8.0 points and especially strong improvements for smaller main models. ZoomBench retains an approximately 3.2-point gain at every tested size. We further develop a supervision pipeline with hint-free verification and a paired role-specific GRPO surrogate for learning the controller and tile-reader roles. SFT improves observed accuracy on all five benchmark splits, including a 4.17-point gain on HR-Bench 4K. Role-specific RL further reshapes search behavior: main-only RL reduces mean tool use from 2.65 to 2.11 with similar pass@1 in an internal four-response evaluation, while external accuracy changes are mixed. Joint training reveals an asymmetry between local evidence reading and global search control. Together, these results support VPS as an effective inference-time scaffold and a trainable decomposition for visual evidence acquisition.
Figures & tables
Figure 1: VPS combines parallel tile inspection with adaptive zoom. The main agent aggregates local evidence and selects a precise or merged region for inspection. The image and question are from TIR-Bench visual search, sample ID 350 (gold answer: 17; Li et al., 2025 ). Reports, selected regions, and the comparison strip illustrate the mechanism rather than reproduce a measured rollout.
Figure 2: Paired role-specific GRPO. Main questions and fixed external tile prompts form separate groups of K responses, each normalized within an identical prompt. Main tokens receive final-answer supervision; fresh tile responses receive local supervision. The two losses update shared weights. The crossed path marks omitted global credit to nested tile-reader tokens.
Main
Recipe
TIR
VSTAR
ZoomBench
HR-4K
HR-8K
27B
Grid + zoom
76.1±0.96
94.4±1.60
67.8±0.94
90.1±0.50
89.3±0.59
Grid + zoom †
77.8±2.93
96.7±1.32
67.9±0.30
90.4±0.07
90.9±0.40
Zoom only
76.9±0.48
93.5±0.60
64.6±1.30
88.8±0.19
88.8±0.40
Δ
−0.8
+0.9
+3.2
+1.3
+0.5
9B
Grid + zoom
66.4±2.93
85.3±2.28
58.4±1.54
84.0±0.80
80.5±0.75
Zoom only
64.7±1.92
77.3±5.45
55.2±1.66
82.0±0.47
78.3±1.45
Table 1: Judge-all accuracy (%) for the two training-free search policies (Appendix C.1 ). Grid + zoom uses a same-size tile reader; † replaces it with Gemini 3.5 Flash. Values are mean ± sample standard deviation over three runs; Δ compares the same-model grid and zoom-only arms using unrounded means. Shading marks same-model grid search; bold marks the higher mean in each same-model pair; blue/orange deltas mark increases/decreases, not significance.
Figure 3: Grid + zoom gain over dedicated zoom-only search. (a) Differences in mean accuracy from Table 1 , in percentage points, with 14 of 15 cells positive. (b) The same gains across model sizes; lines connect observations, without a fitted trend. ZoomBench retains an approximately +3.2 -point gain. Colors encode differences, not statistical significance. Per-arm replicate settings are in Appendix C.1 .
Main
Search
4K
8K
4B
Grid + zoom
59.0
52.8
4B
Dedicated zoom
48.2
40.2
9B
Grid + zoom
68.3
62.0
9B
Dedicated zoom
65.3
57.2
27B
Grid + zoom
81.2
79.3
27B
Dedicated zoom
80.8
78.8
Table 2: HR-Bench CircularEval means (%), computed over 200 questions per subset. A question counts as correct only when all four option rotations are correct. These values are separate from the 800-row accuracy in Table 1 . Shading marks grid search; bold marks the higher mean within each model size.
(a) Cross-benchmark evaluations
Main policy
Sub policy
TIR
VSTAR
ZoomBench
HR-4K
HR-8K
Pretrained
Pretrained
59.44
78.36
51.01
78.58
76.12
SFT
SFT
72.50
87.96
60.00
82.75
82.25
Main-only RL
SFT
74.17
86.91
61.66
see (b)
82.75
Table 3: Pretrained, SFT, and RL policies (accuracy, %). (a) Descriptive cross-benchmark results: three-run pretrained means and single post-training runs. (b) Main-only / early mixed RL role matrix, first of three passes. (c) Selected joint RL role matrix, one pass with its own SFT baseline. VSTAR uses all 191 gold examples; failures count as wrong. Appendix C.6 gives per-run counts and evaluation settings. Shading identifies SFT references; bold marks the highest observed accuracy per column within each panel, not significance.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Main
Search
Replicates
TIR /120
VSTAR /191
ZoomBench /845
4B
Grid + zoom
seeds 1–3
74/75/65
146/155/148
427/441/425
4B
Dedicated zoom
reps 1–3
62/65/62
137/131/135
411/391/409
9B
Grid + zoom
seeds 1–3
83/80/76
161/160/168
493/481/507
9B
Dedicated zoom
base, seeds 2–3
79/75/79
151/156/136
462/455/482
27B
Grid + zoom
seeds 1–3
90/92/92
183/181/177
582/570/567
27B
Dedicated zoom
seeds 1–3
92/93/92
180/178/178
552/552/533
Appendix
Table 4: Training-free run numerators. Within each cell the three counts share the denominator in the column heading. Grid runs on these three benchmarks are seeded; the 4B dedicated zoom-only runs are unseeded. The 9B zoom-only first run is an unseeded baseline followed by two seeded runs. Shading identifies the grid-search reference rows.
Main
Search
4K /800
8K /800
4B
Grid + zoom
627/638/621
610/607/610
4B
Dedicated zoom
592/587/587
553/549/554
9B
Grid + zoom
679/669/667
650/638/645
9B
Dedicated zoom
653/654/660
614/628/637
27B
Grid + zoom
717/725/721
716/709/718
27B
Dedicated zoom
709/710/712
713/712/707
Appendix
Table 5: HR-Bench run numerators for the same training-free arms. All repetitions are unseeded and each subset has 800 option-rotation rows. Shading identifies the grid-search reference rows.
Main
Prompt
TIR /120
VSTAR /191
ZoomBench /845
4B
V1
84
168
527
9B
V1
89
166
496
27B
V1
97
188
599
27B
V2, seeds 1–3
93/97/90
187/185/182
574/571/576
Appendix
Table 6: Gemini 3.5 Flash as the grid tile reader. V1 entries are single-run correct/total values; V2 27B entries show three seeded numerators with the denominator in the heading. All cells use judge-all scoring and count failed gold examples as wrong. Shading identifies the replicated V2 setting; V1 and V2 are not matched comparisons.
Correct /800
Complete /800
Main
Sub
Pass 1
Pass 2
Pass 3
Pass 3
SFT
SFT
679
663
659
786
SFT
Early mixed
667
667
680
788
Main-only RL
SFT
686
670
688
785
Main-only RL
Early mixed
686
682
687
784
Appendix
Table 7: Three unseeded HR-Bench 4K passes of the early mixed RL role matrix. Correct counts share an 800-row denominator; failures count as wrong. All use V2 prompts, original-image crops, and Qwen3.6-27B judge-all. The last column gives completed rows in pass 3. The mixed reader is from the early mixed RL lineage, not the selected joint policy. Shading marks SFT main rows; Table 3 (b) retains pass 1.
SFT sub
Joint RL sub
Main
Accuracy
Complete
Accuracy
Complete
SFT
84.50 (676/800)
789
85.88 (687/800)
790
Joint RL
82.12 (657/800)
720
82.00 (656/800)
722
Appendix
Table 8: HR-Bench 4K deployment-role matrix for the selected joint RL policy. Each cell is a single unseeded run reporting judge-all accuracy and completed trajectories over the same 800 option-rotation rows; failures count as wrong. Shading marks the SFT main; bold marks the higher accuracy within each sub-agent setting, not significance.
Gold
SFT main
RL main
Benchmark
Correct
Completed
Correct
Completed
TIR
120
87
119
89
117
VSTAR
191
168
191
166
188
ZoomBench
845
507
845
521
844
HR-Bench 4K
800
662
784
–
–
HR-Bench 8K
800
658
781
662
761
Appendix
Table 9: Single-run SFT and later main-only RL results with the SFT tile reader fixed. Entries give correct and completed counts; all gold examples remain in the accuracy denominator. The HR-4K SFT row is the historical comparison, not the later role-matrix baseline; HR-8K uses the clean SFT/SFT replacement. Shading identifies benchmarks with both runs. Dashes denote unavailable same-campaign cells; the separate HR-4K role matrix is in Table 7 .
Metric
SFT
RL
Δ
95% CI
Pass@1 (%)
71.25
72.50
+1.25
[−3.25,+5.75]
Pass@4 (%)
89.00
83.00
−6.00
[−12.0,0.0]
Tool calls
2.645
2.105
−0.540
[−0.87,−0.21]
Appendix
Table 10: Four-response evaluation on 100 prompts. Accuracy differences are percentage points; tool-call differences are calls per trajectory. Intervals are 95% prompt-paired bootstrap intervals. Shading marks the tool-use contrast. Blue/orange indicate numerical increases/decreases, not benefit or significance.
Matrix
Contrast
Δ
95% CI
Main-only/early mixed
Main RL − SFT, SFT sub
+0.88
[−1.38,+3.25]
Early mixed − SFT sub, SFT main
−1.50
[−3.88,+1.00]
Early mixed − SFT sub, RL main
+0.00
[−2.12,+2.25]
Main RL − SFT, early mixed sub
+2.38
[+0.25,+4.50]
Joint RL
Joint RL − SFT sub, SFT main
+1.38
[−1.12,+3.88]
Joint RL − SFT main, SFT sub
−2.38
[−5.75,+0.75]
Appendix
Table 11: Paired HR-Bench 4K contrasts in percentage points. Intervals are 95% question-paired bootstrap intervals, preserving option rotations. Blue/orange indicate positive/negative estimates. Shading marks the joint-main contrast whose question-paired interval excludes zero.
Main
Sub
Correct/completed
Mean calls
Rows with ≥10 calls
SFT
SFT
676/789
2.22
18
SFT
Joint RL
687/790
2.11
14
Joint RL
SFT
657/720
4.35
107
Joint RL
Joint RL
656/722
4.37
105
Appendix
Table 12: Completion sensitivity and tool use for the joint RL role matrix. Conditional accuracy uses only that cell’s completed rows; all end-to-end scores in Table 8 use 800 rows. Shading identifies the joint-main arms with longer tool-call tails.
Policy
Accuracy
Grid
Zoom
Unfinished
Forced-grid warm start
67.5%
9.5%
44.5%
0.0%
Continuation 1
71.5%
10.0%
47.5%
7.0%
Continuation 2
69.5%
14.0%
47.5%
5.0%
Continuation 3
67.5%
15.0%
49.0%
8.0%
Continuation 4
64.5%
21.0%
47.5%
10.0%
Appendix
Table 13: Exploratory default-environment evaluations on 200 RL-held-out internal inputs (greedy, one response each). The inputs are disjoint from the current RL training pool by sample identity and image content, but 196 exact main-agent rows appeared in the earlier SFT training pool. “Grid” is the rate of at least one successful grid call. Shading marks the final continuation, where increased grid use accompanies lower accuracy.
Accuracy
Unfinished
Evaluation
Intervention
Control
Intervention
Control
Initial
71.0
69.0
15.5
15.0
Quarter
69.5
69.5
9.5
13.0
Half
69.5
71.5
9.5
10.0
Three-quarter
72.0
73.5
8.0
8.5
Final
75.5
69.5
7.0
11.0
Appendix
Table 14: Completion-aware reward pilot on 200 SFT-exposed internal prompts. All values are percentages; the initial evaluation precedes training, followed by four equally spaced evaluations. Shading marks the final evaluation; it does not indicate a resolved or externally validated gain.
Source pool
TIR
VSTAR
ZoomBench
HR-4K
HR-8K
SFT main
0/0
0/0
0/0
0/0
0/0
SFT sub
0/0
0/0
3/11
0/0
0/0
Main-only RL
0/0
0/0
4/8
0/0
0/0
Early mixed RL
0/0
0/0
3/10
0/0
0/0
Joint stage 1
0/0
0/0
3/8
0/0
0/0
Joint stage 2
0/0
0/0
4/10
0/0
0/0
Appendix
Table 15: Unique benchmark images with direct training-source overlap. Each cell reports exact byte / exact decoded-RGB matches separately. Counts are per pool, not cumulative checkpoint exposure; all nonzero exact matches are on ZoomBench.
Benchmark
Main
Original
Exact RGB
+ pHash
TIR
SFT
87/120 (72.50%)
87/120 (72.50%) Δ=+0.00
86/119 (72.27%) Δ=−0.23
TIR
Main-only RL
89/120 (74.17%)
89/120 (74.17%) Δ=+0.00
88/119 (73.95%) Δ=−0.22
VSTAR
SFT
168/191 (87.96%)
168/191 (87.96%) Δ=+0.00
168/191 (87.96%) Δ=+0.00
VSTAR
Main-only RL
166/191 (86.91%)
166/191 (86.91%) Δ=+0.00
166/191 (86.91%) Δ=+0.00
ZoomBench
SFT
507/845 (60.00%)
502/834 (60.19%) Δ=+0.19
495/825 (60.00%) Δ=+0.00
ZoomBench
Main-only RL
521/845 (61.66%)
507/826 (61.38%) Δ=−0.28
507/825 (61.45%) Δ=−0.20
Appendix
Table 16: Lineage-specific exclusion sensitivity. Cells show correct/total (accuracy, %); exclusion cells also give Δ from the original accuracy in percentage points. The pHash column includes exact-RGB exclusions. SFT and main-only RL both use the SFT reader; HR-4K here is the historical SFT run. All counts retain failures as wrong.
Visual DeepSearch requires multimodal large reasoning model (MLRM) agents to answer complex visual queries by repeatedly inspecting image regions, grounding intermediate reasoning in visual evidence, and connecting fine-grained clues across long reasoning chains. However, existing benchmarks mainly focus on single-step visual understanding or static image-question answering, offering limited evaluation of iterative image inspection, visual-anchor grounding, and multi-hop evidence integration. In this work, we introduce VistaHop, a benchmark for evaluating vision-centric search and multi-hop visual reasoning in Visual DeepSearch. VistaHop contains 300 high-resolution images, 25 visual search scenarios, and 350 multi-hop QA tasks that require models to follow evidence chains from visual anchors or fuse information across multiple image-grounded reasoning paths. We further develop VistaArena, a unified evaluation environment that supports tool-augmented reasoning with text search, image search, image cropping, and evidence-based answer validation. Experiments on seven representative MLRMs show that current models remain far from solving VistaHop: the best model, SenseNova-MARS-32B, achieves only 24.31% Pass@1. These results reveal persistent limitations in visual grounding, evidence revisiting, long-chain reasoning, and multi-anchor information fusion, highlighting the need for stronger benchmarks and training methods for Visual DeepSearch.
Hang He, Chuhuai Yue, Chengqi Dong +6
1East China Normal University · 2Meituan · 3Shanghai Innovation Institute
High-resolution (HR) image perception presents a key bottleneck for multimodal large language models (MLLMs). While visual search offers a promising solution, existing methods struggle with the trade-off between coverage and efficiency. Visual expert-assisted search is efficient but prone to blind spots when proposals fail, whereas scan-based search guarantees coverage at the cost of computational redundancy and semantic fragmentation. To address this dilemma, we introduce CVSearch, a training-free adaptive framework that dynamically schedules search strategies via an Assess-then-Search workflow. Specifically, CVSearch first invokes expert-assisted search when global information is insufficient, and only triggers a novel semantic-aware scanning mechanism upon failure. Distinct from rigid grid partitioning, this efficient scanning paradigm incorporates Semantic Guided Adaptive Patching to decompose images into semantically consistent regions, effectively mitigating object fragmentation. Furthermore, we devise a Dynamic Bottom-Up Search strategy driven by a Visual Complexity prior to enable efficient and precise iterative exploration of local details. Extensive experiments on HR benchmarks demonstrate that CVSearch achieves state-of-the-art accuracy while substantially improving search efficiency. Code is released at https://github.com/liliupeng28/ICML26-CVSearch.
Liupeng Li, Haoqian Kang, Zhenyu Lu +4
Harbin Institute of Technology, Shenzhen, China · Peng Cheng Laboratory, Shenzhen, China · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China.
Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such scores do not necessarily imply faithful use of visual evidence. Prior studies have identified three shortcuts that inflate benchmark performance. First, linguistic priors and lexical cues in questions often enable models to infer plausible answers without seeing the image. Second, coarse global semantics from the visual encoder can bypass fine-grained local details. Third, in some ``think-with-images'' benchmarks, corrupting the intermediate images returned by visual tools barely affects the final answer. These findings suggest that higher input resolution or larger question pools alone do not elicit genuine active visual search. To address this, we introduce VisualNeedle, a challenging, information-dense, and fine-grained benchmark for scenes where critical evidence is spatially constrained to minute regions and not discernible at a glance. We further propose a counterfactual crop-black setting, which replaces crops returned by tools with black images of the same size, to test whether tool-enabled performance truly relies on intermediate visual evidence.We evaluate 9 prominent MLLMs across four settings: text-only, without tools, with tools, and crop-black. Text-only accuracy stays below 10%, while accuracy without tools remains below 20%. The best tool-enabled model reaches only 56.00%, still trailing the 63.00% human majority-vote accuracy. These results reveal persistent limitations in fine-grained visual search, while the crop-black ablation confirms that success on VisualNeedle hinges on genuine intermediate visual evidence.
Jingru Chen, Yiming Liu, Mingtao Chen +5
Hunyuan, Tencent · Peking University · Zhejiang University