Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or require knowledge-intensive and up-to-date information. We term such cases \textit{perception under insufficient evidence} and formulate perception as an agentic process that can obtain information beyond a single glance. To address the absence of data for this setting, we design two dedicated data generation pipelines, yielding EviRover-SFT-5K and EviRover-RL-12K for training. We further construct EviLens, a human-verified benchmark comprising 688 instances across five perception categories. Building on these data, we present EviRover, to our knowledge the first perception agent explicitly trained to resolve perceptual queries through interaction, using supervised fine-tuning followed by agentic reinforcement learning. Experiments show that the 4B EviRover outperforms its backbone by 30 points on average on EviLens, reaching performance comparable to advanced proprietary models. The gains transfer beyond EviLens to WebEyes, conventional perception benchmarks, and general multimodal benchmarks, including a 15-point improvement on BrowseComp-VL. All code, models, and data are released.
Figures & tables
Figure 1: Representative examples from our training data and the EviLens benchmark.
Figure 2: Overview of our data generation pipeline.
Figure 3: Distribution of target area relative to image area.
Figure 4: Tool-use composition of EviRover-SFT-5K.
Benchmark
Task
Required ability
Grounding
Segmentation
Counting
Visual acquisition
Visual comparison
External knowledge
RefCOCOg ( Mao et al., 2016 )
✓
✓
✗
✗
✗
✗
ReasonSeg ( Lai et al., 2024 )
✗
✓
✗
✗
✗
✓
Ref-Adv ( Dong et al., 2026 )
✓
✗
✗
✗
✓
✗
SOREC ( Goto et al., 2025 )
✓
✗
✗
✓
✗
✗
OK-VOS ( Liang et al., 2026 )
✗
✓
✗
✗
✗
✓
Table 1: Comparison of EviLens with existing perception benchmarks.
Model
Grounding
Segmentation
Counting
Localization
Recognition
Spot Diff
IoU
R@.5
IoU
R@.5
F1 mi
F1 ma
gIoU
cIoU
Acc
Closed-source Models
GPT-5.6-Sol ( OpenAI, 2026a )
0.358
0.407
0.615
0.660
0.112
0.087
0.627
0.699
51.3
GPT-5.6-Luna ( OpenAI, 2026a )
0.264
0.250
0.533
0.571
0.192
0.224
0.544
0.564
35.5
GPT-5.6-Terra ( OpenAI, 2026a )
0.304
0.321
0.528
0.548
0.231
0.313
0.539
0.567
43.4
Table 2: Evaluation results on EviLens.
Model
Grounding
Segmentation
VQA
IoU
R@.5
gIoU
cIoU
Doubao-Seed-2.0-Pro
35.7
44.4
61.2
43.3
65.4
Gemini-3.1-Pro
30.5
35.1
54.6
38.8
63.8
Qwen3VL-4B-Inst
25.6
29.9
24.8
21.7
36.0
Qwen3VL-8B-Inst
26.8
32.6
35.8
25.9
36.3
Pixel-Searcher-8B
34.1
41.3
39.1
32.4
42.2
Table 3: Evaluation results on the WebEyes benchmark.
Model
ReasonSeg
RefCOCOg Grounding
RefCOCOg Segmentation
Val
Test
Val
Test
Val
Test
gIoU
cIoU
gIoU
cIoU
@50
@75
@95
@50
@75
@95
gIoU
cIoU
gIoU
cIoU
Seg-Zero-7B
62.6
62.0
57.5
52.0
-
-
-
-
-
-
68.3
65.3
68.8
66.8
Perception-R1
-
-
-
-
85.6
75.6
31.7
85.3
75.9
32.7
-
-
-
-
Qwen3VL-4B-Inst
63.9
58.9
60.9
53.9
85.7
74.9
28.6
85.3
75.4
29.4
71.6
68.8
72.3
69.9
EviRover (Ours)
68.7
62.3
62.6
56.5
86.1
76.4
31.5
85.4
76.7
32.5
73.7
71.9
73.7
72.3
Table 4: Evaluation results on traditional perception benchmarks.
Model
MMMU
MMMU-Pro
MathVerse
BrowseComp-VL
Acc.
Std.
Vision
Overall
L1
L2
Overall
Qwen3-VL-4B-Inst
63.1
44.5
46.9
62.4
10.6
8.5
9.5
EviRover (Ours)
66.2
48.5
48.4
64.5
33.7
16.0
24.8
Table 5: Generalization to unseen benchmarks.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Distribution of description hops in the RL data for anime and real-person questions.
Tool
Function
Parameters
Visual
text_search
Web text search returning titles, snippets, and links
query , top_k (default 6)
text_search_image
Image search from a text query
query , top_k (default 5, range 1–10)
✓
image_search
Reverse image search over the full image or a specified region
image_id , boxes (optional), top_k (default 3)
browse
Opens a URL and extracts content relevant to a query with a summary model
url , query
crop
Crops and enlarges a region for closer inspection
image_id , boxes
✓
verify
Renders proposed boxes on the original image
image_id , boxes
✓
Appendix
Table 6: Tools available to the agent. Visual indicates whether the tool returns an image to the model.
Figure 6: Example localization trajectory generated by EviRover on the training set.
Figure 7: Example segmentation trajectory generated by EviRover on the training set.
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.
Visual perception connects high-level semantic understanding to pixel-level perception, but most existing settings assume that the decisive evidence for identifying a target is already in the image or frozen model knowledge. We study a more practical yet harder open-world case where a visible object must first be resolved from external facts, recent events, long-tail entities, or multi-hop relations before it can be localized. We formalize this challenge as Perception Deep Research and introduce WebEye, an object-anchored benchmark with verifiable evidence, knowledge-intensive queries, precise box/mask annotations, and three task views: Search-based Grounding, Search-based Segmentation, and Search-based VQA. WebEyes contains 120 images, 473 annotated object instances, 645 unique QA pairs, and 1,927 task samples. We further propose Pixel-Searcher, an agentic search-to-pixel workflow that resolves hidden target identities and binds them to boxes, masks, or grounded answers. Experiments show that Pixel-Searcher achieves the strongest open-source performance across all three task views, while failures mainly arise from evidence acquisition, identity resolution, and visual instance binding.
Bokang Yang, Xinyi Sun, Kaituo Feng +3
1Shenzhen Loop Area Institute · 2Wuhan University · 3CUHK MMLab
Fine-grained visual reasoning remains challenging for vision-language models, especially when small but critical visual cues are buried in high-resolution images. Existing approaches rely on repeated cropping or test-time visual search to introduce local evidence, but they typically do not explicitly distinguish perception from reasoning. In this paper, we propose Perceive-to-Reason (P2R), a unified framework that formulates fine-grained visual reasoning as a two-stage process: the model first localizes question-relevant evidence as a Perceiver, and then answers the question as a Reasoner based on the annotated image and cropped regions. To better align training with this decoupled formulation, we further introduce Perception-Reasoning Alternating GRPO (PRA-GRPO), a role-aware reinforcement learning strategy that alternates between perception-focused and reasoning-focused updates using only final-answer supervision. Built on top of Qwen3-VL-Instruct-2B/4B/8B, P2R consistently improves performance across model scales. In particular, P2R-4B achieves 93.2% on V-Star, 81.9% on HR-Bench-4K, and 80.5% on HR-Bench-8K, substantially outperforming its corresponding backbone. Further experiments show that the benefits of P2R extend beyond high-resolution benchmarks to broader multimodal reasoning tasks. These results suggest that explicitly decoupling perception from reasoning provides an effective framework for fine-grained visual reasoning.