Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or require knowledge-intensive and up-to-date information. We term such cases \textit{perception under insufficient evidence} and formulate perception as an agentic process that can obtain information beyond a single glance. To address the absence of data for this setting, we design two dedicated data generation pipelines, yielding EviRover-SFT-5K and EviRover-RL-12K for training. We further construct EviLens, a human-verified benchmark comprising 688 instances across five perception categories. Building on these data, we present EviRover, to our knowledge the first perception agent explicitly trained to resolve perceptual queries through interaction, using supervised fine-tuning followed by agentic reinforcement learning. Experiments show that the 4B EviRover outperforms its backbone by 30 points on average on EviLens, reaching performance comparable to advanced proprietary models. The gains transfer beyond EviLens to WebEyes, conventional perception benchmarks, and general multimodal benchmarks, including a 15-point improvement on BrowseComp-VL. All code, models, and data are released.
Figures & tables
Figure 1: Representative examples from our training data and the EviLens benchmark.
Figure 2: Overview of our data generation pipeline.
Figure 3: Distribution of target area relative to image area.
Figure 4: Tool-use composition of EviRover-SFT-5K.
Benchmark
Task
Required ability
Grounding
Segmentation
Counting
Visual acquisition
Visual comparison
External knowledge
RefCOCOg ( Mao et al., 2016 )
✓
✓
✗
✗
✗
✗
ReasonSeg ( Lai et al., 2024 )
✗
✓
✗
✗
✗
✓
Ref-Adv ( Dong et al., 2026 )
✓
✗
✗
✗
✓
✗
SOREC ( Goto et al., 2025 )
✓
✗
✗
✓
✗
✗
OK-VOS ( Liang et al., 2026 )
✗
✓
✗
✗
✗
✓
Table 1: Comparison of EviLens with existing perception benchmarks.
Model
Grounding
Segmentation
Counting
Localization
Recognition
Spot Diff
IoU
R@.5
IoU
R@.5
F1 mi
F1 ma
gIoU
cIoU
Acc
Closed-source Models
GPT-5.6-Sol ( OpenAI, 2026a )
0.358
0.407
0.615
0.660
0.112
0.087
0.627
0.699
51.3
GPT-5.6-Luna ( OpenAI, 2026a )
0.264
0.250
0.533
0.571
0.192
0.224
0.544
0.564
35.5
GPT-5.6-Terra ( OpenAI, 2026a )
0.304
0.321
0.528
0.548
0.231
0.313
0.539
0.567
43.4
Table 2: Evaluation results on EviLens.
Model
Grounding
Segmentation
VQA
IoU
R@.5
gIoU
cIoU
Doubao-Seed-2.0-Pro
35.7
44.4
61.2
43.3
65.4
Gemini-3.1-Pro
30.5
35.1
54.6
38.8
63.8
Qwen3VL-4B-Inst
25.6
29.9
24.8
21.7
36.0
Qwen3VL-8B-Inst
26.8
32.6
35.8
25.9
36.3
Pixel-Searcher-8B
34.1
41.3
39.1
32.4
42.2
Table 3: Evaluation results on the WebEyes benchmark.
Model
ReasonSeg
RefCOCOg Grounding
RefCOCOg Segmentation
Val
Test
Val
Test
Val
Test
gIoU
cIoU
gIoU
cIoU
@50
@75
@95
@50
@75
@95
gIoU
cIoU
gIoU
cIoU
Seg-Zero-7B
62.6
62.0
57.5
52.0
-
-
-
-
-
-
68.3
65.3
68.8
66.8
Perception-R1
-
-
-
-
85.6
75.6
31.7
85.3
75.9
32.7
-
-
-
-
Qwen3VL-4B-Inst
63.9
58.9
60.9
53.9
85.7
74.9
28.6
85.3
75.4
29.4
71.6
68.8
72.3
69.9
EviRover (Ours)
68.7
62.3
62.6
56.5
86.1
76.4
31.5
85.4
76.7
32.5
73.7
71.9
73.7
72.3
Table 4: Evaluation results on traditional perception benchmarks.
Model
MMMU
MMMU-Pro
MathVerse
BrowseComp-VL
Acc.
Std.
Vision
Overall
L1
L2
Overall
Qwen3-VL-4B-Inst
63.1
44.5
46.9
62.4
10.6
8.5
9.5
EviRover (Ours)
66.2
48.5
48.4
64.5
33.7
16.0
24.8
Table 5: Generalization to unseen benchmarks.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Distribution of description hops in the RL data for anime and real-person questions.
Tool
Function
Parameters
Visual
text_search
Web text search returning titles, snippets, and links
query , top_k (default 6)
text_search_image
Image search from a text query
query , top_k (default 5, range 1–10)
✓
image_search
Reverse image search over the full image or a specified region
image_id , boxes (optional), top_k (default 3)
browse
Opens a URL and extracts content relevant to a query with a summary model
url , query
crop
Crops and enlarges a region for closer inspection
image_id , boxes
✓
verify
Renders proposed boxes on the original image
image_id , boxes
✓
Appendix
Table 6: Tools available to the agent. Visual indicates whether the tool returns an image to the model.
Figure 6: Example localization trajectory generated by EviRover on the training set.
Figure 7: Example segmentation trajectory generated by EviRover on the training set.