Single-image, multi-image, and video deep research require different visual operations but share a workflow of visual grounding, external retrieval, and fact composition. A key challenge is to preserve the dependencies linking localized visual anchors, entity relations, source-supported facts, and answer-producing operations. We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation. Our VGEG-based data engine constructs and verifies multi-image and video questions and filters expert trajectories. Using these data, we assemble OneSearch-VL-SFT-110K and OneSearch-VL-RL-10K for SFT and RL, respectively. We further derive the Evidence-aware Visual-Grounded Rubric reward (EVGR) from VGEG annotations to supervise evidence traceability and visual grounding during RL. For fine-grained evaluation, we construct OneSearch-MI-Bench and OneSearch-Video-Bench, organizing questions by the research operations encoded in their VGEGs. Experiments show that OneSearch-VL-8B improves over Qwen3-VL-8B with tool access by 20.2 and 17.6 percentage points on the two new benchmarks, respectively, while also achieving substantial gains across 7 image benchmarks and VideoDR. Project repository: https://github.com/appletea233/OneSearch-VL
Figures & tables
Tool Set
Tool
Description
Arguments
Visual Input Type
Tvis
Crop
Crop an image region.
Image + Coordinates
I,I,V
OCR
Extract visible text and layout.
Image
PerspectiveCorrect
Correct perspective distortion.
Image
SuperResolution
Upscale low-resolution images.
Image + Scale
Sharpen
Reduce blur and enhance details.
Image + Amount
Tret
ImageSearch
Identify entities through image search.
Image
I,I,V
Table 1: Tool suite of OneSearch-VL.
Figure 1: Overview of the VGEG-centered multimodal data engine. (a) Visual source curation retains videos suitable for external-knowledge research. (b) Dense visual anchor discovery organizes each video into events, key frames, and localized objects. (c) Web entity graph construction links these anchors to real-world entities, source-supported facts, and webpages, forming the evidence graph GX . (d) VGEG-based task construction produces VGEG Γi for question generation, evidence verification, and visual-reference rewriting. (e) Expert trajectories in the tool environment are filtered by answer correctness and process quality to yield multi-turn training trajectories.
Figure 2: Operation-oriented design of the OneSearch benchmarks. Representative questions and reference VGEGs are shown for six principal research operations. Numbered markers bind visual evidence to graph nodes; green nodes denote retrieved entity/fact and purple nodes denote answer-producing operations. The center summarizes the operation distribution and the sizes of the multi-image and video benchmarks.
Figure 3: Overview of the Evidence-aware Visual-Grounded Rubric reward (EVGR). A task-level VGEG and evidence ledger align each multimodal rollout with the required visual anchors and source-supported facts. Claim-and-evidence and visual-grounding judges produce rtrace and rground , whose combination forms REVGR for policy optimization.
Model
SimpleVQA
VDR
MMSearch
LiveVQA
BrowseComp-VL
FVQA
InfoSeek
Avg.
Direct Reasoning
GPT-4o ( OpenAI Team, 2024 )
51.7
1.7
18.7
28.1
5.5
48.0
52.9
29.5
GPT-5 ( OpenAI, 2025b )
61.6
9.8
35.1
44.4
48.6
54.4
61.7
45.1
Gemini-2.5-Flash ( Comanici et al., 2025 )
57.9
6.2
30.4
51.0
37.1
47.7
44.1
39.2
Gemini-2.5-Pro ( Comanici et al., 2025 )
63.0
8.0
39.8
60.3
43.1
60.7
46.9
46.0
Claude-4-Sonnet ( Anthropic Team, 2025b )
50.9
2.0
18.7
38.5
29.3
35.3
57.3
33.1
Table 2: Results on single-image deep-research benchmarks.
VideoDR
OneSearch-MI-Bench
OneSearch-Video-Bench
Model
All
SA
MH
Count
Join
Arith.
Comp.
All
SA
MH
Count
Join
Arith.
Comp.
All
Direct Reasoning
GPT-4o ( OpenAI Team, 2024 )
42.0
35.7
31.3
56.1
46.6
23.0
44.9
39.9
36.0
25.5
29.5
30.8
13.3
29.9
27.4
Gemini-2.5-Flash ( Comanici et al., 2025 )
39.0
32.1
33.3
61.4
46.6
21.3
42.9
40.2
44.0
27.7
32.8
23.1
20.0
33.8
29.6
Gemini-2.5-Pro ( Comanici et al., 2025 )
50.0
42.9
39.6
61.4
46.6
34.4
44.9
45.2
44.0
34.0
37.7
23.1
17.8
37.7
32.3
Qwen2.5-VL-3B ( Bai et al., 2025 )
9.0
10.7
12.5
35.1
13.8
1.6
14.3
15.0
12.0
6.4
18.0
5.8
4.4
3.9
8.1
Table 3: Results on multi-image and video deep-research benchmarks.
(a) SFT data mixture
Data Mixture
Image
Multi-Image
Video
SimpleVQA
InfoSeek
FVQA
OneSearch- MI
OneSearch- Video
VideoDR
Avg.
Qwen3-VL-8B
×
×
×
52.0
50.3
58.7
35.5
17.9
30.0
40.7
I
✓
×
×
66.1
62.4
65.3
44.5
24.8
47.0
51.7
M
×
✓
×
67.0
62.0
67.1
48.5
25.0
49.0
53.1
V
×
×
✓
68.6
62.5
68.5
51.2
28.0
52.0
55.1
I + M
✓
✓
×
68.4
64.5
67.1
50.8
28.6
50.0
54.9
Table 4: Ablation studies on the SFT data mixture and RL reward.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Retained records
Web video source pool
2,500,000
Upload date filter
1,295,220
Metadata coarse filter
377,955
Category balance & exclusion
127,562
Video snippet LLM filter
72,403
Video duration balance
70,781
Appendix
Table 5: Source-video curation. Counts refer to candidate video records.
Operation
Definition
Multi-image
Video
Total
Single-anchor lookup
Retrieve one external attribute for one anchor.
28
25
53
Multi-hop retrieval
Follow an external relation chain from one anchor.
48
47
95
Knowledge-conditioned count
Test a retrieved condition per anchor, then count.
57
61
118
Multi-anchor join
Combine retrieved facts across anchors.
58
52
110
Multi-anchor arithmetic
Compute over retrieved numeric facts.
61
45
106
Multi-anchor comparison
Compare the same retrieved attribute across anchors.
49
77
126
Appendix
Table 6: Definitions and distribution of the six principal research operations.
Figure 4: Benchmark distributions. (a) Smoothed structural-difficulty score distributions, with circles marking observed values. Dashed lines mark the empirical boundaries at 14.4 and 22.2; the easy, medium, and hard groups contain 100/100/101 multi-image and 102/102/103 video questions. (b) Operation composition within six visual domains. Bars are normalized within each domain, and n denotes the number of questions. Every domain covers all six principal research operations.
Category
Hyperparameter
Value / Setting
Model
Base Model
Qwen3-VL-8B-Instruct
Image Max Pixels
262,144 ( ≈512×512 )
Video Max Pixels
100,352 per frame
Video Sampling Rate
2 fps
Maximum Video Frames
128
Trust Remote Code
True
Appendix
Table 7: Agentic SFT configuration for OneSearch-VL-8B.
Score
Evidence traceability
Visual grounding
1.00
All required fact hops are directly supported.
All required visual anchors are correctly grounded and used.
0.75
Most facts are supported, with one minor gap.
Grounding is mostly correct, with one minor localization gap.
0.50
Useful evidence is present, but a major hop is indirect or missing.
Only part of the required visual evidence is grounded.
0.25
Evidence is mostly noisy, speculative, or contradicted.
Identification is weak or incorrect, with limited recovery.
0.00
No supporting evidence is established.
No grounding is established, or the visual identity remains wrong.
Appendix
Table 8: Anchored EVGR scoring criteria used by the trajectory judge.
Category
Hyperparameter
Value / Setting
Model
Initial Policy
OneSearch-VL-8B-SFT
Training Dtype
bfloat16
Data
Train Batch Size (prompts)
256
Val Batch Size
64
Max Prompt Length
40,960 tokens
Max Response Length
16,384 tokens
Appendix
Table 9: GRPO configuration for OneSearch-VL-8B.
Figure 5: Multi-image arithmetic. Separate crops and image searches link the car and brake caliper to Ford and Brembo. Subtracting their retrieved founding years yields 58 years.
Figure 6: Knowledge-conditioned counting. Three input images and four crops support brand and model identification, followed by counting under a retrieved ownership condition. Image handles match the tool calls.
Figure 7: Cross-frame attribute comparison. Blue borders mark selected frames F9 and F116; adjacent frames provide context. Each object undergoes cropping, image retrieval, and text retrieval before release-year comparison. All eight calls and arguments are retained.
Figure 8: Three-anchor cross-frame retrieval. Blue borders mark selected frames F0, F12, and F15 amid neighboring views. The three buildings are independently image-searched before height comparison. All fourteen calls, crop refinement, follow-up searches, and conflicting observations are retained.
Video understanding is moving beyond closed-context perception toward open-world evidence exploration, a paradigm formalized as Video Deep Research (VDR). However, existing multimodal search agents primarily target static images, and the current VDR benchmark relies on text-centric retrieval that discards crucial visual information. To address these limitations, we propose VideoSearcher, a closed-loop agentic framework that empowers Vision-Language Models with multi-tool reasoning for VDR. VideoSearcher unifies temporal localization, spatial focusing, and multimodal search within a single reasoning trajectory, enabling agents to progressively ground visual clues, retrieve relevant evidence, and synthesize answers. To optimize knowledge-intensive reasoning trajectories, we propose Bi-branch Sequence Policy Optimization (BiSPO), a reinforcement learning algorithm that decouples tool-invocation optimization from answer-accuracy optimization. This design provides stable learning signals for both evidence-grounded reasoning and purposeful tool use. Furthermore, we construct VideoSearch-QA, the first benchmark designed to evaluate open-world video information grounding and multimodal search-based reasoning. Extensive experiments demonstrate that VideoSearcher significantly outperforms prior open-source agentic baselines across various search-oriented and multimodal understanding benchmarks.
Zhenkun Gao, Yicheng Bao, Jinlong Peng +13
East China Normal University · Tencent Youtu Lab · HFIPS, Chinese Academy of Sciences +3
Visual DeepSearch requires multimodal large reasoning model (MLRM) agents to answer complex visual queries by repeatedly inspecting image regions, grounding intermediate reasoning in visual evidence, and connecting fine-grained clues across long reasoning chains. However, existing benchmarks mainly focus on single-step visual understanding or static image-question answering, offering limited evaluation of iterative image inspection, visual-anchor grounding, and multi-hop evidence integration. In this work, we introduce VistaHop, a benchmark for evaluating vision-centric search and multi-hop visual reasoning in Visual DeepSearch. VistaHop contains 300 high-resolution images, 25 visual search scenarios, and 350 multi-hop QA tasks that require models to follow evidence chains from visual anchors or fuse information across multiple image-grounded reasoning paths. We further develop VistaArena, a unified evaluation environment that supports tool-augmented reasoning with text search, image search, image cropping, and evidence-based answer validation. Experiments on seven representative MLRMs show that current models remain far from solving VistaHop: the best model, SenseNova-MARS-32B, achieves only 24.31% Pass@1. These results reveal persistent limitations in visual grounding, evidence revisiting, long-chain reasoning, and multi-anchor information fusion, highlighting the need for stronger benchmarks and training methods for Visual DeepSearch.
Hang He, Chuhuai Yue, Chengqi Dong +6
1East China Normal University · 2Meituan · 3Shanghai Innovation Institute
We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool-augmented execution. To address these challenges, we propose Video-DR, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval. Our framework adopts a two-stage training recipe: supervised fine-tuning followed by Group Relative Policy Optimization (GRPO), enabling autonomous exploration that breaks the imitation-learning ceiling. Furthermore, we curate Video-DR-Bench, a human-AI collaborative benchmark comprising 200 complex, multi-hop VQA instances. Empirical results demonstrate that our Video-DeepResearch-35B-A3B establishes a new state-of-the-art of 64.0% average accuracy, surpassing proprietary Claude-4.5-Sonnet (59.0%) by 5.0 points and significantly outperforming GPT-5 (52.5%) and Gemini 2.5 Pro (57.5%). The 30B-A3B variant achieves 59.3%, competitive with Claude-4.5-Sonnet and demonstrating the effectiveness of our training paradigm even at compact scale. Code: https://github.com/Osilly/Vision-DeepResearch.