Single-image, multi-image, and video deep research require different visual operations but share a workflow of visual grounding, external retrieval, and fact composition. A key challenge is to preserve the dependencies linking localized visual anchors, entity relations, source-supported facts, and answer-producing operations. We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation. Our VGEG-based data engine constructs and verifies multi-image and video questions and filters expert trajectories. Using these data, we assemble OneSearch-VL-SFT-110K and OneSearch-VL-RL-10K for SFT and RL, respectively. We further derive the Evidence-aware Visual-Grounded Rubric reward (EVGR) from VGEG annotations to supervise evidence traceability and visual grounding during RL. For fine-grained evaluation, we construct OneSearch-MI-Bench and OneSearch-Video-Bench, organizing questions by the research operations encoded in their VGEGs. Experiments show that OneSearch-VL-8B improves over Qwen3-VL-8B with tool access by 20.2 and 17.6 percentage points on the two new benchmarks, respectively, while also achieving substantial gains across 7 image benchmarks and VideoDR. Project repository: https://github.com/appletea233/OneSearch-VL
Figures & tables
Tool Set
Tool
Description
Arguments
Visual Input Type
Tvis
Crop
Crop an image region.
Image + Coordinates
I,I,V
OCR
Extract visible text and layout.
Image
PerspectiveCorrect
Correct perspective distortion.
Image
SuperResolution
Upscale low-resolution images.
Image + Scale
Sharpen
Reduce blur and enhance details.
Image + Amount
Tret
ImageSearch
Identify entities through image search.
Image
I,I,V
Table 1: Tool suite of OneSearch-VL.
Figure 1: Overview of the VGEG-centered multimodal data engine. (a) Visual source curation retains videos suitable for external-knowledge research. (b) Dense visual anchor discovery organizes each video into events, key frames, and localized objects. (c) Web entity graph construction links these anchors to real-world entities, source-supported facts, and webpages, forming the evidence graph GX . (d) VGEG-based task construction produces VGEG Γi for question generation, evidence verification, and visual-reference rewriting. (e) Expert trajectories in the tool environment are filtered by answer correctness and process quality to yield multi-turn training trajectories.
Figure 2: Operation-oriented design of the OneSearch benchmarks. Representative questions and reference VGEGs are shown for six principal research operations. Numbered markers bind visual evidence to graph nodes; green nodes denote retrieved entity/fact and purple nodes denote answer-producing operations. The center summarizes the operation distribution and the sizes of the multi-image and video benchmarks.
Figure 3: Overview of the Evidence-aware Visual-Grounded Rubric reward (EVGR). A task-level VGEG and evidence ledger align each multimodal rollout with the required visual anchors and source-supported facts. Claim-and-evidence and visual-grounding judges produce rtrace and rground , whose combination forms REVGR for policy optimization.
Model
SimpleVQA
VDR
MMSearch
LiveVQA
BrowseComp-VL
FVQA
InfoSeek
Avg.
Direct Reasoning
GPT-4o ( OpenAI Team, 2024 )
51.7
1.7
18.7
28.1
5.5
48.0
52.9
29.5
GPT-5 ( OpenAI, 2025b )
61.6
9.8
35.1
44.4
48.6
54.4
61.7
45.1
Gemini-2.5-Flash ( Comanici et al., 2025 )
57.9
6.2
30.4
51.0
37.1
47.7
44.1
39.2
Gemini-2.5-Pro ( Comanici et al., 2025 )
63.0
8.0
39.8
60.3
43.1
60.7
46.9
46.0
Claude-4-Sonnet ( Anthropic Team, 2025b )
50.9
2.0
18.7
38.5
29.3
35.3
57.3
33.1
Table 2: Results on single-image deep-research benchmarks.
VideoDR
OneSearch-MI-Bench
OneSearch-Video-Bench
Model
All
SA
MH
Count
Join
Arith.
Comp.
All
SA
MH
Count
Join
Arith.
Comp.
All
Direct Reasoning
GPT-4o ( OpenAI Team, 2024 )
42.0
35.7
31.3
56.1
46.6
23.0
44.9
39.9
36.0
25.5
29.5
30.8
13.3
29.9
27.4
Gemini-2.5-Flash ( Comanici et al., 2025 )
39.0
32.1
33.3
61.4
46.6
21.3
42.9
40.2
44.0
27.7
32.8
23.1
20.0
33.8
29.6
Gemini-2.5-Pro ( Comanici et al., 2025 )
50.0
42.9
39.6
61.4
46.6
34.4
44.9
45.2
44.0
34.0
37.7
23.1
17.8
37.7
32.3
Qwen2.5-VL-3B ( Bai et al., 2025 )
9.0
10.7
12.5
35.1
13.8
1.6
14.3
15.0
12.0
6.4
18.0
5.8
4.4
3.9
8.1
Table 3: Results on multi-image and video deep-research benchmarks.
(a) SFT data mixture
Data Mixture
Image
Multi-Image
Video
SimpleVQA
InfoSeek
FVQA
OneSearch- MI
OneSearch- Video
VideoDR
Avg.
Qwen3-VL-8B
×
×
×
52.0
50.3
58.7
35.5
17.9
30.0
40.7
I
✓
×
×
66.1
62.4
65.3
44.5
24.8
47.0
51.7
M
×
✓
×
67.0
62.0
67.1
48.5
25.0
49.0
53.1
V
×
×
✓
68.6
62.5
68.5
51.2
28.0
52.0
55.1
I + M
✓
✓
×
68.4
64.5
67.1
50.8
28.6
50.0
54.9
Table 4: Ablation studies on the SFT data mixture and RL reward.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Retained records
Web video source pool
2,500,000
Upload date filter
1,295,220
Metadata coarse filter
377,955
Category balance & exclusion
127,562
Video snippet LLM filter
72,403
Video duration balance
70,781
Appendix
Table 5: Source-video curation. Counts refer to candidate video records.
Operation
Definition
Multi-image
Video
Total
Single-anchor lookup
Retrieve one external attribute for one anchor.
28
25
53
Multi-hop retrieval
Follow an external relation chain from one anchor.
48
47
95
Knowledge-conditioned count
Test a retrieved condition per anchor, then count.
57
61
118
Multi-anchor join
Combine retrieved facts across anchors.
58
52
110
Multi-anchor arithmetic
Compute over retrieved numeric facts.
61
45
106
Multi-anchor comparison
Compare the same retrieved attribute across anchors.
49
77
126
Appendix
Table 6: Definitions and distribution of the six principal research operations.
Figure 4: Benchmark distributions. (a) Smoothed structural-difficulty score distributions, with circles marking observed values. Dashed lines mark the empirical boundaries at 14.4 and 22.2; the easy, medium, and hard groups contain 100/100/101 multi-image and 102/102/103 video questions. (b) Operation composition within six visual domains. Bars are normalized within each domain, and n denotes the number of questions. Every domain covers all six principal research operations.
Category
Hyperparameter
Value / Setting
Model
Base Model
Qwen3-VL-8B-Instruct
Image Max Pixels
262,144 ( ≈512×512 )
Video Max Pixels
100,352 per frame
Video Sampling Rate
2 fps
Maximum Video Frames
128
Trust Remote Code
True
Appendix
Table 7: Agentic SFT configuration for OneSearch-VL-8B.
Score
Evidence traceability
Visual grounding
1.00
All required fact hops are directly supported.
All required visual anchors are correctly grounded and used.
0.75
Most facts are supported, with one minor gap.
Grounding is mostly correct, with one minor localization gap.
0.50
Useful evidence is present, but a major hop is indirect or missing.
Only part of the required visual evidence is grounded.
0.25
Evidence is mostly noisy, speculative, or contradicted.
Identification is weak or incorrect, with limited recovery.
0.00
No supporting evidence is established.
No grounding is established, or the visual identity remains wrong.
Appendix
Table 8: Anchored EVGR scoring criteria used by the trajectory judge.
Category
Hyperparameter
Value / Setting
Model
Initial Policy
OneSearch-VL-8B-SFT
Training Dtype
bfloat16
Data
Train Batch Size (prompts)
256
Val Batch Size
64
Max Prompt Length
40,960 tokens
Max Response Length
16,384 tokens
Appendix
Table 9: GRPO configuration for OneSearch-VL-8B.
Figure 5: Multi-image arithmetic. Separate crops and image searches link the car and brake caliper to Ford and Brembo. Subtracting their retrieved founding years yields 58 years.
Figure 6: Knowledge-conditioned counting. Three input images and four crops support brand and model identification, followed by counting under a retrieved ownership condition. Image handles match the tool calls.
Figure 7: Cross-frame attribute comparison. Blue borders mark selected frames F9 and F116; adjacent frames provide context. Each object undergoes cropping, image retrieval, and text retrieval before release-year comparison. All eight calls and arguments are retained.
Figure 8: Three-anchor cross-frame retrieval. Blue borders mark selected frames F0, F12, and F15 amid neighboring views. The three buildings are independently image-searched before height comparison. All fourteen calls, crop refinement, follow-up searches, and conflicting observations are retained.