We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with native tool use. Rather than passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence acquisition part of the reasoning process: it dynamically decides whether to look or listen, and over which temporal window, to retrieve sparse but critical evidence across different modalities within long contexts. Through an iterative multi-turn protocol, the retrieved raw audio or visual segments are appended back into the context to support subsequent reasoning. To cold-start this capability, we build a data engine that synthesizes OmniTraj-170K, a corpus of multi-hop Chain-of-Thought trajectories with interleaved audio and visual evidence. We first supervise the model on these trajectories to instill multi-turn tool-use behavior, and then further optimize the policy via a two-stage reinforcement learning with verifiable rewards. Moreover, we introduce an Audio-Visual Necessity objective that explicitly rewards successful trajectories whose reasoning depends on both modalities, discouraging single-modality shortcuts. Extensive experiments across a wide range of benchmarks demonstrate that OmniSeek learns adaptive cross-modal evidence seeking and consistently improves audio-visual reasoning performance.
Figures & tables
Figure 1 : Interleaved Multi-turn Audio-Visual Reasoning. OmniSeek acts as an active agent operating through an iterative <think> → <tool_call> → <observe> loop. The agent dynamically alternates between fetching audio cues and extracting visual evidence from different time spans to answer a complex multi-hop question, avoiding the pitfalls of single-modality shortcuts.
Figure 2 : Overview of OmniTraj-170K data engine. The pipeline extracts timestamp-aligned audio-visual contexts (Stage-1), synthesizes questions grounded in strict cross-modal evidence chains (Stage-2), and translates them into multi-turn reasoning trajectories (Stage-3).
Figure 3 : OmniTraj-170K Statistics: The dataset contains 169,725 multi-turn trajectories over 39,797 videos and covers 19 cross-modal question types. We report the distributions of (a) question types, (b) source-video durations, (c) evidence-span durations, (d) tool calls per trajectory, and (e) normalized temporal positions of retrieved evidence spans.
Figure 4 : Audio-Visual Necessity. We measure audio-visual dependence through modality-specific attention masking.
Model
Size
Daily-Omni
AVUT
WorldSense
FutureOmni
OmniVideoTest
VideoHolmes
JointAV
OmniVideoBench
MMOU
LVOmni
(44s)
(69s)
(141s)
(166s)
(168s)
(184s)
(212s)
(409s)
(757s)
(2049s)
Gemini-3.0-Pro [ 45 ]
-
81.1
-
66.4
-
-
67.0
-
61.8
-
65.8
Gemini-2.0-Flash [ 11 ]
-
67.8
-
56.2
-
-
30.6
-
41.5
-
42.9
Qwen3.5-Omni-Flash [ 46 ]
-
81.8
81.4
57.9
-
-
57.3
-
-
-
-
VideoLLaMA2 [ 9 ]
7B
35.2
44.9
25.4
40.8
-
35.2
46.8
29.2
28.4
27.2
VITA-1.5 [ 22 ]
7B
52.6
-
36.9
48.7
41.0
-
-
36.4
-
-
Table 1 : Performance comparison of different methods on audio-visual benchmarks. Best open-source results are highlighted in bold.
Model / Variant
P1
P2
P3
ravn
Daily-Omni
WorldSense
FutureOmni
OmniVideoTest
VideoHolmes
OmniVideoBench
LVOmni
Video-MME
(a) Base Model
–
–
–
–
71.9
55.1
53.6
54.5
59.1
43.6
35.8
76.8
(b) + Phase 1 SFT
✓
–
–
–
69.1
50.3
50.1
55.8
55.9
38.4
35.0
77.3
(c) + Phase 2 RL
✓
✓
–
–
75.5
55.0
57.2
63.0
69.6
45.2
41.6
76.8
(d) + Phase 3 RL ( G =8)
✓
✓
✓
–
75.9
58.6
56.4
65.9
72.6
46.8
43.4
77.7
(e) + Phase 3 RL ( G =16)
✓
✓
✓
–
78.2
59.2
58.1
67.3
71.7
47.0
43.6
77.9
(f) + Phase 3 RL ( G =16)
✓
✓
✓
✓
80.0
62.4
58.3
69.5
74.6
47.7
44.2
78.5
Table 2: Ablation study on training strategy. P1, P2, and P3 denote Phase-1 SFT, Phase-2 RL, and Phase-3 RL; G is the rollout size.
Model
Size
Video-MME
LongVideoBench
MLVU
LVBench
( w/o sub , 1059s)
(730s)
( m-avg , 705s)
(4038s)
Gemini-3.0-Pro [ 45 ]
-
88.6
75.9
75.7
77.0
Gemini-2.0-Flash [ 11 ]
-
72.4
-
71.0
57.9
Qwen3.5-Omni-Flash [ 46 ]
-
77.0
-
81.9
65.7
Visual-only inputs
SlowFast-LLaVA-1.5 [ 57 ]
7B
63.9
62.5
71.5
45.3
Table 3 : Performance on general video benchmarks.
Figure 5 : Accuracy vs. Number of tool calls.
Model / Variant
Reasoning Paradigm
Tool Calling
Daily-Omni
WorldSense
FutureOmni
OmniVideoTest
OmniVideoBench
LVOmni
(a) Base Model
-
✗
71.9
55.1
53.6
54.5
43.6
35.8
(b) Text-only CoT
Single-turn text
✗
73.0
54.3
55.7
56.0
44.1
40.5
(c) OmniSeek (ours)
Multi-turn multi-modal
✓
80.0
62.4
58.3
69.5
47.7
44.2
Δ vs. Text-only CoT
-
-
+7.0
+8.1
+2.6
+13.5
+3.6
+3.7
Table 4: Ablation study comparing different reasoning paradigms.
Training Data
Training Format
Daily-Omni
WorldSense
FutureOmni
OmniVideoTest
OmniVideoBench
LVOmni
Video-MME
(a) Base Model
Zero-shot
71.9
55.1
53.6
54.5
43.6
35.8
76.8
(b) OmniVideo-100K
SFT + Single-turn QA
73.7
56.0
56.1
61.7
43.3
40.9
78.0
(c) OmniTraj-170K
SFT + Single-turn QA
75.8
56.0
57.1
59.2
45.0
41.6
77.6
Δ vs. Base Model
-
+3.9
+0.9
+3.5
+4.7
+1.4
+5.8
+0.8
Table 5: Analysis on the utility of the OmniTraj-170K corpus.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Phase 1 (SFT)
Phase 2 (GSPO)
Phase 3 (GSPO)
Training Data & Hardware
Dataset
OmniTraj-170K
OmniVideo100K + VideoHolmes
OmniTraj-170K
Number of Samples
170K (90% QA + 10% Tool)
30K (only Multi-Choice) + 1K
8K (Hard)
Epochs
1
1
1
Train Batch Size
128
256
256
Number of GPUs
32 × H200
128 × H200
128 × H200
Appendix
Table 6 : Implementation and Hyperparameter Details for the Three-Phase Training Strategy.