We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with native tool use. Rather than passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence acquisition part of the reasoning process: it dynamically decides whether to look or listen, and over which temporal window, to retrieve sparse but critical evidence across different modalities within long contexts. Through an iterative multi-turn protocol, the retrieved raw audio or visual segments are appended back into the context to support subsequent reasoning. To cold-start this capability, we build a data engine that synthesizes OmniTraj-170K, a corpus of multi-hop Chain-of-Thought trajectories with interleaved audio and visual evidence. We first supervise the model on these trajectories to instill multi-turn tool-use behavior, and then further optimize the policy via a two-stage reinforcement learning with verifiable rewards. Moreover, we introduce an Audio-Visual Necessity objective that explicitly rewards successful trajectories whose reasoning depends on both modalities, discouraging single-modality shortcuts. Extensive experiments across a wide range of benchmarks demonstrate that OmniSeek learns adaptive cross-modal evidence seeking and consistently improves audio-visual reasoning performance.
Figures & tables
Figure 1 : Interleaved Multi-turn Audio-Visual Reasoning. OmniSeek acts as an active agent operating through an iterative <think> → <tool_call> → <observe> loop. The agent dynamically alternates between fetching audio cues and extracting visual evidence from different time spans to answer a complex multi-hop question, avoiding the pitfalls of single-modality shortcuts.
Figure 2 : Overview of OmniTraj-170K data engine. The pipeline extracts timestamp-aligned audio-visual contexts (Stage-1), synthesizes questions grounded in strict cross-modal evidence chains (Stage-2), and translates them into multi-turn reasoning trajectories (Stage-3).
Figure 3 : OmniTraj-170K Statistics: The dataset contains 169,725 multi-turn trajectories over 39,797 videos and covers 19 cross-modal question types. We report the distributions of (a) question types, (b) source-video durations, (c) evidence-span durations, (d) tool calls per trajectory, and (e) normalized temporal positions of retrieved evidence spans.
Figure 4 : Audio-Visual Necessity. We measure audio-visual dependence through modality-specific attention masking.
Model
Size
Daily-Omni
AVUT
WorldSense
FutureOmni
OmniVideoTest
VideoHolmes
JointAV
OmniVideoBench
MMOU
LVOmni
(44s)
(69s)
(141s)
(166s)
(168s)
(184s)
(212s)
(409s)
(757s)
(2049s)
Gemini-3.0-Pro [ 45 ]
-
81.1
-
66.4
-
-
67.0
-
61.8
-
65.8
Gemini-2.0-Flash [ 11 ]
-
67.8
-
56.2
-
-
30.6
-
41.5
-
42.9
Qwen3.5-Omni-Flash [ 46 ]
-
81.8
81.4
57.9
-
-
57.3
-
-
-
-
VideoLLaMA2 [ 9 ]
7B
35.2
44.9
25.4
40.8
-
35.2
46.8
29.2
28.4
27.2
VITA-1.5 [ 22 ]
7B
52.6
-
36.9
48.7
41.0
-
-
36.4
-
-
Table 1 : Performance comparison of different methods on audio-visual benchmarks. Best open-source results are highlighted in bold.
Model / Variant
P1
P2
P3
ravn
Daily-Omni
WorldSense
FutureOmni
OmniVideoTest
VideoHolmes
OmniVideoBench
LVOmni
Video-MME
(a) Base Model
–
–
–
–
71.9
55.1
53.6
54.5
59.1
43.6
35.8
76.8
(b) + Phase 1 SFT
✓
–
–
–
69.1
50.3
50.1
55.8
55.9
38.4
35.0
77.3
(c) + Phase 2 RL
✓
✓
–
–
75.5
55.0
57.2
63.0
69.6
45.2
41.6
76.8
(d) + Phase 3 RL ( G =8)
✓
✓
✓
–
75.9
58.6
56.4
65.9
72.6
46.8
43.4
77.7
(e) + Phase 3 RL ( G =16)
✓
✓
✓
–
78.2
59.2
58.1
67.3
71.7
47.0
43.6
77.9
(f) + Phase 3 RL ( G =16)
✓
✓
✓
✓
80.0
62.4
58.3
69.5
74.6
47.7
44.2
78.5
Table 2: Ablation study on training strategy. P1, P2, and P3 denote Phase-1 SFT, Phase-2 RL, and Phase-3 RL; G is the rollout size.
Model
Size
Video-MME
LongVideoBench
MLVU
LVBench
( w/o sub , 1059s)
(730s)
( m-avg , 705s)
(4038s)
Gemini-3.0-Pro [ 45 ]
-
88.6
75.9
75.7
77.0
Gemini-2.0-Flash [ 11 ]
-
72.4
-
71.0
57.9
Qwen3.5-Omni-Flash [ 46 ]
-
77.0
-
81.9
65.7
Visual-only inputs
SlowFast-LLaVA-1.5 [ 57 ]
7B
63.9
62.5
71.5
45.3
Table 3 : Performance on general video benchmarks.
Figure 5 : Accuracy vs. Number of tool calls.
Model / Variant
Reasoning Paradigm
Tool Calling
Daily-Omni
WorldSense
FutureOmni
OmniVideoTest
OmniVideoBench
LVOmni
(a) Base Model
-
✗
71.9
55.1
53.6
54.5
43.6
35.8
(b) Text-only CoT
Single-turn text
✗
73.0
54.3
55.7
56.0
44.1
40.5
(c) OmniSeek (ours)
Multi-turn multi-modal
✓
80.0
62.4
58.3
69.5
47.7
44.2
Δ vs. Text-only CoT
-
-
+7.0
+8.1
+2.6
+13.5
+3.6
+3.7
Table 4: Ablation study comparing different reasoning paradigms.
Training Data
Training Format
Daily-Omni
WorldSense
FutureOmni
OmniVideoTest
OmniVideoBench
LVOmni
Video-MME
(a) Base Model
Zero-shot
71.9
55.1
53.6
54.5
43.6
35.8
76.8
(b) OmniVideo-100K
SFT + Single-turn QA
73.7
56.0
56.1
61.7
43.3
40.9
78.0
(c) OmniTraj-170K
SFT + Single-turn QA
75.8
56.0
57.1
59.2
45.0
41.6
77.6
Δ vs. Base Model
-
+3.9
+0.9
+3.5
+4.7
+1.4
+5.8
+0.8
Table 5: Analysis on the utility of the OmniTraj-170K corpus.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Phase 1 (SFT)
Phase 2 (GSPO)
Phase 3 (GSPO)
Training Data & Hardware
Dataset
OmniTraj-170K
OmniVideo100K + VideoHolmes
OmniTraj-170K
Number of Samples
170K (90% QA + 10% Tool)
30K (only Multi-Choice) + 1K
8K (Hard)
Epochs
1
1
1
Train Batch Size
128
256
256
Number of GPUs
32 × H200
128 × H200
128 × H200
Appendix
Table 6 : Implementation and Hyperparameter Details for the Three-Phase Training Strategy.
Multi-hop audio-visual reasoning remains challenging for Omni-LLMs, as relevant evidence is often sparse, temporally dispersed, and distributed across both audio and visual streams. Existing benchmarks provide limited investigation of this setting, typically involving only a limited number of modalities, relevant temporal segments, or reasoning steps. In this work, we introduce MOV-Bench, a benchmark containing 519 carefully curated questions that require multi-hop reasoning over temporally dispersed audio-visual evidence. Evaluations on MOV-Bench reveal that current Omni-LLMs still struggle with multi-hop cross-modal reasoning. To address this challenge, we further propose AOP-Agent, an efficient agentic framework built on open-source Omni-LLMs for active omni-modal perception. By combining a hierarchical omni-modal memory with a collaborative observe-reflect-replan loop, AOP-Agent enables open-source Omni-LLMs to perform active perception without additional training or proprietary models. Experiments on MOV-Bench and OmniVideoBench demonstrate that AOP-Agent consistently improves reasoning performance, with particularly notable gains on long videos and reasoning-intensive questions.
Long audio-video reasoning is difficult for omnimodal LLMs because the decisive evidence is often sparse, cross-modal, and too expensive to preserve with uniformly high-fidelity inputs. We introduce OmniReasoner, a tool-use post-training framework for Thinking with Long Audio-Video: omni-modal LLMs learn, via supervised fine-tuning and reinforcement learning, to decide whether and where to call a zoom-in tool before answering. OmniReasoner first builds a low-cost global preview of the full stream and then, when needed, calls the zoom-in tool with a requested temporal interval for higher-fidelity visual and audio inspection before answering. Because the model observes different sampling granularities before and after this call -- a sparse global preview and a denser local clip -- we introduce TimeAnchor, which keeps the tool's temporal argument valid and round-trip-consistent across these granularities, rather than tied to frame indices from a particular sampling rate. To make this tool-use behavior trainable without expensive manual interval annotation, we build a Temporal Augmented Data Engine that synthesizes tool-use post-training trajectories by video editing and composition. Experiments across omnimodal and video benchmarks show that OmniReasoner improves both answer accuracy and temporal grounding while concentrating high-fidelity computation on informative regions. Code is available at https://github.com/RockyChen0205/OmniReasoner.
Yu Chen, Caorui Li, Ziyu Xiong +8
University of Chinese Academy of Sciences · Institute of Automation, CAS · Southeast University +3
Omnimodal understanding entails a massive, highly redundant search space of cross-modal interactions, demanding focused and deliberative reasoning. Current reasoning paradigms rely on either sequential step-by-step generation or parallel sample-by-sample rollouts, leading to isolated reasoning trajectories. This inability to share promising intermediate paths severely limits exploration efficiency and causes compounding errors in complex audio-visual tasks. To break this bottleneck, we introduce Omni-o3, a novel framework driven by a deep nested deduction policy. By formulating reasoning as a dynamic recursive search, Omni-o3 inherently shares reasoning prefixes across branches, enabling the iterative execution of four atomic cognitive actions: expansion, selection, simulation, and backpropagation. To empower this framework, we propose a robust two-stage training paradigm: (1) cold-start supervised fine-tuning on 101K high-quality, long-chain trajectories distilled from 3.5M diverse omnimodal samples, enabling necessary recursive search patterns; and (2) nested group rollout-driven exploratory reinforcement learning on 18K complex multi-turn samples, explicitly guided by a novel multi-step reward model to stimulate deep nested reasoning. Extensive experiments demonstrate that Omni-o3 achieves competitive performance across 11 benchmarks, unlocking advanced capabilities in comprehensive audio-visual, visual-centric, and audio-centric reasoning tasks.
Zhicheng Zhang, Wentao Gu, Weicheng Wang +5
Nankai University · Kuaishou Technology · Pengcheng Laboratory +1