Organizations: School of Intelligence Science and Technology, Peking University · State Key Laboratory of General Artificial Intelligence, Peking University · Alibaba Token Hub, Alibaba Group
Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities independently, leaving the capability of audio-visual joint reasoning poorly evaluated and insufficiently elicited. We address this gap with a benchmark, data engine, and learning method. First, we introduce OmniReasoningBench, a benchmark where both audio and visual evidence are indispensable. It comprises 1,150 multiple-choice and open-ended questions across two tasks, reasoning over video and reasoning beyond video. Second, we develop a data engine OmniQA. It automatically constructs evidence-grounded QA pairs that explicitly necessitate audio-visual joint reasoning, together with time-stamped clue chains that guide the annotation of thinking process. Besides our benchmark, this engine produces training data OmniReasoning-SFT-112K and OmniReasoning-RL-19K. Finally, we propose an on-policy self-distillation method Modality-Factored Self-Distillation (MFSD). It evaluates each sampled response under modality-specific clue contexts, disentangling the contributions of individual clues and their cross-modal interactions for token-level credit assignment. With our training data and learning method, our model OmniReasoning-30B-A3B achieves 50.0% on OmniVideoBench and 42.5% on OmniReasoningBench, improving the base model Qwen3-Omni-30B-A3B-Thinking by 12.8 and 9.3 percentage points, respectively. Moreover, it delivers substantial gains on general and long-video benchmarks, including Video-MME-v2. We hope our work offers a solid step for facilitating future research in omni-modal joint reasoning.
Figures & tables
Figure 1: OmniReasoning: a benchmark, data engine and learning method for audio-visual joint reasoning. Unlike previous benchmarks, our benchmark OmniReasoningBench truly requires both audio and visual inputs for joint reasoning. Besides this benchmark, our data engine OmniQA additionally produces large-scale training data with evidence-grounded questions. Further, our learning method MFSD leverages the gain from cross-modality joint clues to assign credit at token level, enabling effective exploration along audio-visual joint reasoning. In comparison, previous method GRPO offers only outcome-level guidance, and RLSD does not consider cross-modality interaction. With our training data and learning method, our model achieves large improvements over base model on audio-visual, long video, and general video benchmarks.
Figure 2: OmniReasoningBench tasks and examples. Reasoning over video connects observations across events; reasoning beyond video applies video-derived knowledge to a new scenario. Orange and blue mark audio and visual clues, with numbered markers linking evidence to timestamps.
Figure 3: OmniQA data engine. Gemini-3.1-Pro annotates timestamped audio-visual descriptions. Qwen3.8 generates QA pairs and reasoning steps, followed by clue validation and shortcut screening. Timestamped captions, verified QA pairs, and evidence chains then guide thinking generation.
Figure 4: OmniReasoning released training-data distributions. The SFT and RL corpora span eight content domains, 25 production task types, and varied video durations.
Figure 5: Modality-Factored Self-Distillation. The actor scores the same response under four clue contexts. Joint-clue support and non-additive audio-visual interaction determine bounded token-level advantage weights.
Reasoning over video
Reasoning beyond video
Model
Modality
MCQ
OE
MCQ
OE
Overall
Table 1: OmniReasoningBench accuracy (%). MCQ and OE denote multiple-choice and open-ended questions. Our model is initialized by Qwen3-Omni-30B-A3B-Thinking.