OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
Organizations: The Chinese University of Hong Kong · Alibaba Token Hub, Alibaba Group · Shanghai Jiao Tong University · Shanghai Innovation Institute · Zhejiang University
Abstract
We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, without a separate text question, external captioning, or speech recognition. Direct audio-visual input reduces external latency and computation while preserving perceptual cues. However, research on OmniVChat faces two constraints: data availability and evaluation. Recordings of people using their own devices are scarce. Furthermore, good replies often depend on multimodal context and can be phrased in many ways, making keyword matching unreliable for evaluation. Recent progress in agent systems and video generation makes generation for comprehension viable, which means using synthesized dialogues for training and evaluation. Therefore, we present OmniVChat-Studio, a multi-agent data engine for synthesizing single- and multi-turn audio-visual dialogues. We use synthesized dialogues to build OmniVChat-Bench, an evaluation benchmark that evaluates omni models' basic dialogue abilities across five ability categories. Replies are judged by a large language model based on explicit scoring criteria. We also present OmniVChat-RL, a reinforcement learning reward design that jointly targets reply correctness, efficiency, and style in OmniVChat. Training Qwen3-Omni-Instruct with OmniVChat-RL on synthesized dialogues improves its performance on both OmniVChat-Bench and the human-recorded OmniVChat-Bench-Human. These gains validate the reward design and show transfer to real-world dialogues in training and evaluation.
Figures & tables
| Closed-Source Models | AH | ER | MSA | MEA | DSLP | Single | Multi | Mean | Human | RE | Style |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemini-3.5-Flash | 0.600 | 0.778 | 0.649 | 0.728 | 0.623 | 0.664 | 0.682 | 0.667 | 0.509 | 10.59 | 0.755 |
| Gemini-3.6-Flash | 0.570 | 0.658 | 0.671 | 0.716 | 0.620 | 0.658 | 0.610 | 0.641 | 0.562 | 15.74 | 0.740 |
| Gemini-3.7-Flash | 0.604 | 0.623 | 0.663 | 0.708 | 0.620 | 0.658 | 0.607 | 0.640 | 0.575 | 16.96 | 0.631 |
| Gemini-3.1-Pro | 0.637 | 0.787 | 0.341 | 0.666 | 0.629 | 0.602 | 0.623 | 0.607 | 0.537 | 11.66 | 0.871 |
| Qwen3.5-Omni-Plus | 0.361 | 0.662 | 0.413 | 0.630 | 0.513 | 0.514 | 0.590 | 0.533 | 0.482 | 9.45 | 0.869 |
| Doubao Seed 2.0 Lite | 0.330 | 0.550 | 0.189 | 0.602 | 0.481 | 0.434 | 0.671 | 0.494 | 0.274 | 6.10 | 0.817 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Judge | AH | ER | MSA | MEA | DSLP | Pooled | Mean | Gate |
|---|---|---|---|---|---|---|---|---|
| gpt-5.6-terra | 0.587 | 0.628 | 0.655 | 0.729 | 0.634 | 0.660 | 0.651 | 99.9% |
| gpt-5.6-luna | 0.585 | 0.620 | 0.653 | 0.694 | 0.626 | 0.646 | 0.638 | 99.8% |
| gpt-5.6-sol | 0.622 | 0.673 | 0.675 | 0.747 | 0.636 | 0.678 | 0.673 | 99.9% |
| qwen3.6-flash | 0.644 | 0.614 | 0.653 | 0.741 | 0.632 | 0.671 | 0.657 | 99.1% |
| qwen3.7-plus | 0.595 | 0.593 | 0.638 | 0.702 | 0.619 | 0.643 | 0.631 | 99.9% |
| qwen3.7-max | 0.605 | 0.614 | 0.663 | 0.712 | 0.619 | 0.653 | 0.641 | 99.9% |
| Policy And Objective | Rollout And Reward | ||
|---|---|---|---|
| Base Model | Qwen3-Omni-Instruct | Prompts Per Iteration | 32 |
| Adaptation | LoRA , , Dropout 0 | Samples Per Prompt | 4 |
| Adapter Targets | Attention + MoE Experts | Completions Per Step | 128 |
| Adapter / Total Params | 325M / 30B (3B Active) | Sampling | , Top- 20, Top- 1 |
| Objective | GSPO-based (Eq. 3 ) | Prompt / Reply Budget | 11,072 / 1,024 Tokens |
| Clip Range | (Lo, Hi) | Video Sampling | 2 FPS, Frames |