Omnimodal large language models (Omni-LLMs) encode audio and visual streams into temporally interleaved token sequences for multimodal reasoning. However, processing long audio-visual token sequences incurs substantial prefill costs. Existing compression methods have made progress, but often overlook temporal changes in audio-visual semantic relevance. Motivated by temporal variation and local continuity, we propose OmniRoute, a training-free, two-stage compression framework. First, Temporal Evidence-Guided Budgeting (TEGB) derives chunk-wise modality preferences and initial leading-modality budgets from semantic relevance and local content variation. Second, Budget-Constrained Semantic Compression (BCSC) compresses the leading modality and then calibrates the follower's retention target using the actual retained fraction. For video, it combines spatiotemporal grouping with query-guided selection; for audio, it selects tokens based on encoder attention and query relevance, then merges residual tokens into context anchors under visual guidance. Experiments on four representative benchmarks demonstrate a better trade-off between inference efficiency and performance than competitive baselines. The code and interface will be released to facilitate further research.
Figures & tables
Figure 1: WorldSense accuracy using Qwen2.5-Omni-3B/7B.
Figure 2: Temporal audio-visual evidence. (a) The Di trajectory and its centered five-chunk moving average for a WorldSense video. (b)-(c) Distributions of Vs and Cs over 500 videos.
Figure 3: OmniRoute overview. TEGB derives chunk-wise modality preferences and leading-modality budgets from semantic relevance and local variation. BCSC calibrates follower targets using actual leading-modality retention. Video compression combines spatial and temporal grouping with query-guided selection; audio compression combines attention- and query-based selection with visually guided merging.
Model
Method
Retained Ratio
FLOPs Ratio
WorldSense
AVUTBench
VideoMME
DailyOmni
Avg.
Qwen2.5- Omni-7B
Full Tokens
100%
100%
46.8
64.5
66.0
62.9
100%
Random
55%
48%
43.6
61.0
65.4
58.6
95.0%
DyCoke (V&A)
50%
44%
44.6
62.0
65.5
55.8
94.8%
FlashVID
45%
39%
46.0
63.0
65.9
57.0
96.6%
OmniZip
45%
39%
45.9
63.0
66.1
59.5
97.6%
OmniRoute (Ours)
45%
39%
46.9
63.4
66.0
60.6
98.7%
Table 1: Comparison with token compression methods across different omni-models.
Method
Mem. ↓
Prefill (ms) ↓
Acc. ↑
Latency (s) ↓
Full Tokens (100%)
44 G
2371 (1.00 × )
46.8
10.99 (1.00 × )
DyCoke (50%)
36 G
1386 (1.71 × )
44.6
8.59 (1.28 × )
FlashVID (45%)
35 G
1073 (2.21 × )
46.0
9.63 (1.14 × )
OmniZip (45%)
32 G
894 (2.65 × )
45.9
7.99 (1.38 × )
Ours (45%)
27 G
936 (2.53 × )
46.9
8.12 (1.35 × )
Table 2: Efficiency on WorldSense with Qwen2.5-Omni-7B. Mem. denotes peak GPU memory.
Method
Mem. ↓
Prefill (ms) ↓
Acc. ↑
Latency (s) ↓
Full Tokens (100%)
44 G
2371 (1.00 × )
46.8
10.99 (1.00 × )
DyCoke (50%)
36 G
1386 (1.71 × )
44.6
8.59 (1.28 × )
FlashVID (45%)
35 G
1073 (2.21 × )
46.0
9.63 (1.14 × )
OmniZip (45%)
32 G
894 (2.65 × )
45.9
7.99 (1.38 × )
Ours (45%)
27 G
936 (2.53 × )
46.9
8.12 (1.35 × )
Table 2: Efficiency on WorldSense with Qwen2.5-Omni-7B. Mem. denotes peak GPU memory.
Setting
mi , ni
Routing
Retained Ratio
Acc.
Full OmniRoute
✓
Dynamic
45%
46.9
w/o mi / ni
✗
Dynamic
45%
46.4
Audio-led only
✓
Audio-led
45%
45.9
Video-led only
✓
Video-led
45%
45.2
Table 3: Ablation of OmniRoute’s core components on WorldSense with Qwen2.5-Omni-7B.
Figure 4: Hyperparameter analysis on WorldSense: modality-specific weights γv and γa , and cross-modal coupling β .
Retained Ratio
FLOPs Ratio
Mem. ↓
Prefill (ms) ↓
Acc. ↑
Latency (s) ↓
45%
39%
27 G
936
46.9
8.12
40%
31%
27 G
788
46.3
8.00
35%
26%
27 G
640
46.1
7.73
30%
18%
27 G
558
45.6
7.59
Table 4: Performance of OmniRoute under different retained ratios on WorldSense with Qwen2.5-Omni-7B.
Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often breaks down: for the same query, audio and video relevance often peaks at different moments. This cross-modal salience mismatch makes unidirectional guidance prone to discarding answer-critical cues under aggressive compression. We propose OmniScope, a training-free token compression framework that uses the query as a shared semantic anchor while estimating relevance separately for audio and video. OmniScope allocates modality-specific token budgets, prunes visual tokens with an anchor-delta strategy that preserves both global context and temporal changes, and merges audio tokens within each second to reduce redundancy while maintaining temporal continuity. Across four audio-video benchmarks and two Qwen2.5-Omni model scales, OmniScope achieves the best average accuracy across all compression settings. At 25% overall token retention, it delivers up to 3.53x prefill speedup and more than 15% GPU memory reduction, with only a 0.35-point drop in average accuracy. These results suggest a simple design principle for OmniLLM inference: share the query across modalities, but not the salience estimates. The code is available at https://github.com/MAC-AutoML/OmniScope.
Jinsen Su, Yongdong Luo, Yuexiao Ma +3
Media Analytics and Computing Lab, Xiamen University, Xiamen, China · Institute of Artificial Intelligence, Xiamen University, Xiamen, China · School of Informatics, Xiamen University, Xiamen, China +1
Omni modal large language models (OmniLLMs) have attracted wide attention for their ability to jointly process audio and video, but they generate large token sequences under audio-visual inputs, leading to substantial inference cost. Existing audio-visual token compression methods often rely on unimodal guidance, overlooking the temporal locality of query-relevant evidence in audio-visual inputs and implicitly assuming that the two modalities share a temporally aligned information density distribution. We propose \textbf{OmniFocus}, a training-free query-guided token compression method for OmniLLMs that performs independent importance estimation for video and audio, enabling a modality-symmetric compression design that preserves modality-specific salient evidence while maintaining audio-visual alignment, thereby mitigating the modality bias issue that can arise from unimodal-guided compression. Experiments on the Qwen2.5-Omni model family across four audio-visual benchmarks show that OmniFocus maintains strong compressed performance at low token retention ratios and outperforms existing baselines on several major benchmark scores at 25% token retention. On DailyOmni with Qwen2.5-Omni-7B at 25% token retention, OmniFocus maintains 59.40 accuracy while delivering up to 1.38× prefill speedup relative to the full-token baseline, highlighting a favorable practical accuracy-efficiency trade-off.
Shijie Cao, Qingyu Zhang, Boxi Yu +6
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences · University of Limerick +1
Omnimodal large language models (Omni-LLMs) show strong capability in audio-video understanding, but their practical deployment remains limited by high inference cost of long video streams and dense audio sequences. Despite recent progress, existing compression methods for Omni-LLMs typically rely on fixed or native compression units, which can disrupt cross-modal correspondence and the complementary information required for audio-video reasoning, making it difficult to improve inference efficiency while stably preserving performance. To address this, we propose OmniRefine, a training-free two-stage framework for efficient audio-visual token compression in Omni-LLMs. First, Correspondence-Preserving Chunk Refinement refines native chunk boundaries into cross-modally aligned compression units through frame-audio similarity and dynamic programming. Second, Modality-Aware Cooperative Compression jointly compresses video and audio tokens within each refined unit to reduce redundancy while preserving critical evidence. Extensive experiments show that OmniRefine achieves a better efficiency-performance trade-off than strong baselines and maintains stable performance under lower compression ratios. On WorldSense, it still reaches 46.7% accuracy at a 44% token retention ratio, nearly matching the full-token baseline. The code and interface will be released to facilitate further research.
Yuchen Deng, Zidang Cai, Hai-Tao Zheng +3
Tsinghua Shenzhen International Graduate School, Tsinghua University · Pengcheng Laboratory