Omni-modal large language models (OmniLLMs) enable unified audio-video understanding, but their long multimodal token sequences make deployment computationally expensive. Token compression reduces this cost, yet aggressive compression often lowers accuracy. Existing works predominantly focus on designing better compression mechanisms; however, adapting the underlying language model to reason effectively over the remaining compressed context remains under-explored. To address this, we propose CAFD (Compressed-Context Adaptation via Full-Context Distillation), a ground-truth-free self-distillation framework that adapts OmniLLMs to fixed compression pipelines without requiring reference answers, rationales, or correctness rewards. CAFD leverages the full-token view of the same multimodal sample as a source of privileged information: a full-context self-teacher provides soft target supervision to a compressed-context student along the student's on-policy trajectory. Evaluated on Qwen2.5-Omni-7B across five audio-video benchmarks, five compression pipelines, and five deployment budgets, CAFD demonstrates consistent gains, improving 120 out of 125 conditions with an average accuracy boost of 1.44 points and recovering 26.9% of the accuracy gap on average. These results demonstrate that the proposed ground-truth-free adaptation offers an effective and practical route to improving the accuracy-efficiency trade-off in deployed OmniLLMs.
Figures & tables
Figure 1: Complementary adaptation and accuracy gains. (a) Compression methods reduce the token count; CAFD complements them by adapting the OmniLLM to compressed tokens while keeping the compressor unchanged. Pre-LLM compression is illustrated; CAFD also supports inner-LLM and hybrid compression. (b) Gray and blue bars show each compressor’s mean accuracy before and after CAFD adaptation, averaging five benchmarks and five deployment budgets. The dashed line denotes the unadapted full tokens baseline (54.86%). Red labels indicate accuracy-gap recovery: the percentage of the full-token–unadapted accuracy gap closed by adaptation.
Figure 2: Overview of CAFD. Self-distillation guides the model under compressed context toward its full-context capabilities. The compressed student generates a response; both paths compute next-token distributions along it using the same sample and instruction, without reference answers or rationales. Gradients update only the student; the teacher parameters track the student via an exponential moving average (EMA). Deployment retains the adapted model and original compressor. The framework supports pre-LLM (illustrated), inner-LLM, and hybrid compression.
Compressor
Budget
WorldSense
DailyOmni
AVUT
OmniVideoBench
Video-MME
Avg.
Full tokens
100%
46.85
62.91
64.65
35.50
64.41
54.86
Uniform
35%
43.28 / 44.96
57.23 / 58.65
60.44 / 62.86
33.80 / 33.70
63.44 / 64.81
51.64 / 53.00
25%
41.96 / 43.66
52.05 / 55.89
56.81 / 59.05
31.40 / 32.60
61.63 / 62.70
48.77 / 50.78
15%
38.84 / 39.97
49.37 / 52.72
52.94 / 55.25
29.90 / 31.50
60.33 / 61.67
46.28 / 48.22
10%
36.44 / 37.64
45.45 / 48.45
50.35 / 51.50
29.50 / 30.80
58.15 / 60.22
43.98 / 45.72
5%
35.18 / 36.66
40.77 / 43.69
46.37 / 47.92
30.10 / 31.00
54.89 / 57.37
41.46 / 43.33
Table 1: Main results across compressors, budgets, and benchmarks. Each cell reports Unadapted / +CAFD accuracy (%). Each +CAFD model is trained under the corresponding compressor’s 5% budget configuration and is evaluated at the five listed budgets. Column Avg. averages benchmarks, block-ending Avg. averages budgets, and their intersection averages all 25 conditions. The higher value in each pair is bold. DivPrune-M denotes DivPrune-Merge.
Figure 3: Effects of training data scale and retention. (a) WorldSense five-budget average accuracy for DivPrune across training-set sizes and training retentions. (b) Equal-weight accuracy-point gain over the matched Unadapted baseline across WorldSense, DailyOmni, and AVUT at five deployment budgets per benchmark. All runs use two epochs; data scale therefore changes both the number of training examples and the number of optimizer updates.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Setting
Model and training data
Backbone
Qwen2.5-Omni-7B
Training set
5,700 QA instances from 1,485 OmniVideo videos
Training retention
5% nominal budget, fixed for each compressor
Video sampling
2 FPS; at most 128 frames
Per-frame pixel bounds
Minimum = maximum = 100,352
Appendix
Table 9: Training configuration for the main CAFD runs.
Compressor
Budget
Audio
Video
Total
Prefill compute
Acc. (%)
Retention (%)
(T)
WorldSense
Full tokens
100%
100.00
100.00
100.00
117.80
46.85
Uniform
35%
34.99 / 34.99
35.00 / 35.00
35.00 / 35.00
34.58 / 34.58
43.28 / 44.96
Uniform
25%
25.00 / 25.00
25.00 / 25.00
25.00 / 25.00
24.17 / 24.17
41.96 / 43.66
Uniform
15%
15.00 / 15.00
15.00 / 15.00
15.00 / 15.00
14.40 / 14.40
38.84 / 39.97
Appendix
Table 11: Detailed results before and after adaptation. Paired entries report Unadapted / +CAFD. All adapted models are trained at 5% nominal retention and evaluated with the same compressor. Retention is measured for pre-LLM methods; SEATS reports modality-wise layer-average targets and an input-token-weighted total estimate. Full tokens denotes the unadapted full-token reference; DivPrune-M denotes DivPrune-Merge.
Training data
Train
Eval35
Eval25
Eval15
Eval10
Eval5
Avg.
100% (5700)
35%
45.87
45.52
44.07
42.34
39.06
43.37
25%
46.15
45.55
44.01
42.81
39.34
43.58
15%
46.31
45.78
44.26
42.91
39.66
43.78
10%
46.15
46.06
44.42
42.97
40.04
43.93
5%
46.37
46.15
44.45
43.35
40.45
44.16
50% (2850)
35%
45.46
45.08
43.66
42.50
38.93
43.13
Appendix
Table 12: WorldSense accuracy (%) for the DivPrune data-scale by training-retention grid.
Compressor
Train
Benchmark
Eval35
Eval25
Eval15
Eval10
Eval5
Avg.
DivPrune
35%
WorldSense
45.87
45.52
44.07
42.34
39.06
43.37
DailyOmni
60.65
57.81
56.31
53.97
50.79
55.91
AVUT
62.57
61.48
59.46
57.79
53.34
58.93
25%
WorldSense
46.15
45.55
44.01
42.81
39.34
43.58
DailyOmni
60.90
58.15
56.64
54.39
50.96
56.21
AVUT
62.98
61.42
59.34
58.02
53.46
59.04
Appendix
Table 13: Accuracy (%) across training and deployment retention.
Variant
Benchmark
35%
25%
15%
10%
5%
Avg.
CAFD
WorldSense
46.37
46.15
44.45
43.35
40.45
44.16
(JSD)
DailyOmni
60.82
58.73
58.06
55.22
52.38
57.04
AVUT
63.21
62.00
60.03
58.30
54.27
59.56
Forward KL
WorldSense
44.80
44.39
42.59
41.58
38.65
42.40
DailyOmni
58.90
57.06
55.22
53.63
49.54
54.87
AVUT
62.00
60.38
58.94
56.92
53.23
58.29
Appendix
Table 14: Per-budget distillation-objective results. All runs use DivPrune at 5% training retention and are evaluated with DivPrune at the five listed budgets. Accuracies (%) are reported for WorldSense, DailyOmni, and AVUT. The bold CAFD entry uses JSD and the same seed-42 checkpoint as Table 1 .
Variant
Benchmark
35%
25%
15%
10%
5%
Avg.
Unadapted
WorldSense
45.49
45.02
43.25
42.06
39.12
42.99
DailyOmni
58.48
56.73
55.30
53.55
49.62
54.74
AVUT
62.28
60.90
58.19
56.69
52.88
58.19
CAFD (fixed-generic)
WorldSense
46.09
46.31
44.80
43.38
40.45
44.21
DailyOmni
61.32
59.06
56.98
55.47
52.30
57.03
AVUT
63.32
61.94
59.52
57.55
53.63
59.19
Appendix
Table 15: Per-budget task-conditioning and supervision results. All adapted variants use DivPrune at 5% training retention and are evaluated with DivPrune at the five listed budgets. Accuracies (%) are reported for WorldSense, DailyOmni, and AVUT. The fixed-generic and question-only rows are reference-free CAFD variants. The bold CAFD entry denotes the main configuration used in Table 1 , with its task input in parentheses.
Figure 4: Training dynamics at 5% training retention. (a) Training loss (clipped JSD). (b) Student rollout length. Curves show 15-step trailing averages for DivPrune, OmniZip*, and SEATS.
Figure 5: Paired repair and regression under compression. Bars show the percentages of matched examples changing from incorrect to correct (positive) or correct to incorrect (negative) after adaptation. Percentages use all evaluated examples as the denominator. Each panel reports all five DivPrune deployment budgets.
Benchmark
Problem type
N
Delta (pp)
Repair (%)
Regress (%)
WorldSense
Action Counting
165
0.24
6.00
13.57
WorldSense
Anomaly Recognition
79
1.27
3.71
2.58
WorldSense
Attribute Reasoning
121
3.97
8.32
1.11
WorldSense
Attribute Recognition
181
0.00
4.10
4.62
WorldSense
Audio Change
83
1.45
6.43
6.90
WorldSense
Audio Counting
90
1.11
7.19
9.73
Appendix
Table 18: Five-budget average behavior by benchmark-native problem type. The model is trained with DivPrune at 5% nominal retention and evaluated with DivPrune across five deployment budgets.
Omnimodal large language models (Omni-LLMs) show strong capability in audio-video understanding, but their practical deployment remains limited by high inference cost of long video streams and dense audio sequences. Despite recent progress, existing compression methods for Omni-LLMs typically rely on fixed or native compression units, which can disrupt cross-modal correspondence and the complementary information required for audio-video reasoning, making it difficult to improve inference efficiency while stably preserving performance. To address this, we propose OmniRefine, a training-free two-stage framework for efficient audio-visual token compression in Omni-LLMs. First, Correspondence-Preserving Chunk Refinement refines native chunk boundaries into cross-modally aligned compression units through frame-audio similarity and dynamic programming. Second, Modality-Aware Cooperative Compression jointly compresses video and audio tokens within each refined unit to reduce redundancy while preserving critical evidence. Extensive experiments show that OmniRefine achieves a better efficiency-performance trade-off than strong baselines and maintains stable performance under lower compression ratios. On WorldSense, it still reaches 46.7% accuracy at a 44% token retention ratio, nearly matching the full-token baseline. The code and interface will be released to facilitate further research.
Yuchen Deng, Zidang Cai, Hai-Tao Zheng +3
Tsinghua Shenzhen International Graduate School, Tsinghua University · Pengcheng Laboratory
Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly redundant visual and audio token sequences incurs substantial computational overhead, demanding aggressive token compression for efficient deployment. Existing methods often degrade at low token budgets: pre-LLM compression may discard structurally important and globally distributed evidence, whereas inner-LLM compression often underexploits query-conditioned audio-visual collaboration. To address these limitations, we propose OmniPack, a training-free framework that coordinates structural compression before the LLM with task-relevant semantic refinement within the LLM. Before the LLM, OmniPack removes structural redundancy through modality-specific importance, global coverage, and similarity-aware merging. After sufficient multimodal interaction, it further consolidates diverse, task-relevant representations through textual guidance and audio-visual collaboration. Extensive experiments on five benchmarks with three Omni-LLM backbones demonstrate that OmniPack consistently achieves the best performance-efficiency trade-off across diverse retention ratios, outperforming all existing methods. Notably, on Qwen2.5-Omni-7B, OmniPack preserves 98.0% of the original performance while reducing FLOPs to 16.7%, and still retains 92.9% of the original performance with only 6.8% of the original FLOPs.
Wanshun Su, Yang Shi, Feihu Liu +10
Northwestern Polytechnical University · Peking University · Alibaba Group +1
Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often breaks down: for the same query, audio and video relevance often peaks at different moments. This cross-modal salience mismatch makes unidirectional guidance prone to discarding answer-critical cues under aggressive compression. We propose OmniScope, a training-free token compression framework that uses the query as a shared semantic anchor while estimating relevance separately for audio and video. OmniScope allocates modality-specific token budgets, prunes visual tokens with an anchor-delta strategy that preserves both global context and temporal changes, and merges audio tokens within each second to reduce redundancy while maintaining temporal continuity. Across four audio-video benchmarks and two Qwen2.5-Omni model scales, OmniScope achieves the best average accuracy across all compression settings. At 25% overall token retention, it delivers up to 3.53x prefill speedup and more than 15% GPU memory reduction, with only a 0.35-point drop in average accuracy. These results suggest a simple design principle for OmniLLM inference: share the query across modalities, but not the salience estimates. The code is available at https://github.com/MAC-AutoML/OmniScope.
Jinsen Su, Yongdong Luo, Yuexiao Ma +3
Media Analytics and Computing Lab, Xiamen University, Xiamen, China · Institute of Artificial Intelligence, Xiamen University, Xiamen, China · School of Informatics, Xiamen University, Xiamen, China +1