Omni-modal large language models (OmniLLMs) enable unified audio-video understanding, but their long multimodal token sequences make deployment computationally expensive. Token compression reduces this cost, yet aggressive compression often lowers accuracy. Existing works predominantly focus on designing better compression mechanisms; however, adapting the underlying language model to reason effectively over the remaining compressed context remains under-explored. To address this, we propose CAFD (Compressed-Context Adaptation via Full-Context Distillation), a ground-truth-free self-distillation framework that adapts OmniLLMs to fixed compression pipelines without requiring reference answers, rationales, or correctness rewards. CAFD leverages the full-token view of the same multimodal sample as a source of privileged information: a full-context self-teacher provides soft target supervision to a compressed-context student along the student's on-policy trajectory. Evaluated on Qwen2.5-Omni-7B across five audio-video benchmarks, five compression pipelines, and five deployment budgets, CAFD demonstrates consistent gains, improving 120 out of 125 conditions with an average accuracy boost of 1.44 points and recovering 26.9% of the accuracy gap on average. These results demonstrate that the proposed ground-truth-free adaptation offers an effective and practical route to improving the accuracy-efficiency trade-off in deployed OmniLLMs.
Figures & tables
Figure 1: Complementary adaptation and accuracy gains. (a) Compression methods reduce the token count; CAFD complements them by adapting the OmniLLM to compressed tokens while keeping the compressor unchanged. Pre-LLM compression is illustrated; CAFD also supports inner-LLM and hybrid compression. (b) Gray and blue bars show each compressor’s mean accuracy before and after CAFD adaptation, averaging five benchmarks and five deployment budgets. The dashed line denotes the unadapted full tokens baseline (54.86%). Red labels indicate accuracy-gap recovery: the percentage of the full-token–unadapted accuracy gap closed by adaptation.
Figure 2: Overview of CAFD. Self-distillation guides the model under compressed context toward its full-context capabilities. The compressed student generates a response; both paths compute next-token distributions along it using the same sample and instruction, without reference answers or rationales. Gradients update only the student; the teacher parameters track the student via an exponential moving average (EMA). Deployment retains the adapted model and original compressor. The framework supports pre-LLM (illustrated), inner-LLM, and hybrid compression.
Compressor
Budget
WorldSense
DailyOmni
AVUT
OmniVideoBench
Video-MME
Avg.
Full tokens
100%
46.85
62.91
64.65
35.50
64.41
54.86
Uniform
35%
43.28 / 44.96
57.23 / 58.65
60.44 / 62.86
33.80 / 33.70
63.44 / 64.81
51.64 / 53.00
25%
41.96 / 43.66
52.05 / 55.89
56.81 / 59.05
31.40 / 32.60
61.63 / 62.70
48.77 / 50.78
15%
38.84 / 39.97
49.37 / 52.72
52.94 / 55.25
29.90 / 31.50
60.33 / 61.67
46.28 / 48.22
10%
36.44 / 37.64
45.45 / 48.45
50.35 / 51.50
29.50 / 30.80
58.15 / 60.22
43.98 / 45.72
5%
35.18 / 36.66
40.77 / 43.69
46.37 / 47.92
30.10 / 31.00
54.89 / 57.37
41.46 / 43.33
Table 1: Main results across compressors, budgets, and benchmarks. Each cell reports Unadapted / +CAFD accuracy (%). Each +CAFD model is trained under the corresponding compressor’s 5% budget configuration and is evaluated at the five listed budgets. Column Avg. averages benchmarks, block-ending Avg. averages budgets, and their intersection averages all 25 conditions. The higher value in each pair is bold. DivPrune-M denotes DivPrune-Merge.
Figure 3: Effects of training data scale and retention. (a) WorldSense five-budget average accuracy for DivPrune across training-set sizes and training retentions. (b) Equal-weight accuracy-point gain over the matched Unadapted baseline across WorldSense, DailyOmni, and AVUT at five deployment budgets per benchmark. All runs use two epochs; data scale therefore changes both the number of training examples and the number of optimizer updates.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Setting
Model and training data
Backbone
Qwen2.5-Omni-7B
Training set
5,700 QA instances from 1,485 OmniVideo videos
Training retention
5% nominal budget, fixed for each compressor
Video sampling
2 FPS; at most 128 frames
Per-frame pixel bounds
Minimum = maximum = 100,352
Appendix
Table 9: Training configuration for the main CAFD runs.
Compressor
Budget
Audio
Video
Total
Prefill compute
Acc. (%)
Retention (%)
(T)
WorldSense
Full tokens
100%
100.00
100.00
100.00
117.80
46.85
Uniform
35%
34.99 / 34.99
35.00 / 35.00
35.00 / 35.00
34.58 / 34.58
43.28 / 44.96
Uniform
25%
25.00 / 25.00
25.00 / 25.00
25.00 / 25.00
24.17 / 24.17
41.96 / 43.66
Uniform
15%
15.00 / 15.00
15.00 / 15.00
15.00 / 15.00
14.40 / 14.40
38.84 / 39.97
Appendix
Table 11: Detailed results before and after adaptation. Paired entries report Unadapted / +CAFD. All adapted models are trained at 5% nominal retention and evaluated with the same compressor. Retention is measured for pre-LLM methods; SEATS reports modality-wise layer-average targets and an input-token-weighted total estimate. Full tokens denotes the unadapted full-token reference; DivPrune-M denotes DivPrune-Merge.
Training data
Train
Eval35
Eval25
Eval15
Eval10
Eval5
Avg.
100% (5700)
35%
45.87
45.52
44.07
42.34
39.06
43.37
25%
46.15
45.55
44.01
42.81
39.34
43.58
15%
46.31
45.78
44.26
42.91
39.66
43.78
10%
46.15
46.06
44.42
42.97
40.04
43.93
5%
46.37
46.15
44.45
43.35
40.45
44.16
50% (2850)
35%
45.46
45.08
43.66
42.50
38.93
43.13
Appendix
Table 12: WorldSense accuracy (%) for the DivPrune data-scale by training-retention grid.
Compressor
Train
Benchmark
Eval35
Eval25
Eval15
Eval10
Eval5
Avg.
DivPrune
35%
WorldSense
45.87
45.52
44.07
42.34
39.06
43.37
DailyOmni
60.65
57.81
56.31
53.97
50.79
55.91
AVUT
62.57
61.48
59.46
57.79
53.34
58.93
25%
WorldSense
46.15
45.55
44.01
42.81
39.34
43.58
DailyOmni
60.90
58.15
56.64
54.39
50.96
56.21
AVUT
62.98
61.42
59.34
58.02
53.46
59.04
Appendix
Table 13: Accuracy (%) across training and deployment retention.
Variant
Benchmark
35%
25%
15%
10%
5%
Avg.
CAFD
WorldSense
46.37
46.15
44.45
43.35
40.45
44.16
(JSD)
DailyOmni
60.82
58.73
58.06
55.22
52.38
57.04
AVUT
63.21
62.00
60.03
58.30
54.27
59.56
Forward KL
WorldSense
44.80
44.39
42.59
41.58
38.65
42.40
DailyOmni
58.90
57.06
55.22
53.63
49.54
54.87
AVUT
62.00
60.38
58.94
56.92
53.23
58.29
Appendix
Table 14: Per-budget distillation-objective results. All runs use DivPrune at 5% training retention and are evaluated with DivPrune at the five listed budgets. Accuracies (%) are reported for WorldSense, DailyOmni, and AVUT. The bold CAFD entry uses JSD and the same seed-42 checkpoint as Table 1 .
Variant
Benchmark
35%
25%
15%
10%
5%
Avg.
Unadapted
WorldSense
45.49
45.02
43.25
42.06
39.12
42.99
DailyOmni
58.48
56.73
55.30
53.55
49.62
54.74
AVUT
62.28
60.90
58.19
56.69
52.88
58.19
CAFD (fixed-generic)
WorldSense
46.09
46.31
44.80
43.38
40.45
44.21
DailyOmni
61.32
59.06
56.98
55.47
52.30
57.03
AVUT
63.32
61.94
59.52
57.55
53.63
59.19
Appendix
Table 15: Per-budget task-conditioning and supervision results. All adapted variants use DivPrune at 5% training retention and are evaluated with DivPrune at the five listed budgets. Accuracies (%) are reported for WorldSense, DailyOmni, and AVUT. The fixed-generic and question-only rows are reference-free CAFD variants. The bold CAFD entry denotes the main configuration used in Table 1 , with its task input in parentheses.
Figure 4: Training dynamics at 5% training retention. (a) Training loss (clipped JSD). (b) Student rollout length. Curves show 15-step trailing averages for DivPrune, OmniZip*, and SEATS.
Figure 5: Paired repair and regression under compression. Bars show the percentages of matched examples changing from incorrect to correct (positive) or correct to incorrect (negative) after adaptation. Percentages use all evaluated examples as the denominator. Each panel reports all five DivPrune deployment budgets.
Benchmark
Problem type
N
Delta (pp)
Repair (%)
Regress (%)
WorldSense
Action Counting
165
0.24
6.00
13.57
WorldSense
Anomaly Recognition
79
1.27
3.71
2.58
WorldSense
Attribute Reasoning
121
3.97
8.32
1.11
WorldSense
Attribute Recognition
181
0.00
4.10
4.62
WorldSense
Audio Change
83
1.45
6.43
6.90
WorldSense
Audio Counting
90
1.11
7.19
9.73
Appendix
Table 18: Five-budget average behavior by benchmark-native problem type. The model is trained with DivPrune at 5% nominal retention and evaluated with DivPrune across five deployment budgets.
Media Analytics and Computing Lab, Xiamen University, Xiamen, China · Institute of Artificial Intelligence, Xiamen University, Xiamen, China · School of Informatics, Xiamen University, Xiamen, China +1