Omni-modal large language models deployed in real-world environments encounter external noise that can interfere with their perception and understanding of multimodal inputs. We study their robustness in audio-visual understanding, focusing on question answering under environmental noise and competing speech. The challenge is to resist acoustic interference while preserving useful audio evidence. On-policy distillation provides dense teacher feedback on student-generated responses, but uniform token weighting does not explicitly prioritize positions affected by acoustic interference. We introduce OP-CAD (On-Policy Clean-Audio Distillation), a curriculum-based privileged self-distillation framework for robust audio-visual understanding. Training progresses from mild to severe environmental noise and competing speech, with selective token-level supervision at each stage. The student generates responses from corrupted audio-visual input, while a frozen teacher uses clean audio and the verified answer to supervise the same response prefixes. To allocate this supervision, OP-CAD compares teacher predictions under clean, corrupted, and visual-only contexts without revealing the answer. These matched comparisons measure sensitivity to audio removal and corruption; a bounded weighting rule emphasizes positions identified by either signal while retaining supervision throughout the response. OP-CAD outperforms the compared methods across all evaluated noise conditions. Paired analyses further show improved preservation of clean-correct answers under strong interference, with no observed aggregate clean-accuracy penalty. These results demonstrate the value of directing clean-teacher supervision toward acoustically sensitive predictions for robust audio-visual reasoning.
Figure 2: Overview of the OP-CAD training framework.
Figure 3: Token-level supervision allocation in OPSD and OP-CAD.
Daily-Omni
OmniVideoBench
Method
Clean
Env.
Speech
Clean
Env.
Speech
Base
67.75
61.38
57.84
37.40
36.70
34.80
SFT warm-up
68.76
61.74
59.82
38.10
36.80
35.93
+ SFT
68.42 ↓ 0.34
61.88 ↑ 0.14
59.12 ↓ 0.70
39.00 ↑ 0.90
36.23 ↓ 0.57
36.20 ↑ 0.27
+ GRPO
68.34 ↓ 0.42
62.91 ↑ 1.17
59.93 ↑ 0.11
41.50 ↑ 3.40
37.87 ↑ 1.07
36.87 ↑ 0.94
+ OPSD
69.34 ↑ 0.58
63.49 ↑ 1.75
60.62 ↑ 0.80
41.40 ↑ 3.30
37.40 ↑ 0.60
35.13 ↓ 0.80
Table 1: Full test accuracy (%). Env./Speech: three-SNR averages. Bold: best result. Arrows: changes from SFT warm-up in percentage points, using displayed scores.
Daily-Omni
Cond.
Method
Question Type
Video Duration
Avg.
AV Align
Comp.
Ctx. Und.
Evt. Seq.
Infer.
Reas.
30s
60s
Clean
Base
58.82
77.86
60.62
60.13
81.82
81.14
68.01
67.45
67.75
+ SFT
61.34 ↑ 2.52
77.10 ↓ 0.76
64.77 ↑ 4.15
57.84 ↓ 2.29
83.77 ↑ 1.95
80.57 ↓ 0.57
68.62 ↑ 0.61
68.18 ↑ 0.73
68.42 ↑ 0.67
+ GRPO
59.24 ↑ 0.42
79.39 ↑ 1.53
63.73 ↑ 3.11
59.80 ↓ 0.33
82.47 ↑ 0.65
80.00 ↓ 1.14
70.48 ↑ 2.47
65.82 ↓ 1.63
68.34 ↑ 0.59
+ OPSD
56.72 ↓ 2.10
77.86 ↔ 0.00
68.91 ↑ 8.29
60.78 ↑ 0.65
84.42 ↑ 2.60
82.29 ↑ 1.15
71.41 ↑ 3.40
66.91 ↓ 0.54
69.34 ↑ 1.59
Table 2: Daily-Omni accuracy (%) by question type and duration. Arrows: changes from Base within each condition in percentage points, using displayed scores. Bold: best per column.
OmniVideoBench
Cond.
Method
Audio Type
Video Duration
Avg.
Music
Sound
Speech
(0,1] min
(1,5] min
(5,10] min
(10,30] min
Clean
Base
30.77
32.65
39.11
49.12
35.29
36.40
33.33
37.40
+ SFT
32.97 ↑ 2.20
37.41 ↑ 4.76
40.03 ↑ 0.92
47.95 ↓ 1.17
39.41 ↑ 4.12
35.96 ↓ 0.44
35.25 ↑ 1.92
39.00 ↑ 1.60
+ GRPO
36.26 ↑ 5.49
36.73 ↑ 4.08
43.04 ↑ 3.93
49.12 ↔ 0.00
40.59 ↑ 5.30
39.47 ↑ 3.07
39.46 ↑ 6.13
41.50 ↑ 4.10
+ OPSD
28.57 ↓ 2.20
40.82 ↑ 8.17
43.04 ↑ 3.93
48.54 ↓ 0.58
42.06 ↑ 6.77
40.79 ↑ 4.39
36.40 ↑ 3.07
41.40 ↑ 4.00
Table 3: OmniVideoBench accuracy (%) by audio type and duration. Arrows: changes from Base within each condition in percentage points, using displayed scores. Bold: best per column.
Daily-Omni
OmniVideoBench
Variant
Clean
Env.
Speech
Clean
Env.
Speech
OPSD
69.34
63.49
60.62
41.40
37.40
35.13
Audio-rem. proxy
69.34
64.47
61.71
40.70
39.63
38.50
Corr.-sens. proxy
68.76
65.41
62.13
43.80
39.73
39.00
OP-CAD
70.09
66.97
63.55
41.60
40.23
38.93
Table 4: Accuracy (%) with single-signal and combined allocation. Bold marks the best result.
Training schedule
Daily-Omni
OmniVideoBench
Clean
Env.
Speech
Clean
Env.
Speech
Randomly mixed
69.93
65.16
62.67
42.30
38.27
37.13
Curriculum learning
70.09
66.97
63.55
41.60
40.23
38.93
Table 5: Curriculum order ablation: accuracy (%) under matched training budgets.
Figure 4: Prediction stability at 0 dB: harm rate (left) and clean-correct retention (right).
Figure 5: Token-level supervision allocation on an OP-CAD training rollout.