Seeing What Should Be Heard: Diagnosing and Repairing Cross-Modal Shortcuts in Omni-Modal LLMs
Organizations: The University of Queensland, Brisbane, Australia · Shenzhen University, Shenzhen, China
Abstract
Omni-modal large language models (LLMs) are expected to answer a question using the modality it explicitly refers to. However, existing training paradigms rarely verify whether models actually follow this modality, because multimodal inputs from the same sample often provide redundant evidence for the same answer. In this work, we uncover a pervasive cross-modal shortcut in omni-modal LLMs: when asked an audio-related question, models rely on the image as much as on the audio, and sometimes even more. To systematically diagnose this behavior, we introduce the Factorized Modality Diagnostic, which independently swaps audio and images between samples to isolate each modality's causal contribution. Across two model families in different settings, we find that this shortcut persists throughout supervised fine-tuning and reinforcement learning post-training, while judge-based RL may further amplify such reliance on irrelevant visual information. Based on this finding, we propose DMC-Repair, which trains models on the same kind of cross-modal swapped samples while assigning supervision according to the modality specified by the question. This prevents models from exploiting the spurious correspondence between modalities within the same clip. Experiments demonstrate that DMC-Repair reduces the image-induced share of the answer effect by 59.9%, effectively suppressing the cross-modal shortcut without compromising audio-question answering performance. The reduction in shortcut reliance generalizes across two model families and zero-shot to an unseen dataset and an unseen benchmark, and persists through subsequent post-training. Code is available at https://anonymous.4open.science/r/DMC-Repair.
Figures & tables
| HumanOmniV2 | SFT clean | LoRA | Full | (Full) [95% CI] | |||
| Audio q., direct | (ideal ) | 0.498 | 0.549 | 0.194 | 0.187 | [ , ] | |
| (non-designated) | [ , ] | ||||||
| , confirmation set | — | 0.559 | 0.208 | 0.224 | [ , ] | ||
| Visual q., direct | (ideal ) | 0.716 | 0.713 | 0.872 | 0.872 | [ , ] | |
| (non-designated) | [ , ] | ||||||
| Audio q., own reasoning | (ideal ) | — | 0.612 | 0.256 | 0.271 | [ , ] | |
| Method | Training data | Base | Audio hall. | AV matching |
| AVCD † ( Jung et al., 2025 ) | none (decoding) | 73.0 | — | |
| MAD ( Chung et al., 2026 ) | none (decoding) | 73.0 | — | |
| ACPO ( Baid et al., 2026 ) | audio-swap preferences | 66.7 | — | |
| OmniDPO ( Chen et al., 2026a ) | audio-visual preferences | 67.6 | — | |
| MoD-DPO++ ( Chaubey et al., 2026 ) | modality-perturbed preferences | 77.4 | ||
| Chen et al. (2026b) | (mis)aligned AudioSet clips | 71.7 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| MUSIC-AVQA | AVQA | |
| Source annotation rows | 42,492 | 6,402 |
| usable answer format | 19,478 | 2,343 |
| Candidate families | 2,521 | 846 |
| fewer than two yes or two no | ||
| media already claimed by another family | — | |
| overlapping the frozen evaluation manifest | — |
| Model | Prompt | [95% CI] | ||
|---|---|---|---|---|
| Base | native | +1.64 | +1.68 | 0.478 [0.416, 0.538] |
| Base | neutral | +1.89 | +1.27 | 0.419 [0.366, 0.475] |
| HumanOmniV2 | native | +1.59 | +1.40 | 0.498 [0.441, 0.553] |
| HumanOmniV2 | neutral | +1.00 | +0.80 | 0.486 [0.431, 0.539] |
| SFT ms | native | +0.89 | +0.89 | 0.484 [0.424, 0.543] |
| SFT ms | neutral | +0.24 | +0.28 | 0.565 [0.504, 0.621] |
| Reward design | Start | [95% CI] | [95% CI] | [95% CI] |
|---|---|---|---|---|
| Graded description score ( ) | ms | [ , ] | [ , ] | [ , ] |
| Description gain ( ) | ms | [ , ] | [ , ] | [ , ] |
| Description and answer ( ) | ms | [ , ] | [ , ] | [ , ] |
| Yes/no description score ( ) | ms | [ , ] | [ , ] | inconclusive (31 fam.) |
| GRPO, no modality reward | clean | [ , ] | [ , ] | [ , ] |
| GRPO + grid cells | clean | [ , ] | [ , ] | [ , ] |
| Configuration | [95% CI] |
|---|---|
| Base, native | [ , ] |
| Base, neutral | [ , ] |
| HumanOmniV2, native | [ , ] |
| HumanOmniV2, neutral | [ , ] |
| SFT ms , native | [ , ] |
| SFT ms , neutral | [ , ] |
| Statistic | SFT clean | Repaired (full) | RL endpoint |
|---|---|---|---|
| , own audio | [ , ] | [ , ] | [ , ] |
| , partner audio | [ , ] | [ , ] | [ , ] |
| [ , ] | [ , ] | [ , ] | |
| mean | [ , ] | [ , ] | [ , ] |
| mean | [ , ] | [ , ] | [ , ] |
| Model | Q-type | |||
|---|---|---|---|---|
| Base | Audio (65) | +1.635 | +1.682 | 0.478 |
| Visual (27) | +1.565 | +5.014 | 0.747 | |
| HumanOmniV2 | Audio (65) | +1.591 | +1.403 | 0.498 |
| Visual (27) | +1.856 | +4.744 | 0.716 | |
| MiniCPM-o-2.6 | Audio (65) | +1.565 | +1.883 | 0.554 |
| Visual (27) | +1.363 | +7.104 | 0.802 |
| Q-type | SFT clean | Repaired (full) | RL endpoint | (full) |
|---|---|---|---|---|
| Audio (65) | 64.2 | 72.3 | 73.8 | [ , ] |
| Visual (27) | 83.3 | 80.6 | 81.5 | [ , ] |
| Audio-Visual (38) | 50.7 | 53.9 | 61.8 | [ , ] |
| Test | Statistic | Criterion met |
|---|---|---|
| H1 (primary): repair | [ , ] | Yes |
| G1: repair audio-following (points) | [ , ] | Yes |
| G2: RL endpoint generation retention (points) | [ , ] | Yes |
| H4: SI retention through RL ( ) | [ , ] | Yes |
| H3: Visual-question | [ , ] | Yes |
| H5: LoRA variant | [ , ] | Yes |
| All conflict cells | Parseable outputs only | Lenient | ||||
| Training cells | Audio | Image | Unparseable | Audio | to grid | to grid |
| SFT clean (no training) | 42.8 | 56.6 | 0.6 | 43.0 | [ , ] | [ , ] |
| Complete grid | 61.3 | 30.5 | 8.2 | 67.0 | — | — |
| Single-direction cells | 45.1 | 36.9 | 18.0 | 55.1 | [ , ] | [ , ] |
| Mismatch labels | 33.6 | 29.9 | 36.5 | 52.8 | [ , ] | [ , ] |
| Modality dropout | 43.8 | 41.6 | 14.6 | 51.1 | [ , ] | [ , ] |
| Statistic | Result | Met | |
| C1 answerable | SFT clean matched-cell accuracy chance | [ , ] | Yes |
| C2 replicates | SFT clean | [ , ] | Yes |
| C3 repair transfers | [ , ] | Yes | |
| per-pair | [ , ] | Yes | |
| matched-cell accuracy drop points | Yes | ||
| descriptive | conflict audio-following (gen.) | — |
| Model | [95% CI] | [95% CI] | [95% CI] |
|---|---|---|---|
| Base | [ , ] | [ , ] | 0.292 [0.253, 0.330] |
| HumanOmniV2 | [ , ] | [ , ] | 0.302 [0.258, 0.347] |
| SFT clean | [ , ] | [ , ] | 0.392 [0.349, 0.436] |
| Repaired (full) | [ , ] | [ , ] | 0.218 [0.186, 0.251] |
| RL endpoint | [ , ] | [ , ] | 0.238 [0.201, 0.274] |
| Model | Readout | Audio hall. | Video hall. | AV matching | Overall |
| Qwen2.5-Omni SFT clean | gen. | 61.9 (37.2) | 81.5 (48.6) | 55.4 (62.3) | 63.8 |
| + repair (full) | gen. | 67.7 (44.1) | 82.0 (47.9) | 62.6 (58.6) | 69.0 |
| + subsequent RL | gen. | 73.2 (52.1) | 82.2 (54.2) | 62.7 (65.8) | 71.4 |
| MiniCPM-o-2.6 (base) | margin | 78.7 (44.2) | 76.0 (31.8) | 64.0 (19.1) | 72.9 |
| + repair (LoRA) | margin | 80.2 (41.1) | 78.0 (35.2) | 64.8 (16.6) | 74.3 |
| Variant | Training cells | [95% CI] | Audio-Following, % (gen.) |
|---|---|---|---|
| SFT clean (no training) | — | — | 43.1 |
| Hinge, LoRA (reference) | complete | [ , ] | 60.8 |
| Hinge, full fine-tuning | complete | [ , ] | 61.2 |
| Cross-entropy, LoRA | complete | [ , ] | 61.9 |
| Hinge, LoRA, four seeds | complete | to | — |
| Cross-entropy, LoRA † | complete | [ , ] | — |
| Set (Audio families) | [95% CI] | |||
|---|---|---|---|---|
| Development (65) | [ , ] | |||
| Confirmation (128) | [ , ] |