Omni-modal large language models (LLMs) are expected to answer a question using the modality it explicitly refers to. However, existing training paradigms rarely verify whether models actually follow this modality, because multimodal inputs from the same sample often provide redundant evidence for the same answer. In this work, we uncover a pervasive cross-modal shortcut in omni-modal LLMs: when asked an audio-related question, models rely on the image as much as on the audio, and sometimes even more. To systematically diagnose this behavior, we introduce the Factorized Modality Diagnostic, which independently swaps audio and images between samples to isolate each modality's causal contribution. Across two model families in different settings, we find that this shortcut persists throughout supervised fine-tuning and reinforcement learning post-training, while judge-based RL may further amplify such reliance on irrelevant visual information. Based on this finding, we propose DMC-Repair, which trains models on the same kind of cross-modal swapped samples while assigning supervision according to the modality specified by the question. This prevents models from exploiting the spurious correspondence between modalities within the same clip. Experiments demonstrate that DMC-Repair reduces the image-induced share of the answer effect by 59.9%, effectively suppressing the cross-modal shortcut without compromising audio-question answering performance. The reduction in shortcut reliance generalizes across two model families and zero-shot to an unseen dataset and an unseen benchmark, and persists through subsequent post-training. Code is available at https://anonymous.4open.science/r/DMC-Repair.
Figures & tables
Figure 1: Conventional training and DMC-Repair training for an audio question. DMC-Repair trains on all four image–audio combinations of two samples and labels each cell by its audio. The bottom rows quote excerpts of the reasoning, with errors in red and audio evidence in green.
Figure 2: The 2×2 factorized design for an audio question. Crossing the image source and the audio source, each with two levels, own (A, blue) and partner (B, orange), gives four cells with answer margins muv for image u and audio v . Shaded cells contain one sample’s image and audio.
Figure 3
HumanOmniV2
SFT clean
LoRA
Full
Δ (Full) [95% CI]
Audio q., direct
SI (ideal =0 )
0.498
0.549
0.194
0.187
−0.362
[ −0.43 , −0.29 ]
EV (non-designated)
+1.40
+1.05
−0.04
+0.07
−0.98
[ −1.39 , −0.59 ]
SI , confirmation set
—
0.559
0.208
0.224
−0.335
[ −0.39 , −0.28 ]
Visual q., direct
SI (ideal =1 )
0.716
0.713
0.872
0.872
+0.159
[ +0.09 , +0.22 ]
EA (non-designated)
+1.86
+1.28
−0.10
+0.02
−1.26
[ −1.65 , −0.87 ]
Audio q., own reasoning
SI (ideal =0 )
—
0.612
0.256
0.271
−0.341
[ −0.42 , −0.27 ]
Table 1: Repair results on held-out development families. The confirmation-set row uses the 128 Audio families of the sealed confirmation set. Full and LoRA ( Hu et al., 2022 ) both start from SFT clean and use the same cells (Appendix A ). Full (bold) updates all LLM parameters, and LoRA trains a rank-16 adapter. The released HumanOmniV2 checkpoint ( Yang et al., 2025 ) is evaluated by direct scoring only. Δ is the change of Full from SFT clean with its paired 95% confidence interval (CI). In the own-reasoning block, SI is rescored after the model’s own reasoning, and the other two rows use free generation with unparseable outputs counted as failures (Appendix L ).
Figure 5: Construction controls on the confirmation set (128 Audio families, 510 conflict cells). The complete grid is the LoRA reference, and paired differences with 95% CIs are in Appendix F .
Method
Training data
Base
Δ Audio hall.
Δ AV matching
AVCD † ( Jung et al., 2025 )
none (decoding)
73.0
+2.8
—
MAD ( Chung et al., 2026 )
none (decoding)
73.0
+5.7
—
ACPO ( Baid et al., 2026 )
audio-swap preferences
66.7
+2.6
—
OmniDPO ( Chen et al., 2026a )
audio-visual preferences
67.6
+9.9
—
MoD-DPO++ ( Chaubey et al., 2026 )
modality-perturbed preferences
77.4
+6.0
+15.0
Chen et al. (2026b)
(mis)aligned AudioSet clips
71.7
+8.2
+0.0
Table 2: Comparison with published methods on AVHBench, all built on Qwen2.5-Omni. Base is the audio-hallucination accuracy that each paper reports for its base model, and each change is measured against that base model on the same task. Other numbers are copied from the papers (—, not reported), and Table 14 gives all tasks for our models. † Evaluated by Chung et al. (2026) .
Figure 6: Column area is proportional to the number of Audio families per 9∘ angle of (∣EA∣,∣EV∣) .
Figure 7: Generated answers on two conflict cells, colored by the modality that they follow.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
MUSIC-AVQA
AVQA
Source annotation rows
42,492
6,402
usable answer format
19,478
2,343
Candidate families
2,521
846
− fewer than two yes or two no
−2,134
−741
− media already claimed by another family
—
−1
− overlapping the frozen evaluation manifest
−0
—
Appendix
Table 3: Construction of the two diagnostic sets, from source annotations to retained families. The first two data rows count annotation rows, and the next five rows count families. Samples that share a question form a family, and we retain only families with at least two yes and two no samples. One Audio pair of the MUSIC-AVQA confirmation set was later dropped for missing media, giving 773 pairs rather than 774. The media-uniqueness constraint of Appendix G leaves six AVQA families with a single pair, giving 202 pairs rather than 208.
Figure 8: The diagnostic of Figure 4 with all seven configurations of Table 4 and both repairs. Filled markers are native-prompt runs and open markers neutral-prompt runs. The MiniCPM-o arrow starts at its base model, since no SFT stage precedes it.
Model
Prompt
EA
EV
SI [95% CI]
Base
native
+1.64
+1.68
0.478 [0.416, 0.538]
Base
neutral
+1.89
+1.27
0.419 [0.366, 0.475]
HumanOmniV2
native
+1.59
+1.40
0.498 [0.441, 0.553]
HumanOmniV2
neutral
+1.00
+0.80
0.486 [0.431, 0.539]
SFT ms
native
+0.89
+0.89
0.484 [0.424, 0.543]
SFT ms
neutral
+0.24
+0.28
0.565 [0.504, 0.621]
Appendix
Table 4: Audio-question diagnostic on 65 development families, where the ideal SI is 0 . Base is Qwen2.5-Omni, HumanOmniV2 is the released checkpoint of Yang et al. (2025) , and SFT ms is an earlier version of SFT clean . The neutral prompt omits the system prompt and output format. The paired native-prompt change in SI from Base to HumanOmniV2 is +0.021 [ −0.019 , +0.061 ].
Reward design
Start
ΔEA [95% CI]
ΔEV [95% CI]
ΔSI [95% CI]
Graded description score ( w=1.0 )
ms
−0.012 [ −0.03 , +0.01 ]
+0.053 [ +0.03 , +0.08 ]
+0.019 [ +0.005 , +0.033 ]
Description gain ( w=0.3 )
ms
+0.051 [ +0.03 , +0.07 ]
+0.069 [ +0.04 , +0.09 ]
−0.002 [ −0.018 , +0.015 ]
Description and answer ( w=0.3 )
ms
+0.028 [ +0.02 , +0.04 ]
+0.035 [ +0.02 , +0.05 ]
+0.005 [ −0.010 , +0.020 ]
Yes/no description score ( w=0.3 )
ms
−0.029 [ −0.28 , +0.22 ]
−0.178 [ −0.40 , +0.04 ]
inconclusive (31 fam.)
GRPO, no modality reward
clean
+0.038 [ +0.02 , +0.06 ]
+0.074 [ +0.05 , +0.10 ]
+0.003 [ −0.011 , +0.016 ]
GRPO + grid cells
clean
+0.104 [ +0.07 , +0.14 ]
+0.058 [ +0.03 , +0.09 ]
−0.021 [ −0.036 , −0.007 ]
Appendix
Table 5: Changes in the Audio-question diagnostic after RL, paired against each run’s own starting checkpoint (65 families unless noted). The Start column gives the starting checkpoint (ms for SFT ms , clean for SFT clean ), and w is the weight of the modality reward. Positive ΔEV indicates a larger image effect. Bold marks the design that increases only the image effect.
Configuration
ψ [95% CI]
Base, native
−0.482 [ −0.718 , −0.267 ]
Base, neutral
−0.556 [ −0.813 , −0.327 ]
HumanOmniV2, native
−0.513 [ −0.769 , −0.277 ]
HumanOmniV2, neutral
−0.353 [ −0.511 , −0.201 ]
SFT ms , native
−0.359 [ −0.505 , −0.229 ]
SFT ms , neutral
−0.052 [ −0.113 , +0.005 ]
Appendix
Table 6: Interaction terms ψ for Audio questions, under the orientation and aggregation stated above.
Statistic
SFT clean
Repaired (full)
RL endpoint
δV(o) , own audio
+0.682 [ +0.414 , +0.966 ]
−0.027 [ −0.149 , +0.083 ]
+0.019 [ −0.113 , +0.146 ]
δV(x) , partner audio
+1.417 [ +0.905 , +1.970 ]
+0.165 [ +0.016 , +0.319 ]
+0.449 [ +0.228 , +0.684 ]
ψ
−0.368 [ −0.521 , −0.233 ]
−0.096 [ −0.181 , −0.015 ]
−0.215 [ −0.338 , −0.096 ]
mean ∣EV∣
1.275 [ 0.939 , 1.651 ]
0.343 [ 0.278 , 0.413 ]
0.506 [ 0.417 , 0.602 ]
mean max∣δV∣
1.687 [ 1.220 , 2.191 ]
0.629 [ 0.524 , 0.742 ]
0.928 [ 0.775 , 1.100 ]
Appendix
Table 7: Image effect on Audio questions decomposed by the audio held fixed, 65 development families. Own audio comes from the item whose answer is yes , and partner audio from the item whose answer is no . The last two rows average over families the absolute image effect and the larger absolute conditional effect of each family, so effects of opposite sign cannot cancel. Intervals use 8,000 resamples.
Model
Q-type
EA
EV
SI
Base
Audio (65)
+1.635
+1.682
0.478
Visual (27)
+1.565
+5.014
0.747
HumanOmniV2
Audio (65)
+1.591
+1.403
0.498
Visual (27)
+1.856
+4.744
0.716
MiniCPM-o-2.6
Audio (65)
+1.565
+1.883
0.554
Visual (27)
+1.363
+7.104
0.802
Appendix
Table 8: Diagnostic by question type on the development set. For Audio questions, the image is non-designated ( SI ideal =0 ). For Visual questions, the audio is non-designated ( SI ideal =1 ).
Q-type
SFT clean
Repaired (full)
RL endpoint
Δ (full)
Audio (65)
64.2
72.3
73.8
+8.1 [ +2.3 , +13.5 ]
Visual (27)
83.3
80.6
81.5
−2.8 [ −10.2 , +4.6 ]
Audio-Visual (38)
50.7
53.9
61.8
+3.3 [ −4.6 , +11.2 ]
Appendix
Table 9: Matched-cell accuracy by question type on the development set, under the direct answer scoring of Section 3.2 . Values are percentages, and the final column gives the repaired model’s paired change against SFT clean with 95% family-cluster bootstrap intervals over 8,000 resamples.
Test
Statistic
Criterion met
H1 (primary): repair ΔSI
−0.335 [ −0.390 , −0.279 ]
Yes
G1: repair audio-following Δ (points)
+13.7 [ +7.4 , +19.7 ]
Yes
G2: RL endpoint generation retention (points)
+11.9 [ +8.1 , +15.6 ]
Yes
H4: SI retention through RL ( ≥80% )
D−0.2B=−0.067 [ −0.089 , −0.046 ]
Yes
H3: Visual-question ΔSI
+0.176 [ +0.110 , +0.241 ]
Yes
H5: LoRA variant ΔSI
−0.352 [ −0.407 , −0.294 ]
Yes
Appendix
Table 10: Results for the pre-specified confirmation hypotheses on the 257-family confirmation set. All seven tests meet their criteria, and H2 is reported in two rows. The Audio-question SI tests use 128 families, and H3 evaluates the 53 Visual families. SFT, full, and RL denote the starting checkpoint, the repaired model, and the RL endpoint. H4 uses the repair gain B=SISFT−SIfull and the change through RL D=SIRL−SIfull , observed as 0.000 [ −0.024 , +0.023 ], and G2 uses the audio-following rates pS , pF , and pR of the three checkpoints. H2 tests whether the repaired image effect lies within ±10% of ∣EV(SFT)∣ around zero, and the interval of its upper row must lie below zero and that of its lower row above zero. The lower block compares each construction control with the complete-pool LoRA reference as control minus reference, and its last column states whether the result goes in the pre-registered direction.
All conflict cells
Parseable outputs only
Lenient
Training cells
Audio
Image
Unparseable
Audio
Δ to grid
Δ to grid
SFT clean (no training)
42.8
56.6
0.6
43.0
−24.0 [ −30.7 , −17.3 ]
−22.1 [ −28.5 , −15.6 ]
Complete 2×2 grid
61.3
30.5
8.2
67.0
—
—
Single-direction cells
45.1
36.9
18.0
55.1
−11.8 [ −17.3 , −6.3 ]
−12.3 [ −17.8 , −7.0 ]
Mismatch labels
33.6
29.9
36.5
52.8
−14.2 [ −19.7 , −8.6 ]
−31.2 [ −36.1 , −26.4 ]
Modality dropout
43.8
41.6
14.6
51.1
−15.8 [ −22.0 , −9.6 ]
−14.5 [ −20.5 , −8.6 ]
Appendix
Table 11: Generated answers on the 510 confirmation conflict cells, as family-averaged shares in percent of answers that follow the audio, answers that follow the image, and unparseable outputs. The first numeric column is the pre-registered audio-following rate, in which unparseable outputs count as failures. The parseable columns exclude unparseable outputs, pooled over cells, and the lenient column accepts an unparseable output whose text after </think> contains only one of yes and no , and counts the other unparseable outputs as failures. The complete grid is the LoRA reference of Figure 5 , and differences to it are paired, with 95% family-cluster bootstrap intervals.
Figure 9: Change in Audio-question SI from the repaired model during subsequent RL post-training, with pointwise 95% paired family-cluster bootstrap intervals.
Statistic
Result
Met
C1 answerable
SFT clean matched-cell accuracy > chance
68.3% [ 64.2 , 72.4 ]
Yes
C2 replicates
SFT clean EV>0
+0.852 [ +0.694 , +1.008 ]
Yes
C3 repair transfers
ΔEV<0
−0.529 [ −0.664 , −0.399 ]
Yes
per-pair ΔSI<0
−0.104 [ −0.167 , −0.040 ]
Yes
matched-cell accuracy drop ≤2 points
68.3%→82.2%
Yes
descriptive
conflict audio-following (gen.)
44.4%→62.5%
—
Appendix
Table 12: Pre-registered AVQA criteria, evaluated once on 104 families after the checkpoints were frozen. Intervals are 95% family-cluster bootstrap CIs over 8,000 resamples, and C3 is paired against SFT clean . The per-pair ΔSI row follows the registered definition, which averages the signed ratio EV/(∣EA∣+∣EV∣) over pairs. The family-level SI of Section 3.2 , which Table 13 reports, changes by −0.174 [ −0.210 , −0.140 ]. The generation rows use 40 of the 104 families (160 conflict cells) on AVQA and the 65 development Audio families on MUSIC-AVQA.
Model
EA [95% CI]
EV [95% CI]
SI [95% CI]
Base
+2.911 [ +2.564 , +3.286 ]
+0.950 [ +0.764 , +1.149 ]
0.292 [0.253, 0.330]
HumanOmniV2
+3.381 [ +2.970 , +3.785 ]
+1.001 [ +0.801 , +1.199 ]
0.302 [0.258, 0.347]
SFT clean
+1.556 [ +1.337 , +1.776 ]
+0.852 [ +0.694 , +1.008 ]
0.392 [0.349, 0.436]
Repaired (full)
+1.522 [ +1.347 , +1.693 ]
+0.323 [ +0.247 , +0.402 ]
0.218 [0.186, 0.251]
RL endpoint
+2.044 [ +1.806 , +2.276 ]
+0.466 [ +0.359 , +0.576 ]
0.238 [0.201, 0.274]
Appendix
Table 13: Audio-question diagnostic on the 104 AVQA families under the native prompt. Intervals are 95% family-cluster bootstrap CIs over 8,000 resamples.
Model
Readout
Audio hall.
Video hall.
AV matching
Overall
Qwen2.5-Omni SFT clean
gen.
61.9 (37.2)
81.5 (48.6)
55.4 (62.3)
63.8
+ repair (full)
gen.
67.7 (44.1)
82.0 (47.9)
62.6 (58.6)
69.0
+ subsequent RL
gen.
73.2 (52.1)
82.2 (54.2)
62.7 (65.8)
71.4
MiniCPM-o-2.6 (base)
margin
78.7 (44.2)
76.0 (31.8)
64.0 (19.1)
72.9
+ repair (LoRA)
margin
80.2 (41.1)
78.0 (35.2)
64.8 (16.6)
74.3
Appendix
Table 14: Zero-shot AVHBench accuracy (%) by task, with the yes-rate (%) in parentheses. Tasks are video-driven audio hallucination (audio hall.), audio-driven video hallucination (video hall.), and audio-visual matching. The table covers all 5,302 questions, and “Overall” is the question-weighted mean over the three tasks. The two model families use different answer readouts and are not directly comparable to each other, so each is compared with its own starting checkpoint. Of the audio-hallucination answers that change from no to yes and from yes to no between the starting checkpoint and the RL endpoint, 69% and 64% , respectively, are correct at the RL endpoint.
Figure 10: Two conflict cells in the format of Figure 7 . All three checkpoints follow the image.
Figure 11: Six conflict cells in the format of Figure 7 , four from development pairs and two, in rows three and four, from the confirmation set. Rows one and two are the two directions of the same pair, and the remaining rows show cells from four other pairs. Green boxes indicate outputs that follow the audio, and red boxes indicate outputs that follow the image.
Variant
Training cells
ΔSI [95% CI]
Audio-Following, % (gen.)
SFT clean (no training)
—
—
43.1
Hinge, LoRA (reference)
complete 2×2
−0.356 [ −0.427 , −0.285 ]
60.8
Hinge, full fine-tuning
complete 2×2
−0.362 [ −0.430 , −0.293 ]
61.2
Cross-entropy, LoRA
complete 2×2
−0.317 [ −0.377 , −0.258 ]
61.9
Hinge, LoRA, four seeds
complete 2×2
−0.356 to −0.296
—
Cross-entropy, LoRA †
complete 2×2
−0.297 [ −0.364 , −0.232 ]
—
Appendix
Table 15: Loss and parameterization variants, development families. The reference row is the LoRA recipe on the complete pool, which is also the reference of Figure 5 . The seed row reports the range over four seeds of that reference, each with an interval that excludes zero, and full fine-tuning differs from the reference by −0.007 [ −0.053 , +0.042 ] in paired ΔSI . The matched-cells-only control differs from SFT clean by −4.6 [ −12.3 , +3.5 ] points in audio-following. † Trained with cross-entropy on an earlier and smaller pool.
Set (Audio families)
EA
EV
SI
ΔSI [95% CI]
Development (65)
+1.57→+1.27
+1.88→+0.20
0.554→0.261
−0.293 [ −0.361 , −0.223 ]
Confirmation (128)
+1.17→+0.92
+1.68→+0.17
0.604→0.297
−0.307 [ −0.365 , −0.248 ]
Appendix
Table 16: Audio-question diagnostic of MiniCPM-o-2.6 before and after the repair, under the neutral prompt and the first-token margin. Arrows go from the base model to the repaired model, and ΔSI is the paired change with its 95% family-cluster bootstrap interval over 10,000 resamples.
Omni-modal language models are intended to jointly understand audio, visual inputs, and language, but benchmark gains can be inflated when visual evidence alone is enough to answer a query. We study whether current omni-modal benchmarks separate visual shortcuts from genuine audio-visual-language evidence integration, and how post-training behaves under a visually debiased evaluation setting. We audit nine omni-modal benchmarks with visual-only probing, remove visually solvable queries, and retain full subsets when filtering is undefined or would make comparisons unstable. This yields OmniClean, a cleaned evaluation view with 8,551 retained queries from 16,968 audited queries. On OmniClean, we evaluate OmniBoost, a three-stage post-training recipe based on Qwen2.5-Omni-3B: mixed bi-modal SFT, mixed-modality RLVR, and SFT on self-distilled data. Balanced bi-modal SFT gives limited and uneven gains, RLVR provides the first broad improvement, and self-distillation reshapes the benchmark profile. After SFT on self-distilled data, the 3B model reaches performance comparable to, and in aggregate slightly above, Qwen3-Omni-30B-A3B-Instruct without using a stronger omni-modal teacher. These results show that omni-modal progress is easier to interpret when evaluation controls visual leakage, and that small omni-modal models can benefit from staged post-training with self-distilled omni-query supervision. Project page: https://cheliu-computation.github.io/omni/
Audio and vision provide complementary evidence for audio-visual question answering, yet current audio-visual large language models may suffer from cross-modal interference: information from one modality misguides the interpretation of another, thereby inducing hallucinations. We attribute this issue to uncontrolled cross-modal interactions during intermediate reasoning. To mitigate this, we propose Separate First, Fuse Later (SFFL), an audio-visual reasoning framework designed to reduce cross-modal interference. SFFL enforces modality-specific chain-of-thought reasoning, producing separate audio and visual reasoning traces and integrating evidence for answering. We construct modality-preference labels via a data pipeline under different modality input settings. We use these labels as an auxiliary reward in reinforcement learning to encourage a instance-dependent preference for modality cues when answering. We further introduce a modality-specific reasoning mechanism that preserves modality isolation during the separated reasoning stage while enabling full access to cross-modal information at the evidence fusion stage. Experiments demonstrate consistent improvements in both accuracy and robustness, yielding an average relative gain of 5.16% on general AVQA benchmarks and 11.17% on a cross-modal hallucination benchmark.
Xuanchen Li, Yuheng Lu, Chenrui Cui +6
Tianjin Key Laboratory of Cognitive Computing and Application, Tianjin University, Tianjin, China · Tencent, China · Huiyan Technology Company, Ltd., Tianjin, China +1
Long audio-video reasoning is difficult for omnimodal LLMs because the decisive evidence is often sparse, cross-modal, and too expensive to preserve with uniformly high-fidelity inputs. We introduce OmniReasoner, a tool-use post-training framework for Thinking with Long Audio-Video: omni-modal LLMs learn, via supervised fine-tuning and reinforcement learning, to decide whether and where to call a zoom-in tool before answering. OmniReasoner first builds a low-cost global preview of the full stream and then, when needed, calls the zoom-in tool with a requested temporal interval for higher-fidelity visual and audio inspection before answering. Because the model observes different sampling granularities before and after this call -- a sparse global preview and a denser local clip -- we introduce TimeAnchor, which keeps the tool's temporal argument valid and round-trip-consistent across these granularities, rather than tied to frame indices from a particular sampling rate. To make this tool-use behavior trainable without expensive manual interval annotation, we build a Temporal Augmented Data Engine that synthesizes tool-use post-training trajectories by video editing and composition. Experiments across omnimodal and video benchmarks show that OmniReasoner improves both answer accuracy and temporal grounding while concentrating high-fidelity computation on informative regions. Code is available at https://github.com/RockyChen0205/OmniReasoner.
Yu Chen, Caorui Li, Ziyu Xiong +8
University of Chinese Academy of Sciences · Institute of Automation, CAS · Southeast University +3