This work aims to improve how audio-visual large language models (AVLLMs) associate speech with the correct visible speaker in multi-speaker scenes. We find that current AVLLMs frequently fail at this task, and analyze the nature of these failures. To this end, we construct a synthetic diagnostic dataset in which multiple visible speakers each utter a single word. Analysis on this corpus reveals a consistent error pattern across three recent open-source AVLLMs: models attribute utterances by simply matching the order of spoken sentences with the left-to-right, top-to-bottom arrangement of visible faces, rather than relying on audio-visual cues such as lip synchronization. We term this behavior \emph{ordinal-matching bias}. We further show that this bias can be substantially mitigated through a simple remedy, Ordinal-Decoupled Fine-Tuning (OD-FT), in which models are fine-tuned on synthetic videos where spatial positions of speakers and speaking order are independently randomized. Despite using only 400 synthetic training videos, OD-FT not only suppresses ordinal-matching bias but also improves audio-visual understanding on real-world videos, yielding average gains of 8.27% for Qwen2.5-Omni and 2.57% for video-SALMONN2+ across three audio-visual benchmarks.
Figures & tables
Figure 1 : Position bias on curated two-speaker clips with video-SALMONN2+. Accuracy on “Who spoke first?” is substantially higher when the first speaker is on the left. This gap persists even for the same videos after horizontal flipping, indicating a strong bias.
Figure 2 : Illustration of ordinal-matching bias. SW: The model predicts that the 1st on-screen speaker (“panda”) produced the 1st utterance (“Egypt”), rather than its true utterance (“Mexico”). WS: The model predicts that the 2nd utterance (“Mexico”) was produced by the 2nd on-screen speaker (“tiger”), rather than its true source (“panda”).
Model
Metric
N=4
N=5
N=6
SW
WS
Chance
SW
WS
Chance
SW
WS
Chance
Qwen2.5-Omni [ 18 ] 7B
Accuracy ↑
26.9
25.9
25.0
23.7
24.1
20.0
18.9
19.3
16.7
OMσrow
70.4
75.0
25.0
54.6
62.2
20.0
48.9
53.7
16.7
OMσcol
57.0
53.1
25.0
43.6
45.4
20.0
47.8
47.2
16.7
video-SALMONN2+ [ 12 ] 7B
Accuracy ↑
34.0
27.8
25.0
27.7
25.0
20.0
27.0
24.7
16.7
OMσrow
67.4
53.0
25.0
60.6
49.4
20.0
49.5
46.3
16.7
Table 1 : Existence of ordinal-matching bias. For open-source AVLLMs, accuracy is low while OMσrow is high, i.e., models attribute the k -th utterance to the speaker at the k -th position in row-major order, confirming ordinal-matching bias.
Qwen2.5-Omni
video-SALMONN2+
Qwen3-Omni
Interv.
Metric
SW
WS
SW
WS
SW
WS
Baseline
OMσrow
70.4
75.0
67.4
53.0
68.4
59.4
Swap
OMσrow
28.0 ↓
32.9 ↓
39.5 ↓
36.4 ↓
34.6 ↓
32.4 ↓
OMσrow†
68.8 ↑
75.5 ↑
68.9 ↑
55.9 ↑
61.5 ↑
58.2 ↑
Rotate
OMσrow
24.8 ↓
15.0 ↓
28.0 ↓
19.8 ↓
29.9 ↓
28.1 ↓
OMσrow†
56.2 ↑
61.4 ↑
55.9 ↑
57.9 ↑
43.1 ↑
38.0 ↑
Table 2 : Causal evidence of ordinal-matching bias ( N=4 ). We intervene on the visual positional indices while preserving the original on-screen order (Swap, Rotate). Relative to the baseline, OMσrow , defined with respect to the original on-screen order σrow , drops, whereas OMσrow† , defined with respect to the injected coordinate order σrow† , stays high.
Synthetic diagnosis corpus
Real-world benchmarks
N=4
N=5
N=6
Social Omni
Daily Omni
AV Speaker
Accuracy ↑
OMσrow↓
Accuracy ↑
OMσrow↓
Accuracy ↑
OMσrow↓
Acc. ↑
Acc. ↑
Acc. ↑
Model
SW
WS
SW
WS
SW
WS
SW
WS
SW
WS
SW
WS
Qwen2.5-Omni
26.9
25.9
70.4
75.0
23.7
24.1
54.6
62.2
18.9
19.3
48.9
53.7
38.9
64.3
44.8
+ OC-FT
25.0
25.8
98.5
98.1
20.3
21.7
78.3
69.6
17.2
18.1
66.6
58.7
39.5
67.2
45.8
+ OD-FT
97.0
97.9
26.5
25.9
94.8
92.4
21.5
20.3
91.9
83.7
19.1
16.4
54.8
69.6
48.4
Table 3 : OD-FT suppresses ordinal-matching bias and improves real-world understanding. OD-FT improves accuracy on the synthetic diagnostic corpus while reducing ordinal-matching bias to near-chance levels, and consistently improves performance across all three real-world benchmarks. In contrast, OC-FT provides little improvement on the synthetic corpus, exacerbates ordinal-matching bias, and produces smaller gains on real-world benchmarks.
Figure 3 : Effect of the decoupled share. SocialOmni accuracy increases with the fraction of decoupled examples in the fine-tuning corpus, from 0% (OC-FT) to 100% (OD-FT).
Despite rapid progress in video-capable MLLMs, we find that their apparent audio understanding in videos is often vision-driven: models rely on visual cues to infer or hallucinate acoustic information, rather than verifying the audio stream. This issue appears across both state-of-the-art open-source omni models and leading closed-source models from providers such as Google and OpenAI. We characterize this failure mode as an audio-visual Clever Hans effect, in which models appear (falsely) audio-grounded, but actually exploit visual-acoustic correlations without verifying whether the audio and visual streams are truly aligned. To systematically study this behavior, we introduce Thud, an intervention-driven probing framework based on three counterfactual audio edits: Shift, which tests temporal synchronization; Mute, which tests sound existence; and Swap, which tests audio-visual consistency. Beyond diagnosis, we further study a two-stage alignment recipe: intervention-derived preference pairs teach audio verification, while event-level general video preferences regularize the model against over-specialization. Our best 10K-sample recipe improves average performance across the three intervention dimensions by 28 percentage points, while slightly improving performance on general video and audio-visual QA benchmarks.
Xiaofei Wen, Wenjie Jacky Mo, Xingyu Fu +6
dUniversity of California, Davis · pPrinceton University · wUniversity of Wisconsin–Madison +1
Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination. Existing benchmarks focus on environmental sounds (e.g., dog barking) to indicate event occurrence. In contrast, human speech carries fundamentally different, rich semantics and temporal structures, yet it remains unexplored whether current models can accurately align speech content with corresponding visual signals. In this work, we show that speech content can induce hallucinations in audio-visual LLMs. To systematically study this, we introduce SVHalluc, the first comprehensive benchmark for evaluating speech-vision hallucination in audio-visual LLMs. Our benchmark diagnoses speech-vision hallucinations from two critical and complementary aspects: semantic and temporal. Experimental results demonstrate that state-of-the-art open-source audio-visual LLMs struggle with aligning speech content with corresponding visual signals, with a near-random accuracy on multiple tasks. In contrast, Gemini 2.5 Pro significantly outperforms the open-source models. Our analysis suggests that their failures stem from limited ability in cross-modality understanding, despite strong performance in single-modality perception. Our work uncovers a new and fundamental limitation of current audio-visual LLMs and highlights the need for speech-grounded video comprehension. Project page: https://chenshuang-zhang.github.io/projects/svhalluc/.
Chenshuang Zhang, Kyeong Seon Kim, Chengxin Liu +1
Large audio-language models (LALMs) are often used in tasks that involve reasoning over ordered options. An open question is whether their predictions are influenced by the order of answer choices, which would indicate a form of position bias and undermine their reliability. In this paper, we identify and analyze this problem in LALMs. We demonstrate that no model is immune to this bias through extensive experiments on six LALMs across three widely used benchmarks and their spoken counterparts. Shuffling the order of answer options can cause performance fluctuations of up to 24% and even change model rankings, raising concerns about the reliability of current evaluation practices. We also study permutation-based strategies and show that they can mitigate bias in most cases. Our work represents the first systematic investigation of this issue in LALMs, and we hope it raises awareness and motivates further research in this direction.