This work aims to improve how audio-visual large language models (AVLLMs) associate speech with the correct visible speaker in multi-speaker scenes. We find that current AVLLMs frequently fail at this task, and analyze the nature of these failures. To this end, we construct a synthetic diagnostic dataset in which multiple visible speakers each utter a single word. Analysis on this corpus reveals a consistent error pattern across three recent open-source AVLLMs: models attribute utterances by simply matching the order of spoken sentences with the left-to-right, top-to-bottom arrangement of visible faces, rather than relying on audio-visual cues such as lip synchronization. We term this behavior \emph{ordinal-matching bias}. We further show that this bias can be substantially mitigated through a simple remedy, Ordinal-Decoupled Fine-Tuning (OD-FT), in which models are fine-tuned on synthetic videos where spatial positions of speakers and speaking order are independently randomized. Despite using only 400 synthetic training videos, OD-FT not only suppresses ordinal-matching bias but also improves audio-visual understanding on real-world videos, yielding average gains of 8.27% for Qwen2.5-Omni and 2.57% for video-SALMONN2+ across three audio-visual benchmarks.
Figures & tables
Figure 1 : Position bias on curated two-speaker clips with video-SALMONN2+. Accuracy on “Who spoke first?” is substantially higher when the first speaker is on the left. This gap persists even for the same videos after horizontal flipping, indicating a strong bias.
Figure 2 : Illustration of ordinal-matching bias. SW: The model predicts that the 1st on-screen speaker (“panda”) produced the 1st utterance (“Egypt”), rather than its true utterance (“Mexico”). WS: The model predicts that the 2nd utterance (“Mexico”) was produced by the 2nd on-screen speaker (“tiger”), rather than its true source (“panda”).
Model
Metric
N=4
N=5
N=6
SW
WS
Chance
SW
WS
Chance
SW
WS
Chance
Qwen2.5-Omni [ 18 ] 7B
Accuracy ↑
26.9
25.9
25.0
23.7
24.1
20.0
18.9
19.3
16.7
OMσrow
70.4
75.0
25.0
54.6
62.2
20.0
48.9
53.7
16.7
OMσcol
57.0
53.1
25.0
43.6
45.4
20.0
47.8
47.2
16.7
video-SALMONN2+ [ 12 ] 7B
Accuracy ↑
34.0
27.8
25.0
27.7
25.0
20.0
27.0
24.7
16.7
OMσrow
67.4
53.0
25.0
60.6
49.4
20.0
49.5
46.3
16.7
Table 1 : Existence of ordinal-matching bias. For open-source AVLLMs, accuracy is low while OMσrow is high, i.e., models attribute the k -th utterance to the speaker at the k -th position in row-major order, confirming ordinal-matching bias.
Qwen2.5-Omni
video-SALMONN2+
Qwen3-Omni
Interv.
Metric
SW
WS
SW
WS
SW
WS
Baseline
OMσrow
70.4
75.0
67.4
53.0
68.4
59.4
Swap
OMσrow
28.0 ↓
32.9 ↓
39.5 ↓
36.4 ↓
34.6 ↓
32.4 ↓
OMσrow†
68.8 ↑
75.5 ↑
68.9 ↑
55.9 ↑
61.5 ↑
58.2 ↑
Rotate
OMσrow
24.8 ↓
15.0 ↓
28.0 ↓
19.8 ↓
29.9 ↓
28.1 ↓
OMσrow†
56.2 ↑
61.4 ↑
55.9 ↑
57.9 ↑
43.1 ↑
38.0 ↑
Table 2 : Causal evidence of ordinal-matching bias ( N=4 ). We intervene on the visual positional indices while preserving the original on-screen order (Swap, Rotate). Relative to the baseline, OMσrow , defined with respect to the original on-screen order σrow , drops, whereas OMσrow† , defined with respect to the injected coordinate order σrow† , stays high.
Synthetic diagnosis corpus
Real-world benchmarks
N=4
N=5
N=6
Social Omni
Daily Omni
AV Speaker
Accuracy ↑
OMσrow↓
Accuracy ↑
OMσrow↓
Accuracy ↑
OMσrow↓
Acc. ↑
Acc. ↑
Acc. ↑
Model
SW
WS
SW
WS
SW
WS
SW
WS
SW
WS
SW
WS
Qwen2.5-Omni
26.9
25.9
70.4
75.0
23.7
24.1
54.6
62.2
18.9
19.3
48.9
53.7
38.9
64.3
44.8
+ OC-FT
25.0
25.8
98.5
98.1
20.3
21.7
78.3
69.6
17.2
18.1
66.6
58.7
39.5
67.2
45.8
+ OD-FT
97.0
97.9
26.5
25.9
94.8
92.4
21.5
20.3
91.9
83.7
19.1
16.4
54.8
69.6
48.4
Table 3 : OD-FT suppresses ordinal-matching bias and improves real-world understanding. OD-FT improves accuracy on the synthetic diagnostic corpus while reducing ordinal-matching bias to near-chance levels, and consistently improves performance across all three real-world benchmarks. In contrast, OC-FT provides little improvement on the synthetic corpus, exacerbates ordinal-matching bias, and produces smaller gains on real-world benchmarks.
Figure 3 : Effect of the decoupled share. SocialOmni accuracy increases with the fraction of decoupled examples in the fine-tuning corpus, from 0% (OC-FT) to 100% (OD-FT).