cs.CVSep 28, 2026

Uncovering Ordinal-Matching Bias in Audio-Visual LLMs

Authors: Jihoo Jung, Youngjoon Jang, Hyebin Cho, Suho Yoo, Joon Son Chung

Organizations: Korea Advanced Institute of Science and Technology, South Korea · University of Oxford, United Kingdom

Abstract

This work aims to improve how audio-visual large language models (AVLLMs) associate speech with the correct visible speaker in multi-speaker scenes. We find that current AVLLMs frequently fail at this task, and analyze the nature of these failures. To this end, we construct a synthetic diagnostic dataset in which multiple visible speakers each utter a single word. Analysis on this corpus reveals a consistent error pattern across three recent open-source AVLLMs: models attribute utterances by simply matching the order of spoken sentences with the left-to-right, top-to-bottom arrangement of visible faces, rather than relying on audio-visual cues such as lip synchronization. We term this behavior \emph{ordinal-matching bias}. We further show that this bias can be substantially mitigated through a simple remedy, Ordinal-Decoupled Fine-Tuning (OD-FT), in which models are fine-tuned on synthetic videos where spatial positions of speakers and speaking order are independently randomized. Despite using only 400 synthetic training videos, OD-FT not only suppresses ordinal-matching bias but also improves audio-visual understanding on real-world videos, yielding average gains of 8.27% for Qwen2.5-Omni and 2.57% for video-SALMONN2+ across three audio-visual benchmarks.

Figures & tables

Explore similar work

CardsList
  1. When Vision Speaks for Sound

    May 13, 2026Xiaofei Wen, Wenjie Jacky Mo, Xingyu Fu +6Audio-Visual ReasoningVideo Multimodal Large Language Models