cs.CLAug 27, 2026

Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict

Authors: Adarsh Sudheer, David Li, Omar Elbanna, Ishaan Kodarapu, Arjun Bahuguna, Vasu Sharma

Organizations: Independent Researcher.

Abstract

We study audio-visual conflict as a compositional generalization test for AV-LLMs: the model must combine synchronized but semantically incompatible audio and video evidence and decide whether the pair matches. On VideoLLaMA 2-7B-AV, three alignment configurations remain nearchance on the scored exact-string Yes/No subset of AVHBench, even though their output priors shift substantially. Similarly, off-the-shelf InternVideo2 experienced a 32.3% accuracy decrease specifically under cross-modal conflict, accompanied by a 17.3% instruction-following failure. We call this failure mode prior dominance: late-layer commitment to an internally preferred answer pattern that is weakly grounded in the conflicting inputs. To explain this behavior, we conduct a mechanistic interpretability analysis and find that commitment remains concentrated at 25.5 ±\pm 1 layers. We show that stronger temporal alignment changes answer bias, but do not improve compositional conflict resolution. Code and data to reproduce our mechanistic audit and behavioral evaluations are available at https://github.com/AdarshSudheer09/AVHBench-dmai.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought

    May 11, 2026Xuanchen Li, Yuheng Lu, Chenrui Cui +6Audio-Visual ReasoningLLM Reasoning Strategies

  2. Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs

    Jun 17, 2026Hyebin Cho, Suho Yoo, Jaehyuk Jang +2Large Audio Language ModelsMultimodal Understanding

  3. When Vision Speaks for Sound

    May 13, 2026Xiaofei Wen, Wenjie Jacky Mo, Xingyu Fu +6Audio-Visual ReasoningVideo Multimodal Large Language Models