Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict
Organizations: Independent Researcher.
Abstract
We study audio-visual conflict as a compositional generalization test for AV-LLMs: the model must combine synchronized but semantically incompatible audio and video evidence and decide whether the pair matches. On VideoLLaMA 2-7B-AV, three alignment configurations remain nearchance on the scored exact-string Yes/No subset of AVHBench, even though their output priors shift substantially. Similarly, off-the-shelf InternVideo2 experienced a 32.3% accuracy decrease specifically under cross-modal conflict, accompanied by a 17.3% instruction-following failure. We call this failure mode prior dominance: late-layer commitment to an internally preferred answer pattern that is weakly grounded in the conflicting inputs. To explain this behavior, we conduct a mechanistic interpretability analysis and find that commitment remains concentrated at 25.5 1 layers. We show that stronger temporal alignment changes answer bias, but do not improve compositional conflict resolution. Code and data to reproduce our mechanistic audit and behavioral evaluations are available at https://github.com/AdarshSudheer09/AVHBench-dmai.
Figures & tables
| Configuration | Scored | Overall Acc | GT=Yes Acc | GT=No Acc | Yes/No Preds |
|---|---|---|---|---|---|
| ACTC | 5165 | 49.8% | 55.5% | 44.1% | 2895 / 2270 |
| ACTC + TATI | 5288 | 49.0% | 46.9% | 51.1% | 2529 / 2759 |
| ACTC + TATI + AMD | 5299 | 50.2% | 35.1% | 65.2% | 1853 / 3446 |
| InternVideo2 | 5302 | 52.3% | 51.3% | 53.3% | 2600 / 2702 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.