Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models
Organizations: Harbin Institute of Technology, Shenzhen, China
Abstract
Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual, auditory, and linguistic information. However, recent studies show that AVLLMs face a critical challenge: , where cues from the unused modality induce responses that the required modality does not support, undermining reliability in real-world applications. Existing methods have made progress in mitigating this failure, yet how it arises from internal cross-modal interactions remains insufficiently understood. To address this gap, we conduct path-intervention and representation analyses, revealing a mechanism: question states carry interfering cues alongside required-source evidence, undermining grounding in required-modality evidence. Cutting pathways from interfering modality to question states yields greater correct-answer logit recovery than cutting those to the generation position. Motivated by these findings, we propose (ourc-onditioned lay seering), a training-free method that mitigates cross-modal interference at the question relay. Using contrasting question representations elicited through different modality-pathway interventions, SECRET steers the original question states toward required-source evidence. Experiments on two widely adopted benchmarks CMM and AVHBench across three AVLLMs show that SECRET consistently outperforms prior training-free methods, substantially mitigating source-confused grounding hallucinations (e.g., up to +18.0 and +7.1 percentage points over base models). Modality-specific captioning further demonstrates its generalizability to open-ended generation.
Figures & tables
| Method | CMM | AVHBench | ||||
| Visual Dom. | Audio Dom. | Overall Acc. | Video-Driven | Audio-Driven | Overall Acc. | |
| Audio Hall. | Video Hall. | |||||
| VideoLLaMA2-AV-7B | 71.8 | 80.0 | 75.9 | 75.7 | 79.0 | 77.4 |
| + VCD (CVPR’24) | 71.3 | 83.3 | 77.3 | 66.0 | 74.8 | 70.4 |
| + AVCD (NeurIPS’25) | 71.8 | 84.0 | 77.9 | 78.3 | 80.3 | 79.3 |
| + MAD (CVPR’26) | 82.3 | 84.3 | 83.3 | 79.7 | 79.1 | 79.4 |
| Method | Audio-target Caption | Video-target Caption | ||||
| T-CIDEr | D-CIDEr | LLM-score | T-CIDEr | D-CIDEr | LLM-score | |
| Qwen2.5-Omni-7B | 17.7 | 5.17 | 3.12 | 29.3 | 3.70 | 3.57 |
| +Gen-Interv. | 17.2 | 5.10 | 3.05 | 30.2 | 3.10 | 3.79 |
| +Removal-Interv. | 16.5 | 5.43 | 3.55 | 31.1 | 2.67 | 3.49 |
| + Secret | 17.4 | 4.59 | 3.76 | 31.8 | 2.66 | 4.01 |
| VideoLLaMA2-AV-7B | 13.9 | 20.2 | 2.93 | 32.9 | 5.43 | 3.27 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.