Selective Listening: Mechanism-Guided Control of Audio Influence in Large Audio-Language Models
Organizations: College of Computer Science and Technology, National University of Defense Technology · State Key Laboratory of Complex & Critical Software Environment
Abstract
Large audio-language models (LALMs) exploit multimodal evidence, yet task-irrelevant audio can alter text-reasoning decisions when listening is unnecessary. Aggregate Accuracy can hide this paired drift because audio-induced repairs and damages may cancel. Paired drift analysis and targeted interventions identify architecture-specific, intervention-sensitive late audio pathways as actionable control points. We introduce ICAP-Gate, which applies mechanism-guided, task-conditioned control to each model's pathway. Across four LALMs, two reasoning benchmarks, and environmental-sound and natural-speech interference, ICAP-Gate has lower point estimates for Influence Rate and Answer Flip than ungated inference in all 16 full-split model--condition evaluations. Fixed suppression degrades automatic speech recognition (ASR) across all four models, whereas ICAP-Gate matches ungated ASR performance by preserving the pathway for explicit audio-demand instructions. ICAP-Gate has lower paired-drift point estimates than mitigation prompting in all four evaluated settings and provides competitive stabilization relative to eight-sample Self-Consistency while using one generation per query; in controlled ARC measurements, Self-Consistency incurs -- ungated latency. These results establish selective modality influence control as a design principle for robust multimodal reasoning.
Figures & tables
| Setting | Acc. C/U/I | IR U I | IR | Flip U I | Flip | Net C2W | Net Ans |
| Qwen2.5-Omni-7B | |||||||
| ARC-FSD | 84.56/83.53/ 84.90 | 8.19 6.14 | 2.05 ∗ | 9.04 7.00 | 2.04 ∗ | +20 | +24 |
| ARC-SIB | 84.56/82.85/ 85.07 | 9.39 7.17 | 2.22 ∗ | 11.09 8.28 | 2.81 ∗ | +26 | +33 |
| MMLU-FSD | 67.42/67.58/ 67.55 | 11.80 9.16 | 2.64 ∗ | 16.06 12.60 | 3.46 ∗ | +183 | +486 |
| MMLU-SIB | 67.42/66.00/ 67.68 | 13.65 11.19 | 2.46 ∗ | 18.89 15.58 | 3.31 ∗ | +291 | +465 |
| Qwen2.5-Omni-3B | |||||||
| Dataset | Method | Gen./q. | Mean (s) | Median (s) | Rel. | VRAM (GiB) | Tok./q. |
|---|---|---|---|---|---|---|---|
| ARC-FSD | Ungated | 1 | 6.8881 | 6.4671 | 1.000 | 17.03 | 85.8 |
| Prompting | 1 | 6.6079 | 5.8928 | 0.959 | 17.04 | 81.2 | |
| ICAP-Gate | 1 | 7.2840 | 6.6174 | 1.057 | 17.03 | 94.3 | |
| Self-Consistency | 8 | 63.4709 | 57.7551 | 9.215 | 17.03 | 819.4 | |
| ARC-SIB | Ungated | 1 | 8.8734 | 8.0748 | 1.000 | 16.85 | 88.8 |
| Prompting | 1 | 7.9270 | 7.1944 | 0.893 | 16.85 | 81.6 |
| Setting | Prompt–ICAP (1 vs. 1 gen.) | SCS–ICAP (8 vs. 1 gen.) | ||
|---|---|---|---|---|
| IR | Flip | IR | Flip | |
| ARC-FSD | ||||
| ARC-SIB | ||||
| MMLU-FSD | ||||
| MMLU-SIB | ||||
| Model | Pathway | Depth (block/total) | Full-split confirmation | |
|---|---|---|---|---|
| Qwen2.5-Omni-7B | audio_tower.layers.25 | 0.81 (26/32) | 0.25 | ARC/MMLU FSD/SIB |
| Qwen2.5-Omni-3B | audio_tower.layers.25 | 0.81 (26/32) | 0.25 | ARC/MMLU FSD/SIB |
| Phi-4-MM | audio_embed.encoder.encoders.23 | 1.00 (24/24) | 0.10 | ARC/MMLU FSD/SIB |
| Voxtral-Mini-3B | audio_tower.layers.30 | 0.97 (31/32) | 0.25 | ARC/MMLU FSD/SIB |
| Audit item | Result |
|---|---|
| MMLU-SIB pairs | 14,042 |
| Unique source utterances | 2,694 |
| Speaker/chapter coverage | 40/97 |
| Mean/max utterance reuse | 5.212/14 |
| Transcripts recovered | 14,042/14,042 |
| IDF-overlap candidates | 28 (0.20%) |
| Model | Ungated | Fixed | ICAP | Recovery |
|---|---|---|---|---|
| Qwen2.5-Omni-7B | 17.74 | 98.60 | 17.74 | 80.86 |
| Qwen2.5-Omni-3B | 60.82 | 184.58 | 60.82 | 123.76 |
| Phi-4-MM (5.6B) | 2.45 | 18.73 | 2.45 | 16.28 |
| Voxtral-Mini-3B | 16.72 | 61.05 | 16.72 | 44.33 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Reasoning dataset | Interference source | Examples |
|---|---|---|---|
| ARC-FSD | ARC-Challenge | FSD50K environmental sound | 1,172 |
| ARC-SIB | ARC-Challenge | LibriSpeech dev-clean speech | 1,172 |
| MMLU-FSD | MMLU | FSD50K environmental sound | 14,042 |
| MMLU-SIB | MMLU | LibriSpeech dev-clean speech | 14,042 |
| Model | Condition | Clean Acc. | Ungated Acc. | Ungated IR | Ungated Flip |
|---|---|---|---|---|---|
| Qwen2.5-Omni-7B | ARC-FSD | 84.56 | 83.53 | 8.19 | 9.04 |
| Qwen2.5-Omni-7B | ARC-SIB | 84.56 | 82.85 | 9.39 | 11.09 |
| Qwen2.5-Omni-7B | MMLU-FSD | 67.42 | 67.58 | 11.80 | 16.06 |
| Qwen2.5-Omni-7B | MMLU-SIB | 67.42 | 66.00 | 13.65 | 18.89 |
| Qwen2.5-Omni-3B | ARC-FSD | 77.39 | 77.56 | 11.77 | 13.82 |
| Qwen2.5-Omni-3B | ARC-SIB | 77.39 | 76.54 | 13.82 | 16.13 |
| Model | Condition | Subject | Index | Gold | Clean | Ungated | ICAP-Gate |
|---|---|---|---|---|---|---|---|
| Qwen2.5-Omni-7B | MMLU-FSD | abstract_algebra | 0 | B | B | A | B |
| Qwen2.5-Omni-7B | MMLU-SIB | anatomy | 100 | A | A | D | A |
| Qwen2.5-Omni-3B | MMLU-FSD | clinical_knowledge | 502 | C | C | A | C |
| Qwen2.5-Omni-3B | MMLU-SIB | astronomy | 243 | D | D | B | D |
| Qwen2.5-Omni-7B | MMLU-FSD | college_physics | 1373 | B | A | D | A |
| Qwen2.5-Omni-3B | MMLU-SIB | college_medicine | 1204 | A | A | B | A |
| Setting | Clean ties | Interference ties | Clean tie rate | Interference tie rate |
|---|---|---|---|---|
| ARC-FSD | 25 | 26 | 2.13% | 2.22% |
| ARC-SIB | 25 | 23 | 2.13% | 1.96% |
| MMLU-FSD ( ) | 57 | 59 | 5.70% | 5.90% |
| MMLU-SIB ( ) | 57 | 56 | 5.70% | 5.60% |
| Setting | Method | Sample | Clean Acc. | Interference Acc. | Acc. | IR | Answer Flip | Repair /Damage /Net |
| ARC-FSD | ICAP-Gate | Full | 84.56% | 84.90% | +0.34 | 6.14% | 7.00% | 38/34/+4 |
| ARC-FSD | Prompting | Full | 84.98% | 83.96% | -1.02 | 7.68% | 8.62% | 39/51/-12 |
| ARC-FSD | Self-Consistency | Full | 87.37% | 87.80% | +0.43 | 4.52% | 5.20% | 29/24/+5 |
| ARC-SIB | ICAP-Gate | Full | 84.56% | 85.07% | +0.51 | 7.17% | 8.28% | 45/39/+6 |
| ARC-SIB | Prompting | Full | 84.98% | 81.31% | -3.67 | 10.32% | 12.12% | 39/82/-43 |
| ARC-SIB | Self-Consistency | Full | 87.37% | 87.12% | -0.26 | 4.18% | 5.38% | 23/26/-3 |
| Site | C2W RMS | Control RMS | RMS ratio | C2W margin | Control margin |
|---|---|---|---|---|---|
| late25 | 1.5023 | 1.4833 | 1.0128 | 0.2098 | 3.7383 |
| late30 | 1.7011 | 1.6731 | 1.0168 | 0.2098 | 3.7383 |
| Condition | Acc. C/U/G | IR U G | Flip U G | Net C2W | Net Answer | Interpretation |
|---|---|---|---|---|---|---|
| ARC-FSD | 81.66/75.51/75.77 | 22.53 21.93 | 25.00 24.66 | +5 | +4 | Small drift reduction |
| ARC-SIB | 81.66/73.21/75.51 | 23.46 21.33 | 26.54 24.66 | +26 | +22 | Positive speech transfer |
| MMLU-FSD | 63.99/57.28/58.98 | 29.38 28.04 | 39.56 38.04 | +213 | +214 | Best full Phi setting |
| MMLU-SIB | 63.99/57.25/57.51 | 30.00 28.96 | 40.51 38.73 | +92 | +249 | Near-neutral accuracy |
| Dataset | Scale | Gated Acc. | Gated IR | Gated Flip | Net C2W | Net Overall / Net Answer |
|---|---|---|---|---|---|---|
| MMLU-FSD | 0.10 | 58.98 | 28.04 | 38.04 | +213 | +239 / +214 |
| MMLU-FSD | 0.25 | 58.21 | 28.66 | 38.83 | +116 | +131 / +102 |
| MMLU-SIB | 0.10 | 57.51 | 28.96 | 38.73 | +92 | +37 / +249 |
| MMLU-SIB | 0.25 | 56.38 | 29.43 | 39.74 | -21 | -122 / +108 |
| Model | Condition | Configuration | C2W repair/damage | Overall repair/damage | Answer restore/break |
|---|---|---|---|---|---|
| Phi-4-MM | MMLU-FSD | ||||
| Phi-4-MM | MMLU-FSD | ||||
| Phi-4-MM | MMLU-SIB | ||||
| Phi-4-MM | MMLU-SIB | ||||
| Voxtral-Mini-3B | ARC-FSD | ||||
| Voxtral-Mini-3B | ARC-SIB |
| Condition | Acc. C/U/G | IR U G | Flip U G | Net C2W | Net Answer | Interpretation |
|---|---|---|---|---|---|---|
| ARC-FSD | 75.77/73.38/71.76 | 21.33 20.90 | 26.02 25.85 | Accuracy trade-off | ||
| ARC-SIB | 75.77/69.88/73.38 | 25.17 23.04 | 30.20 27.22 | Positive speech transfer | ||
| MMLU-FSD | 57.79/58.00/58.21 | 27.44 26.26 | 37.53 36.00 | Positive stability transfer | ||
| MMLU-SIB | 57.79/56.62/58.76 | 28.41 25.88 | 39.52 36.28 | Strongest full transfer |
| Scope | TP | TN | FP | FN | Accuracy (95% CI) | Precision | Recall | FPR | FNR | High-scale rate |
|---|---|---|---|---|---|---|---|---|---|---|
| Instruction | 40 | 80 | 40 | 80 | 50.00% [43.72, 56.28] | 50.00% | 33.33% | 33.33% | 66.67% | 33.33% |
| Full query | 40 | 40 | 80 | 80 | 33.33% [27.67, 39.52] | 33.33% | 33.33% | 66.67% | 66.67% | 50.00% |
| Model | Selected path | Ungated WER | Fixed WER | ICAP-Gate WER | Fixed | ICAP | High-scale routing |
|---|---|---|---|---|---|---|---|
| Qwen2.5-Omni-7B | layers.25 | 17.74 | 98.60 | 17.74 | +80.86 | 0.00 | 400/400 |
| Qwen2.5-Omni-3B | layers.25 | 60.82 | 184.58 | 60.82 | +123.76 | 0.00 | 400/400 |
| Phi-4-MM | encoders.23 | 2.45 | 18.73 | 2.45 | +16.28 | 0.00 | 400/400 |
| Voxtral-Mini-3B | layers.30 | 16.72 | 61.05 | 16.72 | +44.33 | 0.00 | 400/400 |
| Statistic | ARC-SIB | MMLU-SIB |
|---|---|---|
| Source speakers | 40 | 40 |
| Source chapters | 97 | 97 |
| Source duration | 5.39 h | 5.39 h |
| Text–speech pairs | 1,172 | 14,042 |
| Unique speech utterances | 1,172 | 2,694 |
| Mean utterance reuse | 1.000 | 5.212 |
| Audit item | Count or rate |
| MMLU-SIB pairs | 14,042 |
| Pairs with recovered transcript | 14,042 |
| Missing transcripts | 0 |
| IDF candidates above threshold | 28 (0.20%) |
| Lexical false positives among candidates | 22 |
| Weakly topical among candidates | 6 |