Selective Listening: Mechanism-Guided Control of Audio Influence in Large Audio-Language Models
Authors: Yulin Sun, Kele Xu, Yong Dou
Organizations: College of Computer Science and Technology, National University of Defense Technology · State Key Laboratory of Complex & Critical Software Environment
Large audio-language models (LALMs) exploit multimodal evidence, yet task-irrelevant audio can alter text-reasoning decisions when listening is unnecessary. Aggregate Accuracy can hide this paired drift because audio-induced repairs and damages may cancel. Paired drift analysis and targeted interventions identify architecture-specific, intervention-sensitive late audio pathways as actionable control points. We introduce ICAP-Gate, which applies mechanism-guided, task-conditioned control to each model's pathway. Across four LALMs, two reasoning benchmarks, and environmental-sound and natural-speech interference, ICAP-Gate has lower point estimates for Influence Rate and Answer Flip than ungated inference in all 16 full-split model--condition evaluations. Fixed suppression degrades automatic speech recognition (ASR) across all four models, whereas ICAP-Gate matches ungated ASR performance by preserving the pathway for explicit audio-demand instructions. ICAP-Gate has lower paired-drift point estimates than mitigation prompting in all four evaluated settings and provides competitive stabilization relative to eight-sample Self-Consistency while using one generation per query; in controlled ARC measurements, Self-Consistency incurs 7.0--9.2× ungated latency. These results establish selective modality influence control as a design principle for robust multimodal reasoning.
Figures & tables
Figure 1: Paired decision drift under task-irrelevant audio. The same text input is evaluated with and without task-irrelevant environmental sound or speech; the downstream decision may change.
Figure 2: Mechanism-guided selective control of irrelevant-audio influence in LALMs. (a) Pathway scaling identifies a late audio pathway through which irrelevant audio changes Qwen2.5-Omni decisions. (b) The actionable control site is architecture-adapted across Qwen2.5-Omni, Phi-4-MM, and Voxtral-Mini-3B. (c) Instruction-conditioned routing keeps the pathway open for explicit audio-demand requests ( g=1 ) and attenuates it otherwise ( g=λm ), followed by one generation.
Figure 3: Activation magnitude and decision margin (Qwen2.5-Omni, FSD). (a) Activation RMS at two late audio sites for clean-correct-to-wrong (C2W) and clean-stable controls; labels show C2W/control ratios. (b) Shared clean gold-answer logit margins for groups. Filled blue/hollow gray markers denote C2W/control aggregates.
Setting
Acc. C/U/I
IR U → I
Δ IR ↑
Flip U → I
Δ Flip ↑
Net C2W
Net Ans
Qwen2.5-Omni-7B
ARC-FSD
84.56/83.53/ 84.90
8.19 → 6.14
2.05 ∗
9.04 → 7.00
2.04 ∗
+20
+24
ARC-SIB
84.56/82.85/ 85.07
9.39 → 7.17
2.22 ∗
11.09 → 8.28
2.81 ∗
+26
+33
MMLU-FSD
67.42/67.58/ 67.55
11.80 → 9.16
2.64 ∗
16.06 → 12.60
3.46 ∗
+183
+486
MMLU-SIB
67.42/66.00/ 67.68
13.65 → 11.19
2.46 ∗
18.89 → 15.58
3.31 ∗
+291
+465
Qwen2.5-Omni-3B
Table 1: Main results across four LALMs. Accuracy is reported as clean/ungated/ICAP (C/U/I); IR and Flip show ungated → ICAP. Δ IR and Δ Flip denote absolute reductions in percentage points. Net C2W and Net Ans report directional recovery toward the clean decision. Asterisks mark paired 95% CIs that exclude zero.
Dataset
Method
Gen./q.
Mean (s)
Median (s)
Rel.
VRAM (GiB)
Tok./q.
ARC-FSD
Ungated
1
6.8881
6.4671
1.000
17.03
85.8
Prompting
1
6.6079
5.8928
0.959
17.04
81.2
ICAP-Gate
1
7.2840
6.6174
1.057
17.03
94.3
Self-Consistency
8
63.4709
57.7551
9.215
17.03
819.4
ARC-SIB
Ungated
1
8.8734
8.0748
1.000
16.85
88.8
Prompting
1
7.9270
7.1944
0.893
16.85
81.6
Table 2: Measured test-time cost. Statistics aggregate 200 measured queries and three sequential repeats after 20 warm-up queries on one NVIDIA A800 80GB GPU (bfloat16; model loading excluded). VRAM is peak allocated memory. Green marks ICAP-Gate; red marks Self-Consistency.
Figure 4: Matched drift relative to ICAP-Gate. Circles and diamonds denote Prompting and Self-Consistency, respectively. Values are comparator minus ICAP-Gate in percentage points: positive values indicate lower drift under ICAP-Gate, whereas negative values indicate lower drift under the comparator. Error bars are paired 95% confidence intervals, with filled markers denoting intervals excluding zero.
Setting
Prompt–ICAP (1 vs. 1 gen.)
SCS–ICAP (8 vs. 1 gen.)
Δ IR
Δ Flip
Δ IR
Δ Flip
ARC-FSD
+1.54
+1.62
−1.62
−1.79
ARC-SIB
+3.16
+3.84
−2.99
−2.90
MMLU-FSD
+3.66
+4.72
−1.80
−1.10
MMLU-SIB
+8.87
+11.76
−0.10
−1.20
Table 3: Matched drift relative to ICAP-Gate (percentage points). Values are comparator minus ICAP-Gate: positive values favor ICAP-Gate, whereas negative values favor the comparator. Bold entries denote paired differences whose 95% confidence intervals exclude zero (20,000 paired bootstrap resamples; seed 0), irrespective of which method is favored.
Model
Pathway ℓm
Depth (block/total)
λm
Full-split confirmation
Qwen2.5-Omni-7B
audio_tower.layers.25
0.81 (26/32)
0.25
ARC/MMLU × FSD/SIB
Qwen2.5-Omni-3B
audio_tower.layers.25
0.81 (26/32)
0.25
ARC/MMLU × FSD/SIB
Phi-4-MM
audio_embed.encoder.encoders.23
1.00 (24/24)
0.10
ARC/MMLU × FSD/SIB
Voxtral-Mini-3B
audio_tower.layers.30
0.97 (31/32)
0.25
ARC/MMLU × FSD/SIB
Table 4: Architecture-adapted ICAP-Gate configurations. The table reports each selected pathway (ℓm) , normalized depth, attenuation scale (λm) , and full-split confirmation scope.
Audit item
Result
MMLU-SIB pairs
14,042
Unique source utterances
2,694
Speaker/chapter coverage
40/97
Mean/max utterance reuse
5.212/14
Transcripts recovered
14,042/14,042
IDF-overlap candidates
28 (0.20%)
Table 5: MMLU-SIB construction and transcript audit. All 14,042 transcripts were recovered; 28 IDF-overlap candidates were manually reviewed, and none was potentially relevant or answer-bearing. Details are in Appendix H .
Figure 5: Selective control preserves evaluated ASR while reducing reasoning drift across architectures. (a) Change in ASR WER relative to ungated inference on 400 LibriSpeech dev-clean utterances per model. Blue circles denote ICAP-Gate, red triangles fixed suppression, and the dashed line the ungated reference. (b) Full-split IR reduction (ungated minus ICAP-Gate); positive values indicate lower drift under ICAP-Gate. In Panel (b), filled and hollow circles indicate whether the paired 95% confidence interval excludes or overlaps zero, respectively. Phi-4-MM uses the canonical s=0.10 condition.
Model
Ungated
Fixed
ICAP
Recovery ↑
Qwen2.5-Omni-7B
17.74
98.60
17.74
80.86
Qwen2.5-Omni-3B
60.82
184.58
60.82
123.76
Phi-4-MM (5.6B)
2.45
18.73
2.45
16.28
Voxtral-Mini-3B
16.72
61.05
16.72
44.33
Table 6: Selectivity of ICAP-Gate. (a) ASR WER (%, ↓ ) and recovery from fixed attenuation. (b) Instruction-scoped routing on selected categories from the n=240 hard-case audit; FP denotes false positives.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Reasoning dataset
Interference source
Examples
ARC-FSD
ARC-Challenge
FSD50K environmental sound
1,172
ARC-SIB
ARC-Challenge
LibriSpeech dev-clean speech
1,172
MMLU-FSD
MMLU
FSD50K environmental sound
14,042
MMLU-SIB
MMLU
LibriSpeech dev-clean speech
14,042
Appendix
Table 7: Evaluation matrix used in the main full-split experiments. The table records the task, interference source, and number of paired examples.
Model
Condition
Clean Acc.
Ungated Acc.
Ungated IR
Ungated Flip
Qwen2.5-Omni-7B
ARC-FSD
84.56
83.53
8.19
9.04
Qwen2.5-Omni-7B
ARC-SIB
84.56
82.85
9.39
11.09
Qwen2.5-Omni-7B
MMLU-FSD
67.42
67.58
11.80
16.06
Qwen2.5-Omni-7B
MMLU-SIB
67.42
66.00
13.65
18.89
Qwen2.5-Omni-3B
ARC-FSD
77.39
77.56
11.77
13.82
Qwen2.5-Omni-3B
ARC-SIB
77.39
76.54
13.82
16.13
Appendix
Table 8: Clean and ungated baseline results for the sixteen full-split model–task combinations. Accuracy is reported as clean/ungated, while IR and Answer Flip are measured relative to clean inference. All values are percentages.
Model
Condition
Subject
Index
Gold
Clean
Ungated
ICAP-Gate
Qwen2.5-Omni-7B
MMLU-FSD
abstract_algebra
0
B
B
A
B
Qwen2.5-Omni-7B
MMLU-SIB
anatomy
100
A
A
D
A
Qwen2.5-Omni-3B
MMLU-FSD
clinical_knowledge
502
C
C
A
C
Qwen2.5-Omni-3B
MMLU-SIB
astronomy
243
D
D
B
D
Qwen2.5-Omni-7B
MMLU-FSD
college_physics
1373
B
A
D
A
Qwen2.5-Omni-3B
MMLU-SIB
college_medicine
1204
A
A
B
A
Appendix
Table 9: Representative Qwen transition examples under task-irrelevant audio. The examples illustrate paired transition categories used by the analysis and are not an additional statistical evaluation.
Setting
Clean ties
Interference ties
Clean tie rate
Interference tie rate
ARC-FSD
25
26
2.13%
2.22%
ARC-SIB
25
23
2.13%
1.96%
MMLU-FSD ( n=1,000 )
57
59
5.70%
5.90%
MMLU-SIB ( n=1,000 )
57
56
5.70%
5.60%
Appendix
Table 10: Self-Consistency tie audit for the matched mitigation comparisons. Tie rates are computed over the corresponding clean or interference records.
Setting
Method
Sample
Clean Acc.
Interference Acc.
Δ Acc.
IR
Answer Flip
Repair /Damage /Net
ARC-FSD
ICAP-Gate
Full
84.56%
84.90%
+0.34
6.14%
7.00%
38/34/+4
ARC-FSD
Prompting
Full
84.98%
83.96%
-1.02
7.68%
8.62%
39/51/-12
ARC-FSD
Self-Consistency
Full
87.37%
87.80%
+0.43
4.52%
5.20%
29/24/+5
ARC-SIB
ICAP-Gate
Full
84.56%
85.07%
+0.51
7.17%
8.28%
45/39/+6
ARC-SIB
Prompting
Full
84.98%
81.31%
-3.67
10.32%
12.12%
39/82/-43
ARC-SIB
Self-Consistency
Full
87.37%
87.12%
-0.26
4.18%
5.38%
23/26/-3
Appendix
Table 11: Matched mitigation comparison. Panel (a) reports method-relative clean/interference outcomes; Panel (b) reports comparator-minus-ICAP-Gate paired differences in IR and Answer Flip with 95% paired percentile-bootstrap confidence intervals from 20,000 resamples (seed 0). Positive differences in Panel (b) indicate lower drift under ICAP-Gate.
Site
C2W RMS
Control RMS
RMS ratio
C2W margin
Control margin
late25
1.5023
1.4833
1.0128
0.2098
3.7383
late30
1.7011
1.6731
1.0168
0.2098
3.7383
Appendix
Table 12: Qwen2.5-Omni FSD activation statistics for C2W and control cases. RMS is the activation root-mean-square at the indicated late audio-tower site; the margin is the clean gold-answer logit margin.
Condition
Acc. C/U/G
IR U → G
Flip U → G
Net C2W
Net Answer
Interpretation
ARC-FSD
81.66/75.51/75.77
22.53 → 21.93
25.00 → 24.66
+5
+4
Small drift reduction
ARC-SIB
81.66/73.21/75.51
23.46 → 21.33
26.54 → 24.66
+26
+22
Positive speech transfer
MMLU-FSD
63.99/57.28/58.98
29.38 → 28.04
39.56 → 38.04
+213
+214
Best full Phi setting
MMLU-SIB
63.99/57.25/57.51
30.00 → 28.96
40.51 → 38.73
+92
+249
Near-neutral accuracy
Appendix
Table 13: Phi-4-MM full-split external validation. Accuracy is clean/ungated/gated; IR and Answer Flip are ungated → gated. Net C2W and Net Answer are paired transition summaries.
Dataset
Scale
Gated Acc.
Gated IR
Gated Flip
Net C2W
Net Overall / Net Answer
MMLU-FSD
0.10
58.98
28.04
38.04
+213
+239 / +214
MMLU-FSD
0.25
58.21
28.66
38.83
+116
+131 / +102
MMLU-SIB
0.10
57.51
28.96
38.73
+92
+37 / +249
MMLU-SIB
0.25
56.38
29.43
39.74
-21
-122 / +108
Appendix
Table 14: Phi-4-MM full MMLU scale comparison. All rows use encoder layer 23 and 14,042 examples. Net Overall denotes repair minus damage over correctness transitions; positive values favor gating.
Model
Condition
Configuration
C2W repair/damage
Overall repair/damage
Answer restore/break
Phi-4-MM
MMLU-FSD
s=0.10
4.31×10−7
6.36×10−6
9.13×10−5
Phi-4-MM
MMLU-FSD
s=0.25
0.00321
0.00866
0.0481
Phi-4-MM
MMLU-SIB
s=0.10
0.0378
0.515
1.35×10−5
Phi-4-MM
MMLU-SIB
s=0.25
0.642
0.0235
0.0527
Voxtral-Mini-3B
ARC-FSD
s=0.25
0.646
0.250
0.951
Voxtral-Mini-3B
ARC-SIB
s=0.25
0.0296
0.0206
0.0498
Appendix
Table 15: Unified transition significance for Phi-4-MM and Voxtral-Mini-3B. Each value is a two-sided exact sign-test p -value over discordant paired cases within the listed model–condition.
Condition
Acc. C/U/G
IR U → G
Flip U → G
Net C2W
Net Answer
Interpretation
ARC-FSD
75.77/73.38/71.76
21.33 → 20.90
26.02 → 25.85
−7
+2
Accuracy trade-off
ARC-SIB
75.77/69.88/73.38
25.17 → 23.04
30.20 → 27.22
+33
+35
Positive speech transfer
MMLU-FSD
57.79/58.00/58.21
27.44 → 26.26
37.53 → 36.00
+98
+215
Positive stability transfer
MMLU-SIB
57.79/56.62/58.76
28.41 → 25.88
39.52 → 36.28
+328
+454
Strongest full transfer
Appendix
Table 16: Voxtral-Mini-3B full-split transfer results. Accuracy is clean/ungated/gated; IR and Answer Flip are ungated → gated. Net C2W and Net Answer are paired transition summaries.
Scope
TP
TN
FP
FN
Accuracy (95% CI)
Precision
Recall
FPR
FNR
High-scale rate
Instruction
40
80
40
80
50.00% [43.72, 56.28]
50.00%
33.33%
33.33%
66.67%
33.33%
Full query
40
40
80
80
33.33% [27.67, 39.52]
33.33%
33.33%
66.67%
66.67%
50.00%
Appendix
Table 17: Instruction-scoped routing audit over 240 manually curated cases. Panel (a) reports scope-level counts and nominal Wilson intervals; Panel (b) reports category-level TP/TN/FP/FN counts. The audit characterizes the deployed instruction-scoped rule and is not a general semantic audio-relevance benchmark.
Model
Selected path
Ungated WER
Fixed WER
ICAP-Gate WER
Fixed Δ
ICAP Δ
High-scale routing
Qwen2.5-Omni-7B
layers.25
17.74
98.60
17.74
+80.86
0.00
400/400
Qwen2.5-Omni-3B
layers.25
60.82
184.58
60.82
+123.76
0.00
400/400
Phi-4-MM
encoders.23
2.45
18.73
2.45
+16.28
0.00
400/400
Voxtral-Mini-3B
layers.30
16.72
61.05
16.72
+44.33
0.00
400/400
Appendix
Table 18: ASR safety and uncertainty on 400 LibriSpeech dev-clean utterances per model. Panel (a) reports point estimates for all four models; Panel (b) reports paired bootstrap intervals for the model-condition pairs with matched utterance-level summaries. WER differences are measured relative to ungated inference.
Statistic
ARC-SIB
MMLU-SIB
Source speakers
40
40
Source chapters
97
97
Source duration
5.39 h
5.39 h
Text–speech pairs
1,172
14,042
Unique speech utterances
1,172
2,694
Mean utterance reuse
1.000
5.212
Appendix
Table 19: SIB source coverage and pairing statistics. Source rows describe the common LibriSpeech pool; pairing rows describe the two reasoning conditions. Processed duration counts paired clips and therefore includes repeated use of an utterance in MMLU-SIB.
Audit item
Count or rate
MMLU-SIB pairs
14,042
Pairs with recovered transcript
14,042
Missing transcripts
0
IDF candidates above threshold
28 (0.20%)
Lexical false positives among candidates
22
Weakly topical among candidates
6
Appendix
Table 20: MMLU-SIB transcript-level validity audit. The audit checks the constructed speech–question pairings and is not part of the ICAP-Gate routing rule.
Imperial College London, UK · Technische Universität München München, Germany · Mohamed bin Zayed University of Artificial Intelligence Abu Dhabi, AE +2