Large audio-language models (LALMs) are sensitive to input perturbations, such as noise, waveform corruption, and adversarial injections. We propose AnchorPrompt, an efficient adaptation method that keeps the model frozen and learns a single block of prompt vectors inserted at the decoder input, between the audio and question embeddings. We train these vectors through self-distillation over diverse audio and text perturbations. To improve answer consistency and mitigate hallucination, we use the model's prediction on the clean recording as the target for answerable inputs, and assign a refusal target when the audio lacks sufficient evidence to answer. Furthermore, AnchorPrompt is perturbation-agnostic at inference, requiring no prior detection of perturbations and enabling zero-shot transfer to unseen distortions. We evaluate three LALMs across three benchmarks and show that AnchorPrompt improves answer consistency in most tested conditions. Clean accuracy improves in six of nine model-benchmark pairs, with minimal impact on the remainder of 1.2% at most. Crucially, AnchorPrompt reduces hallucinations under severe audio corruption while keeping false refusals on clean audio rare. Finally, these consistency gains transfer to unseen perturbations, such as choice permutations and reverberation.
Figures & tables
Figure 1: Overview of AnchorPrompt framework. A student LALM on perturbed input is distilled from a frozen teacher on clean input. When audio lacks sufficient clues to answer, the target is set to CANNOT DETERMINE to prevent hallucination. A soft prompt P ( L learned vectors), inserted between audio and question embeddings, is the only trainable parameter. At inference, the student runs independently with P .
Figure 2: Qwen2.5-Omni validation sweeps: clean-answer consistency on perturbed items assigned clean-answer targets (solid) and refusal under the two refusal-target conditions (dashed) against (a) prompt length with full data and (b) training items per dataset at L=8 . One seed per point.
Clean
Accuracy / clean-answer consistency
Model
Benchmark
Prompt
acc.
Noise
Mask
Text inj.
Audio inj.
Qwen2.5
SAKURA
Base
65.9
58.8/81.0
58.8/78.2
27.4/37.7
29.5/38.7
Manual
60.6
52.1/81.2
52.1/79.0
17.2/29.1
26.8/37.7
AnchorPrompt
75.1
68.6 / 84.5
68.5 / 83.1
73.2 / 87.6
68.2 / 76.3
MMAU
Base
73.4
67.8/84.2
65.6/80.4
31.1/43.1
–
Manual
68.1
60.3/84.3
56.6/79.8
21.2/31.2
–
Table 1: Test results (%): accuracy / clean-answer consistency (defined in Sec. 3.4 ). A refusal counts as incorrect; a question is consistent when the setting keeps its clean-input behavior, i.e. selects the same option or refuses on both inputs. Noise pools 10 and 0 dB SNR; Mask pools 40 and 60% masking; Audio inj. is adversarial audio injection (SAKURA only). Manual is the hand-written instruction at the soft-prompt position.
Answer rate ↓
Refusal ↓
Silence
−20 dB
Clean
Model
Bench.
B
M
AP
B
M
AP
B
M
AP
Qwen2.5
SAKURA
59.1
41.5
0.5
32.4
20.1
3.7
16.2
24.7
0.0
MMAU
73.3
47.8
0.0
59.4
38.9
8.3
3.0
15.5
0.0
MMAR
35.8
17.4
0.0
19.5
9.1
10.2
14.4
33.4
0.3
Qwen3
SAKURA
26.0
26.8
0.1
11.7
11.6
3.3
10.3
14.6
0.0
Table 2: Answer rates under 100% masking and −20 dB SNR noise (percentage of questions answered with an option), and over-refusal on clean audio (percentage of clean questions declined), for B = Base, M = Manual, and AP = AnchorPrompt .
Δ Accuracy vs. clean / consistency
Choice permutation
Reverberation
Model
Bench.
Base
AP
Base
AP
Qwen2.5
SAKURA
+0.2 /82.0
−1.2 /85.6
−8.5 /75.9
−9.0 /79.8
MMAU
+1.4 /81.6
−0.3 /81.9
−5.3 /82.8
−5.9 /78.6
MMAR
−2.3 /69.5
−1.4 /72.4
−7.4 /67.7
−4.1 /71.7
Qwen3
SAKURA
−1.1 /86.0
−0.6 /89.5
−4.3 /84.0
−3.4 /86.6
Table 3: Zero-shot transfer to OOD perturbations (test split). Each entry is the accuracy under the perturbation minus the clean accuracy of the same setting (points) / clean-answer consistency (%), where a refusal on both inputs counts as consistent. AP = AnchorPrompt .
Audio Large Language Models (ALLMs) are highly vulnerable to real-world noise, which often induces severe semantic drift and hallucinations. Existing robustness methods primarily rely on waveform-level acoustic enhancement, answer-level supervision, or the internal suppression of noise representations. To address these issues, we propose echodistill, an alignment-based noisy-to-clean self-distillation framework. Echodistill leverages a frozen clean-audio teacher to provide semantic references for an inference-time noisy-audio student. Specifically, the student samples candidate responses under noisy conditions to expose its test-time behavior. These trajectories are then optimized via group-relative policy optimization (GRPO), where the token-level consistency with the teacher acts as a reward bonus. By aligning the noisy student's candidate responses with clean semantic evidence, and applying audio-aware reward shaping, our method encourages reasoning trajectories that are both correct and genuinely acoustically grounded. Echodistill significantly improves the semantic reliability and task performance of Audio LLMs under complex noise, without introducing any additional inference costs. Extensive experiments show that: (I) Compared with the strongest baseline, echodistill achieves average improvements of 4.18%↑ in GSR under strong noise. (II) Ablation results on Qwen-Omni further show that echodistill improves over the GRPO-only variant by 3.02%↑ in Acc, 3.89%↑ in Noisy, and 4.53%↑ in GSR on average. Our codes are available at https://anonymous.4open.science/r/echodistill-10DE.
Large audio-language models (LALMs) frequently hallucinate by overriding acoustic evidence with language priors. While contrastive decoding (CD) offers training-free mitigation, existing methods rely on blunt perturbations like masking or noise, leaving structured audio transformations unexplored. We explore this design space by evaluating a diverse library of targeted audio perturbations and adaptively selecting the optimal negative branch for each task and example. First, we improve upon earlier prompt engineering by showing that a simple binary yes/no constraint reduces the model's tendency to falsely confirm absent audio features. Second, evaluating our library across temporal, spectral, frequency, and amplitude domains reveals that optimal transformations are highly task-dependent; for instance, reversing the audio array disrupts temporal coherence, raising accuracy on the temporal order task from 74.7% to 81.4%. Finally, we trained a light-weight perturbation selector on model hidden states to dynamically route negative branches, yielding an additional +4.3% gain on the existence task.
Aaron Isidore Grace, Zhouyuan Huo, Weiran Wang
Department of Computer Science University of Iowa Iowa City, IA, USA · Google Mountain View, CA, USA
Large Audio-Language Models (LALMs) excel in Audio QA but often suffer from hallucinations ungrounded in the audio. To our knowledge, we are the first to propose applying vector steering to the audio domain to mitigate this. Unlike text-based steering, our silence-anchored contrastive approach steers the model away from hallucinations by contrasting active audio against a silent baseline. Probing internal states reveals a strong correlation between specific layer representations and output correctness. Leveraging this, we introduce Layer-Weighted Vector Steering (LWVS), a training-free intervention that increases steering strength at influential layers. On the Audio Hallucination QA dataset, LWVS significantly outperforms baselines, boosting Recall on the Gemma model by 15.6% (53.4% to 69.0%). Crucially, MMAU benchmark tests confirm LWVS preserves and even enhances general audio understanding, achieving an 8% relative accuracy increase on the Qwen model (54.8% to 59.2%).
Tsung-En Lin, Kuan-Yi Lee, Hung-Yi Lee
National Taiwan University, Taipei, Taiwan · ASUS Open Cloud Infrastructure Software Center, Taipei, Taiwan