Large audio-language models (LALMs) are sensitive to input perturbations, such as noise, waveform corruption, and adversarial injections. We propose AnchorPrompt, an efficient adaptation method that keeps the model frozen and learns a single block of prompt vectors inserted at the decoder input, between the audio and question embeddings. We train these vectors through self-distillation over diverse audio and text perturbations. To improve answer consistency and mitigate hallucination, we use the model's prediction on the clean recording as the target for answerable inputs, and assign a refusal target when the audio lacks sufficient evidence to answer. Furthermore, AnchorPrompt is perturbation-agnostic at inference, requiring no prior detection of perturbations and enabling zero-shot transfer to unseen distortions. We evaluate three LALMs across three benchmarks and show that AnchorPrompt improves answer consistency in most tested conditions. Clean accuracy improves in six of nine model-benchmark pairs, with minimal impact on the remainder of 1.2% at most. Crucially, AnchorPrompt reduces hallucinations under severe audio corruption while keeping false refusals on clean audio rare. Finally, these consistency gains transfer to unseen perturbations, such as choice permutations and reverberation.
Figures & tables
Figure 1: Overview of AnchorPrompt framework. A student LALM on perturbed input is distilled from a frozen teacher on clean input. When audio lacks sufficient clues to answer, the target is set to CANNOT DETERMINE to prevent hallucination. A soft prompt P ( L learned vectors), inserted between audio and question embeddings, is the only trainable parameter. At inference, the student runs independently with P .
Figure 2: Qwen2.5-Omni validation sweeps: clean-answer consistency on perturbed items assigned clean-answer targets (solid) and refusal under the two refusal-target conditions (dashed) against (a) prompt length with full data and (b) training items per dataset at L=8 . One seed per point.
Clean
Accuracy / clean-answer consistency
Model
Benchmark
Prompt
acc.
Noise
Mask
Text inj.
Audio inj.
Qwen2.5
SAKURA
Base
65.9
58.8/81.0
58.8/78.2
27.4/37.7
29.5/38.7
Manual
60.6
52.1/81.2
52.1/79.0
17.2/29.1
26.8/37.7
AnchorPrompt
75.1
68.6 / 84.5
68.5 / 83.1
73.2 / 87.6
68.2 / 76.3
MMAU
Base
73.4
67.8/84.2
65.6/80.4
31.1/43.1
–
Manual
68.1
60.3/84.3
56.6/79.8
21.2/31.2
–
Table 1: Test results (%): accuracy / clean-answer consistency (defined in Sec. 3.4 ). A refusal counts as incorrect; a question is consistent when the setting keeps its clean-input behavior, i.e. selects the same option or refuses on both inputs. Noise pools 10 and 0 dB SNR; Mask pools 40 and 60% masking; Audio inj. is adversarial audio injection (SAKURA only). Manual is the hand-written instruction at the soft-prompt position.
Answer rate ↓
Refusal ↓
Silence
−20 dB
Clean
Model
Bench.
B
M
AP
B
M
AP
B
M
AP
Qwen2.5
SAKURA
59.1
41.5
0.5
32.4
20.1
3.7
16.2
24.7
0.0
MMAU
73.3
47.8
0.0
59.4
38.9
8.3
3.0
15.5
0.0
MMAR
35.8
17.4
0.0
19.5
9.1
10.2
14.4
33.4
0.3
Qwen3
SAKURA
26.0
26.8
0.1
11.7
11.6
3.3
10.3
14.6
0.0
Table 2: Answer rates under 100% masking and −20 dB SNR noise (percentage of questions answered with an option), and over-refusal on clean audio (percentage of clean questions declined), for B = Base, M = Manual, and AP = AnchorPrompt .
Δ Accuracy vs. clean / consistency
Choice permutation
Reverberation
Model
Bench.
Base
AP
Base
AP
Qwen2.5
SAKURA
+0.2 /82.0
−1.2 /85.6
−8.5 /75.9
−9.0 /79.8
MMAU
+1.4 /81.6
−0.3 /81.9
−5.3 /82.8
−5.9 /78.6
MMAR
−2.3 /69.5
−1.4 /72.4
−7.4 /67.7
−4.1 /71.7
Qwen3
SAKURA
−1.1 /86.0
−0.6 /89.5
−4.3 /84.0
−3.4 /86.6
Table 3: Zero-shot transfer to OOD perturbations (test split). Each entry is the accuracy under the perturbation minus the clean accuracy of the same setting (points) / clean-answer consistency (%), where a refusal on both inputs counts as consistent. AP = AnchorPrompt .