Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by long-term memory in human listening, we propose Long-Term Memory-Guided Audio Enhancement (LTM-AE) to improve selective target perception by refining the audio representations of ALLMs without training. LTM-AE extracts representations in hidden states from separate clean reference recordings as long-term memory for each category, guiding enhancement toward a user-specified listening target. We reconstruct incoming audio tokens in the selected category long-term memory and interpolate the reconstructions with the original tokens before language backbone decoding. This interpolation controls the influence of stored auditory experience while keeping all ALLM parameters fixed. Diagnostic readouts across twenty sound categories and three ALLMs show that LTM-AE strengthens responses to a specified target amid three interfering sources. Averaged over constrained and free-form classification, accuracy gains over raw mixtures range from 29.53 to 46.15 percentage points across multiple open source models. For speech content recovery, LTM-AE with an additional learned token-level gate reduces Qwen2-Audio's word error rate from 23.07% to 14.77%. This work takes an initial step toward using principles of human long-term memory to enhance ALLMs for real-world listening. Our code is available at https://github.com/aynlp/ltm-audio-code
Figures & tables
Figure 1: Overview of LTM-AE. Clean audio hidden states form category long-term memory, while paired validation recordings calibrate the retained rank and interpolation weight. For an incoming mixture, the specified target category selects a memory to reconstruct audio tokens and interpolate them with the original tokens before language backbone decoding, without training the ALLM. The speech transcription branch adds a token-level gate to adjust this guidance at each position.
Figure 2: Diagnostic readouts under acoustic interference. Macro-average accuracy of Qwen2-Audio , Kimi-Audio , and Step-Audio-2-mini on clean audio, raw mixtures, and LTM-AE-enhanced mixtures, evaluated through three readouts: constrained classification, free-form classification, and prompt-free retrieval.
Model
Readout
Clean
Mixed
LTM-AE
Qwen2-Audio
C
66.90
9.60
47.20
F
32.05
4.80
26.25
R
87.90
8.25
86.75
Kimi-Audio
C
58.20
8.30
59.70
F
38.45
4.90
45.80
R
80.40
11.25
75.60
Table 1: Macro-average accuracy (%) over the twenty audio categories. C denotes constrained classification, F denotes free-form classification, and R denotes prompt-free retrieval.
Figure 3: MultiPerception class-wise accuracy gains. Each cell reports the best accuracy across the evaluated memory sizes minus the raw-mixture accuracy, in percentage points (Appendix A ). Columns correspond to Qwen2-Audio , Kimi-Audio , and Step-Audio-2-mini under constrained classification, free-form classification, and prompt-free retrieval.
Readout
Qwen2- Audio
Kimi- Audio
Step-Audio- 2-mini
Constrained classification
+37.60
+51.40
+31.80
Free-form classification
+21.45
+40.90
+35.30
Mean language- output gain
+29.53
+46.15
+33.55
Prompt-free retrieval
+78.50
+64.35
+68.05
Table 2: Accuracy gains in percentage points relative to the raw-mixture condition. The mean language-output gain averages constrained and free-form classification.
Figure 4: Validation-based selection of the token-level gate. WER on the 20-example selection split for λtok∈{0.0,0.1,…,1.0} .
Condition
WER (%)
Δ raw
Raw mixture
23.07
0.00
Ungated LTM-AE
33.75
+10.68
Selected gate
14.77
−8.30
Table 3: Speech transcription results. WER on the matched held-out test set. Lower is better.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
85
0
38
49
44
Dog vocalization
86
14
100
100
99
Human speech
94
64
100
100
100
Flute
43
1
5
9
5
Organ
61
1
38
50
35
Classical guitar
87
3
63
77
64
Appendix
Table 4: Class-wise constrained-classification accuracy (%) of Qwen2-Audio.
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
0
1
0
0
0
Dog vocalization
83
31
100
100
100
Human speech
3
0
0
0
0
Flute
2
0
0
0
0
Organ
38
0
0
0
0
Classical guitar
5
1
1
0
0
Appendix
Table 5: Class-wise free-form classification accuracy (%) of Qwen2-Audio.
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
99
5
100
100
100
Dog vocalization
75
0
44
20
6
Human speech
100
21
100
100
100
Flute
76
0
100
100
100
Organ
87
0
61
65
64
Classical guitar
91
23
100
100
100
Appendix
Table 6: Class-wise prompt-free retrieval accuracy (%) of Qwen2-Audio.
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
100
28
100
100
100
Dog vocalization
90
14
100
91
84
Human speech
84
10
29
29
26
Flute
2
0
0
0
0
Organ
13
4
5
10
11
Classical guitar
91
4
96
98
99
Appendix
Table 7: Class-wise constrained-classification accuracy (%) of Kimi-Audio.
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
46
0
90
92
85
Dog vocalization
76
17
97
73
62
Human speech
100
12
68
65
68
Flute
1
0
0
0
0
Organ
1
0
0
0
0
Classical guitar
62
13
97
95
96
Appendix
Table 8: Class-wise free-form classification accuracy (%) of Kimi-Audio.
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
96
3
100
100
99
Dog vocalization
63
0
71
26
32
Human speech
100
53
97
89
91
Flute
35
0
3
2
2
Organ
81
1
43
42
62
Classical guitar
88
25
86
83
78
Appendix
Table 9: Class-wise prompt-free retrieval accuracy (%) of Kimi-Audio.
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
100
0
29
58
69
Dog vocalization
90
23
96
97
92
Human speech
98
65
98
97
95
Flute
32
4
0
0
1
Organ
30
2
1
1
2
Classical guitar
92
7
48
46
46
Appendix
Table 10: Class-wise constrained-classification accuracy (%) of Step-Audio-2-mini.
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
78
2
54
49
29
Dog vocalization
80
26
87
92
89
Human speech
80
19
55
63
62
Flute
22
0
1
1
8
Organ
2
0
0
0
0
Classical guitar
66
10
44
37
30
Appendix
Table 11: Class-wise free-form classification accuracy (%) of Step-Audio-2-mini.
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
99
2
100
100
100
Dog vocalization
74
0
1
0
0
Human speech
100
60
100
100
100
Flute
72
0
50
51
72
Organ
87
0
32
41
53
Classical guitar
89
32
61
100
100
Appendix
Table 12: Class-wise prompt-free retrieval accuracy (%) of Step-Audio-2-mini.
Label
Sound category
Label
Sound category
A
Piano music
K
Concert drums
B
Dog vocalization
L
Cat meow
C
Human speech
M
Cow moo
D
Flute music
N
Frog croak
E
Organ music
O
Crow call
F
Classical guitar music
P
Glass shatter
Appendix
Table 13: Candidate labels used for constrained classification.
Memory size
Randomized PCA
Full SVD
Difference (pp)
N=20
86.75
86.75
0.00
N=50
86.75
86.75
0.00
N=100
86.10
86.10
0.00
Appendix
Table 14: Solver cross-check for Qwen2-Audio prompt-free retrieval. Accuracies are macro averages over the twenty categories; the final column reports the absolute difference in percentage points.
Target category
Source collections
Piano
MAESTRO ( Hawthorne et al., 2018 )
Dog vocalization
Dog play/pant ( Cuaya et al., 2026 ) , FSD50K ( Fonseca et al., 2021 ) , and UrbanSound8K ( Salamon et al., 2014 )
Human speech
LibriSpeech ( Panayotov et al., 2015 )
Flute, organ
NSynth ( Engel et al., 2017 )
Classical guitar
GAPS ( Riley et al., 2024 )
Violin, trumpet, clarinet, alto saxophone, drums
RWC Instruments Database ( Goto et al., 2003 )
Appendix
Table 15: Source collections used to construct the twenty target categories. Some categories combine several collections.
Figure 5: MultiPerception results summarized by the highest macro accuracy over N∈{20,50,100} for each model and protocol. “Best PCA” denotes this maximum among the evaluated LTM-AE memory sizes. The retained rank and interpolation weight are calibrated on validation data for each size. The numbers above brackets give gains over the raw mixture in percentage points.
“A violin is playing a single note.” → violin music
“A flute is playing a long note.” → flute music
Appendix
Table 16: Representative language-output cases under acoustic interference. The arrows for free-form classification indicate the normalized category mapping.
Model
Original audio
Piano LTM-AE
Qwen2-Audio
0/100
0/100
Appendix
Table 17: Piano mentions in free-form responses to 100 recordings without piano sounds. Each entry gives the number of responses mentioning piano out of the 100 inputs.
Large audio language models (LALMs) are a class of foundation models for audio understanding. Existing LALMs tend to degrade significantly in real-world noisy acoustic conditions where speech and non-speech sounds interfere. While noise-aware fine-tuning can improve robustness, it requires task-specific noisy data and expensive retraining, limiting scalability. To address this issue, we propose Focus-Then-Listen (FTL), a plug-and-play audio enhancer that improves LALMs' noise robustness. Specifically, FTL first separates the input waveform into speech and non-speech, and a modality router is applied to predict the target audio modality (e.g., speech) based on the user's instruction. Finally, a modality-aware fusion block generates a task-adaptive enhanced signal for improved downstream perception and reasoning. Experiments across multiple LALMs and tasks show that FTL improves performance across different noise levels without fine-tuning on LALMs.
Han Yin, Yang Xiao, Younghoo Kwon +2
School of Electrical Engineering, KAIST, Daejeon, Republic of Korea · University of Melbourne, Australia
Audio Large Language Models (ALLMs) are highly vulnerable to real-world noise, which often induces severe semantic drift and hallucinations. Existing robustness methods primarily rely on waveform-level acoustic enhancement, answer-level supervision, or the internal suppression of noise representations. To address these issues, we propose echodistill, an alignment-based noisy-to-clean self-distillation framework. Echodistill leverages a frozen clean-audio teacher to provide semantic references for an inference-time noisy-audio student. Specifically, the student samples candidate responses under noisy conditions to expose its test-time behavior. These trajectories are then optimized via group-relative policy optimization (GRPO), where the token-level consistency with the teacher acts as a reward bonus. By aligning the noisy student's candidate responses with clean semantic evidence, and applying audio-aware reward shaping, our method encourages reasoning trajectories that are both correct and genuinely acoustically grounded. Echodistill significantly improves the semantic reliability and task performance of Audio LLMs under complex noise, without introducing any additional inference costs. Extensive experiments show that: (I) Compared with the strongest baseline, echodistill achieves average improvements of 4.18%↑ in GSR under strong noise. (II) Ablation results on Qwen-Omni further show that echodistill improves over the GRPO-only variant by 3.02%↑ in Acc, 3.89%↑ in Noisy, and 4.53%↑ in GSR on average. Our codes are available at https://anonymous.4open.science/r/echodistill-10DE.
Large audio-language models (LALMs) are sensitive to input perturbations, such as noise, waveform corruption, and adversarial injections. We propose AnchorPrompt, an efficient adaptation method that keeps the model frozen and learns a single block of prompt vectors inserted at the decoder input, between the audio and question embeddings. We train these vectors through self-distillation over diverse audio and text perturbations. To improve answer consistency and mitigate hallucination, we use the model's prediction on the clean recording as the target for answerable inputs, and assign a refusal target when the audio lacks sufficient evidence to answer. Furthermore, AnchorPrompt is perturbation-agnostic at inference, requiring no prior detection of perturbations and enabling zero-shot transfer to unseen distortions. We evaluate three LALMs across three benchmarks and show that AnchorPrompt improves answer consistency in most tested conditions. Clean accuracy improves in six of nine model-benchmark pairs, with minimal impact on the remainder of 1.2% at most. Crucially, AnchorPrompt reduces hallucinations under severe audio corruption while keeping false refusals on clean audio rare. Finally, these consistency gains transfer to unseen perturbations, such as choice permutations and reverberation.
Pooneh Mousavi, Amir Ivry, Mirco Ravanelli +1
Concordia University · Mila – Quebec AI Institute · Technion – IIT +1