Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by long-term memory in human listening, we propose Long-Term Memory-Guided Audio Enhancement (LTM-AE) to improve selective target perception by refining the audio representations of ALLMs without training. LTM-AE extracts representations in hidden states from separate clean reference recordings as long-term memory for each category, guiding enhancement toward a user-specified listening target. We reconstruct incoming audio tokens in the selected category long-term memory and interpolate the reconstructions with the original tokens before language backbone decoding. This interpolation controls the influence of stored auditory experience while keeping all ALLM parameters fixed. Diagnostic readouts across twenty sound categories and three ALLMs show that LTM-AE strengthens responses to a specified target amid three interfering sources. Averaged over constrained and free-form classification, accuracy gains over raw mixtures range from 29.53 to 46.15 percentage points across multiple open source models. For speech content recovery, LTM-AE with an additional learned token-level gate reduces Qwen2-Audio's word error rate from 23.07% to 14.77%. This work takes an initial step toward using principles of human long-term memory to enhance ALLMs for real-world listening. Our code is available at https://github.com/aynlp/ltm-audio-code
Figures & tables
Figure 1: Overview of LTM-AE. Clean audio hidden states form category long-term memory, while paired validation recordings calibrate the retained rank and interpolation weight. For an incoming mixture, the specified target category selects a memory to reconstruct audio tokens and interpolate them with the original tokens before language backbone decoding, without training the ALLM. The speech transcription branch adds a token-level gate to adjust this guidance at each position.
Figure 2: Diagnostic readouts under acoustic interference. Macro-average accuracy of Qwen2-Audio , Kimi-Audio , and Step-Audio-2-mini on clean audio, raw mixtures, and LTM-AE-enhanced mixtures, evaluated through three readouts: constrained classification, free-form classification, and prompt-free retrieval.
Model
Readout
Clean
Mixed
LTM-AE
Qwen2-Audio
C
66.90
9.60
47.20
F
32.05
4.80
26.25
R
87.90
8.25
86.75
Kimi-Audio
C
58.20
8.30
59.70
F
38.45
4.90
45.80
R
80.40
11.25
75.60
Table 1: Macro-average accuracy (%) over the twenty audio categories. C denotes constrained classification, F denotes free-form classification, and R denotes prompt-free retrieval.
Figure 3: MultiPerception class-wise accuracy gains. Each cell reports the best accuracy across the evaluated memory sizes minus the raw-mixture accuracy, in percentage points (Appendix A ). Columns correspond to Qwen2-Audio , Kimi-Audio , and Step-Audio-2-mini under constrained classification, free-form classification, and prompt-free retrieval.
Readout
Qwen2- Audio
Kimi- Audio
Step-Audio- 2-mini
Constrained classification
+37.60
+51.40
+31.80
Free-form classification
+21.45
+40.90
+35.30
Mean language- output gain
+29.53
+46.15
+33.55
Prompt-free retrieval
+78.50
+64.35
+68.05
Table 2: Accuracy gains in percentage points relative to the raw-mixture condition. The mean language-output gain averages constrained and free-form classification.
Figure 4: Validation-based selection of the token-level gate. WER on the 20-example selection split for λtok∈{0.0,0.1,…,1.0} .
Condition
WER (%)
Δ raw
Raw mixture
23.07
0.00
Ungated LTM-AE
33.75
+10.68
Selected gate
14.77
−8.30
Table 3: Speech transcription results. WER on the matched held-out test set. Lower is better.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
85
0
38
49
44
Dog vocalization
86
14
100
100
99
Human speech
94
64
100
100
100
Flute
43
1
5
9
5
Organ
61
1
38
50
35
Classical guitar
87
3
63
77
64
Appendix
Table 4: Class-wise constrained-classification accuracy (%) of Qwen2-Audio.
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
0
1
0
0
0
Dog vocalization
83
31
100
100
100
Human speech
3
0
0
0
0
Flute
2
0
0
0
0
Organ
38
0
0
0
0
Classical guitar
5
1
1
0
0
Appendix
Table 5: Class-wise free-form classification accuracy (%) of Qwen2-Audio.
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
99
5
100
100
100
Dog vocalization
75
0
44
20
6
Human speech
100
21
100
100
100
Flute
76
0
100
100
100
Organ
87
0
61
65
64
Classical guitar
91
23
100
100
100
Appendix
Table 6: Class-wise prompt-free retrieval accuracy (%) of Qwen2-Audio.
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
100
28
100
100
100
Dog vocalization
90
14
100
91
84
Human speech
84
10
29
29
26
Flute
2
0
0
0
0
Organ
13
4
5
10
11
Classical guitar
91
4
96
98
99
Appendix
Table 7: Class-wise constrained-classification accuracy (%) of Kimi-Audio.
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
46
0
90
92
85
Dog vocalization
76
17
97
73
62
Human speech
100
12
68
65
68
Flute
1
0
0
0
0
Organ
1
0
0
0
0
Classical guitar
62
13
97
95
96
Appendix
Table 8: Class-wise free-form classification accuracy (%) of Kimi-Audio.
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
96
3
100
100
99
Dog vocalization
63
0
71
26
32
Human speech
100
53
97
89
91
Flute
35
0
3
2
2
Organ
81
1
43
42
62
Classical guitar
88
25
86
83
78
Appendix
Table 9: Class-wise prompt-free retrieval accuracy (%) of Kimi-Audio.
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
100
0
29
58
69
Dog vocalization
90
23
96
97
92
Human speech
98
65
98
97
95
Flute
32
4
0
0
1
Organ
30
2
1
1
2
Classical guitar
92
7
48
46
46
Appendix
Table 10: Class-wise constrained-classification accuracy (%) of Step-Audio-2-mini.
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
78
2
54
49
29
Dog vocalization
80
26
87
92
89
Human speech
80
19
55
63
62
Flute
22
0
1
1
8
Organ
2
0
0
0
0
Classical guitar
66
10
44
37
30
Appendix
Table 11: Class-wise free-form classification accuracy (%) of Step-Audio-2-mini.
Class
Clean
Mixed
PCA ( N=20 )
PCA ( N=50 )
PCA ( N=100 )
Piano
99
2
100
100
100
Dog vocalization
74
0
1
0
0
Human speech
100
60
100
100
100
Flute
72
0
50
51
72
Organ
87
0
32
41
53
Classical guitar
89
32
61
100
100
Appendix
Table 12: Class-wise prompt-free retrieval accuracy (%) of Step-Audio-2-mini.
Label
Sound category
Label
Sound category
A
Piano music
K
Concert drums
B
Dog vocalization
L
Cat meow
C
Human speech
M
Cow moo
D
Flute music
N
Frog croak
E
Organ music
O
Crow call
F
Classical guitar music
P
Glass shatter
Appendix
Table 13: Candidate labels used for constrained classification.
Memory size
Randomized PCA
Full SVD
Difference (pp)
N=20
86.75
86.75
0.00
N=50
86.75
86.75
0.00
N=100
86.10
86.10
0.00
Appendix
Table 14: Solver cross-check for Qwen2-Audio prompt-free retrieval. Accuracies are macro averages over the twenty categories; the final column reports the absolute difference in percentage points.
Target category
Source collections
Piano
MAESTRO ( Hawthorne et al., 2018 )
Dog vocalization
Dog play/pant ( Cuaya et al., 2026 ) , FSD50K ( Fonseca et al., 2021 ) , and UrbanSound8K ( Salamon et al., 2014 )
Human speech
LibriSpeech ( Panayotov et al., 2015 )
Flute, organ
NSynth ( Engel et al., 2017 )
Classical guitar
GAPS ( Riley et al., 2024 )
Violin, trumpet, clarinet, alto saxophone, drums
RWC Instruments Database ( Goto et al., 2003 )
Appendix
Table 15: Source collections used to construct the twenty target categories. Some categories combine several collections.
Figure 5: MultiPerception results summarized by the highest macro accuracy over N∈{20,50,100} for each model and protocol. “Best PCA” denotes this maximum among the evaluated LTM-AE memory sizes. The retained rank and interpolation weight are calibrated on validation data for each size. The numbers above brackets give gains over the raw mixture in percentage points.
“A violin is playing a single note.” → violin music
“A flute is playing a long note.” → flute music
Appendix
Table 16: Representative language-output cases under acoustic interference. The arrows for free-form classification indicate the normalized category mapping.
Model
Original audio
Piano LTM-AE
Qwen2-Audio
0/100
0/100
Appendix
Table 17: Piano mentions in free-form responses to 100 recordings without piano sounds. Each entry gives the number of responses mentioning piano out of the 100 inputs.