Latent audio generative models are typically trained in two stages: an audio codec is learned first, followed by a latent generative model. This decomposition leads to a decoder train-generation mismatch: the codec decoder is trained on encoder-induced latents but deployed on generator-produced latents at inference time. Across diverse datasets and latent generative models, we observe clear reconstruction-generation gaps under both FD and FAD, showing that strong reconstruction quality does not necessarily translate into strong end-to-end generation quality. A natural remedy is to adapt the decoder on generation-produced latents, but generated latents lack correspondence with source audio and therefore cannot directly provide the paired supervision used for decoder fine-tuning. We introduce \textbf{AudioGAR}, which constructs intermediate latents by perturbing encoder latents and denoising them through the frozen latent diffusion model. These latents form a trajectory from reconstruction toward generation, with lower-noise latents retaining source correspondence and supporting paired decoder fine-tuning. We fine-tune only the codec decoder on these latents, while keeping the codec encoder and latent generative model frozen. When applied to AudioX, AudioGAR substantially improves generative performance. It requires only 1.5% of the original training audio hours and 0.26% of the original training cost.
Figures & tables
Methods
Training Cost
Training Data
AudioCaps
MusicCaps
Wall-Clock (Hours)
Rel. Cost
Audio Hours
Rel. Data
FD ↓
FAD ↓
FD ↓
FAD ↓
AudioX
4000
–
33052.3
–
11.83
1.61
9.56
1.60
AudioX + AudioGAR
10.3
0.26%
508.3
1.5%
11.62
1.34
8.33
1.10
Table 1 : AudioGAR improves AudioX with less training cost and data. Relative cost and data are computed w.r.t. the original AudioX training. Training cost is reported as single-H100-equivalent wall-clock hours.
Figure 1: LDA projections of ze and zg in AudioX. Rows correspond to AudioCaps, MusicCaps, and VGGSound-Omni datasets; columns show latent space and decoder features at Conv In, Block 2, and Block 4. Curves show Gaussian-smoothed, normalized histogram densities of the one-dimensional LDA scores.
Dataset
Models
rFD ( ↓ )
gFD ( ↓ )
rFAD ( ↓ )
gFAD ( ↓ )
AudioCaps
AudioLDM-2-Base [ Liu2024AudioLDM2L ]
3.19
11.54
1.31
2.04
AudioLDM-2-Large [ Liu2024AudioLDM2L ]
3.20
11.53
1.31
1.86
AudioLDM-S-Full [ Liu2023AudioLDMTG ]
3.05
21.85
1.27
4.79
AudioX [ Tian2026AudioXTurboAU ]
7.53
11.83
3.49
1.61
Stable Audio Open [ Evans2025StableAO ]
7.53
30.36
3.49
3.41
Tango [ Ghosal2023TexttoAudioGU ]
2.97
14.47
1.14
1.52
Table 2 : Reconstruction-generation performance gap across latent audio generative models. We report rFD and rFAD for reconstructed audio and gFD and gFAD for generated audio across multiple models on AudioCaps, MusicCaps, and VGGSound-Omni datasets. Reconstruction metrics are generally substantially lower than their generation counterparts. Lower values are better for all metrics.
Figure 2: Quantile comparison of rFD and gFD on AudioCaps. The two curves remain separated throughout.
Figure 3: Illustration of the AudioGAR process.
Figure 4 : AudioGAR bridges encoder-induced and generator-produced latent distributions. Latent-FD between the intermediate AudioGAR distribution Pgt and the two endpoints, Pe and Pg , across noise levels on (a) AudioCaps and (b) MusicCaps. As the noise level increases, Pgt moves farther from Pe and closer to Pg , tracing a trajectory from reconstruction toward generation.
Dataset
Method
Task
KL ( ↓ )
IS ( ↑ )
FD ( ↓ )
FAD ( ↓ )
PC ( ↑ )
PQ ( ↑ )
AudioCaps
AudioGen †
T2A
1.39
10.22
13.29
1.72
3.26
5.25
AudioLDM-L-Full †
T2A
2.00
6.51
37.27
8.37
2.82
5.67
AudioLDM-S-Full ‡
T2A
1.52
7.59
21.85
4.79
3.21
5.81
Tango 2 †
T2A
1.11
10.37
12.22
3.20
3.63
5.82
Tango (AF-AC, FT AudioCaps) ‡
T2A
1.24
10.66
12.08
2.55
3.44
6.13
MAGNET-large †
T2A
1.62
7.46
24.88
2.99
3.25
5.15
Table 3: Performance evaluation across tasks and datasets. T2A and T2M denote Text-to-Audio and Text-to-Music, respectively. † denotes results reported in AudioX [ Tian2026AudioXTurboAU ] , while ‡ denotes results from our experiments, including reproduced baselines and AudioGAR evaluations. The best result for each metric within each task is shown in bold.
Dataset
Decoder Adaptation
Decoder Latents
KL ( ↓ )
IS ( ↑ )
FD ( ↓ )
FAD ( ↓ )
PC ( ↑ )
PQ ( ↑ )
AudioCaps
✗
-
1.31
12.47
11.83
1.61
3.16
5.73
✔
ze
1.29
12.19
11.72
1.55
3.14
5.70
✔
zgt
1.29
12.06
11.62
1.34
3.11
5.73
MusicCaps
✗
-
1.00
3.65
9.56
1.60
4.78
6.61
✔
ze
0.99
3.62
8.87
1.38
4.76
6.50
✔
zgt
0.99
3.59
8.33
1.10
4.75
6.56
Table 4: Effect of latent choice during decoder adaptation. We compare the pretrained AudioX baseline with decoder adaptation using either encoder-induced latents ze or AudioGAR latents zgt on AudioCaps and MusicCaps. ✗and ✔indicate whether decoder adaptation is applied.
Figure 5 : Sensitivity of FD to the AudioGAR noise level. FD throughout decoder adaptation with AudioX for different noise levels ηt on (a) AudioCaps and (b) MusicCaps.
Figure 6 : Sensitivity of FAD to the AudioGAR noise level. FAD throughout decoder adaptation with AudioX for different noise levels ηt on (a) AudioCaps and (b) MusicCaps.
Figure 7 : Training dynamics of AudioGAR on AudioX. FD and FAD over 100k decoder-adaptation steps on AudioCaps and MusicCaps with ηt=0.2 .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: LDA projections of ze and zg in Stable Audio Open. Rows correspond to AudioCaps, MusicCaps, and VGGSound-Omni datasets; columns show latent space and decoder features at Conv In, Block 2, and Block 4. Curves show Gaussian-smoothed, normalized histogram densities of the one-dimensional LDA scores.
Figure 9: Quantile comparison of rFD and gFD on MusicCaps and VGGSound-Omni. The rFD and gFD curves are separated across the full quantile range on both datasets, with median gFD-rFD gaps of 18.88 on MusicCaps and 11.81 on VGGSound-Omni.
Dataset
AudioX
Ours
Clips
Hours
Clips
Share
Hours
AudioCaps
45k
125.1
67554
36.9%
187.4
WavCaps
108.3k
300.8
46866
25.6%
130.0
IF-caps
1268k
3,524.4
58324
32.9%
167.1
AudioTime
20k
355.5
9648
4.7%
23.8
Private T2M dataset
175.2k
11679.3
0
0.0%
0.0
Appendix
Table 5: Comparison of the original AudioX training data and the data used for our decoder adaptation. AudioX training statistics are reported in Appendix A.1 of [ Tian2026AudioXTurboAU ] . Our adaptation uses only four text-to-audio datasets and 508.3 hours of audio in total.