Latent audio generative models are typically trained in two stages: an audio codec is learned first, followed by a latent generative model. This decomposition leads to a decoder train-generation mismatch: the codec decoder is trained on encoder-induced latents but deployed on generator-produced latents at inference time. Across diverse datasets and latent generative models, we observe clear reconstruction-generation gaps under both FD and FAD, showing that strong reconstruction quality does not necessarily translate into strong end-to-end generation quality. A natural remedy is to adapt the decoder on generation-produced latents, but generated latents lack correspondence with source audio and therefore cannot directly provide the paired supervision used for decoder fine-tuning. We introduce \textbf{AudioGAR}, which constructs intermediate latents by perturbing encoder latents and denoising them through the frozen latent diffusion model. These latents form a trajectory from reconstruction toward generation, with lower-noise latents retaining source correspondence and supporting paired decoder fine-tuning. We fine-tune only the codec decoder on these latents, while keeping the codec encoder and latent generative model frozen. When applied to AudioX, AudioGAR substantially improves generative performance. It requires only 1.5% of the original training audio hours and 0.26% of the original training cost.
Figures & tables
Methods
Training Cost
Training Data
AudioCaps
MusicCaps
Wall-Clock (Hours)
Rel. Cost
Audio Hours
Rel. Data
FD ↓
FAD ↓
FD ↓
FAD ↓
AudioX
4000
–
33052.3
–
11.83
1.61
9.56
1.60
AudioX + AudioGAR
10.3
0.26%
508.3
1.5%
11.62
1.34
8.33
1.10
Table 1 : AudioGAR improves AudioX with less training cost and data. Relative cost and data are computed w.r.t. the original AudioX training. Training cost is reported as single-H100-equivalent wall-clock hours.
Figure 1: LDA projections of ze and zg in AudioX. Rows correspond to AudioCaps, MusicCaps, and VGGSound-Omni datasets; columns show latent space and decoder features at Conv In, Block 2, and Block 4. Curves show Gaussian-smoothed, normalized histogram densities of the one-dimensional LDA scores.
Dataset
Models
rFD ( ↓ )
gFD ( ↓ )
rFAD ( ↓ )
gFAD ( ↓ )
AudioCaps
AudioLDM-2-Base [ Liu2024AudioLDM2L ]
3.19
11.54
1.31
2.04
AudioLDM-2-Large [ Liu2024AudioLDM2L ]
3.20
11.53
1.31
1.86
AudioLDM-S-Full [ Liu2023AudioLDMTG ]
3.05
21.85
1.27
4.79
AudioX [ Tian2026AudioXTurboAU ]
7.53
11.83
3.49
1.61
Stable Audio Open [ Evans2025StableAO ]
7.53
30.36
3.49
3.41
Tango [ Ghosal2023TexttoAudioGU ]
2.97
14.47
1.14
1.52
Table 2 : Reconstruction-generation performance gap across latent audio generative models. We report rFD and rFAD for reconstructed audio and gFD and gFAD for generated audio across multiple models on AudioCaps, MusicCaps, and VGGSound-Omni datasets. Reconstruction metrics are generally substantially lower than their generation counterparts. Lower values are better for all metrics.
Figure 2: Quantile comparison of rFD and gFD on AudioCaps. The two curves remain separated throughout.
Figure 3: Illustration of the AudioGAR process.
Figure 4 : AudioGAR bridges encoder-induced and generator-produced latent distributions. Latent-FD between the intermediate AudioGAR distribution Pgt and the two endpoints, Pe and Pg , across noise levels on (a) AudioCaps and (b) MusicCaps. As the noise level increases, Pgt moves farther from Pe and closer to Pg , tracing a trajectory from reconstruction toward generation.
Dataset
Method
Task
KL ( ↓ )
IS ( ↑ )
FD ( ↓ )
FAD ( ↓ )
PC ( ↑ )
PQ ( ↑ )
AudioCaps
AudioGen †
T2A
1.39
10.22
13.29
1.72
3.26
5.25
AudioLDM-L-Full †
T2A
2.00
6.51
37.27
8.37
2.82
5.67
AudioLDM-S-Full ‡
T2A
1.52
7.59
21.85
4.79
3.21
5.81
Tango 2 †
T2A
1.11
10.37
12.22
3.20
3.63
5.82
Tango (AF-AC, FT AudioCaps) ‡
T2A
1.24
10.66
12.08
2.55
3.44
6.13
MAGNET-large †
T2A
1.62
7.46
24.88
2.99
3.25
5.15
Table 3: Performance evaluation across tasks and datasets. T2A and T2M denote Text-to-Audio and Text-to-Music, respectively. † denotes results reported in AudioX [ Tian2026AudioXTurboAU ] , while ‡ denotes results from our experiments, including reproduced baselines and AudioGAR evaluations. The best result for each metric within each task is shown in bold.
Dataset
Decoder Adaptation
Decoder Latents
KL ( ↓ )
IS ( ↑ )
FD ( ↓ )
FAD ( ↓ )
PC ( ↑ )
PQ ( ↑ )
AudioCaps
✗
-
1.31
12.47
11.83
1.61
3.16
5.73
✔
ze
1.29
12.19
11.72
1.55
3.14
5.70
✔
zgt
1.29
12.06
11.62
1.34
3.11
5.73
MusicCaps
✗
-
1.00
3.65
9.56
1.60
4.78
6.61
✔
ze
0.99
3.62
8.87
1.38
4.76
6.50
✔
zgt
0.99
3.59
8.33
1.10
4.75
6.56
Table 4: Effect of latent choice during decoder adaptation. We compare the pretrained AudioX baseline with decoder adaptation using either encoder-induced latents ze or AudioGAR latents zgt on AudioCaps and MusicCaps. ✗and ✔indicate whether decoder adaptation is applied.
Figure 5 : Sensitivity of FD to the AudioGAR noise level. FD throughout decoder adaptation with AudioX for different noise levels ηt on (a) AudioCaps and (b) MusicCaps.
Figure 6 : Sensitivity of FAD to the AudioGAR noise level. FAD throughout decoder adaptation with AudioX for different noise levels ηt on (a) AudioCaps and (b) MusicCaps.
Figure 7 : Training dynamics of AudioGAR on AudioX. FD and FAD over 100k decoder-adaptation steps on AudioCaps and MusicCaps with ηt=0.2 .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: LDA projections of ze and zg in Stable Audio Open. Rows correspond to AudioCaps, MusicCaps, and VGGSound-Omni datasets; columns show latent space and decoder features at Conv In, Block 2, and Block 4. Curves show Gaussian-smoothed, normalized histogram densities of the one-dimensional LDA scores.
Figure 9: Quantile comparison of rFD and gFD on MusicCaps and VGGSound-Omni. The rFD and gFD curves are separated across the full quantile range on both datasets, with median gFD-rFD gaps of 18.88 on MusicCaps and 11.81 on VGGSound-Omni.
Dataset
AudioX
Ours
Clips
Hours
Clips
Share
Hours
AudioCaps
45k
125.1
67554
36.9%
187.4
WavCaps
108.3k
300.8
46866
25.6%
130.0
IF-caps
1268k
3,524.4
58324
32.9%
167.1
AudioTime
20k
355.5
9648
4.7%
23.8
Private T2M dataset
175.2k
11679.3
0
0.0%
0.0
Appendix
Table 5: Comparison of the original AudioX training data and the data used for our decoder adaptation. AudioX training statistics are reported in Appendix A.1 of [ Tian2026AudioXTurboAU ] . Our adaptation uses only four text-to-audio datasets and 508.3 hours of audio in total.
The performance of audio latent diffusion models is primarily governed by generator expressivity and the modelability of the underlying latent space. While recent research has focused primarily on the former, as well as improving the reconstruction fidelity of audio codecs, we demonstrate that latent modelability can be significantly improved through explicit factor disentanglement. We present PoDAR (Power-Disentangled Audio Representation), a framework that utilizes a randomized power augmentation and latent consistency objective to decouple signal power from invariant semantic content. This factorization makes the latent space easier to model, which both accelerates the convergence of downstream generative models and improves final overall performance. When applied to a Stable Audio 1.0 VAE with an F5-TTS generator, PoDAR achieves about a 2× acceleration in convergence to match baseline performance, while increasing final speaker similarity by 0.055 and UTMOS by 0.22 on the LibriSpeech-PC dataset. Furthermore, isolating power into dedicated channels enables the application of CFG exclusively to power-invariant content, effectively extending the stable guidance regime to higher scales.
Stable Audio 3 is a family of fast latent diffusion models (small, medium, large) for variable-length audio generation and editing. Since our models can generate several minutes of audio, variable-length generations are key to avoid the cost of producing full-length generations for short sounds. We also support inpainting, enabling targeted audio editing and the continuation of short recordings. Our latent diffusion models operate on top of a novel semantic-acoustic autoencoder that projects audio into a compact latent space, enabling efficient diffusion-based generation while preserving audio fidelity and encouraging semantic structure in the latent. Finally, we run adversarial post-training to both accelerate inference and improve generation quality, reducing the number of inference steps while improving fidelity and prompt adherence. Stable Audio 3 models are trained on licensed and Creative Commons data to generate music and sounds in less than a 2s on an H200 GPU and less than a few seconds on a MacBook Pro M4. We release the weights of small and medium, that can run on consumer-grade hardware, together with their training and inference pipeline.
Continuous audio autoencoders reconstruct waveforms well but often produce latents with weak structure for understanding, while self-supervised audio encoders capture semantics but are not directly decodable. This mismatch complicates a single audio tokenizer that must support both understanding and generation. We adapt continuous autoencoder latents to this setting with two components: a noise-regularized autoencoder bottleneck and a latent-side representation encoder. The bottleneck uses channel normalization and stochastic perturbation instead of KL-based variational training, yielding scale-controlled continuous latents for reconstruction and autoregressive generation. The representation encoder is trained on frozen autoencoder latents with RQ-MTP and frozen-LLM supervision. The resulting tokenizer provides high-dimensional representations for understanding while preserving normalized continuous latents as generation targets
Dinghao Zhou, Xingchen Song, Di Wu +3
2WeNet Open Source Community · ∗Equal contribution · 1Nanjing University, China