cs.SDOct 5, 2026

AudioGAR: Bridging Reconstruction and Generation in Latent Audio Generative Models

Authors: Xianghong Fang, Geeyang Tay, Wentao Ma, Tim G. J. Rudner, Dehan Kong

Organizations: University of Toronto · Vijil

Abstract

Latent audio generative models are typically trained in two stages: an audio codec is learned first, followed by a latent generative model. This decomposition leads to a decoder train-generation mismatch: the codec decoder is trained on encoder-induced latents but deployed on generator-produced latents at inference time. Across diverse datasets and latent generative models, we observe clear reconstruction-generation gaps under both FD and FAD, showing that strong reconstruction quality does not necessarily translate into strong end-to-end generation quality. A natural remedy is to adapt the decoder on generation-produced latents, but generated latents lack correspondence with source audio and therefore cannot directly provide the paired supervision used for decoder fine-tuning. We introduce \textbf{AudioGAR}, which constructs intermediate latents by perturbing encoder latents and denoising them through the frozen latent diffusion model. These latents form a trajectory from reconstruction toward generation, with lower-noise latents retaining source correspondence and supporting paired decoder fine-tuning. We fine-tune only the codec decoder on these latents, while keeping the codec encoder and latent generative model frozen. When applied to AudioX, AudioGAR substantially improves generative performance. It requires only 1.5% of the original training audio hours and 0.26% of the original training cost.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PoDAR: Power-Disentangled Audio Representation for Generative Modeling

    May 11, 2026Alejandro Luebs, Mithilesh Vaidya, Ishaan Kumar +5Generative Models

  2. Stable Audio 3

    May 18, 2026Zach Evans, Julian D. Parker, Matthew Rice +4Modern Generative Audio ModelsAudio Understanding

  3. F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation

    Jun 4, 2026Dinghao Zhou, Xingchen Song, Di Wu +3Audio TokenizersNeural Audio