cs.SDOct 8, 2026

How Much Audio Is Left In An Embedding? An Inversion Audit Of Audio Encoders

Authors: Marios Glytsos, Brian McFee

Organizations: Music and Audio Research Laboratory, New York University, NY, USA

Abstract

Pretrained audio encoders are reused for downstream tasks that are often unknown when the encoder is trained, so their usefulness depends partly on which signal properties survive the pretext objective. We study this retained information through paired source reconstruction. Using a shared Stable Audio Open latent diffusion decoder, we reconstruct five-second, 44.1-kHz stereo music from frozen representations produced by supervised classifiers (VGGish, ConvNeXt), an audio-text contrastive model (CLAP), and a waveform-reconstruction model (EnCodec). These objectives impose different pressures to preserve source detail, while their exposed interfaces vary substantially in temporal and spectral resolution. Evaluating on the Million Song Dataset (MSD), we find clear differences in reconstructability across encoder families, while within encoder comparisons show improved recovery when finer temporal or spectral structure is exposed. Even compressed task oriented embeddings support reconstructions that preserve measurable source specificity and high level musical content.

Figures & tables

Explore similar work

CardsList
  1. SAME: A Semantically-Aligned Music Autoencoder

    May 18, 2026Julian D. Parker, Zach Evans, CJ Carr +4Audio Representation LearningNeural Audio Codecs

  2. An Empirical Analysis of Task-Induced Encoder Bias in Fréchet Audio Distance

    Feb 27, 2026Wonwoo JeongText-to-Audio GenerationAudio Representation Learning