Pretrained audio encoders are reused for downstream tasks that are often unknown when the encoder is trained, so their usefulness depends partly on which signal properties survive the pretext objective. We study this retained information through paired source reconstruction. Using a shared Stable Audio Open latent diffusion decoder, we reconstruct five-second, 44.1-kHz stereo music from frozen representations produced by supervised classifiers (VGGish, ConvNeXt), an audio-text contrastive model (CLAP), and a waveform-reconstruction model (EnCodec). These objectives impose different pressures to preserve source detail, while their exposed interfaces vary substantially in temporal and spectral resolution. Evaluating on the Million Song Dataset (MSD), we find clear differences in reconstructability across encoder families, while within encoder comparisons show improved recovery when finer temporal or spectral structure is exposed. Even compressed task oriented embeddings support reconstructions that preserve measurable source specificity and high level musical content.
Figures & tables
Condition
Temporal granularity
Input shape
VGGish
0.96 s
128×(6–7)
CLAP 5s
5 s
512×1
CLAP 1s
1 s
512×5
ConvNeXt pooled
≈0.32 s
768×(16–17)
ConvNeXt flattened
≈0.32 s
5376×(16–17)
EnCodec zq
6.7 ms
128×750
Table 1: Frozen interfaces exposed for one 5 second target. Shapes are feature dimension × number of conditioning tokens.
Figure 1: Qualitative reconstruction of an MTG-Jamendo track using the standard interfaces.
Encoder
STFT ↓
VAE ↑
R@1 ↑
R@10 ↑
Tag ↑
VGGish
.21 ± .08
.15 ± .07
.004
.021
.733 ± .144
CLAP 5s
.20 ± .08
.16 ± .07
.006
.034
.751 ± .134
CLAP 1s
.19 ± .08
.18 ± .07
.024
.091
.771 ± .128
ConvNeXt pool.
.18 ± .08
.19 ± .07
.094
.201
.799 ± .121
ConvNeXt flat.
.18 ± .07
.24 ± .07
.352
.539
.812 ± .116
EnCodec
.13 ± .06
.47 ± .16
.894
.913
.847 ± .107
Table 2: Reconstruction on our five-second MSD test crops. Values are mean ± standard deviation where applicable. R@1/R@10 denote paired-source retrieval. Tag is normalized MusicNN top-10 tag overlap. EnCodec native uses its original decoder.
Figure 2: Distribution of paired source–reconstruction SAO-VAE cosine similarity across encoder conditions, compared with 200k randomly sampled distinct source–source pairs.
Latent representations are at the heart of the majority of modern generative models. In the audio domain they are typically produced by a neural-audio-codec autoencoder. In this work we introduce SAME (Semantically-Aligned Music autoEncoder), an autoencoder for stereo music and general audio that reaches a 4096× temporal compression ratio while maintaining reconstruction quality and downstream generative performance. We achieve this by combining a tranformer-based backbone with set of semantic regularisation approaches, phase-aware reconstruction losses and improved discriminator designs. The architecture delivers substantial computational cost benefits, through both its high compression ratio and its reliance on well-optimised transformer primitives. Two variants (a large SAME-L and a CPU-deployable SAME-S) are released in open-weights form.
End-to-end neural audio models achieve high-fidelity compression and generation. We might read that performance as evidence they directly represent interpretable features such as pitch and timbre, but a model can produce plausible outputs without doing so. A model may encode these features in any reachable basis, but regardless of which, the features are well described as compositions of time-frequency-localized primitives. Whether state-of-the-art encoders preserve access to these primitives, and thus to compositions of them, remains unclear. Through theoretical analysis and controlled experiments, we show that several state-of-the-art strided convolutional encoders impose two structural bottlenecks, both predictable from architecture and signal structure, on access to these primitives: (1) they collapse primitives into alias equivalence classes, establishing a bound on representational capacity, and (2) they limit the frequency resolution available to learned filters, restricting separability. For well structured data, we find collapse rates of 31-35% and filter bandwidths 10-35x above the theoretical resolution bound, confirming that both bottlenecks arise under realistic signal conditions. We then introduce Gabor Latent Refactorization (GLRF), a lightweight post-hoc intervention that re-expresses encoder latents in a frequency-localized basis, reducing filter bandwidths from 10-35x to 1.5-3x of the theoretical resolution bound while preserving reconstruction fidelity and improving control over attributes like pitch. These results show that the encoders in question predictably degrade access to frequency-localized primitives, entangling the features that depend on them, and that a lightweight, retraining-free intervention can recover much of that access, improving steerability and interpretability.
Fréchet Audio Distance (FAD) is the de facto standard for evaluating text-to-audio generation, yet its scores depend on the underlying encoder's embedding space. An encoder's training task dictates which acoustic features are preserved or discarded, causing FAD to inherit systematic task-induced biases. We decompose evaluation into Recall, Precision, and Alignment (split into semantic and structural dimensions), using log-scale normalization for fair cross-encoder comparison. Controlled experiments on six encoders across two datasets reveal a four-axis trade-off: reconstruction-based AudioMAE leads precision sensitivity; ASR-trained Whisper dominates structural detection but is blind to signal degradation; classification-trained VGGish maximizes semantic detection but penalizes legitimate intra-class variation. Since no single encoder is a universal evaluator, future metrics must shift toward evaluation-native encoders intrinsically aligned with human perception.
Wonwoo Jeong
Dept. of Computer Science and Engineering, Sogang University, South Korea