Neural audio synthesis models like the Realtime Audio Variational autoEncoder (RAVE) achieve impressive genera tion quality, yet how their internal representations encode musical features remains poorly understood. We present a systematic layer-wise and cross-layer cluster analysis of RAVE decoder activations across three models trained on different musical domains, tested with four stimulus types. We then evaluate architectural generalization with a general purpose EnCodec model. For RAVE, we find that synthetic stimuli are encoded well across models and audio features (pitch |\r{ho}|=0.45, 5.1x the null, BPM |\r{ho}| = 0.76, 8.6x the null). These results are reduced but still substantively apparent when using natural audio (mean across features |\r{ho}|=0.25, 2.8x the null). Natural audio sees a stronger encoding when nonlinear probes are used (mean across features R2=0.56, 18x the null, +0.152 nonlinear gain over the linear probe R2). Encoding strength varies throughout the layers of the decoder and an increased ability to joint-encode in the middle layers is seen across all audio features (\b{eta}2 all negative, p < 0.05). The general purpose EnCodec decoder also sees similar strong synthetic responses across audio features, similar nonlinear gains for natural audio joint encoding and similar depth profiles. We find the best cross-layer cluster improves the strength (r = 0.65, p = 0.006) and prevalence (r = 0.75, p = 0.001) of BPM encoding when compared against the best whole layers within the same section, with no effect for joint encoding. These findings advance the interpretability of neural audio models and inform targeted control strategies for neural synthesis.
Figures & tables
Fig. 1: Diagram of RAVE Decoder architecture
Pitch (Hz)
BPM
Spectral Centroid (Hz)
Spectral Bandwidth (Hz)
Dataset
Min
Max
Mean ( σ )
Min
Max
Mean ( σ )
Min
Max
Mean ( σ )
Min
Max
Mean ( σ )
Strings
80
788
293 (200)
70
140
104 (21)
306
7357
1976 (1123)
969
6082
2601 (813)
Drums
—
—
—
60
175
115 (34)
652
7902
3961 (1711)
948
6074
3553 (1105)
Vocals
80
761
284 (126)
65
123
90 (17)
918
7905
3938 (995)
1803
6608
3799 (779)
Synthetic
80
800
313 (200)
60
180
120 (35)
—
—
—
—
—
—
TABLE I: Acoustic feature distributions per dataset after per-feature balanced sampling. All cells use N=500 samples, stratified to span the full feature range. Exclusions for inappropriate feature / data combinations apply
Fig. 2: Diagram representing the three metrics used for comparing the relationship of acoustic feature values and neuron activations. ∣ρ∣ can be further averaged across layers or models. Null p95 derived from permutation tests. Predictive model for R2 can be linear or MLP.
Per-neuron
Joint probe R2
Feature
Condition
n
Mean ∣ρ∣
× null
% > p95
Linear
Nonlinear
× null
NL gain
Pitch
Natural
6
0.20
2.24
56.4
0.27
0.38
13.18
0.11
Synthetic
3
0.45
5.12
82.3
0.99
1.00
35.31
0.00
ID
2
0.18
2.10
50.5
0.17
0.43
14.41
0.11
OOD
4
0.20
2.32
59.3
0.25
0.36
12.55
0.11
BPM
Natural
9
0.20
2.29
59.7
0.27
0.42
13.56
0.15
TABLE II: Encoding metrics aggregated by feature and condition, reported mean across cells. Synthetic spectral values not analyzed. × null is ratio of observed mean against permutation derived null mean.
Fig. 3: Plots demonstrating what percentage of neurons encode multiple features across all cells. Left shows how many features are >p95 for each neuron; synthetic can have max. 2 features as this is all that is measured. Right shows what percentage of eligible neurons are >p95 for each feature (N=121,472 for BPM (all data), N=91,104 for pitch (no drums), N=91,104 for spectral centroid and spectral bandwidth (no synthetic)).
ICC
Meas.
Feature
N
β2 [IQR]
r
model
data
padj
Mean ∣ρ∣
Pitch
6
0.11 [0.09, 0.17]
1.00
0.37
0.00
0.042*
BPM
9
-0.09 [-0.11, 0.03]
-0.38
0.00
0.00
0.324
Spectral Centroid
9
0.25 [0.16, 0.55]
0.78
0.74
0.00
0.042*
Spectral Bandwidth
9
0.43 [0.17, 0.57]
0.87
0.32
0.00
0.042*
% > p95
Pitch
6
26.16 [12.34, 43.61]
0.81
0.06
0.01
0.125
TABLE III: Median per-cell quadratic coefficient β2 over normalized depth (positive = U-shaped, negative = inverted-U) with [IQR]. r is the matched-pairs rank-biserial effect size. Between-model and between-dataset intraclass correlations (ICC) are provided, along with the Benjamini–Hochberg-adjusted exact sign-permutation p . All natural cells; synthetic omitted (descriptive, N=3 ).
Fig. 4: Plots of per layer mean ∣ρ∣ , neuron prevalence and R2 across four audio features for natural audio, reported as median across cells. Only per layer mean ∣ρ∣ displayed for synthetic audio as other features are near ceiling throughout (e.g. flat). Significance values come from testing curvature coefficient ( β2 ) across cells with a one-sample Wilcoxon signed-rank test. Not possible for synthetic as N=3 (too low).
Fig. 5: Best cluster against best layer per section for each cell for all three metrics ( ∣ρ∣ , % >p95 and R2 ); only natural BPM ∣ρ∣ and >p95 show improvement for clusters.
Per-neuron
Joint probe R2
Feature
Condition
n
Mean ∣ρ∣
× null
% > p95
Linear
Nonlinear
× null
NL gain
Pitch
Natural
2
0.24
2.76
64.6
0.44
0.45
18.03
0.05
Synthetic
1
0.56
6.41
85.7
0.99
0.99
35.91
0.01
BPM
Natural
3
0.20
2.26
55.6
0.46
0.59
21.85
0.20
Synthetic
1
0.66
7.45
94.0
1.00
0.99
32.81
0.00
Spectral Centroid
Natural
3
0.33
3.71
65.9
0.83
0.85
33.80
0.07
TABLE IV: EnCodec encoding metrics aggregated by feature and condition, results are mean across cells. Multiples ( × null) are the observed value divided by the corresponding null baseline ( p95 for per-neuron, mean null R2 for the joint probe).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Per-neuron ∣ρ∣
Joint probe R2
Model
Dataset
Feature
Null p95
Obs. mean [95% CI]
Obs. max
% > p95
Null p95
Linear [95% CI]
Nonlinear [95% CI]
NL gain
Strings
Strings
Pitch
0.087
0.221 [0.210, 0.233]
0.729
58.6
0.023
0.460 [0.436, 0.475]
0.619 [0.584, 0.635]
0.159
BPM
0.087
0.153 [0.148, 0.158]
0.548
45.1
0.037
0.377 [0.323, 0.418]
0.512 [0.480, 0.541]
0.134
Spec C.
0.088
0.367 [0.349, 0.383]
0.904
77.4
0.038
0.817 [0.787, 0.838]
0.908 [0.893, 0.921]
0.091
Spec B.
0.088
0.256 [0.239, 0.277]
0.838
65.2
0.026
0.751 [0.706, 0.786]
0.836 [0.784, 0.866]
0.086
Drums
BPM
0.085
0.172 [0.166, 0.176]
0.431
64.5
0.027
0.246 [0.201, 0.277]
0.329 [0.286, 0.361]
0.083
Appendix
TABLE S1: Per-cell encoding metrics for all (model, dataset, feature) combinations. Per-neuron statistics from Spearman correlations; joint statistics from MLP probes. Null % >p95 values are from 500 permutations of shuffled feature labels. Linear R2 from ridge regression probes; nonlinear R2 from MLP probes; NL gain = nonlinear − linear R2 . Bracketed values are 95% bootstrap confidence intervals calculated at a layer level.
Arch.
Feature
Measure
N
Nat
Syn
Δ (HL) [95% CI]
r
ICC
padj
RAVE
Pitch
Mean ∣ρ∣
84
0.199
0.436
+0.248 [+0.232, +0.263]
+1.00
0.450
< .001 ***
% > p95
84
74.6
90.1
+15.76 [+14.35, +17.19]
+1.00
0.507
< .001 ***
R2
84
0.365
0.997
+0.612 [+0.594, +0.630]
+1.00
0.765
< .001 ***
BPM
Mean ∣ρ∣
84
0.197
0.752
+0.562 [+0.524, +0.625]
+1.00
0.835
< .001 ***
% > p95
84
77.8
97.4
+18.27 [+16.67, +19.88]
+1.00
0.754
< .001 ***
R2
84
0.432
0.996
+0.566 [+0.551, +0.580]
+1.00
0.097
< .001 ***
Appendix
TABLE S2: Natural versus synthetic encoding across architectures. Layer-matched Wilcoxon signed-rank tests (unit = layer pair, N pairs) for RAVE and a pretrained EnCodec decoder. Nat and Syn give medians [IQR]. Δ (HL) is the Hodges–Lehmann synthetic − natural difference with 95% CI (positive = synthetic higher); r is rank-biserial. padj is the Benjamini–Hochberg-adjusted exact sign-permutation p within each measure family. Spectral features excluded (no synthetic spectral analysis).
Measure
Feature
N
β2 [IQR]
r
Mean ∣ρ∣
Pitch
2
-0.14 [-0.16, -0.12]
-1.00
BPM
3
0.16 [0.08, 0.21]
0.67
Spectral Centroid
3
0.29 [0.03, 0.30]
0.67
Spectral Bandwidth
3
0.15 [0.10, 0.20]
1.00
% > p95
Pitch
2
-26.59 [-33.70, -19.49]
-1.00
BPM
3
19.72 [6.30, 40.94]
0.67
Appendix
TABLE S3: Depth curvature, EnCodec (descriptive). Median per-cell quadratic coefficient β2 over normalized depth (positive = U-shaped, negative = inverted-U) with [IQR] and rank-biserial r . As EnCodec is a single model, each feature contributes only N=2 – 3 cells; the across-cell Wilcoxon cannot reach significance ( p≥0.25 throughout), so results are descriptive.
Fig. S1: Plots of per layer mean ∣ρ∣ , neuron prevalence and R2 across four audio features for natural audio with the EnCodec model. Only per layer mean ∣ρ∣ displayed for synthetic audio as other features are near ceiling throughout (e.g. flat). Significance tests not possible as N=3 for natural and N=1 for synthetic.
ICC
Avg. best n
Arch.
Meas.
Feat.
N
Layer
Cluster
HL [95% CI]
r
mod
dat
padj
cluster
layer
RAVE
Mean ∣ρ∣
Pitch
18
0.239
0.237
− 0.004 [ − 0.015, +0.012]
− 0.05
0.154
0.339
0.769
779
480
BPM
27
0.233
0.257
+0.019 [+0.007, +0.033]
+0.65
0.000
0.411
0.006**
360
309
Spec. Centroid
27
0.377
0.362
− 0.018 [ − 0.036, +0.009]
− 0.29
0.268
0.008
0.400
1102
513
Spec. Bandwidth
27
0.348
0.335
− 0.022 [ − 0.044, +0.004]
− 0.34
0.000
0.000
0.605
831
466
% > p95
Pitch
18
82.6
86.9
+3.91 [+0.13, +8.37]
+0.50
0.183
0.153
0.061
514
420
Appendix
TABLE S4: Cross-layer cluster versus best single layer, natural audio only. Hodges–Lehmann estimate of the median per-cell difference (best cross-layer cluster − best single layer within a section) with 95% bootstrap CI, for RAVE and a pretrained EnCodec decoder; positive values indicate the cross-layer cluster outperforms the best single layer. Layer/cluster mean are the raw (non-paired) means; best cluster/layer n are the average number of units in the selected cluster vs. the best single layer. r is the matched-pairs rank-biserial effect size; padj is the Benjamini–Hochberg-adjusted exact sign-permutation p within each measure family. Between-model (mod) and between-dataset (dat) intraclass correlations index dependence among cells. EnCodec is a single model, so between-model ICC is undefined (–).
Fig. S2: Plots of mean silhouette score, cluster imbalance and mean and max feature correlation across values of k .
Fig. S3: Plot of mean cluster size (ranked by size) across values of k .
Fig. S4: Plot of mean ARI for all clusters, and mean Jaccard for best cluster across values of k .
End-to-end neural audio models achieve high-fidelity compression and generation. We might read that performance as evidence they directly represent interpretable features such as pitch and timbre, but a model can produce plausible outputs without doing so. A model may encode these features in any reachable basis, but regardless of which, the features are well described as compositions of time-frequency-localized primitives. Whether state-of-the-art encoders preserve access to these primitives, and thus to compositions of them, remains unclear. Through theoretical analysis and controlled experiments, we show that several state-of-the-art strided convolutional encoders impose two structural bottlenecks, both predictable from architecture and signal structure, on access to these primitives: (1) they collapse primitives into alias equivalence classes, establishing a bound on representational capacity, and (2) they limit the frequency resolution available to learned filters, restricting separability. For well structured data, we find collapse rates of 31-35% and filter bandwidths 10-35x above the theoretical resolution bound, confirming that both bottlenecks arise under realistic signal conditions. We then introduce Gabor Latent Refactorization (GLRF), a lightweight post-hoc intervention that re-expresses encoder latents in a frequency-localized basis, reducing filter bandwidths from 10-35x to 1.5-3x of the theoretical resolution bound while preserving reconstruction fidelity and improving control over attributes like pitch. These results show that the encoders in question predictably degrade access to frequency-localized primitives, entangling the features that depend on them, and that a lightweight, retraining-free intervention can recover much of that access, improving steerability and interpretability.
We argue that training autoencoders to reconstruct inputs from noised versions of their encodings, when combined with perceptually motivated losses, yields encodings that are structured according to a perceptual hierarchy. We demonstrate the emergence of this hierarchy by showing that, after training an audio autoencoder in this manner, perceptually salient information is captured in coarser representation structures than with conventional training. Furthermore, we show that such perceptual hierarchies improve latent diffusion decoding in the context of estimating pitch surprisal in music and predicting EEG-brain responses to music listening. In both cases, our results surpass those of previous methods. Pretrained weights are available on github.com/CPJKU/pa-audioic.
Mathias Rose Bjare, Giorgia Cantisani, Marco Pasini +2
Johannes Kepler University, Linz, AT · STMS, CNRS, IRCAM, Sorbonne Universit´e, Paris, FR · Queen Mary University, London, UK +1
Latent representations are at the heart of the majority of modern generative models. In the audio domain they are typically produced by a neural-audio-codec autoencoder. In this work we introduce SAME (Semantically-Aligned Music autoEncoder), an autoencoder for stereo music and general audio that reaches a 4096× temporal compression ratio while maintaining reconstruction quality and downstream generative performance. We achieve this by combining a tranformer-based backbone with set of semantic regularisation approaches, phase-aware reconstruction losses and improved discriminator designs. The architecture delivers substantial computational cost benefits, through both its high compression ratio and its reliance on well-optimised transformer primitives. Two variants (a large SAME-L and a CPU-deployable SAME-S) are released in open-weights form.