Continuous audio encoders compress audio along feature and time axes through latent width and frame rate, but their joint effect on downstream performance remains unclear. We train sixteen encoders spanning four widths and four frame rates, with downstream adapters and probes, using matched training protocols. Despite generally improved reconstruction at larger widths, automatic speech recognition (ASR) and spoken question answering (SQA) favor moderate widths at higher rates, with the best observed widths shifting toward larger values under stronger temporal compression. Frozen-model PCA interventions reveal distinct reconstruction and recognition sensitivities: removing the trailing half of the components substantially degrades ASR in high-rate 512-dimensional encoders with comparatively small reconstruction penalties, whereas 1024-dimensional encoders largely preserve both. Yet the projected 1024-dimensional model underperforms unmodified narrower models on ASR at 12.5Hz. These findings identify a width--rate interaction in downstream utility and suggest that how representations are organized during training matters beyond reconstruction fidelity and compressibility.
Figures & tables
PESQ ↑
SI-SDR (dB) ↑
Hz
128d
256d
512d
1024d
128d
256d
512d
1024d
3.125
2.79
3.07
3.12
3.09
0.19
0.38
0.43
0.76
6.25
3.24
3.52
3.58
3.62
1.80
4.13
4.42
4.40
12.5
3.61
3.83
3.79
4.41
5.13
5.51
6.06
11.04
25.0
4.43
4.53
4.53
4.54
11.54
14.36
14.80
15.17
Table 1: Speech reconstruction results across all sixteen encoders. PESQ and SI-SDR are higher when better. Bold indicates the best result in each row.
ASR WER (%) ↓
SQA Cover-EM (%) ↑
Spk-Vox Acc (%) ↑
Spk-Libri Acc (%) ↑
Hz
128d
256d
512d
1024d
128d
256d
512d
1024d
128d
256d
512d
1024d
128d
256d
512d
1024d
3.125
8.67
7.82
7.66
7.62
7.22
13.05
13.53
13.74
19.12
22.36
23.43
22.14
94.28
97.14
97.14
96.02
6.25
4.76
4.47
4.37
4.58
19.10
22.37
21.17
26.58
27.94
30.03
30.50
27.37
98.95
99.41
99.13
99.09
12.5
3.91
3.77
3.83
4.22
32.64
33.39
34.58
31.38
35.75
35.89
36.31
23.63
99.58
99.83
99.72
98.67
25.0
3.89
3.83
3.89
3.91
28.95
32.08
29.67
31.62
25.04
25.12
26.09
25.78
97.84
97.66
98.50
98.95
Table 2: Downstream results for sixteen configurations. Bold marks the best result per metric in each row. Paired-bootstrap 95% CIs have ≈0.2 -pp half-widths for WER differences and include zero for SQA differences <1 pp; test-set sampling only.
ASR WER (%) ↓
Spk-Vox Acc (%) ↑
Hz
Setting
128d
256d
512d
1024d
128d
256d
512d
1024d
3.125
Half
22.15
12.28
9.71
7.53
16.87
20.98
22.37
21.49
Δ%
−60.8
−36.3
−21.2
+1.3
−11.8
−6.2
−4.5
−2.9
6.25
Half
13.10
6.17
4.89
4.56
25.24
28.61
29.02
26.38
Δ%
−63.6
−27.7
−10.6
+0.5
−9.7
−4.7
−4.9
−3.6
12.5
Half
8.04
7.37
30.48
4.19
33.06
34.17
34.98
22.97
Table 3: Fixed-weight top-half PCA ablation. “Half” reports performance after retaining the leading half of the PCA components. Δ denotes the relative performance change: 100(WERFull/WERHalf−1) for ASR and 100(AccHalf/AccFull−1) for speaker identification. Darker red indicates a larger absolute change.
PESQ Half ↑
PESQ Δ
SI-SDR Half (dB) ↑
SI-SDR Δ (dB)
Hz
128d
256d
512d
1024d
128d
256d
512d
1024d
128d
256d
512d
1024d
128d
256d
512d
1024d
3.125
1.98
2.48
3.09
3.09
−0.81
−0.59
−0.03
+0.00
−6.51
−1.42
0.30
0.76
−6.70
−1.80
−0.13
+0.00
6.25
2.27
2.83
3.55
3.62
−0.97
−0.69
−0.03
+0.00
−1.52
2.49
4.24
4.40
−3.32
−1.64
−0.18
+0.00
12.5
2.47
2.89
3.72
4.41
−1.14
−0.94
−0.07
+0.00
1.66
3.68
5.92
11.04
−3.47
−1.83
−0.14
+0.00
25.0
1.85
2.50
4.32
4.54
−2.58
−2.03
−0.21
+0.00
3.06
8.60
13.95
15.17
−8.48
−5.76
−0.85
+0.00
Table 4: Codec reconstruction after the ASR top-half PCA intervention. “Half” reports reconstruction performance after retaining the leading half of the PCA components. Δ denotes the difference between Half and the corresponding full-representation result reported in Table 1 . For each reconstruction metric, blue intensity is normalized across all sixteen configurations, with darker blue indicating better performance. Darker red indicates a larger absolute degradation.
End-to-end neural audio models achieve high-fidelity compression and generation. We might read that performance as evidence they directly represent interpretable features such as pitch and timbre, but a model can produce plausible outputs without doing so. A model may encode these features in any reachable basis, but regardless of which, the features are well described as compositions of time-frequency-localized primitives. Whether state-of-the-art encoders preserve access to these primitives, and thus to compositions of them, remains unclear. Through theoretical analysis and controlled experiments, we show that several state-of-the-art strided convolutional encoders impose two structural bottlenecks, both predictable from architecture and signal structure, on access to these primitives: (1) they collapse primitives into alias equivalence classes, establishing a bound on representational capacity, and (2) they limit the frequency resolution available to learned filters, restricting separability. For well structured data, we find collapse rates of 31-35% and filter bandwidths 10-35x above the theoretical resolution bound, confirming that both bottlenecks arise under realistic signal conditions. We then introduce Gabor Latent Refactorization (GLRF), a lightweight post-hoc intervention that re-expresses encoder latents in a frequency-localized basis, reducing filter bandwidths from 10-35x to 1.5-3x of the theoretical resolution bound while preserving reconstruction fidelity and improving control over attributes like pitch. These results show that the encoders in question predictably degrade access to frequency-localized primitives, entangling the features that depend on them, and that a lightweight, retraining-free intervention can recover much of that access, improving steerability and interpretability.
Low frame rates in neural audio codecs are attractive for autoregressive speech synthesis, where the generation cost scales linearly with the sequence length. Recent work has demonstrated that codecs can operate at 12.5 Hz and below, but the mechanisms underlying low frame rate degradation remain insufficiently understood. We investigate these mechanisms through a controlled frame rate ablation. We reproduce a quality cliff at 6.25 Hz reported in previous works and evaluate candidate explanations: phonemic collisions and codebook saturation, neither of which shows evidence of a fundamental barrier. The cliff is instead caused by suboptimal training configuration: fixed clip duration during training yields too few tokens at low frame rates, starving the decoder of inter-token context. Once corrected, WER degrades smoothly with phonemic load down to 3.1 Hz and 1.6 Hz, suggesting the inference-time efficiency gains of low frame rate codecs are more accessible than previously assumed.
Neural audio autoencoders have become a core component of compression, feature extraction, and generation. However, while existing systems support variable bitrate, the vast majority of models still operate at a fixed latent frame-rate, allocating equal temporal budget to regions with very different information density, which can result in unnecessarily long sequences. We introduce Elastic Time, a dynamic frame-rate bottleneck that converts fixed-frame-rate autoencoders to dynamic ones. Our method learns a lightweight latent predictor used to decide which frames can be skipped and later reconstructed, enabling efficient greedy boundary selection at inference. Experiments show our method enables deployment-time rate control while improving efficiency-quality tradeoffs relative to baselines. Overall, we provide a flexible mechanism for adjusting temporal resolution in audio autoencoders, potentially facilitating more efficient downstream modeling for generation and long-context tasks.
Dimitrios Bralios, Paris Smaragdis, Minje Kim
University of Illinois Urbana-Champaign, Urbana, IL, USA · Massachusetts Institute of Technology, Cambridge, MA, USA