Dynamic-frame-rate neural speech codecs replace a uniform frame grid with variable-duration tokens, making boundary placement part of the representation itself. Yet it is unclear what these boundaries encode and whether interpretable boundaries are also useful for neural speech reconstruction. This work combines boundary interpretation and controlled reconstruction analysis by comparing predicted boundaries with linguistic and acoustic references. We find that boundary meaning depends on the underlying speech representation. For the ASR-oriented SenseVoice and Whisper encoders, shallow layers emphasize phonetic, voicing, and acoustic transitions, whereas deeper layers shift toward syllable and subword structure. In a comparison of six dynamic-frame-rate algorithms and a uniform (fixed-frame-rate) baseline, higher-level linguistic alignment is associated with lower pooling distortion and better reconstruction from semantic tokens. Frame-rate-matched boundary tests make the distinction concrete: a syllable-derived partition improves over both Uniform and Similarity, while a phoneme-derived partition does not improve over Uniform. We infer that for semantically rich speech representations, useful codec boundaries are best understood as allocation decisions organized around the syllable scale.
Figures & tables
Figure 1: EXP1: Visualization of reference boundaries and rate-matched segmentations for one utterance.
Boundary type
Uniform
Similarity
CodecSlime
PLE
TADPC
Entropy-Mass
A-ToMe
Syllable
0.38 / 0.00
0.56 / +0.18
0.53 / +0.16
0.48 / +0.10
0.49 / +0.11
0.38 / +0.00
0.40 / +0.02
BPE
0.36 / 0.00
0.51 / +0.15
0.48 / +0.12
0.45 / +0.09
0.45 / +0.09
0.36 / +0.00
0.38 / +0.02
Word
0.29 / 0.00
0.45 / +0.16
0.41 / +0.12
0.40 / +0.11
0.38 / +0.09
0.29 / +0.00
0.30 / +0.01
Phoneme
0.50 / 0.00
0.52 / +0.03
0.52 / +0.02
0.53 / +0.03
0.50 / 0.00
0.50 / +0.00
0.50 / 0.00
Acoustic event
0.47 / 0.00
0.48 / +0.01
0.49 / +0.01
0.50 / +0.02
0.46 / -0.02
0.48 / +0.01
0.48 / +0.01
V/UV
0.44 / 0.00
0.43 / -0.01
0.43 / -0.01
0.46 / +0.01
0.41 / -0.03
0.45 / +0.01
0.44 / 0.00
Table 1: EXP1: Boundary F1 at nominal 6.25 Hz on TIMIT TEST with a 40-ms matching tolerance. Each cell reports F1 and its difference from Uniform, F1/ΔU , rounded to two decimals. Boldface highlights the higher-level linguistic-track deltas for Similarity, CodecSlime, PLE, and TADPC.
Method
Syllable F1/Rank
q1 ASR WER ↓
q1 recon WER ↓
q1:8 recon WER ↓
NMSE ↓
PESQ ↑
SpkSim ↑
Runtime (ms) ↓
Similarity
0.557 / 1
7.762
13.304
5.979
0.0759
2.294
0.9247
0.277
CodecSlime
0.535 / 2
7.586
10.411
5.394
0.0627
2.343
0.9246
19.546
TADPC
0.485 / 3
8.438
16.328
5.731
0.0830
2.316
0.9248
1.945
PLE
0.482 / 4
8.397
12.967
5.415
0.0805
2.330
0.9238
0.283
A-ToMe
0.402 / 5
8.269
16.953
6.116
0.0852
2.360
0.9225
1.290
Entropy-Mass
0.383 / 6
8.417
20.293
6.968
0.0943
2.349
0.9207
31.017
Table 2: EXP2: Reconstruction quality and boundary-selection runtime for seven methods at 6.25 Hz on TIMIT TEST. WER is corpus-level; runtime excludes input loading and encoder extraction.
Rate
q1 recon WER
NMSE
SpkSim
6.25 Hz
0.79/0.79/0.86
0.89/0.89/0.93
0.86/0.86/0.75
8.33 Hz
0.89/0.82/0.82
0.71/0.61/0.61
0.54/0.57/0.57
10 Hz
0.82/0.82/0.79
0.79/0.79/0.75
0.64/0.64/0.61
Table 3: EXP2: Cross-rate absolute Spearman correlations ∣ρ∣ between syllable/BPE/word boundary F1 and reconstruction utility across seven methods.
Segmentation
Phoneme-rate WER ↓
Syllable-rate WER ↓
Uniform
6.47
17.61
Reference-derived
7.10
13.85
CodecSlime
5.11
11.28
Similarity
5.05
14.99
Table 4: EXP2: Corpus WER (%) for reference-derived and rate-matched segmentations on TIMIT TEST at phoneme (8.06 Hz) and syllable (4.47 Hz) rates. Bold/underline mark the best/second-best result within each rate.
Neural audio codecs are a key component in speech language modeling. However, their high frame rates lead to long sequence lengths, increasing computational costs. Dynamic frame rate codecs mitigate this by reducing the effective frame rate using a compression step to merge multiple frames together. However, most prior methods either operate on single-codebook codecs or apply a single compression step before multi-layer quantization. This forces all quantization layers to share the same segmentation boundaries, despite the residual embeddings at different quantization layers exhibiting different rates of change over time. We propose LACE (Layer-Adaptive Codec Encoding), a dynamic frame rate codec that applies an independent compression step at each quantization layer, enabling layer-specific segmentation boundaries. To use LACE tokens in downstream text-to-speech (TTS), we further introduce union alignment and boundary anchor mechanisms to make durations consistent across layers while preserving compression benefits. Experiments on LibriTTS show that LACE offers a better rate-quality tradeoff than prior dynamic frame rate methods on the reconstruction task and improves TTS inference efficiency while maintaining competitive synthesis quality. Our code is released as part of the ESPnet3 codec recipe.
Thanapat Trachu, Samuele Cornell, William Chen +1
Language Technologies Institute, Carnegie Mellon University, Pittsburgh, USA
Variable frame rate (VFR) coding has recently emerged in neural speech codecs, allocating fewer frames to redundant regions and more frames to rapidly changing speech. VFR must transmit side information about retained time steps, but prior gains are either not rigorously addressed or often minor once these overhead bits are included in total bitrate. We present Dynamic Token Masking (DTM)-Codec, a neural speech codec that demonstrates clear gains over fixed-frame-rate baselines under a strict matched-total-bitrate protocol. DTM keeps selected encoder tokens, fills masked positions with a learned <MASK> embedding, and transmits a binary keep-mask for position-aware decoding. We further introduce Path Length Equalization (PLE), a linear-time boundary selector for VFR coding that yields well-spread adaptive segments with negligible overhead. Across operating points, DTM-Codec broadly improves reconstruction quality and intelligibility over fixed-frame-rate baselines.
Hoyeol Sohn, Juhan Nam
Graduate School of Cultural Technology, KAIST, South Korea
Low frame rates in neural audio codecs are attractive for autoregressive speech synthesis, where the generation cost scales linearly with the sequence length. Recent work has demonstrated that codecs can operate at 12.5 Hz and below, but the mechanisms underlying low frame rate degradation remain insufficiently understood. We investigate these mechanisms through a controlled frame rate ablation. We reproduce a quality cliff at 6.25 Hz reported in previous works and evaluate candidate explanations: phonemic collisions and codebook saturation, neither of which shows evidence of a fundamental barrier. The cliff is instead caused by suboptimal training configuration: fixed clip duration during training yields too few tokens at low frame rates, starving the decoder of inter-token context. Once corrected, WER degrades smoothly with phonemic load down to 3.1 Hz and 1.6 Hz, suggesting the inference-time efficiency gains of low frame rate codecs are more accessible than previously assumed.