Dynamic-frame-rate neural speech codecs replace a uniform frame grid with variable-duration tokens, making boundary placement part of the representation itself. Yet it is unclear what these boundaries encode and whether interpretable boundaries are also useful for neural speech reconstruction. This work combines boundary interpretation and controlled reconstruction analysis by comparing predicted boundaries with linguistic and acoustic references. We find that boundary meaning depends on the underlying speech representation. For the ASR-oriented SenseVoice and Whisper encoders, shallow layers emphasize phonetic, voicing, and acoustic transitions, whereas deeper layers shift toward syllable and subword structure. In a comparison of six dynamic-frame-rate algorithms and a uniform (fixed-frame-rate) baseline, higher-level linguistic alignment is associated with lower pooling distortion and better reconstruction from semantic tokens. Frame-rate-matched boundary tests make the distinction concrete: a syllable-derived partition improves over both Uniform and Similarity, while a phoneme-derived partition does not improve over Uniform. We infer that for semantically rich speech representations, useful codec boundaries are best understood as allocation decisions organized around the syllable scale.
Figures & tables
Figure 1: EXP1: Visualization of reference boundaries and rate-matched segmentations for one utterance.
Boundary type
Uniform
Similarity
CodecSlime
PLE
TADPC
Entropy-Mass
A-ToMe
Syllable
0.38 / 0.00
0.56 / +0.18
0.53 / +0.16
0.48 / +0.10
0.49 / +0.11
0.38 / +0.00
0.40 / +0.02
BPE
0.36 / 0.00
0.51 / +0.15
0.48 / +0.12
0.45 / +0.09
0.45 / +0.09
0.36 / +0.00
0.38 / +0.02
Word
0.29 / 0.00
0.45 / +0.16
0.41 / +0.12
0.40 / +0.11
0.38 / +0.09
0.29 / +0.00
0.30 / +0.01
Phoneme
0.50 / 0.00
0.52 / +0.03
0.52 / +0.02
0.53 / +0.03
0.50 / 0.00
0.50 / +0.00
0.50 / 0.00
Acoustic event
0.47 / 0.00
0.48 / +0.01
0.49 / +0.01
0.50 / +0.02
0.46 / -0.02
0.48 / +0.01
0.48 / +0.01
V/UV
0.44 / 0.00
0.43 / -0.01
0.43 / -0.01
0.46 / +0.01
0.41 / -0.03
0.45 / +0.01
0.44 / 0.00
Table 1: EXP1: Boundary F1 at nominal 6.25 Hz on TIMIT TEST with a 40-ms matching tolerance. Each cell reports F1 and its difference from Uniform, F1/ΔU , rounded to two decimals. Boldface highlights the higher-level linguistic-track deltas for Similarity, CodecSlime, PLE, and TADPC.
Method
Syllable F1/Rank
q1 ASR WER ↓
q1 recon WER ↓
q1:8 recon WER ↓
NMSE ↓
PESQ ↑
SpkSim ↑
Runtime (ms) ↓
Similarity
0.557 / 1
7.762
13.304
5.979
0.0759
2.294
0.9247
0.277
CodecSlime
0.535 / 2
7.586
10.411
5.394
0.0627
2.343
0.9246
19.546
TADPC
0.485 / 3
8.438
16.328
5.731
0.0830
2.316
0.9248
1.945
PLE
0.482 / 4
8.397
12.967
5.415
0.0805
2.330
0.9238
0.283
A-ToMe
0.402 / 5
8.269
16.953
6.116
0.0852
2.360
0.9225
1.290
Entropy-Mass
0.383 / 6
8.417
20.293
6.968
0.0943
2.349
0.9207
31.017
Table 2: EXP2: Reconstruction quality and boundary-selection runtime for seven methods at 6.25 Hz on TIMIT TEST. WER is corpus-level; runtime excludes input loading and encoder extraction.
Rate
q1 recon WER
NMSE
SpkSim
6.25 Hz
0.79/0.79/0.86
0.89/0.89/0.93
0.86/0.86/0.75
8.33 Hz
0.89/0.82/0.82
0.71/0.61/0.61
0.54/0.57/0.57
10 Hz
0.82/0.82/0.79
0.79/0.79/0.75
0.64/0.64/0.61
Table 3: EXP2: Cross-rate absolute Spearman correlations ∣ρ∣ between syllable/BPE/word boundary F1 and reconstruction utility across seven methods.
Segmentation
Phoneme-rate WER ↓
Syllable-rate WER ↓
Uniform
6.47
17.61
Reference-derived
7.10
13.85
CodecSlime
5.11
11.28
Similarity
5.05
14.99
Table 4: EXP2: Corpus WER (%) for reference-derived and rate-matched segmentations on TIMIT TEST at phoneme (8.06 Hz) and syllable (4.47 Hz) rates. Bold/underline mark the best/second-best result within each rate.