cs.SDSep 29, 2026

Interpreting and Evaluating Dynamic-Rate Speech Codec Boundaries

Authors: Han Wang, Jiaqi Li, Yingda Shen, Yuxiang Wang, Zhizheng Wu

Organizations: The Chinese University of Hong Kong, Shenzhen · Zhejiang University · Amphion Technology Co., Ltd.

Abstract

Dynamic-frame-rate neural speech codecs replace a uniform frame grid with variable-duration tokens, making boundary placement part of the representation itself. Yet it is unclear what these boundaries encode and whether interpretable boundaries are also useful for neural speech reconstruction. This work combines boundary interpretation and controlled reconstruction analysis by comparing predicted boundaries with linguistic and acoustic references. We find that boundary meaning depends on the underlying speech representation. For the ASR-oriented SenseVoice and Whisper encoders, shallow layers emphasize phonetic, voicing, and acoustic transitions, whereas deeper layers shift toward syllable and subword structure. In a comparison of six dynamic-frame-rate algorithms and a uniform (fixed-frame-rate) baseline, higher-level linguistic alignment is associated with lower pooling distortion and better reconstruction from semantic tokens. Frame-rate-matched boundary tests make the distinction concrete: a syllable-derived partition improves over both Uniform and Similarity, while a phoneme-derived partition does not improve over Uniform. We infer that for semantically rich speech representations, useful codec boundaries are best understood as allocation decisions organized around the syllable scale.

Figures & tables

Explore similar work

CardsList
  1. LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs

    Sep 15, 2026Thanapat Trachu, Samuele Cornell, William Chen +1Frame RateVideo Coding

  2. DTM-Codec: Dynamic Token Masking for VFR Speech Coding with Efficient Boundary Selection

    Jun 28, 2026Hoyeol Sohn, Juhan NamNeural Speech CodecsFrame Rate

  3. Probing Low Frame Rate Degradation in Neural Audio Codecs

    Jun 15, 2026Alex Gichamba, Moise BusogiNeural Audio CodecsFrame Rate