cs.SDSep 15, 2026

LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs

Authors: Thanapat TrachuSamuele CornellWilliam ChenShinji Watanabe

Organizations: Language Technologies Institute, Carnegie Mellon University, Pittsburgh, USA

Abstract

Neural audio codecs are a key component in speech language modeling. However, their high frame rates lead to long sequence lengths, increasing computational costs. Dynamic frame rate codecs mitigate this by reducing the effective frame rate using a compression step to merge multiple frames together. However, most prior methods either operate on single-codebook codecs or apply a single compression step before multi-layer quantization. This forces all quantization layers to share the same segmentation boundaries, despite the residual embeddings at different quantization layers exhibiting different rates of change over time. We propose LACE (Layer-Adaptive Codec Encoding), a dynamic frame rate codec that applies an independent compression step at each quantization layer, enabling layer-specific segmentation boundaries. To use LACE tokens in downstream text-to-speech (TTS), we further introduce union alignment and boundary anchor mechanisms to make durations consistent across layers while preserving compression benefits. Experiments on LibriTTS show that LACE offers a better rate-quality tradeoff than prior dynamic frame rate methods on the reconstruction task and improves TTS inference efficiency while maintaining competitive synthesis quality. Our code is released as part of the ESPnet3 codec recipe.

Explore similar work

CardsList
  1. ZipCodec: Ultra-Low-Frame-Rate Streaming Speech Coding

    Sep 12, 2026Luca Della Libera, Cem Subakan, Mirco RavanelliNeural Audio CodecsVideo Coding

  2. LILAC: An Idempotent Neural Speech Codec

    Aug 6, 2026June Young Yi, Dongwook Lee, Jiheum Yeom +1Neural Audio CodecsLibrispeech