Neural speech codecs increasingly serve as tokenizers for speech language models (SLMs). Lowering the frame rate reduces the computational and memory costs of SLMs, but makes it difficult to preserve both linguistic information and acoustic detail. Existing approaches rely on rule-based compression: average pooling can discard linguistic information, whereas similarity-based merging uses a fixed threshold on adjacent-frame similarity and applies the resulting boundaries to the acoustic stream. We propose Q-SPT, a low-frame-rate dual-stream speech tokenizer with separate, context-aware, learnable query-based compressors specialized for semantic and acoustic representations. In particular, queries at a fixed rate independently attend to the semantic and acoustic streams as separate key-value sources, enabling stream-specific, context-aware aggregation through two separately learned compressors. In addition, an autoregressive text loss explicitly supervises the semantic compressor to preserve linguistic information. Experimental results show that Q-SPT achieves the best reconstruction among the evaluated codecs at the same frame rate. In downstream SLMs, it yields the best speech recognition accuracy and text-to-speech perceptual quality with competitive intelligibility.
Figures & tables
Figure 1: Comparison of temporal compression strategies for low-frame-rate speech tokenization: (a) average pooling in DualCodec [ 22 ] , (b) similarity-based merging in FlexiCodec [ 23 ] , and (c) learnable query-based compression in Q-SPT. (d) Detailed architecture and training objectives of Q-SPT. Dashed elements indicate components and connections used only during training.
Bitrate (kbps)
nq=1
nq=8
Model
Hz
nq=1
nq=8
WER ↓
SIM ↓
WER ↓
PESQ-NB ↑
PESQ-WB ↑
MCD ↓
SIM ↑
UTMOS ↑
Mimi [ 8 ]
12.5
0.138
1.100
76.78
0.09
2.96
2.90
2.25
4.28
0.79
3.34
OmniCodec [ 13 ]
12.5
0.138
1.100
69.05
0.07
2.53
2.83
2.15
3.75
0.85
3.36
LM-SPT [ 16 ]
12.5
0.175
1.225
3.38
0.26
2.21
3.41
2.76
3.12
0.91
3.84
DualCodec [ 22 ]
12.5
0.175
1.225
5.19
0.48
2.25
3.42
2.82
3.01
0.87
3.97
DualCodec ∗
6.25
0.088
0.613
48.95
0.27
4.26
2.80
2.12
3.87
0.72
3.86
Table 1: Speech reconstruction results on LibriSpeech test-clean. ∗ denotes our stride-modified DualCodec. Bold indicates the best result within each frame-rate group.
ASR
TTS
Model
WER ↓
WER ↓
UTMOS ↑
DNSMOS ↑
DualCodec ∗
120.16
14.72
3.07
3.18
FlexiCodec
6.18
10.65
3.05
3.10
Q-SPT (w/o text loss)
7.04
12.31
3.16
3.24
Q-SPT
5.88
11.82
3.44
3.28
Table 2: SLM results on LibriSpeech test-clean at 6.25 Hz. ASR uses the first code stream, and TTS uses all eight streams. ∗ denotes our stride-modified DualCodec. Best results are shown in bold.
nq=1
nq=8
Configuration
LAR
WER ↓
WER ↓
PESQ-WB ↑
UTMOS ↑
Shared weights
✗
6.16
2.86
2.16
4.02
Shared boundaries
✓
4.29
3.83
2.14
3.99
Separate compressors
✗
4.97
2.59
2.25
4.03
Separate compressors
✓
4.29
2.52
2.26
4.07
Table 3: Ablation study of Q-SPT on speech reconstruction. Shared boundaries are applied only at inference. Best results are shown in bold.
Neural audio codecs are a fundamental component of modern speech generation systems. While recent codecs achieve increasingly low bitrates, reducing frame rate remains challenging, as each token must preserve more information while maintaining reconstruction quality. We present ZipCodec, a streaming neural speech codec operating at 6.25 Hz and 0.80 kbps with a theoretical latency of 160 ms. Our approach combines large-scale WavLM distillation with a redesigned transformer-based architecture, a scalar spherical quantizer, and a latency-aware streaming decoder. Experiments show that ZipCodec substantially outperforms existing streaming codecs at comparable bitrates in both reconstruction and downstream tasks, while operating at a significantly lower frame rate. Despite its 842M parameters, ZipCodec achieves real-time single-stream inference on a consumer-grade CPU. Demo samples, code and checkpoints are available at https://lucadellalib.github.io/zipcodec-web/.
Luca Della Libera, Cem Subakan, Mirco Ravanelli
Concordia University · Mila-Quebec AI Institute · Universit´e Laval
Speech codecs serve as bridges between continuous speech signals and large language models, yet face an inherent conflict between acoustic fidelity and semantic preservation. To mitigate this conflict, recent works increasingly adopt dual-tower architectures to decouple semantic and acoustic modeling with separate encoders. However, these dual-tower designs incur substantial architectural overhead. To avoid such complexity, we revisit the single-tower paradigm and propose BiMTokenizer, a low-bitrate speech codec (around 1.1 kbps) combining a bidirectional state-space backbone with Residual Spherical Leech Quantization (RSLQ). The bidirectional backbone strengthens temporal modeling, while RSLQ offers a fixed, well-separated lattice bottleneck for robust semantic and acoustic tokenization without learned-codebook collapse. Experiments show that BiMTokenizer achieves superior acoustic reconstruction and the lowest WER among low-bitrate codec baselines across both clean and noisy environments, while using less than half the parameters of recent dual-tower baselines. Furthermore, its robust semantic representations yield strong performance on downstream speech understanding tasks, confirming that a well-designed single-tower codec can preserve the semantic-acoustic balance at low bitrates. The code and model weights are available at https://github.com/ZhangXinWhut/BiMTokenizer.
Xin Zhang, Lin Li, Chuanbo Liu +2
Wuhan University of Technology, Wuhan, China · NEC Laboratories Asia Pacific, Singapore · NEC Corporation, Japan +1
System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency. In this work we study how to compress such a tokenizer by distillation, taking as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes - the last representation the two token interfaces share. We train only the student encoder to regress the teacher's per-frame latent under a squared-error objective, with a single affine layer absorbing the teacher-student width mismatch. Because the target precedes both the quantizer and the language-model bridge, one recipe covers both token interfaces we support, and applies both to a tokenizer pretrained alone and to one jointly trained with a language model. At 2.8x compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher-student pairs without any fine-tuning, and improves on an independently trained tokenizer of identical capacity by 3.9% relative.
Prasanth Yadla, Mohammad Samragh, Dongseong Hwang +7