Neural speech codecs increasingly serve as tokenizers for speech language models (SLMs). Lowering the frame rate reduces the computational and memory costs of SLMs, but makes it difficult to preserve both linguistic information and acoustic detail. Existing approaches rely on rule-based compression: average pooling can discard linguistic information, whereas similarity-based merging uses a fixed threshold on adjacent-frame similarity and applies the resulting boundaries to the acoustic stream. We propose Q-SPT, a low-frame-rate dual-stream speech tokenizer with separate, context-aware, learnable query-based compressors specialized for semantic and acoustic representations. In particular, queries at a fixed rate independently attend to the semantic and acoustic streams as separate key-value sources, enabling stream-specific, context-aware aggregation through two separately learned compressors. In addition, an autoregressive text loss explicitly supervises the semantic compressor to preserve linguistic information. Experimental results show that Q-SPT achieves the best reconstruction among the evaluated codecs at the same frame rate. In downstream SLMs, it yields the best speech recognition accuracy and text-to-speech perceptual quality with competitive intelligibility.
Figures & tables
Figure 1: Comparison of temporal compression strategies for low-frame-rate speech tokenization: (a) average pooling in DualCodec [ 22 ] , (b) similarity-based merging in FlexiCodec [ 23 ] , and (c) learnable query-based compression in Q-SPT. (d) Detailed architecture and training objectives of Q-SPT. Dashed elements indicate components and connections used only during training.
Bitrate (kbps)
nq=1
nq=8
Model
Hz
nq=1
nq=8
WER ↓
SIM ↓
WER ↓
PESQ-NB ↑
PESQ-WB ↑
MCD ↓
SIM ↑
UTMOS ↑
Mimi [ 8 ]
12.5
0.138
1.100
76.78
0.09
2.96
2.90
2.25
4.28
0.79
3.34
OmniCodec [ 13 ]
12.5
0.138
1.100
69.05
0.07
2.53
2.83
2.15
3.75
0.85
3.36
LM-SPT [ 16 ]
12.5
0.175
1.225
3.38
0.26
2.21
3.41
2.76
3.12
0.91
3.84
DualCodec [ 22 ]
12.5
0.175
1.225
5.19
0.48
2.25
3.42
2.82
3.01
0.87
3.97
DualCodec ∗
6.25
0.088
0.613
48.95
0.27
4.26
2.80
2.12
3.87
0.72
3.86
Table 1: Speech reconstruction results on LibriSpeech test-clean. ∗ denotes our stride-modified DualCodec. Bold indicates the best result within each frame-rate group.
ASR
TTS
Model
WER ↓
WER ↓
UTMOS ↑
DNSMOS ↑
DualCodec ∗
120.16
14.72
3.07
3.18
FlexiCodec
6.18
10.65
3.05
3.10
Q-SPT (w/o text loss)
7.04
12.31
3.16
3.24
Q-SPT
5.88
11.82
3.44
3.28
Table 2: SLM results on LibriSpeech test-clean at 6.25 Hz. ASR uses the first code stream, and TTS uses all eight streams. ∗ denotes our stride-modified DualCodec. Best results are shown in bold.
nq=1
nq=8
Configuration
LAR
WER ↓
WER ↓
PESQ-WB ↑
UTMOS ↑
Shared weights
✗
6.16
2.86
2.16
4.02
Shared boundaries
✓
4.29
3.83
2.14
3.99
Separate compressors
✗
4.97
2.59
2.25
4.03
Separate compressors
✓
4.29
2.52
2.26
4.07
Table 3: Ablation study of Q-SPT on speech reconstruction. Shared boundaries are applied only at inference. Best results are shown in bold.