Discrete speech tokens provide a compact interface between speech encoders and large language models for automatic speech recognition, but single-tokenization systems remain sensitive to the chosen encoder. We propose multi-view discrete token augmentation, a simple strategy that augments each training utterance by generating alternative token sequences from fixed SSL encoders, such as HuBERT, WavLM, and MMS-300M. These tokenizations are treated as complementary training views for a shared LLM decoder, exposing it to more diverse discrete speech representations without requiring multi-encoder inference. At test time, the model can operate with a single encoder. On LibriSpeech, the approach consistently improves all encoders over independently trained baselines, with WavLM reaching 3.30% WER on test-clean and 8.13% on test-other. Budget-matched controls show that the gains come from encoder diversity rather than data volume. ROVER over multi-view hypotheses further improves WER to 3.03% and 7.38%.
Figures & tables
Fig. 1: Discrete speech tokenization. A frozen SSL encoder extracts frame-level features, which are quantized with K-means ( K=2000 ), de-duplicated, and BPE-encoded (vocabulary 6000) into the discrete tokens that are then fed to the LLM decoder.
Encoder
LS clean
LS other
Loq. dev †
WER ↓
CER ↓
WER ↓
CER ↓
WER ↓
CER ↓
Single-encoder baselines
HuBERT
3.71
1.51
9.70
4.84
21.96
12.70
WavLM
4.08
1.81
9.55
4.82
22.17
12.91
MMS-300M
6.08
2.89
15.10
8.34
25.55
14.83
Oracle
2.30
0.95
6.68
3.48
16.02
9.35
TABLE I: Single-encoder baselines and multi-view token augmentation results. The shared-decoder multi-view model is trained on H, W, and M encoder-specific token views and evaluated with one encoder at a time. Oracle selects the best hypothesis per utterance. † YODAS excluded. Best encoder row per block in bold. Values in %.
Configuration
LS clean
LS other
Loq. dev †
WER ↓
CER ↓
WER ↓
CER ↓
WER ↓
CER ↓
Best single-encoder baseline
Best single enc.
3.71
1.51
9.55
4.82
21.96
12.70
Early fusion
H ⊕ W
3.55
1.57
8.48
4.22
23.24
14.21
H ⊕ M
4.97
2.19
12.62
6.54
24.29
13.68
TABLE II: Fusion baselines using multiple SSL encoders. H, W, and M denote HuBERT-Large, WavLM-Large, and MMS-300M. ⊕ is early embedding fusion; +tag , +pipe , and +str are late token-fusion variants. All systems use Qwen 2.5-0.5B with LoRA, K -means ( K=2000 ), BPE 6000, and LS-960 training. † YODAS excluded. Best per column in bold. Values in %.
System
Common Voice
LibriSpeech
VoxPopuli
YODAS
Single-encoder baselines
H
37.34 / 21.74
9.27 / 4.73
26.13 / 14.90
51.39 / 33.72
W
39.00 / 22.41
9.12 / 4.79
25.78 / 14.97
64.26 / 45.17
M
37.12 / 20.62
13.92 / 7.48
31.07 / 18.55
65.51 / 42.67
Shared-decoder multi-view model
H
47.59 / 27.78
8.13 / 4.03
24.96 / 14.70
121.17 / 77.53
TABLE III: Loquacious dev WER/CER (%) by source corpus. H, W, and M denote HuBERT-Large, WavLM-Large, and MMS-300M. Each cell reports WER/CER. YODAS diverges strongly and is excluded from the main Loq. dev † scores.
Model
LS clean
LS other
Loq. dev †
WER ↓
CER ↓
WER ↓
CER ↓
WER ↓
CER ↓
Same- N/3 utterances across H/W/M
H
4.57
1.88
12.11
6.52
31.22
19.89
W
7.28
4.03
10.35
5.92
29.80
18.91
M
6.82
3.31
16.17
8.59
31.80
19.90
Oracle
3.16
1.30
7.77
3.90
17.75
10.32
TABLE IV: Data-budget and repetition controls. H/W/M denote HuBERT-Large, WavLM-Large, and MMS-300M. Same- N/3 uses the same utterances across encoders; split- N/3 uses different utterances per encoder. The 3-stream control repeats one encoder three times. † YODAS excluded. Best per block and column in bold. Values in %.
Pair
View
LS clean
LS other
Loq. dev †
WER ↓
CER ↓
WER ↓
CER ↓
WER ↓
CER ↓
H+M
H
3.45
1.30
8.74
4.23
22.79
13.42
M
5.13
2.37
12.92
6.75
26.06
15.14
Oracle
2.82
1.06
7.53
3.70
17.24
9.63
H+W
H
3.72
1.66
9.28
4.71
29.90
18.31
W
6.04
3.74
8.97
4.55
25.28
15.83
TABLE V: Pairwise shared-decoder results. H/W/M denote HuBERT-Large, WavLM-Large, and MMS-300M. Oracle selects the best hypothesis per utterance from the two encoder views. † YODAS excluded. Best oracle result per column in bold. Values in %.
System
LS clean
LS other
Loq. dev †
WER ↓
CER ↓
WER ↓
CER ↓
WER ↓
CER ↓
Single-encoder hypotheses
Best single view
3.71
1.51
9.55
4.82
21.96
12.70
ROVER H+W+M
3.14
1.52
7.88
4.54
17.39
11.18
Shared-decoder multi-view hypotheses
Best decoded view
3.30
1.23
8.13
3.84
20.28
11.51
TABLE VI: Output-level combination with confidence-averaged ROVER ( avgconf ). H/W/M denote HuBERT-Large, WavLM-Large, and MMS-300M. † YODAS excluded. Best ROVER result per column in bold. Values in %.
We propose AugCodec, a low-bitrate disentangled neural speech codec that leverages data augmentation to decompose speech into three distinct components: semantic, speaker, and prosody tokens. Specifically, we employ tailored augmenta tion strategies to transform speech into distinct variants, each serving as input for extracting tokens that preserve the target attribute while suppressing others. This disentanglement strategy enables substantial reduction in token rate. Further more, we introduce an augmentation loss that aligns semantic encoder outputs between source and voice-converted speech, encouraging speaker-agnostic embeddings while mitigating the acoustic mismatch induced by voice conversion. Experiments on LibriSpeech test-clean demonstrate that AugCodec significantly outperforms state-of-the-art methods in both reconstruction quality and disentanglement, while operating at only 12.5Hz with three token streams.
Neural audio codecs are widely used as tokenizers for spoken language models, but they are optimized for waveform reconstruction rather than autoregressive prediction. This mismatch injects acoustically driven uncertainty into the discrete token space and increases language-model perplexity. We propose \ours, which augments codec training with language-model-facing objectives while keeping both codec and LLM architectures unchanged. \ours introduces (i) future token prediction with Medusa-style multi-step heads to encourage multi-step predictability, and (ii) semantic alignment that matches audio and text representations via a memory-bank contrastive loss. A differentiable Gumbel bridge enables end-to-end gradients from these objectives to the codec encoder. On SALMon speech coherence, token LMs trained on \ours reach 61.6% accuracy (+12.1 points over AUV) while reducing perplexity 35. On Codec-SUPERB-tiny, \ours improves speech Mel distance by 5.0% over AUV while simultaneously achieving the learnability gains, demonstrating that reconstruction fidelity and token predictability can be improved together.
Ho-Lam Chung, Yiming Chen, Hung-yi Lee
Graduate Institute of Communication Engineering, National Taiwan University · ASUS Intelligent Cloud Services · NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)
Discrete audio representations have become increasingly popular for building multimodal text-audio systems and integrating audio capabilities into Large Language Models (LLMs). However, numerous studies report performance degradation on various downstream tasks due to information loss during discretization. To address this, we propose a novel approach combining temporally compressed discrete tokens with dimensionality-reduced continuous residuals. Our framework consists of a hybridized discrete-continuous focal modulation codec and a hybrid Transformer. This architecture performs autoregressive inference in the discrete domain, coupled with non-autoregressive prediction and continuous residual upsampling. Experimental results show that our approach significantly improves the retention of speaker characteristics compared to discrete-only methods, while simultaneously reducing the number of required autoregressive steps.
Artem Ploujnikov, Francesco Verdini, Samir Sadok +1
Mila, Quebec AI Institute, Canada · Concordia University, Canada · Sapienza University of Rome, Italy +1