Discrete speech tokens provide a compact interface between speech encoders and large language models for automatic speech recognition, but single-tokenization systems remain sensitive to the chosen encoder. We propose multi-view discrete token augmentation, a simple strategy that augments each training utterance by generating alternative token sequences from fixed SSL encoders, such as HuBERT, WavLM, and MMS-300M. These tokenizations are treated as complementary training views for a shared LLM decoder, exposing it to more diverse discrete speech representations without requiring multi-encoder inference. At test time, the model can operate with a single encoder. On LibriSpeech, the approach consistently improves all encoders over independently trained baselines, with WavLM reaching 3.30% WER on test-clean and 8.13% on test-other. Budget-matched controls show that the gains come from encoder diversity rather than data volume. ROVER over multi-view hypotheses further improves WER to 3.03% and 7.38%.
Figures & tables
Fig. 1: Discrete speech tokenization. A frozen SSL encoder extracts frame-level features, which are quantized with K-means ( K=2000 ), de-duplicated, and BPE-encoded (vocabulary 6000) into the discrete tokens that are then fed to the LLM decoder.
Encoder
LS clean
LS other
Loq. dev †
WER ↓
CER ↓
WER ↓
CER ↓
WER ↓
CER ↓
Single-encoder baselines
HuBERT
3.71
1.51
9.70
4.84
21.96
12.70
WavLM
4.08
1.81
9.55
4.82
22.17
12.91
MMS-300M
6.08
2.89
15.10
8.34
25.55
14.83
Oracle
2.30
0.95
6.68
3.48
16.02
9.35
TABLE I: Single-encoder baselines and multi-view token augmentation results. The shared-decoder multi-view model is trained on H, W, and M encoder-specific token views and evaluated with one encoder at a time. Oracle selects the best hypothesis per utterance. † YODAS excluded. Best encoder row per block in bold. Values in %.
Configuration
LS clean
LS other
Loq. dev †
WER ↓
CER ↓
WER ↓
CER ↓
WER ↓
CER ↓
Best single-encoder baseline
Best single enc.
3.71
1.51
9.55
4.82
21.96
12.70
Early fusion
H ⊕ W
3.55
1.57
8.48
4.22
23.24
14.21
H ⊕ M
4.97
2.19
12.62
6.54
24.29
13.68
TABLE II: Fusion baselines using multiple SSL encoders. H, W, and M denote HuBERT-Large, WavLM-Large, and MMS-300M. ⊕ is early embedding fusion; +tag , +pipe , and +str are late token-fusion variants. All systems use Qwen 2.5-0.5B with LoRA, K -means ( K=2000 ), BPE 6000, and LS-960 training. † YODAS excluded. Best per column in bold. Values in %.
System
Common Voice
LibriSpeech
VoxPopuli
YODAS
Single-encoder baselines
H
37.34 / 21.74
9.27 / 4.73
26.13 / 14.90
51.39 / 33.72
W
39.00 / 22.41
9.12 / 4.79
25.78 / 14.97
64.26 / 45.17
M
37.12 / 20.62
13.92 / 7.48
31.07 / 18.55
65.51 / 42.67
Shared-decoder multi-view model
H
47.59 / 27.78
8.13 / 4.03
24.96 / 14.70
121.17 / 77.53
TABLE III: Loquacious dev WER/CER (%) by source corpus. H, W, and M denote HuBERT-Large, WavLM-Large, and MMS-300M. Each cell reports WER/CER. YODAS diverges strongly and is excluded from the main Loq. dev † scores.
Model
LS clean
LS other
Loq. dev †
WER ↓
CER ↓
WER ↓
CER ↓
WER ↓
CER ↓
Same- N/3 utterances across H/W/M
H
4.57
1.88
12.11
6.52
31.22
19.89
W
7.28
4.03
10.35
5.92
29.80
18.91
M
6.82
3.31
16.17
8.59
31.80
19.90
Oracle
3.16
1.30
7.77
3.90
17.75
10.32
TABLE IV: Data-budget and repetition controls. H/W/M denote HuBERT-Large, WavLM-Large, and MMS-300M. Same- N/3 uses the same utterances across encoders; split- N/3 uses different utterances per encoder. The 3-stream control repeats one encoder three times. † YODAS excluded. Best per block and column in bold. Values in %.
Pair
View
LS clean
LS other
Loq. dev †
WER ↓
CER ↓
WER ↓
CER ↓
WER ↓
CER ↓
H+M
H
3.45
1.30
8.74
4.23
22.79
13.42
M
5.13
2.37
12.92
6.75
26.06
15.14
Oracle
2.82
1.06
7.53
3.70
17.24
9.63
H+W
H
3.72
1.66
9.28
4.71
29.90
18.31
W
6.04
3.74
8.97
4.55
25.28
15.83
TABLE V: Pairwise shared-decoder results. H/W/M denote HuBERT-Large, WavLM-Large, and MMS-300M. Oracle selects the best hypothesis per utterance from the two encoder views. † YODAS excluded. Best oracle result per column in bold. Values in %.
System
LS clean
LS other
Loq. dev †
WER ↓
CER ↓
WER ↓
CER ↓
WER ↓
CER ↓
Single-encoder hypotheses
Best single view
3.71
1.51
9.55
4.82
21.96
12.70
ROVER H+W+M
3.14
1.52
7.88
4.54
17.39
11.18
Shared-decoder multi-view hypotheses
Best decoded view
3.30
1.23
8.13
3.84
20.28
11.51
TABLE VI: Output-level combination with confidence-averaged ROVER ( avgconf ). H/W/M denote HuBERT-Large, WavLM-Large, and MMS-300M. † YODAS excluded. Best ROVER result per column in bold. Values in %.
Graduate Institute of Communication Engineering, National Taiwan University · ASUS Intelligent Cloud Services · NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)