German is a morphologically rich language whose syllable structure is exceptionally well-predicted by the Knuth--Liang hyphenation algorithm. We ask whether phonologically informed tokenization can serve as a competitive target for end-to-end speech recognition. We compare three tokenizer families on the Omnilingual ASR wav2vec 2.0 backbone fine-tuned with CTC: the pretrained multilingual character inventory, a data-driven Byte-Pair Encoding (BPE) over orthography, and phonologically informed units from Pyphen syllabification and grapheme-to-phoneme conversion. Across 40 fine-tunes, we evaluate on three German test sets spanning orthogonal shifts: in-domain read speech, dialectal spontaneous speech, and standard-German spontaneous speech. In-domain, all phonologically informed tokenizers match BPE and the multilingual character baseline on both WER and CER. Under domain shift the picture splits along vocabulary size rather than the linguistic axis of variation: at small vocabularies, syllable-aware tokenization improves on dialectal speech, where phonetic surface forms vary but syllable structure is preserved, and stays ahead on spontaneous speech, where new word-forms violate vocabulary closure. A phoneme-level confusion analysis further shows that all tokenizers commit the same canonical function-word errors, indicating that the acoustic encoder, not the tokenizer, dominates the error topology. Our findings suggest that tokenizer choice may depend on the vocabulary budget as much as on the distribution shift expected at deployment rather than reducing to a single universal optimum.
Figures & tables
Corpus
Split
n utterances
hours
Role / OOD axis
CV + SWC
train
550 747
960.5
training mix
CV 24 DE ( Ardila et al., 2020 )
(part of train)
457 512
726.5
speakers worldwide
SWC DE ( Köhn et al., 2016 )
(part of train)
93 235
234.0
Spoken Wikipedia
CV + SWC
test
24 604
50.2
in-domain hold-out
RVG 1 ( Burger and Schiel, 1998 )
all, pooled
442
6.6
dialectal-phonetic OOD
Verbmobil VM1 ( Burger et al., 2000 )
all, pooled (German)
13 633
33.8
spontaneous-lexical OOD
Table 1: Corpora. Training uses the CV 24 + SWC mixture; the in-domain test split is held out from the same mixture; RVG 1 and Verbmobil are entirely unseen and pooled across their train/dev/test splits (§ 3.8 ). Both OOD corpora carry parallel ORT (canonical orthography) and TR2 (dialect-/prosody-annotated transliteration) tiers. Verbmobil VM1 also contains English negotiation dialogues; we evaluate on its German subset only (§ 3.8 ).
Family
Variant
Vocab
ntok
cpt
Tokenization
Orthographic track: segmenting the German word “ unterhaltung ” (12 chars, 4 canonical syllables un-ter-hal-tung)
Multilingual char
omni-vocab
10 288
12
1.00
u | n | t | e | r | h | a | l | t | u | n | g
Data-driven BPE
german-BPE-512
512
5
2.40
un | ter | hal | t | ung
Syllable (Pyphen)
syllable-512
512
4
3.00
un | ter | hal | tung
Multisyl (cross-syllable)
multisyl-512
512
6
2.00
un | ter | ha | lt | un | g
Multisyl-nosep
multisyl-nosep-512
512
6
2.00
(identical to multisyl-512 on a single word † )
Table 2: Smallest variant of each tokenizer family on the German word Unterhaltung (orthographic track) and its X-SAMPA transcription ?Un-t6-hal-tUN (phoneme track). Bold rows mark the family that recovers all four canonical syllables (or SAMPA syllables) as individual tokens. cpt = characters per token; on the phoneme track, atomic SAMPA symbols are counted, excluding syllable-separator dashes. † multisyl and multisyl-nosep produce identical outputs on a single word; the boundary marker affects only vocabulary construction.
Tokenizer
WER %
95 % CI
CER %
95 % CI
Multilingual character inventory
omni-vocab
15.64
[15.38, 15.90]
6.09
[5.89, 6.29]
Data-driven BPE (German)
German-BPE-512
14.93
[14.66, 15.18]
6.14
[5.94, 6.35]
German-BPE-1024
16.88
[16.61, 17.13]
6.89
[6.69, 7.09]
German-BPE-2048
16.11
[15.84, 16.36]
6.65
[6.45, 6.85]
Table 3: In-domain WER and CER on CV+SWC test ( n=24,604 ), 300 M wav2vec 2.0 with CTC, on normalized text (lowercase, ß → ss, punctuation stripped). Brackets: paired-bootstrap 95 % CIs ( nboot=10,000 , seed= 42 ); bold marks the per-metric best across the 13 orthographic tokenizers.
Phone Tokenizer
Vocab
PER %
95 % CI
Atomic / raw
phone-char-raw-v2
43
18.05
[17.78, 18.33]
phone-raw-v2
50
18.35
[18.08, 18.64]
Phone-BPE
phone-bpe-v2-128
128
18.04
[17.78, 18.32]
phone-bpe-v2-192
192
18.96
[18.69, 19.24]
Table 4: Phoneme-track in-domain PER on CV+SWC test ( 300 M). Bold marks the best. Spread across the 17 convergent variants is ∼2.5 pp; phone-char-raw-v2 ( 43 tokens) is within reach of phone-bpe-v2- 1024 with a 24× smaller vocabulary. † phone-multisyl-v2- 3840 has not converged at 20 k steps and is excluded from comparisons.
WER %
CER %
Tokenizer
1 B
300 M
1 B
Multilingual character inventory
omni-vocab
11.65
15.64
5.08
Data-driven BPE (German)
German-BPE-512
14.34
14.93
5.93
German-BPE-1024
15.90
16.88
6.47
Table 5: In-domain WER/CER on CV+SWC test ( n=24,604 ) for the nine tokenizers also fine-tuned at 1 B (continued-training checkpoints), alongside their 300 M WER from Table 3 . Scoring as in Table 3 ; bold marks the best per column.
wav2vec 2.0, 1 B (continued-training checkpoints)
RVG 1 (dialectal spontaneous, n=442 )
Verbmobil (spontaneous, German, n≈13,600 )
ORT
TR2
ORT
TR2
Tokenizer
WER
CER
WER
CER
WER
CER
WER
CER
Multilingual character inventory
omni-vocab
36.83 ± 2.1
16.53 ± 1.3
45.35 ± 1.8
18.23 ± 1.1
22.16 ± 0.3
14.08 ± 0.2
22.68 ± 0.3
14.10 ± 0.2
Data-driven BPE (German)
Table 6: Cross-domain WER/CER (%) at 1 B on normalized text (lowercase, ß → ss, punctuation stripped), by tokenizer family. RVG 1: pooled train+dev+test ( n=442 , audio chunked at silence to ≤60 s); Verbmobil: German subset, pooled ( n=13,613 ORT / 13,618 TR2 scored for all nine systems). ± = paired-bootstrap 95 % CI half-width ( nboot=10,000 , seed= 42 ); bold = best per column. Significance tests: Table 7 .
ID
Comparison
Setup
WER A / B
Δ (95 % CI)
pboot
P1-ort
msy-512 vs omni @ 1B
RVG 1 ORT
35.75 / 36.83
−1.1 [ −2.0,−0.2 ]
0.019
P1-tr2
msy-512 vs omni @ 1B
RVG 1 TR2
45.62 / 45.35
+0.3 [ −0.5,+1.0 ]
0.518
P2-300m-cv
syl-512 vs gBPE-512 @ 300M
CV+SWC
15.16 / 14.93
+0.2 [ +0.1,+0.3 ]
<0.001∗
P2-1b-cv
syl-512 vs gBPE-512 @ 1B
CV+SWC
12.27 / 14.34
−2.1 [ −2.2,−2.0 ]
<0.001∗
P2-300m-ort
syl-512 vs gBPE-512 @ 300M
RVG 1 ORT
46.23 / 42.30
+3.9 [ +3.3,+4.7 ]
<0.001∗
P2-1b-ort
syl-512 vs gBPE-512 @ 1B
RVG 1 ORT
38.63 / 47.12
−8.5 [ −9.8,−7.2 ]
<0.001∗
Table 7: Paired-bootstrap tests on Δ -WER (10,000 resamples, seed=42; normalized text as in Table 3 ). Top: nine focal comparisons. Bottom: interaction contrasts at 1 B; C1 on RVG 1 is P1-ort, and for C2 the WER column lists the two constituent differences Δ512 / Δ2048 (syllable minus German-BPE) whose difference is the reported Δ . Verbmobil contrasts use the German subset (Table 6 ). ∗ = passes per-dataset Bonferroni correction ( α=0.05/k ). 1 B values use the continued-training checkpoints, 300 M and phoneme-track values their single-stage runs. omni = omni-vocab, msy = multisyl, syl = syllable, gBPE = German-BPE, p- = phone track.
RVG 1-ORT
Verbmobil-ORT
∣V∣
gBPE / Syl ( Δ )
gBPE / Syl ( Δ )
512
47.12 / 38.63 ( −8.5 )
25.47 / 21.27 ( −4.2 )
1024
45.07 / 42.37 ( −2.7 )
26.95 / 21.83 ( −5.1 )
2048
42.80 / 44.31 ( +1.5 )
24.17 / 25.60 ( +1.4 )
Δ2048−Δ512
+10.0∗
+5.6∗
Table 8: WER (%) at 1 B by vocabulary size for German-BPE (gBPE) and Syllable (Syl) on the two OOD sets (ORT tier; Verbmobil: German subset). Δ = Syllable − German-BPE; the last row is the paired-bootstrap interaction contrast (C2 in Table 7 ; ∗p<0.001 ).
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
wav2vec 2.0, 1 B (Table 7 )
mHuBERT-147, 95 M
Contrast
Set
Δ (95 % CI)
pboot
Δ (95 % CI)
pboot
C1: msy-512 − omni
RVG 1 ORT
−1.1 [ −2.0,−0.2 ]
0.019
−1.0 [ −1.7,−0.3 ]
0.006
Verbmobil ORT
−0.1 [ −0.3,+0.1 ]
0.214
−0.3 [ −0.5,−0.1 ]
<0.001
C2: (syl − gBPE)@2048 − @512
RVG 1 ORT
+10.0 [ +8.4,+11.6 ]
<0.001
+8.7 [ +7.8,+9.7 ]
<0.001
Verbmobil ORT
+5.6 [ +5.4,+5.9 ]
<0.001
−0.6 [ −0.8,−0.4 ]
<0.001
C3: syl-1024 − msy-nosep-512
RVG 1 ORT
−2.8 [ −3.9,−1.6 ]
<0.001
−2.1 [ −2.6,−1.5 ]
<0.001
Appendix
Table 9: Interaction contrasts C1–C3 ( Δ -WER in pp, paired bootstrap, 10,000 resamples, seed=42) on the two backbones; scoring and abbreviations as in Table 7 ; Verbmobil on the German subset.
RVG 1
Verbmobil
CV+SWC
Tokenizer
ORT
ORT
test
omni-vocab
69.50
45.36
22.35
German-BPE-512
70.36
45.23
24.38
German-BPE-2048
63.27
46.57
28.24
syllable-512
63.32
43.44
23.26
syllable-1024
64.04
44.51
24.20
Appendix
Table 10: WER (%) of the eight mHuBERT-147 fine-tunes ( 30 k steps, 10 k of them with the encoder frozen); test sets and scoring as in Tables 3 and 6 (Verbmobil: German subset). Bold marks the best per column.
Reference
omni-vocab
multisyl-512
German-BPE-512
RVG 1
und dann habe ich gedacht
und ann ab ch gedacht
und dann habe ich gedacht
und dann habe ich gedacht
da konnte man etwas erzählen
da konte an etas erzählen
da konnte man etwas erzählen
da konn manwas erzählen
und dann haben wir jetzt
und dan aben ir jetzt
und dann haben wir jetzt
und dann haben wir jetzt
tag verdienen hektik gab es auch
tag erdienen hekti gab s auch
tag verdienen hektik gab es auch
tag verdiener hektik gab es auch
und die alarmanlage muss man †
und die alarmanlage muss man
und der lamanlage mussss man
und der lamanlage muss man
Appendix
Table 11: Automatically selected reference/hypothesis windows ( 1 B systems, normalized text). Bold marks hypothesis words that do not match the reference. Selection: RVG 1 windows in which multisyl-512 is error-free and every omni-vocab error is a near-miss substitution (character distance ≤2 ); Verbmobil windows in which omni-vocab is error-free and both subword systems err on a word absent from the training text. † Counterexample with the opposite pattern.
System
Most frequent word substitutions (reference → hypothesis, count)
RVG 1-ORT
omni-vocab
habe → hab (296), das → es (76), es → s (69), das → des (61), dann → dan (45), ist → is (44), dass → das (42), und → un (35)
multisyl-512
habe → ha (232), habe → hab (100), das → es (60), halt → ha (55), das → des (41), dass → dasss (38), noch → nur (38), dass → das (36)
German-BPE-512
ja → j (70), was → w (50), das → es (46), bisschen → bis (42), habe → hab (38), habe → h (32), bin → b (32), das → des (29)
Verbmobil-ORT (German)
omni-vocab
wäre → wär (767), es → s (637), habe → hab (469), das → es (327), einen → ein (211), würde → wird (174), ginge → ging (160), mal → nochmal (159)
Appendix
Table 12: Eight most frequent word substitutions per system ( 1 B, normalized text) on the two OOD sets. On RVG 1 the substitutions are dominated by reduced dialectal realizations of function words; on Verbmobil by colloquial contractions shared by all systems.
Tokenizer
∣V∣
tok/wd
ch/tok
tokens
per
(M)
entry (k)
Multilingual character inventory
omni-vocab
10,288
6.96
0.87
42.7
4.2
Data-driven BPE (German)
German-BPE-512
512
2.72
2.23
16.7
32.6
German-BPE-1024
1024
2.32
2.61
14.2
13.9
Appendix
Table 13: Target-token statistics for the 13 orthographic tokenizers, from corpus-level fertility on the training text ( 550,747 sentences, 6,141,326 words). tokens = tokens per word × words; per entry = tokens / ∣V∣ , i.e., how often an average vocabulary entry occurs in one pass over the training text; this is distinct from gradient steps, which are fixed at 20 k for every run (§ 4.1 ). The phoneme track is tokenized on its own SAMPA transcripts and is not included.
Unsupervised syllabic tokenization aims to learn discrete syllabic tokens that capture latent linguistic content-related structure from raw speech. Recent syllabic tokenization methods employ teacher-student distillation of the pretrained HuBERT to organize latent speech frame representations into syllabic segments. However, when trained with an utterance-level cross-entropy objective, the model predicts speaker identity rather than linguistic content, thereby compromising the purity of syllabic tokens. To address this problem, we propose a speaker-disentangled syllabic tokenizer that regresses speaker-perturbed student representations toward clean teacher targets within fixed-length chunks. Experimental results demonstrate that our proposed method achieves state-of-the-art performance in syllable boundary detection and syllabic segment clustering. Moreover, a speech language model trained on our syllabic tokens achieves a 7% relative improvement in syntactic and semantic understanding over the phone-level SpiRit-LM.
Ryota Komatsu, Kota Kawakita, Takuma Okamoto +1
Institute of Science Tokyo, Meguro, Tokyo 152-8550, Japan · National Institute of Information and Communications Technology, Kyoto 619-0289, Japan
Tokenization is the first step in every language model (LM), yet it never takes the sounds of words into account. We investigate how tokenization influences text-only LMs' ability to represent phonological knowledge. Through a series of probing experiments, we show that subword-based tokenization systematically weakens the encoding of both local (e.g., rhyme) and global (e.g., syllabification) phonological features. To quantify this effect, we introduce the syllabification-tokenization alignment distance (STAD), a metric that measures the misalignment between a model's tokenization and the natural syllable boundaries of words, and find that higher misalignment correlates with poorer phonological representations, providing a simple diagnostic for phonology-aware tokenization. To address these limitations, we propose a lightweight IPA-based fine-tuning method that infuses phonological awareness into LMs, leading to consistent improvements across three phonology-related tasks while largely preserving math and general reasoning ability, with 1.1% and 0.9% drops on GSM8K and MMLU, respectively.
Speech tokenizers are essential for connecting speech to large language models (LLMs) in multimodal systems. Speech tokenizers are expected to preserve both semantic and acoustic information for downstream understanding and generation tasks. However, emerging evidence suggests that the term "semantic" in speech processing does not align with linguistic lexical-semantic, leading to a mismatch between speech and text modality. In this paper, we systematically analyze the information encoded by several widely used speech tokenizers, evaluating their lexical-semantic and phonetic content through three tasks. Our results show that current tokenizers primarily capture phonetic rather than lexical-semantic structure, deriving practical implications for the design of next-generation speech tokenization methods. Code is released to public at https://github.com/Alexuan/codec_probing_release.
Xuan Shi, Chang Zeng, Tiantian Feng +3
University of Southern California, USA · Dolby Laboratories, USA