Tokenizers are commonly optimized for compression, but a compact vocabulary does not necessarily distribute its capacity evenly across languages. We introduce the Latent Core Tokenizer (LCT), a language-agnostic approach that separates structural discovery from vocabulary construction. LCT uses Minimum Description Length, entropy-based boundary signals, and morphotactic constraints to identify reusable linguistic units before constructing a shared vocabulary. Across 104 languages with a 200K-token vocabulary, LCT achieves lower fertility and higher MorphScore than BPE, Unigram, and parity-aware BPE, while maintaining comparable cross-lingual disparity in tokenization cost. Across four multilingual downstream benchmarks, LCT improves aggregate score by 1.48, 1.83, and 2.00 points over BPE, Unigram, and parity-aware BPE, respectively. Our findings show that compression alone does not predict representation quality and highlight the importance of morphology-driven structural discovery and how frequency is used to allocate the final vocabulary across languages.
Figures & tables
Comp. Rate ↑
TTR ↑
Vocab. Util. ↑
Fertility ↓
Rényi ( α=2.5 ) ↑
MorphScore F1 ↑
Gini ↓
LCT-MEM
4.866
0.1781
0.2526
1.999
0.4472
0.5209
0.1341
BPE
4.295
0.1349
0.2168
2.241
0.4613
0.4410
0.1342
Parity-Aware BPE
3.340
0.0630
0.1301
2.742
0.4879
0.3186
0.1563
Unigram
3.195
0.0878
0.1898
2.920
0.3142
0.3109
0.1435
Table 1: Intrinsic tokenizer quality at 200k vocabulary. LCT-MEM denotes the LCT configuration combining MDL, branching entropy, and morphotactic refinement. Per-language results are displayed in Table 8 and Table 9 .
Figure 1: Intrinsic tokenizer performance across low-, mid-, and high-resource language tertiles. Bars show macro-averages and error bars indicate 95% confidence intervals across languages. MorphScore uses 65 languages, while fertility and vocabulary utilization use the 93 languages covered by TokEval. All configurations use a 200k vocabulary.
Rank
Fertility ↓
MorphScore ↑
Gini ↓
# 200k
# 256k
System
200k
256k
200k
256k
200k
256k
1
1
LCT-MEM τ=0.2
2.023
1.935
0.5249
0.5556
0.1369
0.1409
2
2
LCT-MEM τ=0
2.017
1.929
0.5209
0.5507
0.1341
0.1384
3
3
LCT-MEM τ=0.4
2.047
1.960
0.5244
0.5543
0.1374
0.1422
4
4
LCT-ME τ=0.8
2.053
1.963
0.5233
0.5549
0.1536
0.1570
5
8
LCT-MDL τ=0
1.962
1.876
0.4880
0.5245
0.1402
0.1439
Table 2: Top 10 tokenizer configurations at 200k and 256k vocabulary sizes. Additional intrinsic evaluation results are provided in Figure 3 .
Benchmark
Aggregate accuracy (%)
Language
XNLI
Belebele
XStoryCloze
PAWS-X
LCT-MDL τ=0
LCT-MDL τ=0.2
LCT-MEM τ=0
LCT-MEM τ=0.2
BPE
Unigram
Parity-Aware BPE
Arabic
✓
✓
✓
–
42.14
42.65
41.90
42.10
42.46
41.92
41.79
Basque
–
✓
✓
–
42.71
42.83
42.78
43.10
42.76
41.34
42.78
Bulgarian
✓
✓
–
–
39.27
40.25
39.68
40.18
38.98
37.84
38.99
Chinese
✓
✓
✓
✓
49.89
50.04
50.11
51.40
47.30
47.77
45.99
English
✓
✓
✓
✓
62.20
62.33
61.63
62.39
61.70
61.45
61.38
Table 3: Aggregate downstream accuracy (%) for languages covered by at least 2 of the four multilingual benchmarks. A checkmark identifies the benchmarks included for each language. For each tokenizer and language, accuracy is averaged equally across the checked tasks within each fine-tuning seed and then averaged over five seeds. All configurations use one pretrained encoder seed and a 200k vocabulary. The best result in each row is bold.
Tokenizer
Score
LCT-MDL τ=0
LCT-MDL τ=0.2
LCT-MEM τ=0
LCT-MEM τ=0.2
BPE
Unigram
Parity-Aware BPE
LCT-MDL τ=0
32.84±0.38
−0.80∗
−0.24
−1.45∗∗,‡
+0.03
+0.38
+0.55
LCT-MDL τ=0.2
33.64±0.32
0.54
+0.56
−0.65∗
+0.83∗
+1.18∗∗
+1.35∗∗
LCT-MEM τ=0
33.08±0.64
0.65
0.78
−1.21∗
+0.27
+0.62
+0.79
LCT-MEM τ=0.2
34.29±0.33
0.23
0.35
0.70
+1.48∗∗,†
+1.83∗∗,†
+2.00∗∗
BPE
32.81±0.26
0.39
0.51
0.40
0.43
+0.35
+0.52
Unigram
32.46±0.22
0.38
0.50
0.77
0.42
0.38
+0.17
Table 4: Pairwise comparison of the aggregated downstream score (accuracy %). Score is each configuration’s mean and standard deviation over five fine-tuning seeds. Rows and columns contain the same seven configurations in the same order. All configurations use a 200k vocabulary and one pretrained encoder each. Values above the diagonal report the paired difference between the row and column configurations, with positive values favouring the row. Differences are paired by seed ( df=4 ) and Holm-corrected across the 21 unordered pairs. ∗ and ∗∗ indicate raw paired-test p<0.05 and p<0.01 , while † and ‡ indicate Holm-adjusted p<0.05 and p<0.01 , respectively. Values that remain significant after correction are bold.
Table 6: The systems compared. All share the 200,024-identifier budget, the byte-level pre-tokenizer, the encoder corpus and the same 41,347,986 lines of tokenizer training text; they differ only in how the 199,763 learned pieces are chosen.
Table 7: LCT’s latent discovery reads at most 100k lines per language.
Figure 3: Intrinsic evaluation results for the 200k and 256k vocabulary configurations.
Fertility ↓
MorphScore F1 ↑
Gini ↓
LCT-M τ=0
BPE
P-base
Uni.
LCT-M τ=0
BPE
P-base
Uni.
LCT-M τ=0
BPE
P-base
Uni.
Afrikaans
1.431
1.376
1.482
1.772
0.397
0.354
0.270
0.332
3.67
-0.23
-4.26
-2.99
Albanian, Tosk
1.785
1.775
1.804
2.254
0.331
0.236
0.250
0.196
5.06
0.48
-4.58
5.74
Amharic †
70.610
50.665
77.457
62.488
16.89
1.23
7.45
-0.93
Arabic, Standard
1.618
1.513
1.651
2.056
-4.49
-5.60
-5.05
-5.00
Armenian
2.002
2.113
3.035
3.041
0.412
0.372
0.319
0.194
-5.02
-5.63
-4.80
-5.01
Appendix
Table 8: Intrinsic evaluation (Fertility, MorphScore and Gini) by language
Comp. Rate ↑
TTR ↑
Vocab. Util. ↑
Rényi ↑
LCT-M τ=0
BPE
P-base
Uni.
LCT-M τ=0
BPE
P-base
Uni.
LCT-M τ=0
BPE
P-base
Uni.
LCT-M τ=0
BPE
P-base
Uni.
Afrikaans
3.917
4.128
3.770
3.435
0.270
0.286
0.252
0.221
3.57
3.58
3.45
3.32
0.356
0.353
0.361
0.322
Albanian, Tosk
3.858
3.958
3.826
2.813
0.260
0.276
0.265
0.176
4.25
4.39
4.35
3.93
0.328
0.325
0.329
0.228
Amharic
2.798
3.899
2.550
3.161
0.097
0.108
0.018
0.136
2.83
2.25
0.57
3.50
0.174
0.197
0.245
0.178
Arabic, Standard
5.487
5.987
5.354
4.548
0.370
0.410
0.327
0.283
4.63
4.69
4.19
4.27
0.407
0.414
0.418
0.323
Armenian
7.368
6.982
4.873
4.837
0.265
0.230
0.104
0.179
4.18
3.82
2.48
4.28
0.389
0.398
0.396
0.256
Appendix
Table 9: Intrinsic evaluation (Comp. Rate, TTR, Vocab. Util. and Rényi) by language
Accuracy
Δ vs BPE
Language
mdl 0
mdl 0.2
mem 0
mem 0.2
BPE
Uni.
mdl 0
mdl 0.2
mem 0
mem 0.2
XNLI , all 15 languages, grouped by family
Indo-European (9)
bulgarian
48.97
50.48
49.56
47.90
48.89
47.28
+0.08
+1.59
+0.67
−0.99
english
68.14
67.64
68.07
68.34
68.59
69.30
−0.45
−0.96
−0.53
−0.26
french
50.18
49.82
50.06
48.75
53.64
46.61
−3.47
−3.82
−3.58
−4.89
Appendix
Table 10: Per-language test accuracy (%) on every evaluated language, averaged over five fine-tuning seeds. mdl and mem are the twoLCT objectives, mem abbreviating mdl_ent_morph , and the number beside each is τ ; Uni. is Unigram. Chance is 33.3 on XNLI, 25.0 on Belebele, 50.0 on XStoryCloze and 50.0 on PAWS-X, so margins are not comparable across the sections and the Belebele ones are small in absolute terms. Δ is the paired difference against BPE over the five shared seeds. Green marks a difference in LCT’s favour and red against it; only differences that clear a two-sided paired t test at p<0.05 across the five seeds are shaded, and saturation scales with the size of the difference up to 5 points. Unshaded cells are within seed-to-seed noise. Languages are grouped by genealogical family, ordered by how many of the 104 pretraining languages each family contributes. Belebele covers the 79 test languages present in the pretraining corpus. All arms share one pretraining seed.
Figure 4: Tokenization examples
Figure 5: Aggregate downstream accuracy across low-, mid-, and high-resource language tertiles. Scores are averaged across the available XNLI, Belebele, XStoryCloze, and PAWS-X evaluations and five fine-tuning seeds. Error bars indicate 95% confidence intervals across languages.
The Unigram tokenizer uses an elegant representation which makes it straightforward to edit vocabularies, but its training is comparatively heavy and complex. We introduce MinGram (Minimalist Unigram), which keeps the token-list representation but simplifies training using a BPE-derived seed vocabulary, Hard EM on a minimum-token path, and a single flat score-pruning step. This removes the suffix array, the forward-backward pass, and the iterative prune loop, leaving a procedure that requires little beyond tokenizer inference itself. By making token count the primary objective and using a Unigram score only as a tiebreak, MinGram keeps the compression of pure token-count methods while retaining much of the morphological alignment and downstream quality of probabilistic ones. Across six languages, MinGram compresses better than both BPE and standard Unigram, and a compression-oriented variant matches the strongest token-count compressors while retaining substantially higher morphological alignment. In controlled downstream language-model training, Unigram-family tokenizers, with MinGram among the best, consistently beat BPE in bits-per-byte.
We introduce Tokenization with Split Trees (ToaST), a subword tokenization method that directly optimizes compression under a new recursive inference procedure. ToaST greedily splits each pretoken into a full binary tree using precomputed byte n-gram counts, independent of any vocabulary. Given a vocabulary, inference recursively descends each split tree and emits the first in-vocabulary node reached on each path. Vocabulary selection is formulated as an Integer Program (IP) that minimizes the total token count over all split trees under this inference procedure. The Linear Programming (LP) relaxation is near-integral in practice, yielding provably near-optimal vocabularies, with training time empirically scaling quadratically in the number of split trees. On English text, ToaST reduces token counts by more than 11% compared to BPE, WordPiece, and UnigramLM at vocabulary sizes of 40,960 and above, reducing the number of inference tokens for models using this tokenizer, thus extending the effective context length. ToaST also uses common single-byte tokens less frequently than these baselines, leading to a substantial improvement in Renyi efficiency. In experiments training 1.5B parameter language models, ToaST achieves the highest CORE score, outperforming baselines by 2.6%--7.6%, with significance for two of three, and scoring best on 13 of 22 individual tasks.
Craig W. Schmidt, Michael Krumdick, Adam Wiemerslage +4
Kensho Technologies · Ben-Gurion University · MIT Cambridge, MA
Language-specific tokenizers improve tokenization quality and the downstream performance of models on those languages. However, using such a tokenizer comes at a cost: either a new model must be trained from scratch, or the vocabulary of an existing pretrained model must be adapted. We propose Language-adaptive Maximum a Posteriori (LangMAP) Tokenization, a tokenization scheme that extends the UnigramLM algorithm to the multilingual setting, producing language-specific tokenization from a single shared vocabulary. Notably, LangMAP can be used when training a multilingual language model from scratch or to adapt a pretrained model's tokenizer to individual languages without changing its vocabulary. While language labels are required at training time, a key feature of the algorithm is that it then performs language-specific tokenization at inference without knowledge of the input's language. Across 14 open-source tokenizers, 9 natural languages, and 9 programming languages, LangMAP improves morphological boundary alignment and, for all coding languages tested, alignment with abstract syntax tree (AST) leaf boundaries. In fine-tuning experiments, results are mixed: LangMAP improves target-language grammatical acceptability (MultiBLiMP) on the languages tested; its benefits are less consistent on knowledge-related tasks (Global-PIQA, Belebele).
Clara Meister, Suchir Salhan, Andrzej Szablewski +3