Tokenizers are commonly optimized for compression, but a compact vocabulary does not necessarily distribute its capacity evenly across languages. We introduce the Latent Core Tokenizer (LCT), a language-agnostic approach that separates structural discovery from vocabulary construction. LCT uses Minimum Description Length, entropy-based boundary signals, and morphotactic constraints to identify reusable linguistic units before constructing a shared vocabulary. Across 104 languages with a 200K-token vocabulary, LCT achieves lower fertility and higher MorphScore than BPE, Unigram, and parity-aware BPE, while maintaining comparable cross-lingual disparity in tokenization cost. Across four multilingual downstream benchmarks, LCT improves aggregate score by 1.48, 1.83, and 2.00 points over BPE, Unigram, and parity-aware BPE, respectively. Our findings show that compression alone does not predict representation quality and highlight the importance of morphology-driven structural discovery and how frequency is used to allocate the final vocabulary across languages.
Figures & tables
Comp. Rate ↑
TTR ↑
Vocab. Util. ↑
Fertility ↓
Rényi ( α=2.5 ) ↑
MorphScore F1 ↑
Gini ↓
LCT-MEM
4.866
0.1781
0.2526
1.999
0.4472
0.5209
0.1341
BPE
4.295
0.1349
0.2168
2.241
0.4613
0.4410
0.1342
Parity-Aware BPE
3.340
0.0630
0.1301
2.742
0.4879
0.3186
0.1563
Unigram
3.195
0.0878
0.1898
2.920
0.3142
0.3109
0.1435
Table 1: Intrinsic tokenizer quality at 200k vocabulary. LCT-MEM denotes the LCT configuration combining MDL, branching entropy, and morphotactic refinement. Per-language results are displayed in Table 8 and Table 9 .
Figure 1: Intrinsic tokenizer performance across low-, mid-, and high-resource language tertiles. Bars show macro-averages and error bars indicate 95% confidence intervals across languages. MorphScore uses 65 languages, while fertility and vocabulary utilization use the 93 languages covered by TokEval. All configurations use a 200k vocabulary.
Rank
Fertility ↓
MorphScore ↑
Gini ↓
# 200k
# 256k
System
200k
256k
200k
256k
200k
256k
1
1
LCT-MEM τ=0.2
2.023
1.935
0.5249
0.5556
0.1369
0.1409
2
2
LCT-MEM τ=0
2.017
1.929
0.5209
0.5507
0.1341
0.1384
3
3
LCT-MEM τ=0.4
2.047
1.960
0.5244
0.5543
0.1374
0.1422
4
4
LCT-ME τ=0.8
2.053
1.963
0.5233
0.5549
0.1536
0.1570
5
8
LCT-MDL τ=0
1.962
1.876
0.4880
0.5245
0.1402
0.1439
Table 2: Top 10 tokenizer configurations at 200k and 256k vocabulary sizes. Additional intrinsic evaluation results are provided in Figure 3 .
Benchmark
Aggregate accuracy (%)
Language
XNLI
Belebele
XStoryCloze
PAWS-X
LCT-MDL τ=0
LCT-MDL τ=0.2
LCT-MEM τ=0
LCT-MEM τ=0.2
BPE
Unigram
Parity-Aware BPE
Arabic
✓
✓
✓
–
42.14
42.65
41.90
42.10
42.46
41.92
41.79
Basque
–
✓
✓
–
42.71
42.83
42.78
43.10
42.76
41.34
42.78
Bulgarian
✓
✓
–
–
39.27
40.25
39.68
40.18
38.98
37.84
38.99
Chinese
✓
✓
✓
✓
49.89
50.04
50.11
51.40
47.30
47.77
45.99
English
✓
✓
✓
✓
62.20
62.33
61.63
62.39
61.70
61.45
61.38
Table 3: Aggregate downstream accuracy (%) for languages covered by at least 2 of the four multilingual benchmarks. A checkmark identifies the benchmarks included for each language. For each tokenizer and language, accuracy is averaged equally across the checked tasks within each fine-tuning seed and then averaged over five seeds. All configurations use one pretrained encoder seed and a 200k vocabulary. The best result in each row is bold.
Tokenizer
Score
LCT-MDL τ=0
LCT-MDL τ=0.2
LCT-MEM τ=0
LCT-MEM τ=0.2
BPE
Unigram
Parity-Aware BPE
LCT-MDL τ=0
32.84±0.38
−0.80∗
−0.24
−1.45∗∗,‡
+0.03
+0.38
+0.55
LCT-MDL τ=0.2
33.64±0.32
0.54
+0.56
−0.65∗
+0.83∗
+1.18∗∗
+1.35∗∗
LCT-MEM τ=0
33.08±0.64
0.65
0.78
−1.21∗
+0.27
+0.62
+0.79
LCT-MEM τ=0.2
34.29±0.33
0.23
0.35
0.70
+1.48∗∗,†
+1.83∗∗,†
+2.00∗∗
BPE
32.81±0.26
0.39
0.51
0.40
0.43
+0.35
+0.52
Unigram
32.46±0.22
0.38
0.50
0.77
0.42
0.38
+0.17
Table 4: Pairwise comparison of the aggregated downstream score (accuracy %). Score is each configuration’s mean and standard deviation over five fine-tuning seeds. Rows and columns contain the same seven configurations in the same order. All configurations use a 200k vocabulary and one pretrained encoder each. Values above the diagonal report the paired difference between the row and column configurations, with positive values favouring the row. Differences are paired by seed ( df=4 ) and Holm-corrected across the 21 unordered pairs. ∗ and ∗∗ indicate raw paired-test p<0.05 and p<0.01 , while † and ‡ indicate Holm-adjusted p<0.05 and p<0.01 , respectively. Values that remain significant after correction are bold.
Table 6: The systems compared. All share the 200,024-identifier budget, the byte-level pre-tokenizer, the encoder corpus and the same 41,347,986 lines of tokenizer training text; they differ only in how the 199,763 learned pieces are chosen.
Table 7: LCT’s latent discovery reads at most 100k lines per language.
Figure 3: Intrinsic evaluation results for the 200k and 256k vocabulary configurations.
Fertility ↓
MorphScore F1 ↑
Gini ↓
LCT-M τ=0
BPE
P-base
Uni.
LCT-M τ=0
BPE
P-base
Uni.
LCT-M τ=0
BPE
P-base
Uni.
Afrikaans
1.431
1.376
1.482
1.772
0.397
0.354
0.270
0.332
3.67
-0.23
-4.26
-2.99
Albanian, Tosk
1.785
1.775
1.804
2.254
0.331
0.236
0.250
0.196
5.06
0.48
-4.58
5.74
Amharic †
70.610
50.665
77.457
62.488
16.89
1.23
7.45
-0.93
Arabic, Standard
1.618
1.513
1.651
2.056
-4.49
-5.60
-5.05
-5.00
Armenian
2.002
2.113
3.035
3.041
0.412
0.372
0.319
0.194
-5.02
-5.63
-4.80
-5.01
Appendix
Table 8: Intrinsic evaluation (Fertility, MorphScore and Gini) by language
Comp. Rate ↑
TTR ↑
Vocab. Util. ↑
Rényi ↑
LCT-M τ=0
BPE
P-base
Uni.
LCT-M τ=0
BPE
P-base
Uni.
LCT-M τ=0
BPE
P-base
Uni.
LCT-M τ=0
BPE
P-base
Uni.
Afrikaans
3.917
4.128
3.770
3.435
0.270
0.286
0.252
0.221
3.57
3.58
3.45
3.32
0.356
0.353
0.361
0.322
Albanian, Tosk
3.858
3.958
3.826
2.813
0.260
0.276
0.265
0.176
4.25
4.39
4.35
3.93
0.328
0.325
0.329
0.228
Amharic
2.798
3.899
2.550
3.161
0.097
0.108
0.018
0.136
2.83
2.25
0.57
3.50
0.174
0.197
0.245
0.178
Arabic, Standard
5.487
5.987
5.354
4.548
0.370
0.410
0.327
0.283
4.63
4.69
4.19
4.27
0.407
0.414
0.418
0.323
Armenian
7.368
6.982
4.873
4.837
0.265
0.230
0.104
0.179
4.18
3.82
2.48
4.28
0.389
0.398
0.396
0.256
Appendix
Table 9: Intrinsic evaluation (Comp. Rate, TTR, Vocab. Util. and Rényi) by language
Accuracy
Δ vs BPE
Language
mdl 0
mdl 0.2
mem 0
mem 0.2
BPE
Uni.
mdl 0
mdl 0.2
mem 0
mem 0.2
XNLI , all 15 languages, grouped by family
Indo-European (9)
bulgarian
48.97
50.48
49.56
47.90
48.89
47.28
+0.08
+1.59
+0.67
−0.99
english
68.14
67.64
68.07
68.34
68.59
69.30
−0.45
−0.96
−0.53
−0.26
french
50.18
49.82
50.06
48.75
53.64
46.61
−3.47
−3.82
−3.58
−4.89
Appendix
Table 10: Per-language test accuracy (%) on every evaluated language, averaged over five fine-tuning seeds. mdl and mem are the twoLCT objectives, mem abbreviating mdl_ent_morph , and the number beside each is τ ; Uni. is Unigram. Chance is 33.3 on XNLI, 25.0 on Belebele, 50.0 on XStoryCloze and 50.0 on PAWS-X, so margins are not comparable across the sections and the Belebele ones are small in absolute terms. Δ is the paired difference against BPE over the five shared seeds. Green marks a difference in LCT’s favour and red against it; only differences that clear a two-sided paired t test at p<0.05 across the five seeds are shaded, and saturation scales with the size of the difference up to 5 points. Unshaded cells are within seed-to-seed noise. Languages are grouped by genealogical family, ordered by how many of the 104 pretraining languages each family contributes. Belebele covers the 79 test languages present in the pretraining corpus. All arms share one pretraining seed.
Figure 4: Tokenization examples
Figure 5: Aggregate downstream accuracy across low-, mid-, and high-resource language tertiles. Scores are averaged across the available XNLI, Belebele, XStoryCloze, and PAWS-X evaluations and five fine-tuning seeds. Error bars indicate 95% confidence intervals across languages.