Tokenizer choice affects multilingual language modeling, but vocabulary capacity is finite and vocabulary size is often constrained: improving representation for some languages often comes at the expense of others. We therefore ask whether tokenizer choice matters equally across languages, a question that the current literature leave unanswered. To this end, we train 123 language models spanning 54 tokenizers. In the main comparison, architecture, training corpus, training-token budget, and optimization are held fixed, so the models differ only in their tokenizer. We find that tokenizer choice matters more for languages with less language-model training data: across the 54 tokenizers, the standard deviation of a language's bits-per-byte (BPB) increases as its model training-data share decreases (Spearman rho = -0.52 over the 31 trained languages and -0.69 over the 28 written with word boundaries). Leaving a language out of tokenizer training raises its BPB in every language we study, and the penalty tends to be larger for languages with less language-model training data. Giving lower-resource languages a larger share of tokenizer-training data, however, does not unconditionally help those languages: both equal weighting and an allocation inverting the shares with respect to the language model training data increase their BPB, particularly when language-model training repeats data. Finally, which intrinsic tokenizer properties are associated with better BPB differs across languages, providing further evidence that what makes a good tokenizer depends on the language. We find that the metrics quantifying these properties can be successfully used to predict downstream models' pairwise BPB rankings, suggesting a practical strategy for screening tokenizer candidates before training language models.
Figures & tables
Figure 1: Left : the standard deviation of a language’s FLORES+ BPB across the 54 tokenizers (BPB spread) vs. the language’s log LM share (Spearman ρ=−0.52 , p=0.003 ). Open diamonds mark the three languages written without word boundaries (Chinese, Japanese, and Thai). Right : the exclusion penalty (defined in the text) against the same axis (Spearman ρ=−0.33 , p=0.068 ). The horizontal line marks a penalty of zero. Dashed lines are ordinary-least-squares fits, and the grey regions are 95% confidence regions for the fitted mean.
Figure 2: Per-language change in FLORES+ BPB relative to reference (the model trained with the GPT-4o BPE reference tokenizer) for the 31 languages of the training mixture, as a function of the language’s log LM share. Filled circles: inverted-allocation tokenizer, in which languages with the smallest nominal sampling weights in the language-model training data receive the largest tokenizer shares. Open triangles: the two permuted allocation tokenizers, in which the tokenizer shares are assigned randomly.
Figure 3: The nine intrinsic metrics we consider and their within-language relationships with FLORES+ BPB (lower ↓ better). Each tile gives the metric’s definition, the direction of the relationship, and the coefficient from the mixed effects model of § 6.1 , in BPB per standard deviation of the metric.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Group
Optimizer
Eff. LR
Betas
WD
Transformer matrices
Muon
0.0283 ∗
mom. schedule
0.050 → 0
Input embeddings
AdamW
0.2121
(0.8, 0.995)
0.001
Output projection
AdamW
0.0057
(0.8, 0.96)
0.01
Value embeddings
AdamW
0.1061
(0.8, 0.995)
0.01
Residual scalars
AdamW
0.0071
(0.8, 0.95)
0.05
Skip scalars
AdamW
0.707
(0.96, 0.95)
0.0
Appendix
Table 2: Effective learning rates and optimizer settings per parameter group of the d24 models. ∗ Muon multiplies the learning rate of a matrix by the square root of its ratio of rows to columns when that ratio exceeds one. The learning rate is therefore 0.0566 for the MLP input projection and 0.0980 for the matrices that project the value embeddings to the attention values. The smear parameters mix the embedding of the previous token into the current position, and the backout parameter subtracts the residual stream of the middle layer before the output layer.
Model
Layers
Width
Heads
Params
nanochat d8
8
512
4
0.22 B
nanochat d12
12
768
6
0.38 B
nanochat d16
16
1024
8
0.60 B
nanochat d24 (main)
24
1536
12
1.27 B
Appendix
Table 3: Smaller proxy models used for ranking transfer ( Table 16 ). All share the d24 training data, tokenizers, optimizer recipe, and a budget of about 3.5 B tokens, which is 3,337 steps for d8 and d12 and 3,337 to 3,370 steps for d16; μ P transfers learning rates across width. Parameters are totals at the 128K vocabulary.
Figure 4: Per-language spread of FLORES+ BPB (top) and MultiBLiMP accuracy (bottom) across the 54 tokenizers, languages in approximately decreasing order of share of the language-model training data (largest at left). Each dot is the standard deviation of the outcome across the 54 tokenizers, and the horizontal tick in the same column, joined to the dot by a thin line, is that language’s seed variation; orange dots are languages where the cross-tokenizer standard deviation does not exceed its seed variation (MultiBLiMP: Catalan, Dutch and Turkish). MultiBLiMP accuracy for a language is estimated from a finite number of minimal pairs, so a fixed model’s measured accuracy varies with a binomial standard error of (p(1−p)/n) for n pairs. An open dot marks a language for which the pooled seed standard deviation does not exceed this standard error: for these languages, run-to-run variation cannot be measured more finely than the benchmark’s own sampling noise. Languages without MultiBLiMP coverage have no dot in the bottom panel.
Contrast (baseline → variant)
Δ FLORES
ratio
Δ Val
ratio
Data
Balanced → English-only (GPT-4o) a
+0.0386
32.69
+0.0332
81.73
Balanced → English-only (Claude) b
+0.0307
26.06
+0.0248
61.00
Balanced → English-only (Punct) b
+0.0238
20.18
+0.0254
62.41
Balanced → Code-heavy (GPT-4o)
+0.0189
16.06
+0.0153
37.56
Balanced → High-resource (GPT-4o)
+0.0202
17.09
+0.00025
0.62
Balanced → High+mid (GPT-4o)
+0.0121
10.29
+0.00059
1.45
Appendix
Table 5: Controlled comparisons: 25 tokenizer pairs that differ in one configuration choice. Δ FLORES: variant minus baseline in the 31-language FLORES+ mean. Δ Val: variant minus baseline in validation BPB. The ratio column after each difference is the difference’s magnitude divided by the unrounded pooled seed standard deviation of that outcome across seed replicates ( § A.5 ). Bold : the difference is larger than run-to-run variation ( § 2.3 ), i.e., the ratio is at least 22≈2.83 ; the factor 2 converts the standard deviation of one run into the standard deviation of a difference between two runs. a full-byte-alphabet rebuild; b tokenizer with the byte-drop defect ( § A.1 )
Variant (vs. balanced baseline)
Δ mean
Δ tail
Spearman ρ
Equal weights (GPT-4o)
+0.0107
+0.0177
−0.61∗
Equal weights (Claude)
+0.0064
+0.0162
−0.52∗
Equal weights (Punct)
+0.0015
+0.0031
−0.07
Equal weights (GPT-4o, NFC)
+0.0054
+0.0099
−0.34
Equal weights, no repeat sampling
+0.0060
+0.0120
−0.44∗
File cap only (control)
−0.0035
−0.0039
+0.17
Appendix
Table 6: Equal vs. proportional tokenizer-data weighting across 30 non-English languages (English fixed at 35%). Δ mean: the 31-language FLORES+ mean of the model with equal weights minus that of the model with proportional weights. Δ tail: the same difference in FLORES+ BPB, averaged over the 15 languages with the smallest proportional weights. Spearman ρ : correlation between a language’s change in FLORES+ BPB and its proportional weight. Bold and a star mark correlations with p<0.05 . Last row: a control that changes only the limit on the number of files.
# of coverage tokenizers that exclude the language
# of languages in group
Difference in 31-language mean
Mean penalty
Ratio
One (English-only)
5
0.0278
0.0671
0.41
Two (English-only, high-resource)
15
0.0233
0.0213
1.09
Three (English-only, high-resource, high-and-mid-resource)
10
0.0236
0.0330
0.72
Appendix
Table 7: For the three groups of languages that the same coverage tokenizers exclude: the difference between the coverage tokenizers that is not specific to a language. Difference in 31-language mean: the average of the 31-language FLORES+ means of the coverage tokenizers that exclude the group’s languages, minus the average of the 31-language FLORES+ means of the coverage tokenizers that include them. Mean penalty: the mean of the exclusion penalties ( § 4 ) of the group’s languages. Ratio: the difference in 31-language mean divided by the mean penalty. English is not included in the analysis below.
Table 8: Spearman correlation, for each reallocated tokenizer, between a language’s change in FLORES+ BPB against the reference run and the language’s log share of the language-model training data (realized share), over the 31 trained languages. A star marks p<0.05 .
Original mixture
Reweighted mixture
Reference tokenizer
0
−0.052
Inverted-allocation tokenizer
+0.267
+0.181
Appendix
Table 9: Tamil’s FLORES+ BPB minus that of the reference model on the original mixture, for the four combinations of tokenizer and language-model training mixture. Reweighted mixture: the mixture on which the model with the inverted-allocation tokenizer is trained on as many tokens of each language as the reference model is on the original mixture.
Repeated
Fresh (single epoch)
Tamil
+0.2674
+0.1386
Bengali
+0.2616
+0.0997
Hebrew
+0.0724
0.0059
Korean
+0.0468
−0.0087
Mean (31 trained languages)
+0.0256
+0.0123
Appendix
Table 10: Per-language change in FLORES+ BPB from the reference run under the inverted-allocation tokenizer, on the repeated training data of the 54 tokenizers’ runs and on single-pass data: the two languages whose fertility ratio fell furthest, two languages whose fertility ratio changed less, and the mean over the 31 training languages. In both settings the inverted-allocation value is the average of two seeds, compared against one reference run; the largest difference between the two seeds in any language’s change is 0.017 on repeated data and 0.019 on single-pass data.
Intrinsic metric
β
SE m
SE cl
padjcl
R2
Fertility
+0.0156
0.0013
0.0027
<10−7
0.082
UTF-8 char split
+0.0113
0.0005
0.0012
<10−16
0.214
Trigram entropy
−0.0112
0.0006
0.0015
<10−13
0.170
Bigram entropy
−0.0083
0.0006
0.0017
<10−5
0.120
Compression rate
−0.0072
0.0007
0.0024
0.003
0.057
Rényi eff. ( α=2 )
−0.0070
0.0005
0.0015
<10−5
0.093
Appendix
Table 11: Within-language relationship between intrinsic metrics and FLORES BPB on the 31 trained languages: β from BPB∼z(m)+(1∣language) , in BPB per standard deviation of the metric ( n=1674 tokenizer–language pairs; 1673 for UTF-8 char split, whose value is unavailable for one (tokenizer, language) entry; 31 languages, 54 tokenizers; null-model language ICC =0.995 ). SE m : model-based standard error. SE cl and padjcl : cluster-robust standard error and BH-adjusted p -value, clustered at the family level ( G=42 families); the clustered p is the inference we rely on. R2 : share of within-language variance explained. Sorted by ∣β∣ . Vocabulary utilization’s raw R2 is marginally negative, a degenerate output reported as 0.000 . All values are for models trained on the same number of tokens; Table 14 reports the coefficients after the adjustment for the number of training bytes.
Figure 5: Per-language Spearman ρ between each intrinsic metric and FLORES+ BPB across the 54 tokenizers, for the 31 languages of the training mixture ordered from the highest to the lowest share of it. Only BH-FDR-significant entries are colored. The boundary-crossing entries for English, Indonesian, and Dutch are undefined (zero variance), and the char-split entry for Korean has n=53 . Hatching marks the 13 entries that are not identified under the measured-composition control of App. E , where the metric and the measured number of training bytes of the language correlate at ∣ρ∣>0.9 across the tokenizer set.
Figure 6: Per-language Spearman ρ between each FineWeb2-measured intrinsic metric and per-language FineWeb2-validation BPB, across the 54 tokenizers, for the 31 trained languages ordered from highest to lowest training-data share. FineWeb2 analogue of Fig. 5 .
Figure 7: Per-language Spearman ρ between each intrinsic metric (measured on FLORES+ text, as in Fig. 5 ) and per-language MultiBLiMP accuracy, across the 54 tokenizers, for the 24 covered trained languages ordered from highest to lowest training-data share. The boundary-crossing entries for English and Dutch are undefined (zero variance).
Intrinsic metric
β
SE m
SE cl
padjcl
R2
Fertility
+0.0322
0.0014
0.0022
<10−47
0.244
Trigram entropy
−0.0168
0.0005
0.0016
<10−26
0.458
Unigram entropy
−0.0151
0.0005
0.0020
<10−13
0.329
Compression rate
−0.0151
0.0007
0.0027
<10−7
0.240
Bigram entropy
−0.0145
0.0004
0.0017
<10−17
0.396
UTF-8 char split
+0.0140
0.0005
0.0008
<10−71
0.344
Appendix
Table 12: Within-language relationship between intrinsic metrics and per-language FineWeb2-validation BPB on the 31 trained languages, with both sides measured on FineWeb2 text: β from BPB∼z(m)+(1∣language) , in BPB per standard deviation of the metric ( n=1674 tokenizer–language pairs; 31 languages, 54 tokenizers; null-model language ICC =0.994 ). Columns as in Table 11 : model-based and family-clustered standard errors ( G=42 ), BH-adjusted clustered p -value, share of within-language variance explained. Sorted by ∣β∣ .
Agg. ρ
Median lang. ρ
Sig.
Intrinsic metric
tr.-31
all-214
trained
untrained
tr.
untr.
Rényi eff. ( α=2 )
−0.455
−0.193
−0.288
−0.072
9
11
Trigram entropy
−0.446
+0.130
−0.328
+0.080
15
34
Compression rate
−0.326
+0.227
−0.226
+0.106
9
56
UTF-8 char split
+0.331
−0.194
+0.348
−0.047
17
34
Bigram entropy
−0.286
+0.052
−0.214
+0.024
6
40
Appendix
Table 13: Per-language decomposition of the full FLORES+ language set, across the 54 tokenizers. Agg. ρ : Spearman between the tokenizer’s global metric value and its mean per-language FLORES+ BPB over the 31 trained languages (tr.) and over all 214 languages. Median lang. ρ : median of the per-language correlations (global metric value against one language’s BPB) over the trained and the untrained languages. Sig.: languages with a significant per-language correlation (BH-FDR within metric), of 31 trained and of 183 untrained.
unadjusted
estimate of β used in the adjustment
Metric
−0.0186
−0.0257
−0.0365
−0.0397
Fertility
+0.0154∗
+0.0154∗
+0.0147∗
+0.0135∗
+0.0135∗
UTF-8 char split
+0.0113∗
+0.0113∗
+0.0108∗
+0.0099∗
+0.0099∗
Trigram entropy
−0.0110∗
−0.0109∗
−0.0095∗
−0.0070∗
−0.0072∗
Bigram entropy
−0.0083∗
−0.0082∗
−0.0067∗
−0.0039∗
−0.0041∗
Compression rate
−0.0071∗
−0.0070∗
−0.0055
−0.0029
−0.0031
Appendix
Table 14: Within-language coefficients of Table 11 after each run’s BPB is adjusted for its number of training bytes ( § E.3 ), for the 51 tokenizers with a measured number of training bytes. In each column the adjustment is made with one estimate of β , the change in the 31-language FLORES+ mean per unit increase in the natural logarithm of the number of training bytes: from left to right, no adjustment, the SuperBPE byte-matched retrain, the thirteen repeated-data continuations, the UnigramLM byte-matched retrain, and the eleven single-pass continuations. ∗ : BH-adjusted family-clustered p<0.05 . The unadjusted column is the fit of Table 11 on the same 51 tokenizers, so its values differ slightly from that table, which uses all 54.
Intrinsic metric
High
Mid
Low
Trigram entropy
−0.0111
−0.0084
−0.0128
UTF-8 char split
+0.0215
+0.0074
+0.0119
Vocab utilization
−0.0052
+0.0012
+0.0112
Rényi eff. ( α=2 )
−0.0002
−0.0065
−0.0105
Fertility
+0.0539
+0.0070
+0.0099
Bigram entropy
−0.0060
−0.0077
−0.0094
Appendix
Table 15: Within-language coefficient β of each intrinsic metric on FLORES+ BPB, fit separately within resource tiers (high n=6 , mid n=15 , low n=10 languages, by share of the language-model training data), in BPB per standard deviation of the metric. Sorted by the low-tier ∣β∣ . The coefficients are the mixed-effects tier fits. Bold : BH-adjusted padj<0.05 under the family-clustered tier fit, an OLS fit with language fixed effects clustered by family ( G=42 ; § A.5 ); the compression-rate high-tier entry is not significant ( padj=0.063 ), and it is also not significant after the adjustment for the number of training bytes ( § E.3 ).
Ranking method
Val BPB
FLORES-tr.
Majority-class baseline
0.601
0.565
All nine metrics (logistic ridge)
0.700
0.735
Best single metric, selected per fold
0.655
0.668
Bigram entropy (fixed)
0.683
0.676
Trigram entropy (fixed)
0.639
0.664
Rényi eff. ( α=2 , fixed)
0.675
0.653
Appendix
Table 16: Pairwise accuracy at ordering held-out tokenizer pairs by the 1.27B models’ aggregate BPB, under leave-one-family-out cross-validation over the 42 tokenizer families ( 40 for the validation-BPB target, which excludes the three SCRIPT tokenizers; § A.5 ). Val BPB: validation BPB. FLORES-tr.: the 31-language FLORES+ mean. Tokenizer-level rows (second block) use one value of each metric per tokenizer. Per-language rows (third block) use the same metrics measured on each of the 31 languages ( 278 features; § E.4 states the excluded column); a single-metric row of that block uses all per-language measurements of that metric as the features. The majority-class baseline is the accuracy of always predicting the more frequent pair order. Each single-metric row uses one metric fixed in advance; selecting the best single-metric row from this table afterward is a choice made in hindsight, so we recommend the fixed nine-metric combination for screening. The three smaller-model rows are direct ranking transfer, not cross-validation: the 17 tokenizers with runs at all four scales are ranked by the smaller model’s validation BPB and the fraction of pairs ordered as by the 1.27B outcome is reported.
Intrinsic metric
β
SE m
SE cl
padjcl
UTF-8 char split
−0.0066
0.0006
0.0014
<10−5
Fertility
−0.0056
0.0007
0.0012
<10−5
UTF-8 boundary crossing
−0.0046
0.0009
0.0009
<10−5
Trigram entropy
+0.0040
0.0007
0.0010
<10−4
Bigram entropy
+0.0035
0.0006
0.0009
<10−3
Rényi eff. ( α=2 )
+0.0031
0.0006
0.0007
<10−4
Appendix
Table 17: Within-language relationship between intrinsic metrics and per-language MultiBLiMP accuracy on the 24 covered trained languages: β from accuracy∼z(m)+(1∣language) , in accuracy per standard deviation of the metric ( n=1296 tokenizer–language pairs; 24 languages, 54 tokenizers). SE m , SE cl , and padjcl as in Table 11 ( G=42 families; the clustered p -value is the inference we rely on). Sorted by ∣β∣ .
Intrinsic metric
β
SE m
SE cl
padjcl
Fertility
−0.0043
0.0012
0.0009
<10−5
Trigram entropy
+0.0042
0.0008
0.0010
<10−4
UTF-8 char split
−0.0030
0.0006
0.0007
<10−4
Compression rate
+0.0023
0.0008
0.0011
0.07
UTF-8 boundary crossing
−0.0020
0.0006
0.0004
<10−5
Unigram entropy
+0.0019
0.0007
0.0010
0.08
Appendix
Table 18: Within-language relationship between intrinsic metrics and per-language XNLI accuracy on the 13 covered trained languages: β from accuracy∼z(m)+(1∣language) , in accuracy per standard deviation of the metric ( n=702 tokenizer–language pairs; 13 languages, 54 tokenizers). SE m , SE cl , and padjcl as in Table 11 ( G=42 families; the clustered p -value is the inference we rely on). Sorted by ∣β∣ .
Tokenizers provide the fundamental basis through which text is represented and processed by language models (LMs). Despite the importance of tokenization, its role in LM performance and behavior is poorly understood due to the challenge of measuring the impact of tokenization in isolation. To address this need, we present TokSuite, a collection of models and a benchmark that supports research into tokenization's influence on LMs. Specifically, we release fourteen pre-trained models that use different off-the-shelf tokenizers but are otherwise identical, using the same architecture, dataset, training budget, and initialization. We also release a multilingual robustness benchmark that measures model performance under real-world perturbations in English, Chinese, Farsi, Italian, and Turkish, curated by native annotators. Together, TokSuite allows robust decoupling of the influence of a model's tokenizer, supporting a series of novel findings that elucidate the respective benefits and shortcomings of a wide range of popular tokenizers.
Gül Sena Altıntaş, Malikeh Ehghaghi, Brian Lester +4
University of Toronto · Vector Institute · Google DeepMind +4
Multilingual large language models (LLMs) depend on subword tokenization to bridge discrete text and continuous neural representation. State-of-the-art multilingual LLMs often use Byte-level Byte-Pair Encoding (BPE) tokenizers that structurally favor high-resource languages and Latin scripts. For speakers of underrepresented languages, particularly those across Southeast Asia, this bias inflates inference costs and widens cross-lingual capability gaps. We present the first systematic comparison of equitable tokenizers on a unified benchmark spanning 11 Southeast Asian languages. Beyond tokenizer-level analysis of compression efficiency and cross-lingual equity, we assess downstream task performance through controlled 1.5B-parameter language model training using the same training data. Our results show that Parity-aware BPE lies on the Pareto frontier of the efficiency-equity trade-off, achieving strong compression parity at competitive cost. Morphology-Driven Byte Encoding delivers the best semantic reasoning performance through morphologically richer representations, albeit at a higher computational expense. Byte Latent Transformer underperforms on downstream tasks, possibly because its architectural assumptions misalign with the constraints of limited low-resource training data. Together, our findings demonstrate that cross-lingual fairness and tokenization efficiency are not fundamentally at odds, and offer practical guidance for designing equitable multilingual models.
Kieron Seven Jun Wei Lee, Muhammad Reza Qorib, Andrew Ivan Soegeng +1
National University of Singapore · Carnegie Mellon University · SAP
Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. We introduce TokEval, a framework of tokenizer evaluation metrics that goes beyond standard measures like fertility and compression rate to capture linguistically and structurally meaningful properties, e.g., UTF-8 character boundary integrity and digit place-value boundary alignment for mathematics. To validate whether these metrics are predictive of downstream model performance, we conduct controlled language model pretraining experiments, varying solely the tokenizers' training data mixture, pretokenization strategy, and training algorithm. We evaluate the resulting models on bits-per-byte (a tokenizer-agnostic version of perplexity) and several benchmarks, spanning linguistic understanding, mathematical reasoning, and code generation. Our experiments suggest that different intrinsic properties have different impacts on model abilities: information-theoretic metrics predict language modeling abilities (Spearman rho up to 0.80), while structure-sensitive metrics, such as those measuring digit and line-break handling, correlate with task accuracy. We hope TokEval enables more principled tokenizer evaluation, replacing pretraining sweeps with intrinsic measurement wherever the two agree.