Language-Specific Effects of Tokenizer Choice in Multilingual Language Models
Organizations: EPFL · University of Toronto, Vector Institute
Abstract
Tokenizer choice affects multilingual language modeling, but vocabulary capacity is finite and vocabulary size is often constrained: improving representation for some languages often comes at the expense of others. We therefore ask whether tokenizer choice matters equally across languages, a question that the current literature leave unanswered. To this end, we train 123 language models spanning 54 tokenizers. In the main comparison, architecture, training corpus, training-token budget, and optimization are held fixed, so the models differ only in their tokenizer. We find that tokenizer choice matters more for languages with less language-model training data: across the 54 tokenizers, the standard deviation of a language's bits-per-byte (BPB) increases as its model training-data share decreases (Spearman rho = -0.52 over the 31 trained languages and -0.69 over the 28 written with word boundaries). Leaving a language out of tokenizer training raises its BPB in every language we study, and the penalty tends to be larger for languages with less language-model training data. Giving lower-resource languages a larger share of tokenizer-training data, however, does not unconditionally help those languages: both equal weighting and an allocation inverting the shares with respect to the language model training data increase their BPB, particularly when language-model training repeats data. Finally, which intrinsic tokenizer properties are associated with better BPB differs across languages, providing further evidence that what makes a good tokenizer depends on the language. We find that the metrics quantifying these properties can be successfully used to predict downstream models' pairwise BPB rankings, suggesting a practical strategy for screening tokenizer candidates before training language models.
Figures & tables
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Group | Optimizer | Eff. LR | Betas | WD |
|---|---|---|---|---|
| Transformer matrices | Muon | 0.0283 ∗ | mom. schedule | 0.050 0 |
| Input embeddings | AdamW | 0.2121 | (0.8, 0.995) | 0.001 |
| Output projection | AdamW | 0.0057 | (0.8, 0.96) | 0.01 |
| Value embeddings | AdamW | 0.1061 | (0.8, 0.995) | 0.01 |
| Residual scalars | AdamW | 0.0071 | (0.8, 0.95) | 0.05 |
| Skip scalars | AdamW | 0.707 | (0.96, 0.95) | 0.0 |
| Model | Layers | Width | Heads | Params |
|---|---|---|---|---|
| nanochat d8 | 8 | 512 | 4 | B |
| nanochat d12 | 12 | 768 | 6 | B |
| nanochat d16 | 16 | 1024 | 8 | B |
| nanochat d24 (main) | 24 | 1536 | 12 | B |
| Contrast (baseline variant) | FLORES | ratio | Val | ratio | |
|---|---|---|---|---|---|
| Data | Balanced English-only (GPT-4o) a | ||||
| Balanced English-only (Claude) b | |||||
| Balanced English-only (Punct) b | |||||
| Balanced Code-heavy (GPT-4o) | |||||
| Balanced High-resource (GPT-4o) | |||||
| Balanced High+mid (GPT-4o) |
| Variant (vs. balanced baseline) | mean | tail | Spearman |
|---|---|---|---|
| Equal weights (GPT-4o) | |||
| Equal weights (Claude) | |||
| Equal weights (Punct) | |||
| Equal weights (GPT-4o, NFC) | |||
| Equal weights, no repeat sampling | |||
| File cap only (control) |
| # of coverage tokenizers that exclude the language | # of languages in group | Difference in 31-language mean | Mean penalty | Ratio |
|---|---|---|---|---|
| One (English-only) | 5 | |||
| Two (English-only, high-resource) | 15 | |||
| Three (English-only, high-resource, high-and-mid-resource) | 10 |
| Label | Variant | Spearman |
|---|---|---|
| L2 | inverted-allocation tokenizer (two seeds averaged) | |
| P1 | permuted allocation 1 | |
| P2 | permuted allocation 2 | |
| E1 | no English web text |
| Original mixture | Reweighted mixture | |
|---|---|---|
| Reference tokenizer | ||
| Inverted-allocation tokenizer |
| Repeated | Fresh (single epoch) | |
|---|---|---|
| Tamil | ||
| Bengali | ||
| Hebrew | ||
| Korean | ||
| Mean (31 trained languages) |
| Intrinsic metric | SE | SE | |||
|---|---|---|---|---|---|
| Fertility | |||||
| UTF-8 char split | |||||
| Trigram entropy | |||||
| Bigram entropy | |||||
| Compression rate | |||||
| Rényi eff. ( ) |
| Intrinsic metric | SE | SE | |||
|---|---|---|---|---|---|
| Fertility | |||||
| Trigram entropy | |||||
| Unigram entropy | |||||
| Compression rate | |||||
| Bigram entropy | |||||
| UTF-8 char split |
| Agg. | Median lang. | Sig. | ||||
|---|---|---|---|---|---|---|
| Intrinsic metric | tr.-31 | all-214 | trained | untrained | tr. | untr. |
| Rényi eff. ( ) | 9 | 11 | ||||
| Trigram entropy | 15 | 34 | ||||
| Compression rate | 9 | 56 | ||||
| UTF-8 char split | 17 | 34 | ||||
| Bigram entropy | 6 | 40 | ||||
| unadjusted | estimate of used in the adjustment | ||||
|---|---|---|---|---|---|
| Metric | |||||
| Fertility | |||||
| UTF-8 char split | |||||
| Trigram entropy | |||||
| Bigram entropy | |||||
| Compression rate | |||||
| Intrinsic metric | High | Mid | Low |
|---|---|---|---|
| Trigram entropy | |||
| UTF-8 char split | |||
| Vocab utilization | |||
| Rényi eff. ( ) | |||
| Fertility | |||
| Bigram entropy |
| Ranking method | Val BPB | FLORES-tr. |
| Majority-class baseline | ||
| All nine metrics (logistic ridge) | ||
| Best single metric, selected per fold | ||
| Bigram entropy (fixed) | ||
| Trigram entropy (fixed) | ||
| Rényi eff. ( , fixed) |
| Intrinsic metric | SE | SE | ||
|---|---|---|---|---|
| UTF-8 char split | ||||
| Fertility | ||||
| UTF-8 boundary crossing | ||||
| Trigram entropy | ||||
| Bigram entropy | ||||
| Rényi eff. ( ) |
| Intrinsic metric | SE | SE | ||
|---|---|---|---|---|
| Fertility | ||||
| Trigram entropy | ||||
| UTF-8 char split | ||||
| Compression rate | ||||
| UTF-8 boundary crossing | ||||
| Unigram entropy |