Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities
Organizations: Samsung Research
Abstract
Byte-level byte-pair encoding (BBPE) tokenizers are attractive for multilingual large language models (LLMs) because they cover all Unicode text. In UTF-8-based BBPE, however, many scripts start from a higher fallback cost than English: when no learned merges can be applied, a multibyte character requires multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor. A higher floor can increase token counts and per-request cost and shrink usable context. Changing the text encoding can reduce this gap, but a single global encoding can make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer that keeps 1-2-byte UTF-8 characters on the UTF-8 path while routing 3-4-byte UTF-8 characters through UTF-16. This lowers the encoding floor for 3-byte Basic Multilingual Plane (BMP) characters in scripts with high token premiums (token counts relative to English) without raising it for already-efficient spans in mixed-script text. UBE changes only the byte representation presented to byte-pair encoding (BPE); the merge rule remains standard, and exact decoding is preserved. UBE also composes with alternative boundary policies and morphology-based representations. In a Unicode 17 audit, UBE exactly round-trips all Unicode scalar values and all inputs in the official normalization, grapheme-break, and emoji test suites. Across intrinsic evaluations, UBE lowers dispersion in English-normalized token-count ratios, reducing cross-lingual token-budget disparity. In multilingual language model (LM) experiments, UBE matches BBPE's LM quality. In the main multilingual settings, UBE reduces token counts most for high-premium scripts and slightly lowers English token counts, yielding more usable context under fixed token budgets and faster prompt processing in content-matched benchmarks.
Figures & tables
| ASCII | Cyrillic/Arabic/Hebrew | CJK, Devanagari, Thai | Non-BMP | |
| BBPE | 1 | 2 | 3 | 4 |
| BBPE16 | 2 | 2 | 2 | 4 |
| UBE | 1 | 2 | 2 | 4 |
| Dataset | Eval. units | BBPE Gini | UBE Gini | Gini | Note |
|---|---|---|---|---|---|
| FLORES-200 | 204 | 0.2183 | 0.1979 | -9.4% | parallel sentences |
| UDHR | 387 | 0.2499 | 0.2221 | -11.1% | public-text stress test |
| SIB-200 | 205 | 0.2297 | 0.2102 | -8.5% | topic classification |
| MGSM | 11 | 0.1841 | 0.1843 | +0.1% | math reasoning |
| Language | Script | BBPE | UBE | % |
|---|---|---|---|---|
| Santali | Ol Chiki | 341.287 | 222.944 | -34.7% |
| Central Atlas Tamazight | Tifinagh | 269.166 | 183.305 | -31.9% |
| Tibetan | Tibetan | 336.804 | 286.737 | -14.9% |
| Arabic | Arabic | 60.600 | 60.407 | -0.3% |
| English | Latin | 31.160 | 31.068 | -0.3% |
| Config | Tok | BPB | BPB Gini | XNLI | Bele | XCOPA | ARC-E | ARC-C | HSwag | PIQA | LAMB |
|---|---|---|---|---|---|---|---|---|---|---|---|
| En-only | BBPE | 1.4909 | 0.1993 | 34.3 | 23.6 | 51.4 | 38.1 | 18.3 | 27.5 | 61.9 | 23.6 |
| UBE | 1.4698 | 0.1989 | 34.4 | 23.0 | 51.5 | 38.5 | 17.0 | 27.8 | 62.7 | 22.5 | |
| 6-lang | BBPE | 1.3680 | 0.1930 | 34.6 | 23.5 | 52.0 | 35.6 | 17.1 | 27.4 | 61.3 | 23.1 |
| UBE | 1.3760 | 0.1880 | 34.9 | 23.2 | 52.2 | 36.7 | 18.4 | 28.0 | 62.0 | 23.1 | |
| 101-lang | BBPE | 1.5174 | 0.1762 | 34.3 | 24.8 | 52.0 | 35.5 | 17.0 | 27.7 | 59.3 | 22.8 |
| UBE | 1.5237 | 0.1739 | 34.4 | 23.7 | 51.8 | 35.9 | 17.8 | 27.9 | 60.1 | 23.1 |
| Fixed component | English | Selected BMP (12) | Overall (204) |
|---|---|---|---|
| Standard | 31.160 31.068 | 178.499 145.838 | 71.889 69.857 |
| SuperBPE-style | 33.143 32.973 | 161.463 130.218 | 74.187 72.137 |
| SCRIPT-style | 31.698 31.652 | 184.743 150.700 | 69.957 68.001 |
| MYTE maps | 109.140 109.094 | 255.737 183.867 | 148.142 141.354 |
| Method | Common-text BPB | BPB Gini | Multi. mean | English mean |
|---|---|---|---|---|
| UBE (ours) | 1.5643 | 0.1937 | 36.8 | 33.0 |
| BBPE | 1.5908 | 0.2006 | 36.6 | 32.7 |
| BBPE16 | 1.5561 | 0.1996 | 36.3 | 32.5 |
| MYTE+BPE (derived 32K) | 1.6474 | 0.2029 | 36.5 | 32.2 |
| SCRIPT-BPE | 1.5401 | 0.1951 | 36.4 | 32.9 |
Appendix figures & tables51 assets
Supplementary material from the paper’s appendix.
Appendix
| Vocab size | BBPE merges | UBE merges | UBE overhead |
|---|---|---|---|
| 16K | 15,739 | 15,483 | 1.6% |
| 32K | 31,739 | 31,483 | 0.8% |
| 64K | 63,739 | 63,483 | 0.4% |
| 128K | 127,739 | 127,483 | 0.2% |
| 256K | 255,739 | 255,483 | 0.1% |
| Dataset | Original Gini | Det. Gini | Gini | mean premium |
|---|---|---|---|---|
| FLORES-200 | 0.2522 | 0.2517 | -0.00050 | -0.0028 |
| UDHR | 0.2473 | 0.2471 | -0.00025 | +0.0003 |
| SIB-200 | 0.2600 | 0.2595 | -0.00050 | -0.0025 |
| MGSM | 0.3120 | 0.3120 | +0.00002 | -0.0004 |
| Setting | Runs | Layers/heads/width | Params | Batch recipe | Global tok/step | Role |
|---|---|---|---|---|---|---|
| 182M | 6 | 14/12/768 | 182M | 32 1 | 524,288 | full coverage matrix (200K steps, 105B tok) |
| 329M | 4 | 14/12/768 | 329M | 8 4 | 524,288 | 128K composition check (200K steps, 105B tok) |
| 1.3B | 6 | 24/16/2048 | 1.35B actual | 8 1 | 131,072 | scale verification (205K steps, 26.9B tok) |
| Tokenizer coverage | BBPE train | UBE train | BBPE eval | UBE eval | Steps/epoch | 200K coverage |
|---|---|---|---|---|---|---|
| En-only on 101-lang mC4 | 184.0B | 166.9B | 7.5B | 7.2B | 350,991 / 318,320 | 57.0% / 62.8% |
| 6-lang on 101-lang mC4 | 120.0B | 119.0B | 7.0B | 6.9B | 228,967 / 227,033 | 87.3% / 88.1% |
| 101-lang on 101-lang mC4 | 108.1B | 107.8B | 5.3B | 5.3B | 206,142 / 205,695 | 97.0% / 97.2% |
| Config | Langs | Train | Eval | |||
|---|---|---|---|---|---|---|
| Shards | GiB | Docs | Shards | Docs | ||
| En-only | 1 | 118 | 101 | 32.1M | 1 | 0.27M |
| 6-lang | 6 | 1,330 | 253 | 62.3M | 6 | 0.46M |
| 101-lang | 101 | 2,273 | 346 | 89.3M | 101 | 4.36M |
| Language | Region | Speakers | BBPE | UBE | % | Script |
|---|---|---|---|---|---|---|
| Amharic | Ethiopia | 57M | 202.775 | 137.539 | -32.2% | Ethiopic |
| Lao | Laos | 30M | 342.100 | 233.445 | -31.8% | Lao |
| Santali | India | 8M | 341.287 | 222.707 | -34.7% | Ol Chiki |
| Telugu | India | 82M | 237.530 | 220.807 | -7.0% | Telugu |
| Kannada | India | 44M | 251.956 | 237.506 | -5.7% | Kannada |
| Tibetan | China/Nepal | 6M | 264.246 | 247.286 | -6.4% | Tibetan |
| Language | Script | 16K | 32K | 64K | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| BBPE | UBE | % | BBPE | UBE | % | BBPE | UBE | % | ||
| Amharic | Ethiopic | 219.321 | 140.388 | -36.0% | 202.775 | 137.539 | -32.2% | 163.896 | 134.837 | -17.7% |
| Lao | Lao | 343.176 | 234.068 | -31.8% | 342.100 | 233.445 | -31.8% | 256.355 | 231.798 | -9.6% |
| Santali | Ol Chiki | 363.850 | 230.685 | -36.6% | 341.287 | 222.707 | -34.7% | 341.274 | 213.132 | -37.5% |
| Telugu | Telugu | 263.163 | 224.362 | -14.7% | 237.530 | 220.807 | -7.0% | 184.658 | 181.593 | -1.7% |
| Kannada | Kannada | 354.337 | 238.641 | -32.7% | 251.956 | 237.506 | -5.7% | 220.944 | 219.524 | -0.6% |
| Language | Script | Vocab | BBPE | UBE | % |
|---|---|---|---|---|---|
| Santali | Ol Chiki | 16K | 341.312 | 237.326 | -30.467% |
| 32K | 341.287 | 222.944 | -34.676% | ||
| 64K | 341.271 | 212.692 | -37.677% | ||
| 128K | 251.125 | 211.712 | -15.695% | ||
| 256K | 228.583 | 207.379 | -9.276% | ||
| Central Atlas Tamazight | Tifinagh | 16K | 269.200 | 192.112 | -28.636% |
| Evaluation suite | Eval sentences | Exact overlaps | Prefix overlaps |
|---|---|---|---|
| FLORES-200 | 206,448 | 0 | 53 |
| UDHR | 35,870 | 1 | 7,102 |
| SIB-200 | 41,820 | 0 | 10 |
| MGSM | 2,750 | 0 | 0 |
| Config | Vocab | UBE wins/ties/losses | Gini |
|---|---|---|---|
| En-only | 16K | 37 / 19 / 148 | -13.9% |
| En-only | 32K | 41 / 23 / 140 | -11.9% |
| En-only | 64K | 37 / 82 / 85 | -7.3% |
| 6-lang | 16K | 202 / 0 / 2 | -8.7% |
| 6-lang | 32K | 191 / 0 / 13 | -7.2% |
| 6-lang | 64K | 169 / 24 / 11 | -6.2% |
| Vocab | BBPE Gini | UBE Gini | Gini | UBE wins |
|---|---|---|---|---|
| 16K | 0.2194 | 0.1927 | -12.1% | 196/204 |
| 32K | 0.2183 | 0.1979 | -9.4% | 189/204 |
| 64K | 0.2225 | 0.2043 | -8.2% | 187/204 |
| 128K | 0.2258 | 0.2140 | -5.2% | 174/204 |
| 256K | 0.2246 | 0.2204 | -1.9% | 150/204 |
| Tokenizer | EN avg | Latin avg | Selected BMP (12) | All avg | Gini |
|---|---|---|---|---|---|
| BBPE 16K | 34.863 | 62.373 | 200.188 | 79.408 | 0.21938 |
| BBPE 32K | 31.160 | 56.743 | 178.499 | 71.889 | 0.21833 |
| BBPE 64K | 28.716 | 52.540 | 161.576 | 66.056 | 0.22254 |
| BBPE 128K | 27.174 | 48.751 | 140.748 | 60.594 | 0.22576 |
| BBPE 256K | 26.272 | 45.100 | 115.565 | 55.389 | 0.22456 |
| UBE 16K | 34.671 | 62.163 | 153.261 | 76.410 | 0.19274 |
| Tokenizer | EN avg | Latin avg | Selected BMP (12) | All avg | Gini |
|---|---|---|---|---|---|
| SuperBPE BBPE 16K | 38.308 | 66.840 | 180.079 | 82.853 | 0.19967 |
| SuperBPE BBPE 32K | 33.143 | 61.177 | 161.463 | 74.187 | 0.19217 |
| SuperBPE BBPE 64K | 29.133 | 55.586 | 151.281 | 66.871 | 0.18996 |
| SuperBPE BBPE 128K | 26.337 | 51.476 | 140.300 | 61.045 | 0.19042 |
| SuperBPE BBPE 256K | 24.357 | 47.479 | 122.457 | 55.220 | 0.18694 |
| SuperBPE UBE 16K | 38.099 | 66.534 | 149.793 | 80.676 | 0.18269 |
| Language | BBPE avg | UBE avg | BBPE16 (128K) avg | UBE/BBPE |
|---|---|---|---|---|
| Top 10: UBE helps most | ||||
| Santali | 341.271 | 212.692 | 222.620 | 0.6232 |
| Central Atlas Tamazight | 269.123 | 182.243 | 183.143 | 0.6772 |
| Tamashek (Tifinagh) | 270.571 | 185.609 | 187.934 | 0.6860 |
| Dzongkha | 286.182 | 265.685 | 267.341 | 0.9284 |
| Tibetan | 264.208 | 247.171 | 247.411 | 0.9355 |
| Script | Language | Codepoint | Floor (bytes/char) | Avg tokens/sent | Save | ||
|---|---|---|---|---|---|---|---|
| UTF-8 | UTF-16 | BBPE | UBE | ||||
| Ol Chiki | Santali | U+1C5E | 3 | 2 | 341.3 | 212.7 | 37.7% |
| Tifinagh | Central Atlas Tamazight | U+2D30 | 3 | 2 | 269.1 | 182.2 | 32.3% |
| Tibetan | Dzongkha | U+0F42 | 3 | 2 | 286.2 | 265.7 | 7.2% |
| Script | Word | Chars | BBPE | UBE | Save | ||
| Base | Tokens | Base | Tokens | ||||
| Ol Chiki | ona | 3 | 9 | 9 | 6 | 6 | 33% |
| Ol Chiki | kemikal | 7 | 21 | 21 | 14 | 12 | 43% |
| Tifinagh | sin | 3 | 9 | 9 | 6 | 6 | 33% |
| Tifinagh | acukula | 7 | 21 | 21 | 14 | 14 | 33% |
| Ol Chiki | (24 chars) | 24 | 66 | 63 | 45 | 39 | 38% |
| Script | Word | Chars | 16K | 32K | 64K | 128K | 256K | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| BBPE | UBE | BBPE | UBE | BBPE | UBE | BBPE | UBE | BBPE | UBE | |||
| Ol Chiki | ona | 3 | 9 | 6 | 9 | 6 | 9 | 6 | 6 | 6 | 6 | 6 |
| Ol Chiki | kemikal | 7 | 21 | 14 | 21 | 13 | 21 | 12 | 14 | 12 | 14 | 10 |
| Tifinagh | sin | 3 | 9 | 6 | 9 | 6 | 9 | 6 | 9 | 6 | 6 | 6 |
| Tifinagh | acukula | 7 | 21 | 14 | 21 | 14 | 21 | 14 | 21 | 14 | 14 | 14 |
| Ol Chiki | (24 chars) | 24 | 63 | 42 | 63 | 40 | 63 | 39 | 45 | 39 | 42 | 38 |
| Script family | Langs | UTF-8 | Metric | 16K | 32K | 64K | 128K | 256K |
|---|---|---|---|---|---|---|---|---|
| Latin | 127 | 1–2 | BBPE | 62.373 | 56.743 | 52.540 | 48.751 | 45.100 |
| UBE | 62.163 | 56.627 | 52.500 | 48.722 | 45.097 | |||
| Arabic | 22 | 2 | BBPE | 78.312 | 69.687 | 62.817 | 56.489 | 51.358 |
| UBE | 77.610 | 69.549 | 62.741 | 56.453 | 51.336 | |||
| Metric | BBPE | UBE | S-BBPE | S-UBE |
|---|---|---|---|---|
| Alphabet tokens | 256 | 512 | 256 | 512 |
| Merge tokens | 31,739 | 31,483 | 31,739 | 31,483 |
| Dead merges | 13 | 9 | 524 | 531 |
| Active utilization | 99.9% | 99.8% | 98.3% | 98.2% |
| Pure PUA–PUA merges | — | 23.6% | — | 24.6% |
| GPT-2–GPT-2 merges | — | 66.6% | — | 64.6% |
| Family | Metric | 16K | 32K | 64K | 128K | 256K |
|---|---|---|---|---|---|---|
| BBPE | Dead merges | 9 | 13 | 20 | 82 | 224 |
| Active util. | 99.825% | 99.900% | 99.939% | 99.921% | 99.905% | |
| Depth-1 | 12.326% | 7.889% | 4.937% | 2.982% | 1.774% | |
| UBE | Dead merges | 6 | 9 | 16 | 78 | 217 |
| Active util. | 99.713% | 99.847% | 99.911% | 99.907% | 99.899% | |
| Pure PUA | 24.758% | 23.587% | 22.297% | 20.924% | 19.641% |
| Dataset | Vocab | BBPE Gini | UBE Gini | Gini |
|---|---|---|---|---|
| FLORES-200 | 16K | 0.2194 | 0.1927 | -12.14% |
| FLORES-200 | 32K | 0.2183 | 0.1979 | -9.36% |
| FLORES-200 | 64K | 0.2225 | 0.2043 | -8.20% |
| FLORES-200 | 128K | 0.2258 | 0.2140 | -5.20% |
| FLORES-200 | 256K | 0.2246 | 0.2204 | -1.86% |
| UDHR | 16K | 0.2449 | 0.2158 | -11.86% |
| Corpus | Vocab | Wilcoxon | Holm | BH | Gini BBPE UBE | Gini 95% CI |
|---|---|---|---|---|---|---|
| En-only | 16K | 0.3495 0.3010 | [-0.0616, -0.0323] | |||
| En-only | 32K | 0.3380 0.2978 | [-0.0537, -0.0244] | |||
| En-only | 64K | 0.3183 0.2950 | [-0.0353, -0.0109] | |||
| 6-lang | 16K | 0.2960 0.2703 | [-0.0376, -0.0139] | |||
| 6-lang | 32K | 0.2717 0.2522 | [-0.0325, -0.0074] | |||
| 6-lang | 64K | 0.2603 0.2442 | [-0.0299, -0.0044] |
| Script cohort | Standard | SuperBPE-style | SCRIPT-style | |
|---|---|---|---|---|
| Latin/Arabic/Cyrillic | 161 | 59.205 59.078 | 64.534 64.262 | 59.832 59.719 |
| South Asian | 21 | 103.401 103.199 | 91.472 91.514 | 77.316 78.385 |
| Han/Japanese | 4 | 43.363 43.364 | 43.455 43.465 | 42.722 42.801 |
| Khmer/Lao/Myanmar/Thai | 5 | 130.013 130.446 | 113.257 115.093 | 126.617 124.854 |
| Remaining | 13 | 164.495 134.332 | 160.236 130.662 | 170.053 139.682 |
| All | 204 | 71.889 69.857 | 74.187 72.137 | 69.957 68.001 |
| Pre-tokenization | Method | English | Selected BMP (12) | Overall | Gini | Lowest |
|---|---|---|---|---|---|---|
| Standard GPT-2 | BBPE | 36.959 | 183.592 | 67.875 | 0.2281 | 6/204 |
| Standard GPT-2 | UBE (ours) | 36.938 | 153.442 | 66.079 | 0.2086 | 162 /204 |
| Official script/category | SCRIPT-BPE | 37.982 | 151.444 | 63.785 | 0.1785 | 25/204 |
| SCRIPT-style | BBPE | 37.569 | 188.228 | 65.677 | 0.2048 | 7/204 |
| SCRIPT-style | UBE (ours) | 37.539 | 154.015 | 63.918 | 0.1832 | 4/204 |
| Evaluation | Script cohort | Comparisons | UBE lower | SCRIPT-BPE lower |
|---|---|---|---|---|
| FLORES-200 | Latin/Arabic/Cyrillic | 161 | 156 | 5 |
| FLORES-200 | South Asian scripts | 21 | 1 | 20 |
| FLORES-200 | Han/Japanese | 4 | 0 | 4 |
| FLORES-200 | Khmer/Lao/Myanmar/Thai | 5 | 0 | 5 |
| FLORES-200 | Remaining scripts | 13 | 12 | 1 |
| FLORES-200 | All | 204 | 169 | 35 |
| Boundary policy | Representation | English | Selected BMP (12) | Overall | Gini |
|---|---|---|---|---|---|
| Standard | BBPE | 36.959 | 183.592 | 67.875 | 0.2281 |
| Standard | UBE (ours) | 36.938 | 153.442 | 66.079 | 0.2086 |
| SuperBPE-style | BBPE | 39.586 | 163.120 | 68.858 | 0.1916 |
| SuperBPE-style | UBE (ours) | 39.524 | 135.050 | 67.208 | 0.1736 |
| Base representation | Alphabet | English | Selected BMP (12) | Overall | Total positions |
|---|---|---|---|---|---|
| UTF-8 bytes | 256 | 130.530 | 292.895 | 196.106 | 40,485,643 |
| UBE (ours) | 512 | 130.481 | 199.638 | 176.088 | 36,353,040 |
| MYTE | 256 | 109.140 | 255.737 | 148.142 | 30,583,555 |
| MYTE-UBE (ours) | 512 | 109.094 | 183.867 | 141.354 | 29,182,333 |
| Unicode 17 input set | Inputs | UBE raw exact | SCRIPT-BPE after NFC | MYTE raw exact |
|---|---|---|---|---|
| Assigned through Unicode 16 (excluding Private-Use) | 155,063 | All recovered | All recovered | 154,023/155,063 |
| Assigned in Unicode 17 | 4,803 | All recovered | Outside released coverage | All recovered |
| Private-Use (Co) | 137,468 | All recovered | Filtered by design | All recovered |
| Unassigned/reserved (Cn) | 814,730 | All recovered | Filtered by design | All recovered |
| NormalizationTest.txt unique inputs | 37,671 | All recovered | 166 outside coverage | 21,373/37,671 |
| GraphemeBreakTest.txt unique inputs | 764 | All recovered | 74 outside coverage | 684/764 |
| Tokenizer Steps Model-input positions Supervised UTF-8 bytes Common-text BPB BPB Gini UBE (ours, primary tokenizer) 200,000 104,857,600,000 361,441,233,445 1.5298 0.1735 BBPE 200,437 105,086,713,856 361,439,652,484 1.5439 0.1740 BBPE16 237,228 124,375,793,664 361,440,525,559 1.5418 0.1882 |
| Tokenizer XNLI XCOPA Bele ARC-E ARC-C HSwag PIQA LAMB English mean UBE (ours, primary tokenizer) 34.4 51.8 23.7 35.9 17.8 27.9 60.1 23.1 33.0 BBPE 34.4 51.5 24.0 35.3 16.8 27.9 60.3 23.3 32.7 BBPE16 34.4 51.9 23.3 35.5 18.0 27.6 60.8 19.5 32.3 |
| Evaluation | Entries | Primary 3-byte BMP | Primary 4-byte non-BMP |
|---|---|---|---|
| XNLI | 15 | 3 | 0 |
| Belebele | 122 | 26 | 0 |
| XCOPA | 11 | 3 | 0 |
| MGSM | 11 | 5 | 0 |
| Each English benchmark | 1 | 0 | 0 |
| Held-out BPB | 101 | 20 | 0 |
| Input content | ASCII-byte share | UBE vs. BBPE | BBPE16 vs. BBPE |
|---|---|---|---|
| Issue text | 73.91% | -0.27% [-0.47%, -0.15%] | +21.92% [+20.56%, +23.19%] |
| Issue text + edit-file contents | 99.24% | -0.18% [-0.21%, -0.15%] | +24.86% [+24.35%, +25.40%] |
| Config | Tok | BPB | BPB Gini | XNLI | Bele | XCOPA | ARC-E | ARC-C | HSwag | PIQA | LAMB |
|---|---|---|---|---|---|---|---|---|---|---|---|
| En-only | BBPE | 1.4909 | 0.1993 | 34.3 | 23.6 | 51.4 | 38.1 | 18.3 | 27.5 | 61.9 | 23.6 |
| En-only | UBE | 1.4698 | 0.1989 | 34.4 | 23.0 | 51.5 | 38.5 | 17.0 | 27.8 | 62.7 | 22.5 |
| 6-lang | BBPE | 1.3680 | 0.1930 | 34.6 | 23.5 | 52.0 | 35.6 | 17.1 | 27.4 | 61.3 | 23.1 |
| 6-lang | UBE | 1.3760 | 0.1880 | 34.9 | 23.2 | 52.2 | 36.7 | 18.4 | 28.0 | 62.0 | 23.1 |
| 101-lang | BBPE | 1.5174 | 0.1762 | 34.3 | 24.8 | 52.0 | 35.5 | 17.0 | 27.7 | 59.3 | 22.8 |
| 101-lang | UBE | 1.5237 | 0.1739 | 34.4 | 23.7 | 51.8 | 35.9 | 17.8 | 27.9 | 60.1 | 23.1 |
| Config | Tok | BPB | BPB Gini | XNLI | Bele | XCOPA | ARC-E | ARC-C | HSwag | PIQA | LAMB |
|---|---|---|---|---|---|---|---|---|---|---|---|
| En-only | BBPE | 1.4193 | 0.1971 | 34.5 | 22.9 | 52.3 | 39.4 | 17.4 | 28.5 | 62.6 | 25.6 |
| En-only | UBE | 1.3668 | 0.1956 | 34.8 | 22.9 | 52.8 | 37.3 | 19.0 | 28.2 | 62.3 | 27.0 |
| 6-lang | BBPE | 1.3071 | 0.1956 | 34.8 | 23.1 | 52.9 | 40.4 | 18.9 | 29.2 | 63.5 | 28.9 |
| 6-lang | UBE | 1.3106 | 0.1973 | 35.0 | 23.9 | 52.6 | 39.6 | 18.2 | 29.2 | 63.0 | 28.6 |
| 101-lang | BBPE | 1.4673 | 0.1972 | 35.0 | 23.0 | 51.1 | 39.1 | 18.1 | 29.1 | 64.0 | 29.9 |
| 101-lang | UBE | 1.4707 | 0.1983 | 34.5 | 23.4 | 51.9 | 39.9 | 18.4 | 29.1 | 63.2 | 29.1 |
| Config | Tok | BPB | BPB Gini | XNLI | Bele | XCOPA | ARC-E | ARC-C | HSwag | PIQA | LAMB |
|---|---|---|---|---|---|---|---|---|---|---|---|
| En-only | BBPE | 1.4879 | 0.2006 | 34.2 | 23.0 | 52.1 | 40.2 | 18.1 | 27.6 | 61.2 | 20.6 |
| En-only | UBE | 1.4588 | 0.2002 | 34.2 | 23.2 | 51.9 | 40.2 | 18.4 | 27.6 | 61.9 | 22.6 |
| 6-lang | BBPE | 1.4329 | 0.1751 | 34.7 | 23.0 | 52.0 | 37.7 | 18.9 | 28.0 | 61.3 | 22.4 |
| 6-lang | UBE | 1.4456 | 0.1812 | 34.5 | 22.9 | 51.6 | 38.1 | 19.3 | 27.8 | 62.1 | 22.8 |
| 101-lang | BBPE | 1.4431 | 0.2037 | 34.8 | 23.2 | 51.5 | 39.5 | 17.6 | 28.1 | 61.4 | 23.2 |
| 101-lang | UBE | 1.4435 | 0.2047 | 34.8 | 23.4 | 51.9 | 38.2 | 18.1 | 27.6 | 61.0 | 24.7 |
| Config | Tok | BPB | BPB Gini | XNLI | Bele | XCOPA | ARC-E | ARC-C | HSwag | PIQA | LAMB |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 101-lang | BBPE | 1.5044 | 0.2029 | 34.6 | 23.0 | 51.4 | 41.0 | 19.4 | 28.3 | 62.4 | 25.5 |
| 101-lang | UBE | 1.5028 | 0.2021 | 34.9 | 23.1 | 51.6 | 40.0 | 19.4 | 28.7 | 62.2 | 26.7 |
| 101-lang | S-BBPE | 1.5064 | 0.1912 | 34.3 | 24.2 | 52.3 | 39.4 | 18.0 | 28.4 | 60.8 | 29.9 |
| 101-lang | S-UBE | 1.4929 | 0.1914 | 34.3 | 23.4 | 52.4 | 39.9 | 17.8 | 28.2 | 62.6 | 26.7 |
| Language modeling and multilingual accuracy | ||||||
|---|---|---|---|---|---|---|
| Tok | Run | BPB | BPB Gini | XNLI | Bele | XCOPA |
| BBPE | A | 1.4674 | 0.1863 | 34.7 | 23.6 | 51.8 |
| B | 1.4715 | 0.1833 | 34.3 | 24.5 | 51.8 | |
| Primary | 1.5174 | 0.1762 | 34.3 | 24.8 | 52.0 | |
| mean | 1.4854 0.0278 | 0.1819 0.0052 | 34.4 0.2 | 24.3 0.7 | 51.9 0.1 | |
| UBE | A | 1.4717 | 0.1879 | 34.9 | 23.5 | 51.8 |
| Language | BBPE tok | UBE tok | tok | BBPE rate (seq/s, 95% CI) | UBE rate (seq/s, 95% CI) | spd |
|---|---|---|---|---|---|---|
| Amharic | 1438 | 976 | -32.1% | 31.8 0.01 | 46.0 0.02 | +45.01% |
| Arabic | 975 | 957 | -1.8% | 46.1 0.01 | 48.1 0.02 | +4.29% |
| English | 226 | 223 | -1.3% | 145.3 0.15 | 151.0 0.41 | +3.90% |
| Hindi | 672 | 672 | 0.0% | 65.2 0.03 | 65.3 0.02 | +0.08% |
| Japanese | 354 | 347 | -2.0% | 107.5 0.05 | 108.7 0.05 | +1.07% |
| Korean | 260 | 253 | -2.7% | 140.1 0.10 | 143.9 0.10 | +2.76% |
| Stage | Configured steps | Runs | Wall-clock/run | GPU-hours/run | Total GPU-hours |
|---|---|---|---|---|---|
| 182M/32K LM pre-training | 200,000 | 12 | 1–2 days | 192–384 | 2,304–4,608 |
| 230M/64K LM pre-training | 200,000 | 6 | 1.5–3 days | 288–576 | 1,728–3,456 |
| 329M/128K LM pre-training | 200,000 | 4 | 2–4 days | 384–768 | 1,536–3,072 |
| 1.3B/32K LM pre-training | 205,000 | 6 | 8–14 days | 1,536–2,688 | 9,216–16,128 |
| Downstream eval. + inference | n/a | agg. | 6 h | 6 | 6 |
| Method | Steps | Training time (hours) |
|---|---|---|
| UBE | 200,000 | 12.94 |
| BBPE | 178,549 | 24.76 |
| BBPE16 | 177,699 | 24.68 |
| MYTE+BPE | 179,526 | 12.06 |
| SCRIPT-BPE | 170,191 | 11.39 |
| Paper | Venue | Tok. langs | Eval langs | LM size | LM tokens | Benchmarks | Vocab |
|---|---|---|---|---|---|---|---|
| Arnett+ | NeurIPS’25 | 97 | 97 | — | — | — | 8–262K |
| Schmidt+ | EMNLP’24 | 1 (en) | 1 | 350M–2.4B (64 runs) | 200B | 10 EN | 33–49K |
| SuperBPE | COLM’25 | 1 (en) | 1 | 680M–11B | 330B (8B baseline) | 30 EN | 200K |
| Abagyan+ | arXiv’25 | 62 | NR | 3.3B | 100B+10.5B | 2 multi+11 EN + Aya | 100–250K |
| Land & Arnett | TokShop’25 | 12 (mono) | 12 (mono) | — | — | — | 64K (mono); 256K merges (multi) |
| MAGNET | NeurIPS’24 | 9 | 9 | 126M | 10B bytes | 4 multi+2 | 50–250K (BPE baselines) |
| Tokenizer | Vocab | EN avg | Total (M) | Gini | Type |
|---|---|---|---|---|---|
| GPT-4o (o200k) | 200,019 | 26.554 | 11.6 | 0.2323 | tiktoken |
| GPT-4 (cl100k) | 100,277 | 26.860 | 18.6 | 0.3287 | tiktoken |
| Llama 3 | 128,256 | 26.849 | 16.6 | 0.3318 | tiktoken-BPE |
| Qwen 2.5 | 151,665 | 27.293 | 15.4 | 0.2724 | BPE |
| Qwen 3.6 multimodal tokenizer † | 248,044 | 27.204 | 12.8 | 0.2420 | BPE |
| Mistral NeMo | 131,072 | 27.497 | 14.2 | 0.3030 | Tekken |
| Model | Lang scope | Params | Vocab | Tok. source | Scale |
| Compact tokenizer (32K–50K) | |||||
| SmolLM2 | English | 135M–1.7B | 49,152 | self-trained | 2T–11T tok |
| MobileLLM | English | 125M–1B | 32,000 | Llama 2 | 1T tok |
| Pythia | English | 160M–2.8B | 50,277 | GPT-NeoX | 300B tok |
| TinyLlama | English + code | 1.1B | 32,000 | Llama 2 | 3T tok |
| Phi-3.5-mini | 23 langs | 3.8B | 32,064 | Llama 2 | 3.4T tok |
| Asset | Access point / version used | License | Terms / handling in this work |
|---|---|---|---|
| mC4 pretraining corpus | Hugging Face allenai/c4 multilingual; deterministic 101-language local Arrow mirror | ODC-BY | Derived from Common Crawl snapshots. Fixed capped subsets are used for tokenizer and LM training; the local Arrow mirror is reconstructed from public sources and not redistributed. |
| FLORES-200 | Hugging Face facebook/flores , devtest split | CC BY-SA 4.0 | Used for the primary parallel intrinsic benchmark. |
| UDHR | NLTK udhr2 corpus (Unicode UDHR) and United Nations declaration text | Public UN text; public corpus distribution | Used only for robustness evaluation. Our loader downloads the public corpus when absent and records the local corpus checksum. |
| SIB-200 | Hugging Face Davlan/sib200 , test split | CC BY-SA 4.0 | Used only for robustness evaluation. Our loader uses a pinned public revision and records a text checksum for the measured split. |
| MGSM | Hugging Face juletxara/mgsm , test split | CC BY-SA 4.0 | Used for robustness evaluation. Our preprocessing script builds the evaluation JSON with the original question text from the public dataset. |
| XNLI | facebookresearch/XNLI ; Hugging Face facebook/xnli , test split | CC BY-NC 4.0 | Used for multilingual downstream evaluation only. We preserve the noncommercial restriction and do not redistribute the data. |
| Asset | Access point / version used | License | Terms / handling in this work |
|---|---|---|---|
| Hugging Face tokenizers | Upstream huggingface/tokenizers ; 0.22.3-dev.778/779 | Apache 2.0 | Primary SCRIPT-style tokenizer training uses dev.779. Our modifications retain the Apache 2.0 license and attribution files. |
| LitGPT | Lightning-AI/litgpt , version 0.5.11 | Apache 2.0 | Our trainer vendors and modifies selected LitGPT modules while preserving the upstream license notice and attribution. |
| lm-evaluation- harness | EleutherAI repo, version 0.4.9.1 | MIT | Used as the evaluation backend for our offline benchmark runner, pinned to this version. |
| SCRIPT-BPE | sanderland/script_tok , commit 54a1058 | Apache 2.0 | Official SCRIPT-BPE tokenizers in Tables E.3 , E.7 , and 6 . |
| MYTE maps and decoder | tomlimi/MYTE , commit 177299c | Not stated in the repository | Used unmodified for Tables E.6 , E.7 , and 6 ; not redistributed. |
| Language | Tokenizer | 16K | 32K | 64K | 128K | 256K |
|---|---|---|---|---|---|---|
| Acehnese (Arabic) | UBE | 85.927 | 79.600 | 75.741 | 70.241 | 67.480 |
| BBPE | 86.626 | 79.774 | 75.897 | 70.374 | 67.620 | |
| Acehnese (Latin) | UBE | 60.001 | 56.095 | 52.427 | 49.271 | 45.328 |
| BBPE | 60.122 | 56.191 | 52.458 | 49.284 | 45.389 | |
| Afrikaans | UBE | 52.071 | 47.458 | 43.605 | 40.014 | 36.589 |
| BBPE | 52.299 | 47.583 | 43.651 | 40.024 | 36.590 |
| Language | BBPE | UBE | |
|---|---|---|---|
| ar | 31.8 | 33.6 | +1.8 |
| bg | 33.2 | 33.4 | +0.2 |
| de | 33.4 | 34.1 | +0.7 |
| el | 33.7 | 33.4 | -0.3 |
| en | 42.0 | 40.4 | -1.6 |
| es | 33.4 | 34.0 | +0.6 |
| Language | BBPE | UBE | |
|---|---|---|---|
| et | 50.8 | 49.8 | -1.0 |
| ht | 52.0 | 52.4 | +0.4 |
| id | 53.4 | 52.2 | -1.2 |
| it | 48.4 | 49.6 | +1.2 |
| qu | 50.0 | 48.6 | -1.4 |
| sw | 54.6 | 53.4 | -1.2 |
| Language | BBPE BPB | UBE BPB | BPB | UBE PPL |
|---|---|---|---|---|
| af | 1.9670 | 1.9795 | +0.0125 | 52.31 |
| am | 1.0390 | 1.0483 | +0.0093 | 9.50 |
| ar | 1.2644 | 1.2544 | -0.0100 | 17.86 |
| az | 1.9380 | 1.9210 | -0.0170 | 36.04 |
| be | 1.4010 | 1.4363 | +0.0353 | 30.66 |
| bg | 1.1531 | 1.1601 | +0.0070 | 21.12 |