cs.CLOct 1, 2026

Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities

Authors: Hyunsik Kim, Youngmoon Jung

Organizations: Samsung Research

Abstract

Byte-level byte-pair encoding (BBPE) tokenizers are attractive for multilingual large language models (LLMs) because they cover all Unicode text. In UTF-8-based BBPE, however, many scripts start from a higher fallback cost than English: when no learned merges can be applied, a multibyte character requires multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor. A higher floor can increase token counts and per-request cost and shrink usable context. Changing the text encoding can reduce this gap, but a single global encoding can make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer that keeps 1-2-byte UTF-8 characters on the UTF-8 path while routing 3-4-byte UTF-8 characters through UTF-16. This lowers the encoding floor for 3-byte Basic Multilingual Plane (BMP) characters in scripts with high token premiums (token counts relative to English) without raising it for already-efficient spans in mixed-script text. UBE changes only the byte representation presented to byte-pair encoding (BPE); the merge rule remains standard, and exact decoding is preserved. UBE also composes with alternative boundary policies and morphology-based representations. In a Unicode 17 audit, UBE exactly round-trips all Unicode scalar values and all inputs in the official normalization, grapheme-break, and emoji test suites. Across intrinsic evaluations, UBE lowers dispersion in English-normalized token-count ratios, reducing cross-lingual token-budget disparity. In multilingual language model (LM) experiments, UBE matches BBPE's LM quality. In the main multilingual settings, UBE reduces token counts most for high-premium scripts and slightly lowers English token counts, yielding more usable context under fixed token budgets and faster prompt processing in content-matched benchmarks.

Figures & tables

Appendix figures & tables51 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Writing-System-Level Tokenizer Adaptation for Byte-Level BPE

    Aug 1, 2026Bohdan DidenkoByte-Pair EncodingTokenizer

  2. Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization

    Aug 6, 2025Negar Foroutan, Clara Meister, Debjit Paul +4Byte-Pair EncodingCross-Lingual Consistency

  3. Equity with Efficiency: An Empirical Study of Tokenizers for Multilingual Large Language Models

    Jun 13, 2026Kieron Seven Jun Wei Lee, Muhammad Reza Qorib, Andrew Ivan Soegeng +1Multilingual Language ModelsTokenizer