cs.CLOct 5, 2026

Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation

Authors: Jie Wang, Shiwei Luo, Qi Zhang, Yuanbin Wu

Organizations: School of Computer Science, East China Normal University, Shanghai, China · School of Computer Science, Fudan University, Shanghai, China

Abstract

Tokenizer-free language models remove the inductive bias of fixed tokenizers by modeling text directly as bytes, but the resulting longer sequences substantially increase computation and eliminate explicit text abstractions. We ask whether this additional computation can be useful, and whether standard Transformers can learn the abstractions that tokenization provides. We study these questions on Transformers without specialized tokenization-related architectures. With token-superposition training and hash embeddings, byte Transformers consistently outperform subword Transformers as model size scales. We further find that byte Transformers build local text abstractions as external tokenizers: a set of segmentation-like positions are used to collect local context representations, and restricting up to 25%25\% of intermediate layers to these local representations preserves downstream performance. Finally, these learned structures induce highly non-uniform generation difficulty, with uncertainty concentrated near local structure boundaries; exploiting them for speculative decoding yields 3.4×3.4\times more accepted tokens than in subword Transformers.

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Fast Byte Latent Transformer

    May 8, 2026Julie Kallini, Artidoro Pagnoni, Tomasz Limisiewicz +5Speculative DecodingLLM Inference Acceleration

  2. Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models

    Sep 14, 2026Kalyani Marathe, Artidoro Pagnoni, Tomasz Limisiewicz +4Decoder-Only Language ModelsLanguage Model Scaling Laws

  3. Distilling Token-Trained Models into Byte-Level Models

    Feb 1, 2026Zishuo Bao, Jiaqi Leng, Junxiong Wang +2LLM Fine-TuningByte-Level Language Model