Byte-Level Language Model

Latest papers 14

All topics
CardsList
  1. Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation

    Oct 5, 2026Jie Wang, Shiwei Luo, Qi Zhang +1Language Model ScalingByte-Level Language Model

  2. Learning to Learn a Language

    Oct 5, 2026Lennart Carstens-Behrens, Holger FröhlichLanguage Model PretrainingMeta-Learning

  3. Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models

    Sep 14, 2026Kalyani Marathe, Artidoro Pagnoni, Tomasz Limisiewicz +4Decoder-Only Language ModelsLanguage Model Scaling Laws

  4. Toppling the Hierarchy in Byte-level Language Modeling

    Aug 31, 2026Lukas Edman, Alexander FraserEfficient Language Model TrainingByte-Level Language Model

  5. Disentangling Language Modeling and Boundaries

    Aug 4, 2026Mykola HaltiukByte-Level Language ModelTransfer Learning

  6. EntropyMoE: Entropy-Aware Sparse Expert Routing for Tokenizer-Free LLMs

    Jul 31, 2026Bo Liu, Muxuab Yu, Yu Zhang +2Sparse Mixture-of-ExpertsByte-Level Language Model

  7. Beyond Perplexity: UTF-8 Validity in Byte-aware Language Models

    Jun 12, 2026Sangwhan Moon, Daisuke Oba, Youmi Ma +2Byte-Level Language ModelLLM Reliability

  8. Large Byte Model: Teaching Language Models About Compiled Code

    Jun 1, 2026Florian Störtz, Catalin-Andrei Stan, Alexandru Dinu +4Malware ClassificationByte-Level Language Model

  9. Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models

    May 10, 2026Lin Zheng, Vasilisa Bashlovkina, Timothy Dozat +3Byte-Level Language ModelLanguage Modeling

  10. Fast Byte Latent Transformer

    May 8, 2026Julie Kallini, Artidoro Pagnoni, Tomasz Limisiewicz +5Speculative DecodingLLM Inference Acceleration

  11. A Systematic Benchmark of Machine Transliteration Models for the Tajik-Farsi Language Pair: A Comparative Study from Rule-Based to Transformer Architectures

    May 4, 2026Mullosharaf K. ArabovMultilingual Language ModelsTransformer

  12. Distilling Token-Trained Models into Byte-Level Models

    Feb 1, 2026Zishuo Bao, Jiaqi Leng, Junxiong Wang +2LLM Fine-TuningByte-Level Language Model