Language Model Scaling

Momentum

6 papers in the last four weeks, up 100% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 37

All topics
CardsList
  1. Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation

    Oct 5, 2026Jie Wang, Shiwei Luo, Qi Zhang +1Language Model ScalingByte-Level Language Model

  2. ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control

    Sep 30, 2026Peng Jin, Zihan Qiu, Zekun Wang +8Expert Load BalancingMixture-of-Experts Language Models

  3. Trajectory Soup: Pushing the Compute-Scaling Frontier of LLM Mid-training via Diverse Trajectories

    Sep 29, 2026Zhehao Huang, Changxin Tian, Qingyuan Yang +5Language Model ScalingModel Merging

  4. When Harness Beats Scale, and When Reading Beats Both

    Sep 28, 2026Ivan Bondarenko, Nikolay O. NikitinNumerical Reasoning in Language ModelsLanguage Model Scaling

  5. Likelihood Ranking doesn't Scale Like Prompting in LLMs

    Sep 24, 2026Alessandro Bondielli, Lucia Passaro, Davide Bacciu +1LLM EvaluationLanguage Model Generation Evaluation

  6. Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs

    Sep 1, 2026Jingtan Wang, Arun Verma, Xiaoqiang Lin +4Supervised Fine-TuningLanguage Model Scaling

  7. When Does Bigger Help? A Controlled Study of LLM Scale for Ontology Learning

    Aug 31, 2026Hamed Babaei Giglou, Sören Auer, Jennifer D'SouzaLanguage Model ScalingOntology Learning

  8. Kathleen Writes: Autoregressive Generation and Data Scaling Without Attention

    Aug 5, 2026George FountzoulasLanguage Model Generation EvaluationAutoregressive Language Modeling

  9. Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory

    Jul 30, 2026Rubin Wei, Jiaqi Cao, Jiarui Wang +4Language Model PretrainingMemory-Augmented Language Models

  10. Understanding Layer Patching in Model Size Interpolation

    Jul 9, 2026Sara Kangaslahti, Jonathan Geuter, Nihal V. Nayak +3Language Model Scaling

  11. Think Big, Search Small: Where Capacity Matters in Hierarchical Search Agents?

    Jul 8, 2026Qinnan Cai, Yibo Zhao, Xiang LiMulti-Hop QAMulti-Agent LLM Systems

  12. Fixed RAG Compression Collapses Measured Reader Scaling

    Jun 20, 2026Sugam Panthi, Rabab AbdelfattahRetrieval-Augmented GenerationLLM Evaluation

  13. Variable-Width Transformers

    Jun 16, 2026Zhaofeng Wu, Oliver Sieberling, Shawn Tan +3TransformerDecoder-Only Language Models

  14. A Dual-Path Architecture for Scaling Compute and Capacity in LLMs

    May 28, 2026Markus Frey, Behzad Shomali, Joachim Koehler +1Recurrent TransformersLanguage Model Scaling

  15. Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory

    May 20, 2026Runxi Cheng, Yuchen Guan, Yongxian Wei +7Language Model PretrainingMemory-Augmented Language Models

  16. Introspective X Training: Feedback Conditioning Improves Scaling Across all LLM Training Stages

    May 19, 2026Brandon Cui, Ximing Lu, Jaehun Jung +7Language Model PretrainingEfficient Language Model Training

  17. Scale Determines Whether Language Models Organize Representation Geometry for Prediction

    May 16, 2026Weilun XuLanguage Model PretrainingNeural Representation Geometry

  18. When is Warmstarting Effective for Scaling Language Models?

    May 13, 2026Neeratyoy Mallik, Maciej Janowski, Johannes Hog +4Language Model Scaling LawsLanguage Model Scaling

  19. Sparse Layers are Critical to Scaling Looped Language Models

    May 9, 2026Ryan Lee, Jacob Biloki, Edward J. Hu +1Mixture-of-Experts Language ModelsRecurrent Transformers

  20. Scaling Categorical Flow Maps

    May 8, 2026Oscar Davis, Anastasiia Filippova, Pierre Ablin +4Language Model ScalingFlow Map Learning

  21. Seeking Information with RAG-Assistants: Does Model Size Matter in Human-AI Collaborations?

    May 1, 2026Lennard C. Froma, Tom Kouwenhoven, Maaike H. T. de Boer +2LLM EvaluationLanguage Model Scaling

  22. Scaling Probabilistic Transformer via Efficient Cross-Scale Hyperparameter Transfer

    Apr 28, 2026Penghao Kuang, Haoyi Wu, Kewei TuTransformerNeural Network Optimization