Language Model Scaling Laws

Momentum

15 papers in the last four weeks, up 200% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 95

All topics
CardsList
  1. LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws

    May 22, 2026Xu Ouyang, Deyi Liu, Yuhang Cai +5Language Model Scaling Laws

  2. Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate

    May 20, 2026Dayal Singh Kalra, Maissam BarkeshliLanguage Model Scaling LawsHyperparameter Transfer

  3. Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency

    May 18, 2026Matthew L. Smith, Jonathan P. Shock, Samuel T. Segun +2LLM EvaluationLanguage Model Scaling Laws

  4. Continuous Diffusion Scales Competitively with Discrete Diffusion for Language

    May 18, 2026Zhihan Yang, Wei Guo, Shuibai Zhang +5Language Model Scaling LawsDiffusion Language Models

  5. Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance

    May 17, 2026Kazuya Horibe, Masaomi Hatakeyama, Gen Masumoto +2Multi-Agent LLM SystemsMulti-Agent Coordination

  6. The Scaling Laws of Skills in LLM Agent Systems

    May 15, 2026Charles Chen, Qiming Yu, Yuhang Gu +12Language Model Scaling LawsLLM Agent Reliability

  7. A Hierarchical Language Model with Predictable Scaling Laws and Provable Benefits of Reasoning

    May 13, 2026Jason Gaitonde, Frederic Koehler, Elchanan Mossel +2Long-Context Language ModelingLanguage Model Scaling Laws

  8. When is Warmstarting Effective for Scaling Language Models?

    May 13, 2026Neeratyoy Mallik, Maciej Janowski, Johannes Hog +4Language Model Scaling LawsLanguage Model Scaling

  9. The Efficiency Gap in Byte Modeling

    May 13, 2026Celine Lee, Jing Nathan Yan, Chen Liang +9Language Model Scaling LawsAutoregressive Language Modeling

  10. Scaling Laws for Mixture Pretraining Under Data Constraints

    May 12, 2026Anastasiia Sedova, Skyler Seto, Natalie Schluter +1Language Model PretrainingLanguage Model Scaling Laws

  11. Predicting Large Model Test Losses with a Noisy Quadratic System

    May 9, 2026Chuning Li, Chris J. MaddisonLanguage Model Scaling LawsEfficient Language Model Training

  12. Where Larger Models Excel: The Primacy of Constraint-Guided Reasoning

    May 9, 2026Guan-Yi Lin, Hen-Hsen HuangLanguage Model Scaling LawsLLM Reasoning

  13. Tokens-per-Parameter Coverage Is Critical for Robust LLM Scaling Law Extrapolation

    May 8, 2026Joshua Shay Kricheli, Alexander Lawrence Reid, Soumajyoti Sarkar +2Parameter EstimationLanguage Model Scaling Laws

  14. On the Invariance and Generality of Neural Scaling Laws

    May 8, 2026Xing Han, Liu Ziyin, Suchi Saria +1Language Model Scaling LawsCross-Domain Generalization

  15. Limits of Reliability and Scaling in Language Models

    May 8, 2026Subhabrata MajumdarLanguage Model Scaling LawsAutoregressive Language Modeling

  16. Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key

    May 7, 2026Tianle Wang, Zhaoyang Wang, Guangchen Lan +4Logical ReasoningRL for Language Model Reasoning

  17. Safety and accuracy follow different scaling laws in clinical large language models

    May 5, 2026Sebastian Wind, Tri-Thien Nguyen, Jeta Sopa +9LLM Safety BenchmarksRetrieval-Augmented Generation

  18. Finite-Size Gradient Transport in Large Language Model Pretraining: From Cascade Size to Intensive Transport Efficiency

    May 3, 2026Ping Wang, Yan-Qi DuLanguage Model PretrainingLanguage Model Scaling Laws

  19. Prescriptive Scaling Laws for Data Constrained Training

    May 2, 2026Justin Lovelace, Christian Belardi, Srivatsa Kundurthy +2Weight DecayLanguage Model Scaling Laws

  20. Compute Optimal Tokenization

    May 2, 2026Tomasz Limisiewicz, Artidoro Pagnoni, Srini Iyer +6Language Model Scaling LawsEfficient Language Model Training

  21. Scaling Properties of Continuous Diffusion Spoken Language Models

    Apr 27, 2026Jason Ramapuram, Eeshan Gunesh Dhekane, Amitis Shidani +6Language Model Scaling LawsSpeech Language Models

  22. How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models

    Apr 22, 2026Kristian Schwethelm, Daniel Rueckert, Georgios KaissisLanguage Model Scaling LawsRecurrent Neural Networks

  23. Nexusformer: Nonlinear Attention Expansion for Stable and Inheritable Transformer Scaling

    Apr 21, 2026Weijie Zhao, Mingquan Liu, Bolun Wang +4TransformerLanguage Model Scaling Laws

  24. Stabilizing Native Low-Rank LLM Pretraining

    Feb 12, 2026Paul Janson, Edouard Oyallon, Eugene BelilovskyLanguage Model PretrainingLow-Rank Matrix Decomposition