Language Model Scaling Laws

Momentum

15 papers in the last four weeks, up 200% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 95

All topics
CardsList
  1. Will Scaling Improve Social Simulation with LLMs?

    Jul 2, 2026Caleb Ziems, William Held, Su Doga Karaca +3Opinion DynamicsLanguage Model Scaling Laws

  2. When Does Generating More Help? Disentangling Fixed-Source Synthesis from Source Expansion in Synthetic Data Scaling

    Jul 2, 2026Xu Guo, Jian Tong, Zhihui Lu +1Synthetic Data GenerationLanguage Model Scaling Laws

  3. How to Allocate Your Tokens? Scaling Laws with Training Steps and Batch Size

    Jul 1, 2026Fabian SchaippLanguage Model Scaling LawsScaling Laws

  4. Two AI Metrics Diverged: Will it Make All the Difference?

    Jul 1, 2026Alex Fogelson, Zachary A. Brown, Hans Gundlach +2Language Model Scaling LawsAI Governance

  5. Smooth Scaling Laws Hide Stepwise Token Learning

    Jun 29, 2026Pingjie Wang, Zechen Hu, Peiru Yang +2Language Model PretrainingLanguage Model Scaling Laws

  6. On the Nonlinearity of Learning Rate Scaling for LLM Training

    Jun 28, 2026Zaiwen Yang, Huaqing Zhang, Jing Xu +1Language Model Scaling LawsHyperparameter Transfer

  7. Scaling limit of the Random Language Model

    Jun 26, 2026Eric De GiuliLanguage Model Scaling LawsLanguage Modeling

  8. Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients

    Jun 23, 2026Yizhou Liu, Jeff GoreLanguage Model Scaling LawsNeural Scaling Laws

  9. Can Scale Save Us From Plasticity Loss in Large Language Models?

    Jun 23, 2026J. Fernando Hernandez-Garcia, Tomás Figliolia, Beren MillidgeContinual Learning for LLMsLoss of Plasticity

  10. Scaling Laws for Task-Specific LLM Distillation

    Jun 23, 2026Lavinia Ghita, Dhruv Desai, Ioana BoierLLM PruningLLM Compression

  11. Internal Data Repetition Destroys Language Models

    Jun 23, 2026Jessica Chudnovsky, Joshua Kazdan, Noam Levi +6Memorization in Language ModelsLanguage Model Scaling Laws

  12. On the Smallness of the Large Language Models Scaling Exponents

    Jun 23, 2026Sauro Succi, Peter V. Coveney, Alex HansenEnergy-Efficient MLLanguage Model Scaling Laws

  13. The Energy Consumption of Transformer Fine-Tuning: A Roofline-Inspired Scaling Model

    Jun 22, 2026Mansour Zoubeirou a MayakiEnergy-Efficient MLTransformer

  14. Decomposing Factual Sycophancy in Language Models: How Size and Instruction Tuning Shape Robustness

    Jun 4, 2026Victor De Marez, Luna De Bruyne, Walter DaelemansLLM EvaluationLLM Sycophancy

  15. Entity tracking emerges in sub-billion parameter language models and exceeds human performance in naturalistic narratives

    Jun 4, 2026Karolina Drożdż, Micha HeilbronLLM EvaluationLanguage Model Scaling Laws

  16. Predictable Scaling Laws of Optimal Hyperparameters for LLM Continued Pre-training

    Jun 4, 2026Yongwei Zhou, Juncheng Diao, Junlin Shang +2Language Model Scaling LawsHyperparameter Transfer

  17. Structure and Scale in Simplicial Sequence Modelling

    May 31, 2026Matthew Farrugia-RobertsRepresentation GeometryTransformer

  18. When Data Is Scarce: Scaling Sparse Language Models with Repeated Training

    May 31, 2026Boqian Wu, Qiao Xiao, Patrik Okanovic +6Language Model Scaling LawsEfficient Language Model Training

  19. Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation

    May 29, 2026Sang Truong, Yuheng Tu, Rylan Schaeffer +1Language Model Scaling LawsTest-Time Scaling

  20. How LoRA Remembers? A Parametric Memory Law for LLM Finetuning

    May 28, 2026Ziwen Xu, Haiwen Hong, Linsong Yu +4Memorization in Language ModelsLLM Fine-Tuning

  21. Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

    May 28, 2026Jing Huang, Daniel Wurgaft, Rachit Bansal +6Long-Tail LearningLanguage Model Scaling Laws

  22. The Curse of Helpfulness: Inverse Scaling Law in Robustness to Distractor Instructions via DistractionIF

    May 28, 2026Zeli Su, Zhankai Xu, Tianlei Chen +4Language Model Scaling LawsLanguage Model Robustness

  23. Law of Neural Interaction: Depth-Width Shape, Interaction Efficiency, and Generalization

    May 27, 2026Wenjie Sun, Jinning Yang, Shuai Zhang +1Neural Network GeneralizationLanguage Model Scaling Laws

  24. Unified Neural Scaling Laws

    May 25, 2026Ethan Caballero, Priyank Jaini, David Krueger +1Language Model Scaling LawsScaling Laws