cs.LGOct 6, 2026

The Best Optimizer Depends on Batch Size

Authors: Xingyu Dang, Kaiyue Wen, Sadhika Malladi

Organizations: Princeton University · Stanford University · University of California, San Diego

Abstract

A plethora of new adaptive optimizers are designed to efficiently estimate and use minibatch gradient statistics to shape parameter updates, but they are typically benchmarked at a single batch size. Hyperparameter scaling rules promise to preserve performance as batch size and gradient noise change, suggesting that the best optimizer at one batch size should remain the best at another. We challenge this approach to developing and evaluating optimizers by showing: (1) no principled scaling rule for Muon works consistently across training settings, and (2) the best optimizer for language model pretraining changes with batch size even after extensive hyperparameter tuning.

Figures & tables

Appendix figures & tables80 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers

    May 9, 2026Aditya RanganathDeep Learning OptimizationMemory-Efficient Optimization

  2. Fantastic Pretraining Optimizers and Where to Find Them II: Hyperball Optimization

    Jun 15, 2026Kaiyue Wen, Xingyu Dang, Kaifeng Lyu +2Language Model PretrainingWeight Decay

  3. Towards joint scaling laws with optimal batch size schedules

    Jul 30, 2026Jiaxiang Li, Zhiqi Bu, Shiyun XuDeep Learning OptimizationLanguage Model Scaling Laws