cs.LGJul 30, 2026

Towards joint scaling laws with optimal batch size schedules

Authors: Jiaxiang LiZhiqi BuShiyun Xu

Organizations: Meta · Independent researcher

Abstract

Modern deep learning typically keeps the batch size static throughout training, thus overlooking the joint effect of learning rate and batch size on the training dynamics. In this paper, we study the deep learning dynamics through the lens of convex optimization and derive a joint characterization of loss in terms of both schedules, applicable to general optimizers and model architectures. This characterization yields a closed-form optimal batch size schedule for any prescribed learning rate schedule, and further leads to joint scaling laws that consistently outperform static batch size baselines, highlighting the significance of dynamic batch size schedule in large language model training.

Explore similar work

CardsList
  1. Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

    Aug 28, 2026Niccolò Ajroldi, Diana Alexandra Onutu, Haider Al-Tahan +4Learning RatesBatch Size