Batch

Momentum

25 papers in the last four weeks, up 127% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 134

All topics
CardsList
  1. Optimal Learning Rate Schedules under Functional Scaling Laws: Power Decay and Warmup-Stable-Decay

    Feb 6, 2026Binghui Li, Zilin Wang, Fengling Chen +3BatchScaling Laws

  2. Gradient Flow Through Diagram Expansions: Learning Regimes and Explicit Solutions

    Feb 4, 2026Dmitry Yarotsky, Eugene Golikov, Yaroslav GusevWasserstein Gradient FlowsGradient

  3. Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent

    Feb 3, 2026Hiroki Naganuma, Shagun Gupta, Youssef Briki +4Stochastic Gradient DescentBatch

  4. SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning

    Feb 2, 2026Qifan Yu, Xinyu Ma, Zhijian Zhuo +7BatchSparsity

  5. Learnability Window in Gated Recurrent Neural Networks

    Dec 5, 2025Lorenzo LiviRecurrent Neural NetworksLearnability

  6. Correctness Forensics for Batch Speculative Decoding: Diagnosing the Ragged Tensor Problem

    Oct 26, 2025Ranran Haoran Zhang, Soumik Dey, Ashirbad Mishra +3Speculative DecodingBatch

  7. Why Do We Need Warm-up? A Theoretical Perspective

    Oct 3, 2025Foivos Alimisis, Rustem Islamov, Aurelien LucchiWarm StartsBatch

  8. Never Skip a Batch: Dense Learning of Temporal GNNs via Adaptive Pseudo-Supervision

    May 18, 2025Alexander Panyshev, Dmitry Vinichenko, Oleg Travkin +2Temporal Graph Neural NetworksBatch

  9. Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation: The Case of Multi-Armed Bandits

    May 6, 2025Max Qiushi Lin, Jincheng Mei, Matin Aghaei +6Policy GradientConvergence

  10. Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses

    Jun 20, 2024Steffen Dereich, Arnulf Jentzen, Adrian RiekertStochastic Gradient DescentAdam

  11. On Regularization via Early Stopping for Least Squares Regression

    Jun 6, 2024Rishi Sonthalia, Jackie Lok, Elizaveta RebrovaEarly StoppingKernel Ridge Regression

  12. Personalized Execution Time Optimization for Billion-Scale Scheduled Jobs

    Mar 11, 2022Yang Liu, Juan Wang, Idris Malik +7SchedulersSmooth Execution

  13. ExpTest: Loss-Curve Hypothesis Testing for Autonomous Learning-Rate Selection in Deep Neural Networks

    Date pendingZan Chaudhry, Naoko MizunoBatchHyperparameter

  14. On the Residual Scaling of Looped Transformers: Stability and Transferability

    Date pendingShaowen Wang, Bingrui Li, Ge Zhang +3Residual NetworksTransformer Architectures