Neural Network Training Dynamics

Latest papers 349

All topics
CardsList
  1. The Implicit Bias of Depth: From Neural Collapse to Softmax Codes

    May 21, 2026Connall Garrod, Jonathan P. Keating, Christos ThrampoulidisDeep Linear NetworksImplicit Bias

  2. Anytime Training with Schedule-Free Spectral Optimization

    May 21, 2026Anuj Apte, Pranav Deshpande, Niraj Kumar +2Deep Learning OptimizationLearning Rate Scheduling

  3. Human-Centered Learning Mechanics: A Dynamical Framework for Entropy-Regulated Representation Learning

    May 21, 2026Kim Phuc TranRepresentation LearningNeural Network Training Dynamics

  4. Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics

    May 21, 2026Igor Ignashin, Anna Radovskaya, Andrew Semenov +7Neural Network Training DynamicsGradient Descent Dynamics

  5. AMUSE: Anytime Muon with Stable Gradient Evaluation

    May 21, 2026Jueun Kim, Baekrok Shin, Jihun Yun +3Deep Learning OptimizationMuon Optimizer

  6. A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification

    May 21, 2026Marcel Kühn, Yoon Thelge, Bernd RosenowStatistical Physics of LearningMulticlass Classification

  7. Uniform-in-Time Weak Propagation-of-Chaos in Shallow Neural Networks

    May 21, 2026Margalit Glasgow, Joan BrunaShallow Neural NetworksMean-Field Theory

  8. Thermodynamic Irreversibility of Training Algorithms

    May 21, 2026Liu Ziyin, Yuanjie Ren, Adam Levine +1Neural Network Training DynamicsGradient Descent Dynamics

  9. Large-Step Training Dynamics of a Two-Factor Linear Transformer Model

    May 20, 2026Krishnakumar BalasubramanianTransformerPhase Transitions

  10. The Devil is in the Condition Numbers: Why is GLU Better than non-GLU Structure?

    May 20, 2026Xingyu Lyu, Qianqian Xu, Zhiyong Yang +2Neural Network OptimizationNeural Network Training Dynamics

  11. StableGrad: Backward Scale Control without Batch Normalization

    May 19, 2026Jose I. Mestre, Alberto Fernández-Hernández, Cristian Pérez-Corral +2Deep Learning OptimizationNeural Network Training Dynamics

  12. Adynamical systems view of training generativemodels and the memorization phenomenon

    May 19, 2026Siva Athreya, Chiranjib Bhattacharya, Vivek S. BorkarNeural Network MemorizationMemorization in Generative Models

  13. Understanding Dynamics of Adam in Zero-Sum Games: An ODE Approach

    May 19, 2026Yi Feng, Weiming Ou, Xiao WangZero-Sum GamesMin-Max Optimization

  14. Canonical Regularisation of Wide Feature-Learning Neural Networks

    May 18, 2026George Whittle, Pranav Vaidhyanathan, Juliusz Ziomek +2Kernel Ridge RegressionRepresentation Learning

  15. How does feature learning reshape the function space?

    May 18, 2026João Lobo, Bruno Loureiro, Long Tran-Than +1Representation LearningKernel Methods

  16. Training Infinitely Deep and Wide Transformers

    May 17, 2026Raphaël Barboni, Maarten V. de Hoop, Takashi Furuya +1TransformerWasserstein Gradient Flows

  17. Bug or Feature2^2: Weight Drift, Activation Sparsity and Spikes

    May 17, 2026Egor Shvetsov, Aleksandr Serkov, Shokorov Viacheslav +3ReLU Neural NetworksActivation Sparsity

  18. The Neural Tangent Kernel for Classification

    May 17, 2026Jonathan Plenk, Sergio Calvo-Ordonez, Alvaro Cartea +3ClassificationNeural Network Training Dynamics

  19. High-dimensional Limit of SGD for Diagonal Linear Networks

    May 16, 2026Begoña García Malaxechebarría, Courtney Paquette, Maryam Fazel +1Stochastic Optimization ConvergenceNeural Network Training Dynamics

  20. DynMuon: A Dynamic Spectral Shaping View of Muon

    May 16, 2026Fangzhou Wu, Rikhav Shah, Sandeep Silwal +1Neural Network OptimizationMuon Optimizer

  21. A Fourier perspective on the learning dynamics of neural networks: from sample complexities to mechanistic insights

    May 16, 2026Fabiola Ricci, Claudia Merger, Sebastian GoldtSample ComplexityFrequency-Domain Feature Learning

  22. Does Weight Decay Enhance Training Stability?

    May 15, 2026Marius Saether, Amir Kolic, Tomaso Poggio +1Edge of StabilityDeep Learning Optimization

  23. Characterizing Learning in Deep Neural Networks using Tractable Algorithmic Complexity Analysis

    May 15, 2026Pedram Bakhtiarifard, Sophia N. Wilson, Mahmoud Afifi +2Neural Network GeneralizationNeural Network Training Dynamics

  24. Rethinking Neural Network Learning Rates: A Stackelberg Perspective

    May 15, 2026Sihan Zeng, Sujay Bhatt, Sumitra GaneshDeep Learning OptimizationNeural Network Optimization