Neural Network Training Dynamics

Latest papers 349

All topics
CardsList
  1. Spectral Lens: Activation and Gradient Spectra as Diagnostics of LLM Optimization

    May 7, 2026Andy Zeyi Liu, Elliot Paquette, John SousRepresentation LearningDecoder-Only Language Models

  2. SPHERE: Mitigating the Loss of Spectral Plasticity in Mixture-of-Experts for Deep Reinforcement Learning

    May 6, 2026Lirui Luo, Guoxi Zhang, Hongming Xu +2Loss of PlasticityReinforcement Learning

  3. Demystifying Manifold Constraints in LLM Pre-training

    May 6, 2026Kang An, Jiaxiang Li, Donald Goldfarb +1Language Model PretrainingRiemannian Optimization

  4. Information Plane Analysis of Binary Neural Networks

    May 5, 2026Maximilian Nothnagel, Bernhard C. GeigerNeural Network GeneralizationBinary Neural Networks

  5. Learning reveals invisible structure in low-rank RNNs

    May 5, 2026Yoav Ger, Omri BarakNeural ProcessesRecurrent Neural Networks

  6. Learning Dynamics of Zeroth-Order Optimization: A Kernel Perspective

    May 5, 2026Zhe Li, Bicheng Ying, Zidong Liu +1Zeroth-Order OptimizationLLM Fine-Tuning

  7. Trust, but Verify: Peeling Low-Bit Transformer Networks for Training Monitoring

    May 4, 2026Arian Eamaz, Farhang Yeganegi, Mojtaba SoltanalianDeep Learning OptimizationNeural Network Training Dynamics

  8. Deciphering Shortcut Learning from an Evolutionary Game Theory Perspective

    May 4, 2026Xiayang Li, Kuo Gai, Shihua ZhangNeural Network Training DynamicsShortcut Learning

  9. A Theory of Saddle Escape in Deep Nonlinear Networks

    May 2, 2026Divit Rawal, Michael R. DeWeeseNeural Network OptimizationNeural Network Training Dynamics

  10. Focus and Dilution: The Multi-stage Learning Process of Attention

    May 2, 2026Zheng-An Chen, Pengxiao Lin, Zhi-Qin John Xu +1Self-AttentionTransformer Attention

  11. Dendritic Neural Networks with Equilibrium Propagation

    May 1, 2026Yoshimasa KuboEquilibrium PropagationNeural Network Training Dynamics

  12. Better Models, Faster Training: Sigmoid Attention for single-cell Foundation Models

    Apr 29, 2026Vijay Sadashivaiah, Georgios Dasoulas, Judith Mueller +1Softmax AttentionSelf-Attention

  13. Self-Abstraction Learning for Effective and Stable Training of Deep Neural Networks

    Apr 27, 2026Wonyong Cho, Taemin Kim, Jungmin Kim +2Neural Network GeneralizationNeural Network Optimization

  14. GFlowState: Visualizing the Training of Generative Flow Networks Beyond the Reward

    Apr 23, 2026Florian Holeczek, Andreas Hinterreiter, Alex Hernandez-Garcia +2Data VisualizationGenerative Flow Networks

  15. There Will Be a Scientific Theory of Deep Learning

    Apr 23, 2026Jamie Simon, Daniel Kunin, Alexander Atanasov +11Neural Network Training DynamicsStatistical Learning Theory

  16. SGD at the Edge of Stability: The Stochastic Sharpness Gap

    Apr 22, 2026Fangshuo Liao, Afroditi Kolomvaki, Anastasios KyrillidisEdge of StabilityNeural Network Optimization

  17. Too Sharp, Too Sure: When Calibration Follows Curvature

    Apr 22, 2026Alessandro Morosini, Matea Gjika, Tomaso Poggio +1Model CalibrationNeural Network Training Dynamics

  18. The Origin of Edge of Stability

    Apr 22, 2026Elon LitmanEdge of StabilityGradient Descent

  19. Neuro-evolutionary stochastic architectures in gauge-covariant neural fields

    Apr 22, 2026Rodrigo Carmo TerinEvolutionary OptimizationEquivariant Neural Networks

  20. Generalization at the Edge of Stability

    Apr 21, 2026Mario Tuci, Caner Korkmaz, Umut Şimşekli +1Intrinsic DimensionalityNeural Network Generalization