Deep Linear Networks

Momentum

3 papers in the last four weeks, with none the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 19

All topics
CardsList
  1. Global Exponential Convergence of Two-Layer Linear Network Training

    Oct 7, 2026Stephen Y Zhang, Gabriel PeyréNeural Network Training DynamicsGradient Descent Dynamics

  2. Learning the identity: a case study of how SGD selects among functional decompositions

    Sep 30, 2026Andy Arditi, Weian Xie, David Bau +1Deep Linear NetworksImplicit Bias

  3. Not all solutions are created equal: An analytical dissociation of functional and representational similarity in deep linear neural networks

    Sep 30, 2026Lukas Braun, Erin Grant, Andrew M. SaxeRepresentational Similarity AnalysisNeural Representation Geometry

  4. Statistical Benefits of Fine-Tuning from Pretrained Initialization in Diagonal Linear Networks

    Sep 28, 2026Alexandre Declèves, Etienne Boursier, Nicolas FlammarionFine-TuningDeep Linear Networks

  5. Cascading Through the Hierarchy: Regularizer-Induced Feature Detection as Phase Transitions in Deep Linear Neural Networks

    Aug 6, 2026Björn Ladewig, Ibrahim Talha Ersoy, Karoline WiesnerStatistical Physics of LearningRepresentation Learning

  6. New Complexity-Theoretic Frontiers of Tractability for Neural Network Training

    Jul 23, 2026Cornelius Brand, Robert Ganian, Mathis RoctonReLU Neural NetworksNeural Network Optimization

  7. How the Hessian-Spectrum of Neural Networks Depends on Data

    Jul 15, 2026Jasraj Singh, Enea Monzio Compagnoni, Antonio OrvietoNeural Network OptimizationDeep Linear Networks

  8. Gradient Flow Dynamics and Implicit Bias of Diagonal Linear Networks under Infinitesimal Initialization

    Jul 14, 2026Jiajie Zhao, Jianxing Wang, Junjie Yang +2Deep Linear NetworksImplicit Bias

  9. How are linear representations learned? Exact solutions to the dynamics of abstraction

    Jul 9, 2026William W. Yang, Andrew M. Saxe, Peter E. LathamLinear Representation HypothesisDeep Linear Networks

  10. Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks

    Jul 8, 2026Yedi Zhang, Peter E. Latham, Leena Chennuru Vankadara +1Learning Rate SchedulingDeep Linear Networks

  11. Deciphering Two Training Clocks in Grokking via Deep Linear Network Theory with Conditional ReLU Reduction

    Jun 4, 2026Hu Tan, Kuo Gai, Shihua ZhangReLU Neural NetworksDeep Linear Networks

  12. Gradient Descent with Large Step Size Restores Symmetry in Deep Linear Networks with Multi-Pathway

    May 29, 2026Hee-Sung Kim, Sungyoon LeeEdge of StabilitySymmetry Breaking

  13. The Implicit Bias of Depth: From Neural Collapse to Softmax Codes

    May 21, 2026Connall Garrod, Jonathan P. Keating, Christos ThrampoulidisDeep Linear NetworksImplicit Bias

  14. Understanding Sample Efficiency in Predictive Coding

    May 12, 2026Gaspard Oliviers, Elene Lominadze, Rafal BogaczBackpropagationPredictive Coding

  15. Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer

    May 8, 2026Clarissa Lauditi, Cengiz Pehlevan, Blake BordelonDeep Linear NetworksNeural Network Training Dynamics

  16. To Use or not to Use Muon: How Simplicity Bias in Optimizers Matters

    Feb 28, 2026Sara Dragutinović, Yedi Zhang, Rajesh RanganathNeural Network GeneralizationMuon Optimizer