Neural Network Training Dynamics

Latest papers 349

All topics
CardsList
  1. Prospective Prediction of OOD Degradation from Source-Side Training Dynamics

    Oct 8, 2026Sasha, MoninOOD GeneralizationNeural Network Training Dynamics

  2. Probability-Signature Dynamics: Unpacking Modular Addition Learning Within Two-Layer Networks

    Oct 8, 2026Yunji Wang, Junjie Yao, Linyu Liu +2Neural Network Training DynamicsFrequency-Domain Feature Learning

  3. Rare Gate Disagreements Can Limit Plasticity: When Gradient Flow Mispredicts Finite-Batch SGD

    Oct 8, 2026Ruoyu Zhao, Mingxuan Zhang, Jianbo Dai +3Neural Network Training DynamicsStochastic Gradient Descent

  4. Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals

    Oct 8, 2026Lei Zhao, Qichao Zhao, Bowen Zuo +1On-Policy DistillationLanguage Model Distillation

  5. When does a network's training history predict its future learning better than its current state? Evidence from a response probe and a forecasting screen

    Oct 7, 2026Martin Hofmann, Patrick MäderNeural Network Training Dynamics

  6. Global Exponential Convergence of Two-Layer Linear Network Training

    Oct 7, 2026Stephen Y Zhang, Gabriel PeyréNeural Network Training DynamicsGradient Descent Dynamics

  7. Neural Fields Encode Adaptation Geometry

    Oct 5, 2026Prateik Sinha, Stefania DrugaImplicit Neural RepresentationsNeural Network Training Dynamics

  8. Mind the Drift: Diagonal Linear Networks Under Large Learning Rates

    Oct 5, 2026Aniket Sanyal, Tom Jacobs, Rebekka BurkholzEdge of StabilityDeep Learning Optimization

  9. Generalization in Neural Networks Through the Lens of Magnitude Potential

    Oct 1, 2026Sahel Torkamani, Henry Gouk, Rik SarkarNeural Network GeneralizationNeural Representation Geometry

  10. Why Does Train-Validation Separation Emerge? Update-Pressure Density Dynamics in Pretrained Backbones

    Oct 1, 2026Yuchen Li, Mingyu Du, Zongqi Fan +2Neural Network GeneralizationTransfer Learning

  11. Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning

    Sep 30, 2026Syon Mansur, Joel ZylberbergDeep Learning OptimizationConvolutional Neural Networks

  12. Disentangling Computation in Multi-Task Neural Networks with the Green's Operator

    Sep 30, 2026James HazeldenRecurrent Neural NetworksNeural Network Interpretability

  13. The Life Cycle of a Massive Activation: Stochastic Birth, Weight-Decay-Driven Growth, and Competitive Consolidation

    Sep 30, 2026S. Aaron McClendon, Jorge Gallego-Feliciano, Antonios SaravanosWeight DecayAttention Sinks

  14. How Does Local Landscape Geometry Evolve in Language Model Pre-Training?

    Sep 30, 2026Zhanpeng Zhou, Yuhan Sun, Bingrui Li +4Language Model PretrainingLearning Rate Scheduling

  15. Stable Transformers for Graph Generation

    Sep 30, 2026Luca Miglior, Alessio Gravina, Davide BacciuRepresentation CollapseGraph Generation

  16. Awakening of the Buddha: Subspace Learning During Population-Loss Plateaus

    Sep 30, 2026Akash KumarRepresentation LearningReLU Neural Networks

  17. A Dynamical Theory of LoRA in Continual Learning

    Sep 30, 2026Théo Marchetta, Filippo Alessandroni, Alessandro Breccia +2Continual LearningLow-Rank Adaptation

  18. Where Does Randomness Matter in Neural Cellular Automata?

    Sep 29, 2026Fei Zuo, Jiaqi Shi, Yujing LiuNeural Cellular AutomataCellular Automata

  19. The Hidden Ratio in Adam: Stable Structure, Compression, and Sign Dynamics

    Sep 28, 2026Yihe Zhou, Tongtian Zhu, Yingxiao Huo +4Stochastic OptimizationSign-Based Optimization

  20. First Learn, Then Memorize: The Spectral Bias of Diffusion Models

    Sep 28, 2026Raphaël Urfin, Tony Bonnaire, Giulio Biroli +1Neural Network GeneralizationMemorization in Generative Models

  21. Weighting Schedules Govern What and When Score-Based Generative Models Learn from Multimodal Data

    Sep 28, 2026Jérémie Klinger, Raphaël Urfin, Giulio Biroli +1Generative ModelingScore-Based Generative Modeling

  22. Fast Learning Rate Transfer in Shallow Linear Networks at Growing Training Horizons

    Sep 28, 2026Mana Sakai, Masaaki ImaizumiHyperparameter TransferNeural Network Training Dynamics

  23. Muon Sublates the Edge of Stability in LLM Pretraining

    Sep 28, 2026Yanzhe Chen, Qifang Zhao, Xiaoxiao Xu +1Stochastic OptimizationLanguage Model Pretraining

  24. When Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy Distillation

    Sep 28, 2026Xinke Jiang, Tao Feng, Zhibang Yang +4Gradient InterferenceReinforcement Learning with Verifiable Rewards

  25. Two-Timescale Fine-tuning Provably Learns New Features for Two-Layer ReLU Networks

    Sep 28, 2026Etienne Boursier, Nicolas FlammarionRepresentation LearningFine-Tuning