Layer-Wise Learning Rate Adaptation

Momentum

0 papers in the last four weeks, against 2 the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 11

All topics
CardsList
  1. What to Preserve, Where to Adapt: A Depth-Wise Analysis of Forgetting in Continual Gynecological Image Segmentation

    Aug 13, 2026Amal Saqib, Tausifa Jan Saleem, Numan Saeed +1Medical Image Segmentation3D Medical Image Segmentation

  2. FedA2L: Adaptive layer-wise learning rate adjustment in decentralized federated learning

    Aug 10, 2026Van Truong Vo, Khoa Nguyen, Taehong KimNon-IID Federated LearningFederated Learning

  3. LionVote: Per-Layer Learning Rate Adaptation for Lion

    Jul 10, 2026Kris AtallahDeep Learning OptimizationLayer-Wise Learning Rate Adaptation

  4. Directional Curvature from Armijo Backtracking: A Low-Cost Sharpness Probe and a Calibration-Free Learning-Rate Safeguard for Adam

    Jul 4, 2026Ashmitha R, Jörg FrochteGradient DescentAdaptive Gradient Methods

  5. Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks

    May 29, 2026Tianyu Pang, Vignesh Kothapalli, Shenyang Deng +3Deep Learning OptimizationNeural Network Training Dynamics

  6. One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs

    May 21, 2026Di He, Songjun Tu, Keyu Wang +2Learning Rate SchedulingEfficient Language Model Training

  7. Training Neural Networks with Optimal Double-Bayesian Learning

    May 19, 2026Vy Bui, Hang Yu, Karthik Kantipudi +2Neural Network OptimizationBayesian Decision Theory

  8. Rethinking Neural Network Learning Rates: A Stackelberg Perspective

    May 15, 2026Sihan Zeng, Sujay Bhatt, Sumitra GaneshDeep Learning OptimizationNeural Network Optimization

  9. Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining

    May 11, 2026Jinchang Zhu, Jindong Li, Yuwen Hao +3Language Model PretrainingSelf-Attention

  10. Learning Rate Engineering: From Coarse Single Parameter to Layered Evolution

    Apr 30, 2026Ming-Hong Yao, Di Wang, Jian Cui +5Deep Learning OptimizationFine-Tuning