cs.LGOct 1, 2026

SoftServe: A Scalable Quasi-Newton Method for Deep Learning

Authors: Joohwan Ko, Tetiana Parshakova, Diana Cai, Robert M. Gower

Organizations: University of Massachusetts Amherst · CCM, Flatiron Institute · Cornell University

Abstract

Quasi-Newton (QN) methods have long been among the most effective methods for large-scale unconstrained convex optimization. Two obstacles have limited their use in deep learning: non-convexity and enormous parameter sizes. We introduce SoftServe, a family of QN methods designed to overcome these obstacles without line searches or ad hoc curvature corrections. SoftServe derives positivedefinite curvature estimates from the variational objective of Berglund et al. (2025), even in the presence of negative curvature. We develop diagonal and Kroneckerfactored variants that preserve positive definiteness by construction and scale to massive neural networks. Finally, SoftServe relies on the stable coupled Newton-Schulz iteration for the required matrix operations, replacing costly matrix decompositions with GPU-friendly matrix multiplications. SoftServe excels on problems that are severely ill-conditioned, including tasks such as recurrent networks, deep autoencoders, physics-informed neural networks, and a 136M-parameter physics-informed diffusion model, often achieving lower losses than established baselines including Adam, Muon, and SOAP.

Figures & tables

Appendix figures & tables23 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Softsign: Smooth Sign in Your Optimizer For Better Parameter Heterogeneity Handling

    May 29, 2026Dmitrii Feoktistov, Timofey Belinsky, Andrey Veprikov +2Adaptive OptimizersStochastic Gradient Descent

  2. Sven: Singular Value Descent as a Computationally Efficient Natural Gradient Method

    Apr 1, 2026Samuel Bright-Thonney, Thomas R. Harvey, Andre Lukas +1

  3. Layerwise LQR for Geometry-Aware Optimization of Deep Networks

    May 5, 2026Simon Dufort-Labbé, Pierre-Luc Bacon, Razvan Pascanu +2Spectral PreconditioningNeural Network Optimization