Sven: Singular Value Descent as a Computationally Efficient Natural Gradient Method
Organizations: Department of Physics, Massachusetts Institute of Technology · The NSF Institute for Artificial Intelligence and Fundamental Interactions · Rudolf Peierls Centre for Theoretical Physics, University of Oxford · Institut des Hautes Études Scientifiques · Institut de Physique Théorique, CEA Paris-Saclay
Abstract
We introduce Sven (Singular Value dEsceNt), a new optimization algorithm for neural networks that exploits the natural decomposition of loss functions into a sum over individual data points, rather than reducing the full loss to a single scalar before computing a parameter update. Sven treats each data point's residual as a separate condition to be satisfied simultaneously, using the Moore-Penrose pseudoinverse of the loss Jacobian to find the minimum-norm parameter update that best satisfies all conditions at once. In practice, this pseudoinverse is approximated via a truncated singular value decomposition, retaining only the most significant directions. We show that Sven can be understood as a natural gradient method generalized to the overparametrized regime, recovering natural gradient descent in the underparametrized limit. We test Sven on a variety of regression and classification tasks, including small-scale language modeling with transformers, and find that it is competitive with leading baselines such as Adam, Muon, and K-FAC. We also discuss Sven's memory overhead, which presents a barrier to scaling under a naive implementation, and introduce an optimized implementation that keeps memory usage on par with standard baselines under mild restrictions on model architecture. Beyond standard machine learning benchmarks, we anticipate that Sven will find natural application in scientific computing settings where custom loss functions decompose into several conditions.
Figures & tables
Appendix figures & tables45 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Loss | Train | Val. | Test | Construction |
|---|---|---|---|---|---|
| 1D regression | 10k | 10k | 10k | independent draws; pool-standardized | |
| Random polynomial | 10k | 10k | 10k | independent draws; pool-standardized | |
| MNIST | label reg./CE | 50k | 10k | 10k | official train split; official test set |
| CIFAR-10 | label reg./CE | 45k | 5k | 10k | official train split; official test set |
| Shakespeare (char) | token CE | 6,971 | 871 | 871 | contiguous 80/10/10 by position |
| FineWeb-edu (BPE) | token CE | 210k | 200 | 200 | val/test from disjoint documents |
| Task | Model | Configuration | |
|---|---|---|---|
| 1D regression | MLP | , GeLU | 593 |
| Random polynomial | MLP | , GeLU | 673 |
| MNIST | MLP | , GeLU | 27,562 |
| CIFAR-10 | ResNet18 | 10 classes, functional BatchNorm | 11,181,642 |
| Shakespeare (char) | nanoGPT | 4 layers, 4 heads, , block 128 | 826,368 |
| FineWeb-edu | GPT-2 small | 12 layers, 12 heads, , block 1024 | 163,109,376 |
| Scan | rtol | Runs | |||
|---|---|---|---|---|---|
| 1D regression | 32 | 1,2,4,8,16,32 | .01,.02,.05,.1,.5,1 | (5) | 900 |
| Random polynomial | 32 | 1,2,4,8,16,32 | .05,.1,.5,1 | (6) | 720 |
| MNIST label reg. | 64 | 1,2,4,8,16,32,48,64 | .05,.1,.5,1 | (4) | 640 |
| MNIST CE | 64 | 1,2,4,8,16,32,48,64 | .05,.1,.5,1 | (5) | 800 |
| CIFAR-10 label reg. | 128 | 64,128 | .1,.5,1 | 90 | |
| CIFAR-10 CE | 128 | 64,128 | .02,.05,.1,.5,1 | 150 |
| Study | rtol | Swept axis | Runs | |||
|---|---|---|---|---|---|---|
| (MNIST lab. reg.) | 64 | 32,64 | .125–1.5 (7) | : 1,2,3 | 210 | |
| Micro-batch (1D, poly.) | 32 | 32 | .05,.1,.5,1 | : 1–32 (6) | 120 | |
| Micro-batch (MNIST) | 64 | 64 | .05,.1,.5,1 | : 1–64 (7) | 140 | |
| Param. frac. (1D, poly.) | 32 | 32 | .05,.1,.5,1 | : .1–1 (5) | 100 | |
| Param. frac. (MNIST) | 64 | 64 | .05,.1,.5,1 | : .1–1 (5) | 100 | |
| Param. frac. (CIFAR-10) | 128 | 128 | .5 | : .05–1 (5) | 15 |
| Scan | Standard optimizers | L-BFGS | JD | HIG |
|---|---|---|---|---|
| 1D regression | (8) | .01,.03,.1,.5,1 | (6) | (7) |
| Random polynomial | (10) | .1,.5,1,2,4 | (6) | (5) |
| MNIST label reg. | (8) | .1,.5,1 | (6) | (5) |
| MNIST CE | (8) | .1,.5,1,2,4 | (6) | (5) |
| CIFAR-10 label reg. | (7) | .1,.5,1,2,4 | — | — |
| CIFAR-10 CE | (9) | .1,.5,1,2,4 | — | — |
| Method | Configuration | Val. loss | Val. (tune) | Test loss | Test acc. | fin./att. | s/ep. | MB |
|---|---|---|---|---|---|---|---|---|
| Random Polynomial | ||||||||
| Sven | =0.5, =16, rtol=0.03 | 0.1388 0.055 | 0.1095 | 0.136 0.055 | – | 15/15 | 1.33 | 18.86 |
| HIG | =0.1, tau=1e-08 | 0.08406 0.024 | 0.07149 | 0.08028 0.02 | – | 15/15 | 1.4 | 21.59 |
| MuonW | =0.01, wd=0.1 | 0.2313 0.07 | 0.2061 | 0.2226 0.06 | – | 15/15 | 0.967 | 18.44 |
| Muon | =0.01 | 0.207 0.062 | 0.1651 | 0.202 0.058 | – | 15/15 | 0.978 | 18.44 |
| SOAP | =0.01 | 0.1818 0.05 | 0.1496 | 0.1756 0.042 | – | 15/15 | 1.15 | 21.29 |
| Method | Val. (pooled) | Val. (inst. mean) | Betw.-inst. std | Within-inst. std | Ratio |
|---|---|---|---|---|---|
| Toy 1D | |||||
| HIG | 0.791 | ||||
| Sven | 0.682 | ||||
| SOAP | 0.318 | ||||
| SGD + momentum | 0.652 | ||||
| AdamW | 0.657 | ||||
| Method | Selected configuration | Val. loss | Test loss | Train (eval) |
|---|---|---|---|---|
| HIG | lr=0.005, tau=1e-06 | |||
| Sven | k=32, lr=0.01, rtol=0.0001 | |||
| SOAP | lr=0.001 | |||
| SGD + momentum | lr=0.1 | |||
| AdamW | lr=0.001, weight_decay=0.01 | |||
| KFAC | lr=0.001 |
| Method | Selected configuration | Val. loss | Test loss | Test acc. | Train (eval) |
|---|---|---|---|---|---|
| MuonW | lr=0.001, weight_decay=0.1 | 0.04982 0.0013 | 0.04959 0.0025 | 97.03% | 0.02206 0.00041 |
| HIG | lr=0.05, tau=0.03 | 0.05298 0.0026 | 0.05324 0.0023 | 96.60% | 0.02293 0.0019 |
| Sven | k=64, lr=0.5, rtol=0.001 | 0.05328 0.0007 | 0.05363 0.0024 | 96.72% | 0.02804 0.0012 |
| Muon | lr=0.001 | 0.05473 0.0017 | 0.05424 0.0038 | 96.72% | 0.01652 0.0011 |
| SGD | lr=0.1 | 0.05524 0.0025 | 0.05439 0.0019 | 96.89% | 0.02968 0.0019 |
| SGD + momentum | lr=0.01 | 0.05524 0.0015 | 0.05454 0.0017 | 96.95% | 0.02863 0.0011 |
| Scan | / | rtol | Binds | Used, first final | Kept, first final |
|---|---|---|---|---|---|
| Toy 1D | / | rtol | 4.2 14.4 | 92.08% 91.53% | |
| Polynomial | / | rtol | 7.0 13.4 | 92.25% 99.59% | |
| MNIST-LR | / | neither ( ) | 64.0 63.6 | 100.00% 100.00% | |
| MNIST-CE | / | rtol | 10.0 2.8 | 31.01% 95.17% | |
| CIFAR-LR | / | rtol | 128.0 81.6 | 100.00% 100.00% | |
| CIFAR-CE | / | rtol | 106.0 36.6 | 90.13% 76.69% |
| Scan | Probe rows | Spectrum width | Resolved rank, final step | |
|---|---|---|---|---|
| Toy 1D | 10,000 | 593 | 593 | 45.0 |
| Random Polynomial | 10,000 | 673 | 673 | 673.0 |
| MNIST (label reg.) | 512 | 27,562 | 512 | 512.0 |
| MNIST (CE) | 512 | 27,562 | 512 | 512.0 |