Standard weight decay treats each weight matrix as a vector and ignores its spectral structure. We introduce spectral weight decay, a post-step decoupled nuclear-norm update that applies additive rather than multiplicative spectral shrinkage. We connect the update to approximate proximal descent and show that its sensitivity to update order can exceed that of conventional ℓ2 weight decay near rank deficiency. Across LLaMA models with 124M to 500M parameters, spectral weight decay lowers effective rank and improves SVD-LLM compression at matched validation loss. At 500M and a 4% distortion budget, it reaches 1.89× compression and 1.18× GPU inference speedup, compared with 1.14× and 1.01× after standard weight decay. Under fixed-horizon training with 60% label noise, it also improves final mean clean-test accuracy over matched ℓ2 regularization by up to 17.8 points on MNIST and 4.6 points across four BERT-base tasks. Code is available at https://github.com/brain-lab-research/SpectralWD.
Figures & tables
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
124 M
257 M
500 M
Transformer blocks
12
16
22
Hidden dimension
768
1,024
1,280
Attention heads
12
16
20
Feed-forward dimension
2,048
2,816
3,584
Embedding vocabulary size
50,304
50,304
50,304
Training steps
19,000
39,000
76,294
Appendix
Table 1: LLaMA architectures and reference pretraining configurations. Token budgets are rounded.
Method
Compression
val loss
Δ val loss
ARC-E
HellaSwag
PIQA
baseline (uncompressed)
–
3.1129
–
0.4947
0.3004
0.6088
truncated SVD
1.43×
3.1488
+0.0359
0.4947
0.2999
0.6099
SliceGPT
1.22×
4.6347
+1.5218
0.4140
0.2815
0.5680
ASVD
1.45×
3.1416
+0.0286
0.4877
0.2998
0.6121
SVD-LLM
1.42×
3.1325
+0.0196
0.4982
0.2991
0.6088
Dobi-SVD
1.21×
3.1336
+0.0207
0.4912
0.3005
0.6110
Appendix
Table 2: Five compression methods applied to the 124 M spectral-weight-decay model ( λ=1 ). Lower losses are better. Higher compression rates and downstream accuracies are better. Bold marks the best compressed value in each column.
Method
10% noise
25% noise
40% noise
60% noise
No WD
87.99±0.93
79.52±1.67
65.29±3.27
45.71±1.79
ℓ2 WD
86.89±1.70
77.90±2.00
67.12±2.71
45.78±1.29
Spectral WD (post-step)
87.99±0.93
81.22±2.36
75.01±3.01
63.56±1.04
Spectral WD (pre-step)
87.31±1.13
81.63±1.77
75.52±2.35
61.63±2.74
Appendix
Table 3: MNIST MLP: final clean-test accuracy, mean ± one sample standard deviation.
Method
10% noise
25% noise
40% noise
60% noise
No WD
91.29±0.64
82.42±1.18
71.50±1.13
51.65±2.12
ℓ2 WD
91.07±1.55
82.98±0.54
72.69±0.84
51.20±2.78
Spectral WD (post-step)
94.06±1.03
89.96±0.97
81.38±3.15
65.60±3.70
Spectral WD (pre-step)
94.83±0.62
90.00±1.17
80.02±2.03
67.33±1.52
Appendix
Table 4: MNIST GRU: final clean-test accuracy, mean ± one sample standard deviation.
Dataset
Noise
No WD
ℓ2 -SP
Spectral WD (post-step)
Spectral WD (pre-step)
AG News
10%
88.49±0.10
89.63±0.03
90.57±0.10
90.55±0.11
25%
79.23±0.47
88.60±0.21
89.77±0.11
89.76±0.13
40%
66.02±0.50
86.65±0.55
88.56±0.33
88.61±0.31
60%
44.05±0.85
78.40±1.91
81.07±1.14
81.08±1.10
DBpedia-14
10%
97.97±0.11
98.72±0.01
98.86±0.03
98.85±0.04
25%
95.19±0.18
98.52±0.13
98.72±0.10
98.72±0.10
Appendix
Table 5: BERT-base: final clean-test accuracy, mean ± one sample standard deviation.
Modern Transformer architectures frequently employ normalization mechanisms such as RMSNorm and Query-Key Normalization, making parts of the model approximately scale-invariant with respect to weight magnitudes. In this regime, standard Frobenius-norm weight decay acts purely along the radial direction of the weight space and cannot directly simplify the function represented by the normalized layer. We study grokking in small algorithmic tasks through this lens and propose \emph{Low-Rank Decay} (LRD), a nuclear-norm-like spectral regularizer whose subgradient -- the polar factor UV⊤ -- retains a tangential component even in the scale-invariant setting. This distinction has a concrete dynamical consequence: after the model memorizes the training set and task gradients vanish, L2 decay can no longer reshape the weight spectrum, whereas LRD continues to compress singular values in an ℓ1-like fashion. On modular arithmetic tasks, we find that LRD induces rapid effective-rank collapse in Query/Key matrices and expands the data-fraction boundary at which delayed generalization (grokking) occurs. We further provide a spectral-geometric interpretation through the ``needle-to-fan'' expansion of the nuclear-norm subdifferential near low-rank strata.
The discovery of scaling laws has motivated training neural networks on ever increasing quantities of data. This is typically done with a constant decoupled weight decay which causes the network weights to shrink steadily over the course of training. Taking inspiration from the Robbins--Monro conditions, we propose to scale weight decay by the fraction of the peak learning rate η/ηmax. We prove that this scaled weight decay preserves the asymptotic stationarity guarantees of the corresponding unregularized methods for both stochastic gradient descent and the non-Euclidean spectral optimizer Muon, thereby avoiding the additional asymptotic bias introduced by constant decoupled weight decay. This retains the stability benefits of weight decay without changing the asymptotic optimization target. Using a steady-state analysis, we explain why under standard weight decay the weight norm shrinks steadily as training proceeds, whereas under scaled weight decay it settles to a roughly constant value. When applied to the training of mixture-of-experts models, Muon with scaled weight decay (Muon-SW) consistently outpaces Muon with identical hyperparameters, reaching the same validation loss 30% faster at our largest scale across models from 72−930 million parameters trained at ∼600 tokens per active parameter. If this trend continues to hold, the method promises to substantially accelerate the pre-training of frontier models while requiring only a few lines of code to implement.
Anuj Apte
Global Technology Applied Research, JPMorganChase, New York, NY 10001
Matrix-level low-rank compression is a promising way to reduce the cost of large language models, but running compression and evaluating the resulting models on language tasks can be prohibitively expensive. Can compression-induced degradation be predicted before committing to this compute? We systematically analyze the Qwen3 and Gemma3 model families across four representative low-rank compression methods: vanilla SVD, two ASVD variants, and SVD-LLM. We find that stable rank and information density, measured in bits per parameter, dominate performance degradation. The interaction term γ⋅ρˉs, defined as compression ratio times stable rank, is a robust predictor of accuracy degradation, achieving leave-one-out cross-validation Pearson correlations of 0.890 for attention layers and 0.839 for MLP layers. We provide theoretical intuition for why this predictor succeeds by connecting it to standard SVD truncation bounds and error composition mechanisms in transformer layers. These findings enable a predict-then-compress workflow: compute γ⋅ρˉs from weights, estimate degradation, and invest compute only in desirable configurations.
Mingxue Xu
Department of Electrical and Electronic Engineering, Imperial College London, London, United Kingdom