Standard weight decay treats each weight matrix as a vector and ignores its spectral structure. We introduce spectral weight decay, a post-step decoupled nuclear-norm update that applies additive rather than multiplicative spectral shrinkage. We connect the update to approximate proximal descent and show that its sensitivity to update order can exceed that of conventional ℓ2 weight decay near rank deficiency. Across LLaMA models with 124M to 500M parameters, spectral weight decay lowers effective rank and improves SVD-LLM compression at matched validation loss. At 500M and a 4% distortion budget, it reaches 1.89× compression and 1.18× GPU inference speedup, compared with 1.14× and 1.01× after standard weight decay. Under fixed-horizon training with 60% label noise, it also improves final mean clean-test accuracy over matched ℓ2 regularization by up to 17.8 points on MNIST and 4.6 points across four BERT-base tasks. Code is available at https://github.com/brain-lab-research/SpectralWD.
Figures & tables
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
124 M
257 M
500 M
Transformer blocks
12
16
22
Hidden dimension
768
1,024
1,280
Attention heads
12
16
20
Feed-forward dimension
2,048
2,816
3,584
Embedding vocabulary size
50,304
50,304
50,304
Training steps
19,000
39,000
76,294
Appendix
Table 1: LLaMA architectures and reference pretraining configurations. Token budgets are rounded.
Method
Compression
val loss
Δ val loss
ARC-E
HellaSwag
PIQA
baseline (uncompressed)
–
3.1129
–
0.4947
0.3004
0.6088
truncated SVD
1.43×
3.1488
+0.0359
0.4947
0.2999
0.6099
SliceGPT
1.22×
4.6347
+1.5218
0.4140
0.2815
0.5680
ASVD
1.45×
3.1416
+0.0286
0.4877
0.2998
0.6121
SVD-LLM
1.42×
3.1325
+0.0196
0.4982
0.2991
0.6088
Dobi-SVD
1.21×
3.1336
+0.0207
0.4912
0.3005
0.6110
Appendix
Table 2: Five compression methods applied to the 124 M spectral-weight-decay model ( λ=1 ). Lower losses are better. Higher compression rates and downstream accuracies are better. Bold marks the best compressed value in each column.
Method
10% noise
25% noise
40% noise
60% noise
No WD
87.99±0.93
79.52±1.67
65.29±3.27
45.71±1.79
ℓ2 WD
86.89±1.70
77.90±2.00
67.12±2.71
45.78±1.29
Spectral WD (post-step)
87.99±0.93
81.22±2.36
75.01±3.01
63.56±1.04
Spectral WD (pre-step)
87.31±1.13
81.63±1.77
75.52±2.35
61.63±2.74
Appendix
Table 3: MNIST MLP: final clean-test accuracy, mean ± one sample standard deviation.
Method
10% noise
25% noise
40% noise
60% noise
No WD
91.29±0.64
82.42±1.18
71.50±1.13
51.65±2.12
ℓ2 WD
91.07±1.55
82.98±0.54
72.69±0.84
51.20±2.78
Spectral WD (post-step)
94.06±1.03
89.96±0.97
81.38±3.15
65.60±3.70
Spectral WD (pre-step)
94.83±0.62
90.00±1.17
80.02±2.03
67.33±1.52
Appendix
Table 4: MNIST GRU: final clean-test accuracy, mean ± one sample standard deviation.
Dataset
Noise
No WD
ℓ2 -SP
Spectral WD (post-step)
Spectral WD (pre-step)
AG News
10%
88.49±0.10
89.63±0.03
90.57±0.10
90.55±0.11
25%
79.23±0.47
88.60±0.21
89.77±0.11
89.76±0.13
40%
66.02±0.50
86.65±0.55
88.56±0.33
88.61±0.31
60%
44.05±0.85
78.40±1.91
81.07±1.14
81.08±1.10
DBpedia-14
10%
97.97±0.11
98.72±0.01
98.86±0.03
98.85±0.04
25%
95.19±0.18
98.52±0.13
98.72±0.10
98.72±0.10
Appendix
Table 5: BERT-base: final clean-test accuracy, mean ± one sample standard deviation.