Joint Affine Spectral Shaping: Coupling Weight and Bias Updates Beyond Weight-Only Muon
Organizations: State Key Laboratory of Robotics and Systems, Harbin Institute of Technology, Shenzhen, Shenzhen, 518055, China.
Abstract
Matrix spectral optimizers reshape weight-update spectra but usually delegate vector-valued biases to a separate optimizer. We study whether this separation is neutral. We formulate each affine layer as a joint momentum matrix and apply a capped regularized-inverse spectral map to the complete matrix, producing both the weight and physical bias updates. A strict five-seed ablation on a four-layer BERT-mini trained from scratch on IMDb compares exact-SVD Muon, weight-only inverse shaping, affine-probe inverse shaping, and the proposed joint regularized inverse (JRI). Weight-only inverse shaping raises validation-loss-selected test accuracy from to and lowers selected test loss from to . Allowing bias to alter the joint SVD while retaining an independent Adam bias update does not improve over weight-only inverse shaping. Using the transformed bias jointly raises selected test accuracy to and lowers test loss to , with all five seeds improving relative to the probe baseline. During the peak-performance window, JRI preserves the eligible weight-update norm while reducing the bias-update norm from to , lowers boundary-function share from to , and changes the cosine between weight-induced boundary motion and explicit bias from to . An independent 22-seed replication yields selected test accuracy. These results identify joint affine spectral allocation as a small but consistent extension to weight-only spectral optimization.