cs.LGSep 27, 2026

Predicting Block-Coordinate Performance via Cross-Curvature

Authors: Shengkun Zhu, Jinshan Zeng, Zhiqiang Kou, Yongxin Tong, Yang Liu

Organizations: The Hong Kong Polytechnic University · School of Management, Xi’an Jiaotong University · Beijing Key Laboratory of AI-Native Data Systems and SKLCCSE Lab, Beihang University

Abstract

Simultaneous and sequential block updates are two basic optimization strategies used across machine learning, such as neural-network training, federated learning, and low-rank adaptation. Choosing between them is difficult because their relative advantage depends on both the objective geometry and the number of iterations. We develop a unified theory for comparing Jacobi (JC), Gauss--Seidel (GS), and partially sequential deterministic block-gradient updates. Our analysis expresses the one-step loss difference through cross-block curvature, with an O(η3)O(η^3) remainder, where ηη is the learning rate. We derive a signed loss comparison after KK iterations with O(Kη3)O(Kη^3) error under regularity conditions and ηK≤TηK\le T for fixed TT, identifying the better method when the predicted difference exceeds this error. We evaluate these formulas along observed training trajectories across different machine learning settings. Over 500 iterations, our theory correctly identifies the lower-loss method in 98.0% of iterations for the neural network, 83.4% for federated learning, and 97.6% for LoRA. Applying the loss recursion at each step using the measured parameter difference raises these rates to 100.0%, 93.2%, and 99.6%, respectively.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

May 29, 2026cs.LG

Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks

We study optimal learning-rate selection in two-layer and three-layer linear neural networks trained to learn linear target functions. In particular, we derive the exact closed-form expressions for the gradients and test loss after one and two steps of gradient descent, enabling a precise characterization of early training dynamics. We characterize how learning rates should scale under the gradient approximation in the first two steps, and prove that performing updates with this approximation yields a tractable surrogate loss with a tight, small approximation error. This formulation enables the theoretical analysis of layer-wise learning rates and reveals a distinct early-training regime: test loss can be minimized by unequal learning rates at the initial step, while equal learning rates become optimal in subsequent steps. Our numerical experiments validate the theory and demonstrate the importance of balancing layer-wise learning rates early during training. The code is available at: https://github.com/TDCSZ327/Layer-Balancing.
Sep 8, 2026cs.LG

Equivariance Breaks the Learning Rate

Equivariant networks are commonly trained with Adam, yet recent work reports that matrix structured optimizers such as Muon can perform better, with the reasons for these gains only partly understood. We identify one source of this difference inside equivariant layers. An equivariant layer learns one channel mixing matrix WlW_l per degree ll, which we call an irrep block, and shares it across the 2l+12l+1 components, giving the expanded map Wl⊗I2l+1W_l \otimes I_{2l+1}. This sharing sums gradient contributions across components and can produce different update scales under SGD. Adam's entrywise normalization reduces sensitivity to gradient scale, but neither optimizer directly controls the effective step size of each block. A single learning rate can therefore produce different effective step sizes across blocks. Muon instead controls the effective step size by approximately equalizing the singular values of each momentum matrix. We normalize each irrep block update by a single scalar, preserving its singular value ratios while letting the learning rate control its size. We implement this with spectral normalization or a simpler root-mean-square normalization. We evaluate spectral normalization in a controlled SO(3)\mathrm{SO}(3)-equivariant model with a matched non-equivariant model. In this setting, the step size mismatch grows with width in the equivariant model but not in the non-equivariant model. We evaluate both variants across molecular force prediction on the rMD17 and MD22 datasets, QM9 molecular property prediction, and charged particle dynamics. Across these applications, block normalization generally improves Adam and closes part of its gap to Muon. These results highlight an overlooked interaction between equivariant architectures and their optimizers. Studying and designing the two together may help explain and address training difficulties often attributed to equivariance itself.
Sep 28, 2026cs.LG

CTP-FL: Common-Trajectory Gradient Prediction for Federated Learning

Communication-efficient federated optimization commonly spends several gradient evaluations between server updates. Existing local-update methods use this computation to advance an independent model on each client. Under heterogeneous data, however, these models evaluate gradients at different locations, making the aggregated update difficult to interpret as a gradient of the global objective. We study an alternative use of the same computation budget: \emph{evaluate the global objective along a shared, predicted path}. We propose Common-Trajectory Predictive Federated Learning (\texttt{CTP-FL}). At each round, all clients construct the same sequence of query points from the current global model and the previous aggregated direction, evaluate KK stochastic gradients along this sequence, and upload their average. The server then performs a single global update. Thus, \texttt{CTP-FL} uses KK mini-batch gradients per client and one model-sized vector in each communication direction, matching the per-round computation and communication of full-participation FedAvg-M. Shared query points make the aggregated direction an unbiased estimator of the average \emph{global} gradient along the predicted path. The remaining discrepancy from the gradient at the current model is controlled by the path length, without assuming bounded client-gradient dissimilarity or bounded gradients. For smooth non-convex objectives, we establish an O ⁣(LΔσ2/(NKR)+LΔ/R)\mathcal{O}\!\left( \sqrt{LΔσ^2/(NKR)}+LΔ/R \right) average-stationarity bound under full participation. The analysis isolates a testable trade-off: extending the prediction path provides more forward-looking gradient information but increases its displacement bias.