cs.LGSep 27, 2026

Predicting Block-Coordinate Performance via Cross-Curvature

Authors: Shengkun Zhu, Jinshan Zeng, Zhiqiang Kou, Yongxin Tong, Yang Liu

Organizations: The Hong Kong Polytechnic University · School of Management, Xi’an Jiaotong University · Beijing Key Laboratory of AI-Native Data Systems and SKLCCSE Lab, Beihang University

Abstract

Simultaneous and sequential block updates are two basic optimization strategies used across machine learning, such as neural-network training, federated learning, and low-rank adaptation. Choosing between them is difficult because their relative advantage depends on both the objective geometry and the number of iterations. We develop a unified theory for comparing Jacobi (JC), Gauss--Seidel (GS), and partially sequential deterministic block-gradient updates. Our analysis expresses the one-step loss difference through cross-block curvature, with an O(η3)O(η^3) remainder, where ηη is the learning rate. We derive a signed loss comparison after KK iterations with O(Kη3)O(Kη^3) error under regularity conditions and ηK≤TηK\le T for fixed TT, identifying the better method when the predicted difference exceeds this error. We evaluate these formulas along observed training trajectories across different machine learning settings. Over 500 iterations, our theory correctly identifies the lower-loss method in 98.0% of iterations for the neural network, 83.4% for federated learning, and 97.6% for LoRA. Applying the loss recursion at each step using the measured parameter difference raises these rates to 100.0%, 93.2%, and 99.6%, respectively.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks

    May 29, 2026Tianyu Pang, Vignesh Kothapalli, Shenyang Deng +3BatchTwo-Layer Neural Networks

  2. Equivariance Breaks the Learning Rate

    Sep 8, 2026Andrei Manolache, Mathias NiepertAdamEquivariant Neural Networks