Predicting Block-Coordinate Performance via Cross-Curvature
Organizations: The Hong Kong Polytechnic University · School of Management, Xi’an Jiaotong University · Beijing Key Laboratory of AI-Native Data Systems and SKLCCSE Lab, Beihang University
Abstract
Simultaneous and sequential block updates are two basic optimization strategies used across machine learning, such as neural-network training, federated learning, and low-rank adaptation. Choosing between them is difficult because their relative advantage depends on both the objective geometry and the number of iterations. We develop a unified theory for comparing Jacobi (JC), Gauss--Seidel (GS), and partially sequential deterministic block-gradient updates. Our analysis expresses the one-step loss difference through cross-block curvature, with an remainder, where is the learning rate. We derive a signed loss comparison after iterations with error under regularity conditions and for fixed , identifying the better method when the predicted difference exceeds this error. We evaluate these formulas along observed training trajectories across different machine learning settings. Over 500 iterations, our theory correctly identifies the lower-loss method in 98.0% of iterations for the neural network, 83.4% for federated learning, and 97.6% for LoRA. Applying the loss recursion at each step using the measured parameter difference raises these rates to 100.0%, 93.2%, and 99.6%, respectively.
Figures & tables
| Task | Estimate | Correct | JC hits | GS hits | MAE | Max. AE |
|---|---|---|---|---|---|---|
| DNN | Cumulative | 98.0% | 45/45 | 445/455 | ||
| Recursive | 100.0% | 45/45 | 455/455 | |||
| FL | Cumulative | 83.4% | 25/28 | 392/472 | ||
| Recursive | 93.2% | 25/28 | 441/472 | |||
| LoRA | Cumulative | 97.6% | 74/75 | 414/425 | ||
| Recursive | 99.6% | 74/75 | 424/425 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Example | Parameter blocks | Update semantics | |
|---|---|---|---|
| Simultaneous | Sequential | ||
| MLP MNIST | One weight–bias pair per layer | JC. Compute all gradients before updating any parameters. | GS-like . Update output-to-input, using updated weights in the backward recursion. |
| ResNet-18 CIFAR-10 | Client-specific input layers ; shared output layers | FedSim. Joint local updates of . | FedAlt. Personal phase, then shared phase using the updated personal weights. |
| LoRA SST-2 | All factors; all factors | JC. One forward–backward pass at . | GS. Update , then recompute the network to update . |