cs.LGSep 30, 2026

Robustifying Asynchronous SGD via Soft Throttling

Authors: Kaoru Otsuka, Maxime Meyer, Yuki Takezawa, Makoto Yamada, Anastasia Koloskova

Organizations: Okinawa Institute of Science and Technology · National University of Singapore · Toyota Motor Corporation · University of Zurich

Abstract

Asynchronous SGD is a popular algorithm for distributed learning where each client's gradient update is applied on arrival. This leads to a speed-up, but also an increased vulnerability to attacks, as fast clients can dominate the total update. We introduce Throttle, a Byzantine-robust generalization of asynchronous SGD where the key idea is to exponentially down-weight updates from faster clients by a factor qq. Both asynchronous SGD (q=1q=1) and synchronous Byzantine-robust SGD (q→∞q\to\infty) correspond to specific settings of Throttle. We provide a theoretical analysis of the convergence rate and validate the robustness to attacks both theoretically and empirically. Remarkably, our experiments show that this down-weighting mechanism can also improve performance over standard asynchronous SGD even in the non-Byzantine setting.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Bringing Order to Asynchronous SGD: Towards Optimality under Data-Dependent Delays with Momentum

    May 3, 2026Tehila Dahan, Roie Reshef, Sharon Goldstein +1Stochastic Gradient DescentAsynchronous Execution

  2. Clipping Makes Distributed and Federated Asynchronous SGD Robust to Stragglers

    Jun 11, 2026Samuel Erickson, Mikael JohanssonStochastic Gradient DescentGradient Clipping