Large Stepsizes Federated Learning on Logistic Regression with Linearly Separable Data: The Case of Heterogeneous Devices
Authors: Hok Fong Wong, Hoi-To Wai, Chung-Yiu Yau
Organizations: Department of CSE, The Chinese University of Hong Kong, Hong Kong SAR · Department of SEEM, The Chinese University of Hong Kong, Hong Kong SAR · Department of ECE, University of Minnesota, Minnesota, USA
This paper revisits the distributed learning problem for training a multinomial logistic regression model with the Federated Averaging (FedAvg) algorithm. We concentrate on a scenario with arbitrarily large stepsizes and heterogeneous update rules where the devices may perform a different number of local updates in each round. We show that, with linearly separable data, FedAvg is stable with any stepsizes and the objective values converge to zero at the rate of O(1/R), where R is the number of communication rounds. Our result also demonstrates that the effects of device heterogeneity vanish asymptotically. For sufficiently large R, the objective values decrease monotonically and is bounded by O(1/(RTavg)), where Tavg is the average number of local update steps per communication round across devices. Numerical experiments support our findings.
Figures & tables
Fig. 1: Convergence of FedAvg with homogeneous (“homo. device”: Ti=27 for all devices) and heterogeneous (“hete. device”: Ti=4 for device 1 - 4 , Ti=50 for device 5 - 8 ) devices. Both configurations share the same average Tavg=27 with global/local stepsizes: (Left) ηg=64,ηl=1 , (Right) ηg=128,ηl=1 .
Fig. 2: Effects of the number of local steps T . All devices perform the same number of local epochs T∈{1,4,16,64} , with global stepsizes ηg∈{1,64,128} and local stepsize ηl=1 .
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Fig. 3: Effect of hard-sample placement under heterogeneous local step counts. All configurations share the same set of data and the same average local work Tavg=57 . The hardest samples are assigned to either straggler clients with Ti=13 ( hardLowT ), fast clients with Ti=101 ( hardHighT ) under the heterogeneous setup, or arbitrarily for the homogeneous setup where all Ti=57 ( homo ). We use the global stepsizes ηg∈{32,64} , local stepsize ηl=1 .
Fig. 4: Effect of heterogeneous local update steps under minibatch SGD. All devices perform local updates using minibatches of size b∈{5,25,125} , with global stepsizes ηg∈{32,64} and local stepsize ηl=1 . Two configurations are compared at a matched average Tavg=57 : homogeneous ( Ti=57 for all devices) and heterogeneous ( Ti=13 for devices 1-4, Ti=101 for devices 5-8).
School of Mathematics, China University of Mining and Technology, Xuzhou, China · School of Mathematical Sciences, Dalian University of Technology, Dalian, China