Fast Convergence through Distributed Augmentation for Class-Imbalanced Federated Learning
Authors: Arathi Nair M, J. Harshan, Anwitaman Datta
Organizations: Amar Nath and Shashi Khosla School of Information Technology, Indian Institute of Technology Delhi, India · Department of Electrical Engineering, Indian Institute of Technology Delhi, India · College of Computing and Data Science, Nanyang Technological University, Singapore
In federated learning, mitigating class imbalance is essential to improve minority-class performance. A common approach to address this problem is to augment minority-class samples to achieve local class balance. Existing approaches treat augmentation as a heuristic and do not establish how the amount of augmentation influences the convergence of federated learning, leading to excessive augmentation and increased training time. To address this limitation, we first establish the relationship between augmentation and the convergence behavior of federated learning. Leveraging this insight, we propose DAFL, a distributed augmentation framework that determines the minimum augmentation required for each client-class pair by jointly minimizing augmentation and training time while constraining global class imbalance, thereby improving minority-class F1-score. Experimental results demonstrate that DAFL consistently improves minority-class F1-score while substantially reducing training time, particularly under severe global class imbalance and high label proportion imbalance.
Figures & tables
Category / References
CI
MS
CD
LA
Limit.
Loss modification
Imbalance [ 6 , 7 ]
✔
✘
⚫
✘
Does not increase diversity of minority-class data
Non-IID [ 10 , 11 , 12 , 13 ]
⚫
✘
✔
✘
Aggregation- based
Imbalance [ 8 , 9 ]
✔
✘
⚫
✘
Non-IID [ 14 , 15 , 16 ]
⚫
✘
✔
✘
Data augmentation [ 17 , 18 , 19 , 20 , 21 , 23 ]
✔
✔
⚫
✘
Heuristic augmentation policies
DAFL (proposed)
✔
✔
✔
✔
TABLE I: Comparison of existing methods with DAFL. ✔, ✘, and ⚫ denote yes, no, and indirect support, respectively. Indirect: addressed by some works or achieved as a byproduct. CI: class imbalance; MS: minority samples; CD: client drift; LA: latency-aware; Limit.: severe classimbalance limitation.
D1
D2
I / J
0.0002 / 0.4097
0.3982 / 0.4097
Desc.
Low I , High J
High I , High J
C1
C2
C3
C4
C5
C6
C7
C8
C1
C2
C3
C4
C5
C6
C7
C8
Client 1
52
18
61
2
30
612
70
5
1
126
1
1
1
1
485
1
Client 2
63
21
74
3
28
701
65
6
1
1
1
10
248
1
1
1
Client 3
48
20
69
4
33
5800
60
4
395
2
1
1
1
1
1
1
TABLE II: Dataset details for D1–D6, including client-class counts (C1–C8 denote Class 1–Class 8) and the corresponding I and J values. D1–D3 isolate extreme I / J settings while D4–D6 cover intermediate settings. Here, the descriptors low/moderate/high are relative rather than specific thresholds. Also, J∈[0,CC−1] .
Dataset
I
J
Added Samples
Time for Augmentation (s)
Time at Server to find A^ (s)
LB
DAFL
LB
DAFL
LB
DAFL
LB
DAFL
q=2
q=3
q=4
q=2
q=3
q=4
q=2
q=3
q=4
q=2
q=3
q=4
q=2
q=3
q=4
D1
0
0.0001
0.0001
0.0001
0.001
0.001
0.001
0.001
24913
24909
24909
24909
0.993
0.390
0.291
0.373
33.666
35.219
34.900
D2
0.0005
0.1479
0.1479
0.0798
0.0005
0.001
0.001
0.001
35571
24913
24913
24911
7.258
0.351
0.343
0.364
34.782
38.081
40.032
D3
0.0009
0.0558
0.2836
0.3492
0.0001
0.001
0.001
0.001
81622
39060
3080
7
5.526
0.682
0.042
0.297
55.485
4.238
0.012
D4
0.0003
0.0795
0.0795
0.0461
0.0007
0.001
0.001
0.001
28020
24309
24309
24311
4.406
0.334
0.395
0.318
32.878
32.897
34.163
TABLE III: Updated I , J , number of augmented samples, and time associated with augmentation for LB and DAFL (λ2=10q) .
Fig. 1: Heatmap for client-class counts for D2 (a) before augmentation, (b) after LB, and (c) after DAFL.
Fig. 2: Performance on datasets D1 (a)–(c), D2 (d)–(f), D3 (g)–(i), D4 (j)–(l), D5 (m)–(o), and D6 (p)–(r), with each triplet showing minority-class F1-score, accuracy, and macro F1-score, respectively.
Federated learning has emerged as the foremost approach for decentralized model training with privacy preservation. The global class imbalance and cross-client data heterogeneity naturally coexist, and the mismatch between local and global imbalances exacerbates the performance degradation of the aggregated model. The agnosticism of global class distribution poses significant challenges for data-level methods, especially under extreme conditions with severe class absence across clients. In this paper, we propose FedReLa, a novel data-level approach that tackles the coexistence of data heterogeneity and class imbalance in federated learning. By re-labeling samples with a feature-dependent label re-allocator, FedReLa corrects biased global decision boundaries without requiring knowledge of the global class distribution. This modular, model-agnostic approach can be integrated with algorithmic methods to deliver consistent improvements without additional communication overhead. Through extensive experiments, our method significantly improves the accuracy of minority classes and the overall accuracy on stepwise-imbalanced and long-tailed datasets, outperforming the previous state of the art.
Guangzheng Hu, Patricia Menéndez, Feng Liu +3
School of Mathematics and Statistics, University of Melbourne, Victoria, Australia · School of Computing and Information Systems, University of Melbourne, Victoria, Australia · School of Statistics and Data Science, LPMC, KLMDASR, and LEBPS, Nankai University
Class imbalance poses a critical challenge in federated learning (FL), where underrepresented classes suffer from poor predictive performance yet cannot be addressed by standard centralized techniques due to privacy and heterogeneity constraints. We propose FedCGNM (Federated Class-Grouped Normalized Momentum), a client-side optimizer in FL that partitions classes into a small number of groups based on minimum within-group variance, maintains a momentum per group, normalizes each group momentum to unit length, and uses the summation of the normalized group momentums as an update direction. This design both equalizes gradient magnitude across majority and minority groups and mitigates the noise inherent in rare-class gradients. We further provide a theoretical convergence analysis explicitly accounting for time-varying resampling-rates. Additionally, to efficiently optimize these rates in small-client regimes, we introduce FedHOO, an X-armed-bandit (XAB) based algorithm that exploits federated parallelism that evaluates many combinations of two candidate rates per client at linear cost. Empirical evaluation on four public long-tailed benchmarks and a proprietary chip-defect dataset demonstrates that FedCGNM consistently outperforms baselines, with FedHOO yielding further gains in small-scale federations.
Haemin Park, Diego Klabjan, Martin W. Braun +2
Department of Industrial Engineering & Management Sciences, Northwestern University, Evanston, IL, USA · Intel Corporation, Chandler, AZ, USA
Label skew in federated learning (FL) causes client drift and degrades global accuracy. Synthetic data augmentation can reduce this imbalance; however, full class balancing requires substantial computation cost. We propose FedEAS, a policy that assigns each client an entropy-adaptive per-class generation budget computed from its local label distribution. The budget jointly decides \emph{how much} each client generates and \emph{WHERE} the samples go. Accordingly, the total generation budget follows from the per-client budgets rather than being fixed in advance. FedEAS recovers most of the accuracy gain of full class balancing while reducing the generation budget by 94.1%. At the same total generation budget, it outperforms Uniform allocation by up to 18.82% across CIFAR-10 and CIFAR-100.