cs.LGOct 4, 2026

Fast Convergence through Distributed Augmentation for Class-Imbalanced Federated Learning

Authors: Arathi Nair M, J. Harshan, Anwitaman Datta

Organizations: Amar Nath and Shashi Khosla School of Information Technology, Indian Institute of Technology Delhi, India · Department of Electrical Engineering, Indian Institute of Technology Delhi, India · College of Computing and Data Science, Nanyang Technological University, Singapore

Abstract

In federated learning, mitigating class imbalance is essential to improve minority-class performance. A common approach to address this problem is to augment minority-class samples to achieve local class balance. Existing approaches treat augmentation as a heuristic and do not establish how the amount of augmentation influences the convergence of federated learning, leading to excessive augmentation and increased training time. To address this limitation, we first establish the relationship between augmentation and the convergence behavior of federated learning. Leveraging this insight, we propose DAFL, a distributed augmentation framework that determines the minimum augmentation required for each client-class pair by jointly minimizing augmentation and training time while constraining global class imbalance, thereby improving minority-class F1-score. Experimental results demonstrate that DAFL consistently improves minority-class F1-score while substantially reducing training time, particularly under severe global class imbalance and high label proportion imbalance.

Figures & tables

Explore similar work

Jun 24, 2026stat.ML

FedReLa: Imbalanced Federated Learning via Re-Labeling

Federated learning has emerged as the foremost approach for decentralized model training with privacy preservation. The global class imbalance and cross-client data heterogeneity naturally coexist, and the mismatch between local and global imbalances exacerbates the performance degradation of the aggregated model. The agnosticism of global class distribution poses significant challenges for data-level methods, especially under extreme conditions with severe class absence across clients. In this paper, we propose FedReLa, a novel data-level approach that tackles the coexistence of data heterogeneity and class imbalance in federated learning. By re-labeling samples with a feature-dependent label re-allocator, FedReLa corrects biased global decision boundaries without requiring knowledge of the global class distribution. This modular, model-agnostic approach can be integrated with algorithmic methods to deliver consistent improvements without additional communication overhead. Through extensive experiments, our method significantly improves the accuracy of minority classes and the overall accuracy on stepwise-imbalanced and long-tailed datasets, outperforming the previous state of the art.
Jul 1, 2026cs.LG

Class-Grouped Normalized Momentum and Faster Hyperparameter Exploration to Tackle Class Imbalance in Federated Learning

Class imbalance poses a critical challenge in federated learning (FL), where underrepresented classes suffer from poor predictive performance yet cannot be addressed by standard centralized techniques due to privacy and heterogeneity constraints. We propose FedCGNM (Federated Class-Grouped Normalized Momentum), a client-side optimizer in FL that partitions classes into a small number of groups based on minimum within-group variance, maintains a momentum per group, normalizes each group momentum to unit length, and uses the summation of the normalized group momentums as an update direction. This design both equalizes gradient magnitude across majority and minority groups and mitigates the noise inherent in rare-class gradients. We further provide a theoretical convergence analysis explicitly accounting for time-varying resampling-rates. Additionally, to efficiently optimize these rates in small-client regimes, we introduce FedHOO, an X-armed-bandit (XAB) based algorithm that exploits federated parallelism that evaluates many combinations of two candidate rates per client at linear cost. Empirical evaluation on four public long-tailed benchmarks and a proprietary chip-defect dataset demonstrates that FedCGNM consistently outperforms baselines, with FedHOO yielding further gains in small-scale federations.
Jul 7, 2026cs.LG

WHERE to Generate Matters: Budget-Aware Synthetic Augmentation for Label Skewed Federated Learning

Label skew in federated learning (FL) causes client drift and degrades global accuracy. Synthetic data augmentation can reduce this imbalance; however, full class balancing requires substantial computation cost. We propose FedEAS, a policy that assigns each client an entropy-adaptive per-class generation budget computed from its local label distribution. The budget jointly decides \emph{how much} each client generates and \emph{WHERE} the samples go. Accordingly, the total generation budget follows from the per-client budgets rather than being fixed in advance. FedEAS recovers most of the accuracy gain of full class balancing while reducing the generation budget by 94.1%. At the same total generation budget, it outperforms Uniform allocation by up to 18.82% across CIFAR-10 and CIFAR-100.