Client and Training Data Selection for Computationally Efficient Synchronized Federated Learning
Authors: Muzaffer Citir, Hiroki Nishikawa, Sangyoung Park
Organizations: Smart Mobility Systems, Technical University of Berlin, Germany · Graduate School of Information Science and Technology, The University of Osaka, Japan
Federated learning (FL) is a promising paradigm of machine learning, which preserves user privacy by enabling learning without sharing raw data with a cloud server. Straggling clients have been a problem for FL as they introduce delays in aggregating the local models and hence, the convergence of the global model. Therefore, it is important to have a mechanism that ensures fast convergence of the global model as well as good FL participation rate. Another issue for the convergence of a model in FL is the non-independent and identically distributed (non-iid) data across the clients. Prior approaches based on probabilistic client selection do not work well under non-iid data especially when the number of clients is small. We show scenarios where such approaches fail and propose a joint client-training data selection algorithm for fast convergence of FL models. Our experiments on CIFAR-100 dataset show that convergence of the FL model can be significantly improved over prior works that can consider non-iid data and heterogeneous computation and higher model accuracy.
Figures & tables
Fig. 1: (a) A deadline miss example of FL devices scheduling, and (b) avoiding it with proper data allocation.
Fig. 2: Training time vs dataset size for CIFAR-10 dataset.
Fig. 3: Time distribution with 5 ×103 images of CIFAR-10.
Component
Specification
CPU
Intel i9-12900KF 3,2 GHz
GPU
NVIDIA GeForce RTX 4090 24 GB (CUDA 12.2)
RAM
64 GB DDR5-SDRAM 4800 MHz
Operating System
Ubuntu 22.04.3 LTS
TABLE I: Hardware specifications for FL Training
Fig. 4: Test accuracy over time for the proposed, probPart and MinCost algorithms for CIFAR-100 dataset distributed across 50 clients. Shaded area denotes 1 std range over 6 experiments with random seeds for data distribution.
Fig. 5: Component ablation of the usefulness metric (Eq. 13 ) as test accuracy over training time, under two degrees of non-iid severity: (a) α=0.3 and (b) α=0.1 . Each variant disables one term of uk ; for w1 , its two sub-factors (the freshness penalty and the client dataset-size preference) are disabled separately. Proposed keeps all terms, while probPart and MinCost are shown for reference. The experiments use the same client, model, and deadline configuration as Fig. 4 , and are reported for a single random seed.
Method
Overall
Tail-10
CV
Gini
acc. (%)
acc. (%)
Proposed (full)
42.3
15.3
0.43
0.24
probPart
38.0
6.8
0.49
0.28
MinCost
31.2
3.5
0.62
0.35
TABLE II: Per-class fairness on the CIFAR-100 test set at α=0.3 for each method’s final trained model (single run): overall accuracy, tail-class accuracy (mean of the 10 lowest-accuracy classes), and the dispersion of the per-class accuracies (coefficient of variation and Gini coefficient; lower is more uniform).
Federated learning (FL) is a popular distributed learning framework where multiple clients perform local training and a server aggregates the locally updated models. FL enables decentralized training while preserving the privacy of clients' datasets. However, non-independent and identically distributed (non-IID) or noisy datasets can lead to low model accuracy or high convergence latency. Precluding these clients through client selection may mitigate the problem, but heavily biased client selections may also degrade the learning performance. In this study, we first experimentally measure the impact of non-IID data (including skews in data quantity and label distribution), noisy data, and fairness in client selection on model accuracy and convergence. We then propose a privacy-preserving scoring method to assess each client's contribution in FL, with experiments conducted to demonstrate the effectiveness of the proposed assessment.
Yuan-Heng Tsai, Li-Hsing Yen, Yan-Wei Chen
Department of Computer Science, National Yang Ming Chiao Tung University, Hsinchu, Taiwan.
Federated Learning enables collaborative model training across decentralized data sources without data transfer. Averaging-based FL is limited by the presence of non-IID data, which negatively impacts convergence speed and final model accuracy. Conventional alternatives suffer from significant inefficiency. Clients with noisy or highly heterogeneous data contribute expensive gradient computations that are either discarded or heavily down-weighted before aggregation. These reactive approaches waste computational resources, require more communication rounds and result in unnecessary privacy exposure. In this paper, we propose a proactive client selection framework that aims to find an optimal federation of clients whose combined data match utility and fairness requirements before training begins. Our method relies on mutual information computed from differentially private contingency tables to quantify the relevance of cross-feature correlations in the union dataset. We introduce a Potential Federation Loss (PFL) over the set of fixed-size federations, which balances two objectives. Maximizing collective data utility while ensuring fair cross-features correlations to prevent group unfairness. Client selection is expressed as an optimal subset search problem over the PFL objective, which we solve using simulated annealing under strong differential privacy guarantees for clients' local statistics. Experimental results on four benchmarks show faster, fairer, and more accurate models trained on optimally found federations, compared to uniform sampling, even when state-of-the-art adaptive aggregation or sampling strategies are employed.
There are two categories of methods in Federated Learning (FL) for joint training across multiple clients: (i) parallel FL (PFL), where clients train models in a parallel manner; and (ii) sequential FL (SFL), where clients train models in a sequential manner. In contrast to that of PFL, the convergence theory of SFL on heterogeneous data is still lacking. In this paper, we establish the convergence guarantees of SFL for strongly/general/non-convex objectives on heterogeneous data. The convergence guarantees of SFL are better than that of PFL on heterogeneous data with both full and partial client participation. Experimental results validate the counterintuitive analysis result that SFL outperforms PFL on extremely heterogeneous data in cross-device settings.
Yipeng Li, Xinchen Lyu
National Engineering Research Center for Mobile Network Technologies Beijing University of Posts and Telecommunications Beijing, 100876, China