Client and Training Data Selection for Computationally Efficient Synchronized Federated Learning
Organizations: Smart Mobility Systems, Technical University of Berlin, Germany · Graduate School of Information Science and Technology, The University of Osaka, Japan
Abstract
Federated learning (FL) is a promising paradigm of machine learning, which preserves user privacy by enabling learning without sharing raw data with a cloud server. Straggling clients have been a problem for FL as they introduce delays in aggregating the local models and hence, the convergence of the global model. Therefore, it is important to have a mechanism that ensures fast convergence of the global model as well as good FL participation rate. Another issue for the convergence of a model in FL is the non-independent and identically distributed (non-iid) data across the clients. Prior approaches based on probabilistic client selection do not work well under non-iid data especially when the number of clients is small. We show scenarios where such approaches fail and propose a joint client-training data selection algorithm for fast convergence of FL models. Our experiments on CIFAR-100 dataset show that convergence of the FL model can be significantly improved over prior works that can consider non-iid data and heterogeneous computation and higher model accuracy.
Figures & tables
| Component | Specification |
|---|---|
| CPU | Intel i9-12900KF 3,2 GHz |
| GPU | NVIDIA GeForce RTX 4090 24 GB (CUDA 12.2) |
| RAM | 64 GB DDR5-SDRAM 4800 MHz |
| Operating System | Ubuntu 22.04.3 LTS |
| Method | Overall | Tail-10 | CV | Gini |
|---|---|---|---|---|
| acc. (%) | acc. (%) | |||
| Proposed (full) | 42.3 | 15.3 | 0.43 | 0.24 |
| probPart | 38.0 | 6.8 | 0.49 | 0.28 |
| MinCost | 31.2 | 3.5 | 0.62 | 0.35 |