Training with a fixed global batch limits how many distributed clients can provide examples in any one step. We examine a way to use additional server workers without increasing the batch processed by an individual workload. Global Clustered Parallel Split Learning (GCPSL) assigns clients to fixed clusters, executes a Parallel Split Learning with Global Sampling (GPSL) workload for each cluster concurrently, and periodically fuses the client and server model segments. In simulations with 256 logical clients, dividing the population across more workloads improves direct data participation, while smaller clusters can incur an accuracy cost. A four-H100 implementation of label-aware GCPSL reaches 85% CIFAR-10 validation accuracy in 6.13±0.15 minutes over three matched runs, versus 19.09±0.45 minutes when the same workloads are serialized. Within the four-GPU allocation, size-balanced and random fixed affiliations reach the target in similar mean times (5.70 and 5.66 minutes); size balancing increases direct participation by 3.25 percentage points. These measurements characterize a trade-off among execution concurrency, assignment information, participation, and accuracy for stable-client split learning.
Figures & tables
Fig. 2: Execution organization. GPSL uses one global workload with batch size B , so only clients represented in that batch supply data. GCPSL assigns clients to stable clusters and runs one GPSL workload with batch size B in each cluster. The cluster workloads execute concurrently, and their split-model replicas are fused at periodic barriers. Clients that supply no data in a round still receive the synchronized client-model update.
Method
Assignment
N
Acc. (%) ↑
R85↓
R88↓
Inactivity (%) ↓
GPSL
global
–
89.40 ± 0.24
6.78 ± 0.45
14.86 ± 2.07
78.23 ± 0.03
GCPSL
label-aware
2
91.36 ± 0.10
4.44 ± 0.22
8.36 ± 0.83
61.69 ± 0.07
4
90.17 ± 0.44
3.85 ± 0.14
6.45 ± 0.64
39.30 ± 0.16
8
88.93 ± 0.15
3.40 ± 0.10
6.76 ± 0.81
17.32 ± 0.15
16
86.64 ± 0.47
3.61 ± 0.02
–
4.35 ± 0.21
size-balanced
2
90.91 ± 0.88
4.69 ± 0.39
7.95 ± 0.45
61.69 ± 0.07
TABLE I: Cluster-count study.
Method
GPUs
t85 (min) ↓
GPU-min ↓
Acc. (%) ↑
Best-val (%) ↑
Part. (%) ↑
Payload (TB) ↓
Fixed-assignment comparison
Random
4
5.66 ± 0.18
22.62
89.69
90.07
56.96
1.461
Size-balanced
4
5.70 ± 0.50
22.81
89.32
89.52
60.21
1.483
Label-aware
4
6.13 ± 0.15
24.51
89.53
89.73
59.43
1.528
Other reference baselines
Single-worker
1
10.88 ± 0.58
10.88
89.59
89.83
21.70
1.654
TABLE II: Real-execution results. Values summarize three matched runs.
Wireless split learning (SL) reduces on-device computation by offloading upper layers to a server, yet transmitting high-dimensional intermediate features at each iteration remains a major communication bottleneck. Existing methods select features at the client side using task-agnostic criteria such as magnitude, statistics, or clustering, which increases client-side processing and often degrades accuracy under non-independent and identically distributed (non-i.i.d.) client data. We propose importance-aware class-balanced sparsification (ICS), a lightweight approach in which the server ranks feature channels using Grad-CAM-based scores obtained from the true-class logit during backpropagation. The per-class scores are aggregated into a class-balanced, label-agnostic importance vector that mitigates head-class bias under label skew, and each client reuses this vector in the next round to retain the top-N feature channels, incurring no additional client-side forward or backward passes. We further derive a non-asymptotic convergence bound that isolates the sparsification-induced error and characterizes how the sparsification ratio and mini-batch size jointly affect convergence under a fixed communication budget, and we analyze the communication and computational overhead of ICS against representative baselines. Beyond sequential CNN-based SL, we extend ICS to parallel split learning and to transformer-based models. Experiments show that ICS consistently outperforms the baselines, with larger gains under severe non-i.i.d. partitions.
Bumjun Kim, Yoon Huh, Wan Choi
Department of Electrical and Computer Engineering, Seoul National University (SNU), and the Institute of New Media and Communications, SNU, Seoul 08826, Korea
Federated learning enables collaborative model training without sharing raw data, but its performance can degrade substantially under heterogeneous client data distributions. A single global model often cannot satisfy diverse client requirements, so personalized federated learning has therefore been explored to improve client specific performance while preserving global generalization. Existing PFL methods often face a fundamental tradeoff in which stronger global sharing can undermine local specialization, whereas stronger local adaptation can lead to overfitting under limited data, label imbalance, and missing class scenarios. In this work, we propose PGFedSplit, a personalized federated learning framework that improves both personalization and global generalization under severe client heterogeneity. PGFedSplit adopts a split architecture and performs adaptive aggregation scheduling tailored to the roles of different model components, enabling stable knowledge sharing while maintaining client specific adaptation. Each client further leverages a mixture of locally extracted representations and synthetic representations generated from server side Gaussian statistics, improving robustness under label imbalance and missing class conditions. Extensive experiments on Fashion MNIST, CIFAR 10, CIFAR 100, and Tiny ImageNet demonstrate consistent improvements over state of the art PFL methods, with stable convergence and superior personalization in highly heterogeneous settings.
Yunseok Kang, Jaeyoung Song
Department of Electronics Engineering, Pusan National University, Republic of Korea
Hierarchical federated learning (HFL) leverages edge servers for partial aggregation in edge computing. Yet existing FL methods lack mechanisms for jointly optimizing cluster assignment and client selection under data heterogeneity. This paper proposes Fed-BAC, which integrates additive cluster personalization with a two-level bandit framework: contextual bandits at the cloud learn server-to-cluster assignments, while Thompson Sampling at each edge server identifies high-contributing clients. The additive decomposition enables the sharing of knowledge between groups through a globally aggregated network, while cluster-specific networks capture distribution variations. Across three classification benchmarks (CIFAR-10, SVHN, Fashion-MNIST) under moderate (α=0.5) and severe (α=0.1) Dirichlet non-IID partitioning, Fed-BAC achieves distributed accuracy gains of up to +35.5pp over HierFAVG and +8.4pp over IFCA, while requiring only 80% client participation, converging 1.5 to 4.8× faster depending on dataset and accuracy target, and improving cross-server fairness. These gains are further validated at 5× deployment scale on CIFAR-10. The advantage of Fed-BAC increases with heterogeneity severity, confirming that additive cluster personalization becomes increasingly valuable as data distributions diverge.
Satwat Bashir, Tasos Dagiuklas, Muddesar Iqbal
Computer Science and Informatics · London South Bank University · London, UK +1