Training with a fixed global batch limits how many distributed clients can provide examples in any one step. We examine a way to use additional server workers without increasing the batch processed by an individual workload. Global Clustered Parallel Split Learning (GCPSL) assigns clients to fixed clusters, executes a Parallel Split Learning with Global Sampling (GPSL) workload for each cluster concurrently, and periodically fuses the client and server model segments. In simulations with 256 logical clients, dividing the population across more workloads improves direct data participation, while smaller clusters can incur an accuracy cost. A four-H100 implementation of label-aware GCPSL reaches 85% CIFAR-10 validation accuracy in 6.13±0.15 minutes over three matched runs, versus 19.09±0.45 minutes when the same workloads are serialized. Within the four-GPU allocation, size-balanced and random fixed affiliations reach the target in similar mean times (5.70 and 5.66 minutes); size balancing increases direct participation by 3.25 percentage points. These measurements characterize a trade-off among execution concurrency, assignment information, participation, and accuracy for stable-client split learning.
Figures & tables
Fig. 2: Execution organization. GPSL uses one global workload with batch size B , so only clients represented in that batch supply data. GCPSL assigns clients to stable clusters and runs one GPSL workload with batch size B in each cluster. The cluster workloads execute concurrently, and their split-model replicas are fused at periodic barriers. Clients that supply no data in a round still receive the synchronized client-model update.
Method
Assignment
N
Acc. (%) ↑
R85↓
R88↓
Inactivity (%) ↓
GPSL
global
–
89.40 ± 0.24
6.78 ± 0.45
14.86 ± 2.07
78.23 ± 0.03
GCPSL
label-aware
2
91.36 ± 0.10
4.44 ± 0.22
8.36 ± 0.83
61.69 ± 0.07
4
90.17 ± 0.44
3.85 ± 0.14
6.45 ± 0.64
39.30 ± 0.16
8
88.93 ± 0.15
3.40 ± 0.10
6.76 ± 0.81
17.32 ± 0.15
16
86.64 ± 0.47
3.61 ± 0.02
–
4.35 ± 0.21
size-balanced
2
90.91 ± 0.88
4.69 ± 0.39
7.95 ± 0.45
61.69 ± 0.07
TABLE I: Cluster-count study.
Method
GPUs
t85 (min) ↓
GPU-min ↓
Acc. (%) ↑
Best-val (%) ↑
Part. (%) ↑
Payload (TB) ↓
Fixed-assignment comparison
Random
4
5.66 ± 0.18
22.62
89.69
90.07
56.96
1.461
Size-balanced
4
5.70 ± 0.50
22.81
89.32
89.52
60.21
1.483
Label-aware
4
6.13 ± 0.15
24.51
89.53
89.73
59.43
1.528
Other reference baselines
Single-worker
1
10.88 ± 0.58
10.88
89.59
89.83
21.70
1.654
TABLE II: Real-execution results. Values summarize three matched runs.
Department of Electrical and Computer Engineering, Seoul National University (SNU), and the Institute of New Media and Communications, SNU, Seoul 08826, Korea