Fine-tuning large models on edge devices is severely hindered by the memory-intensive backpropagation (BP) in standard frameworks like federated learning and split learning. While substituting BP with zeroth-order optimization can significantly reduce memory footprints, it typically suffers from prohibitively degraded convergence speed. To resolve this dilemma, we propose Hybrid-Order Split Federated Learning (HO-SFL). By reformulating the split learning process within a Lagrangian framework, HO-SFL decouples the optimization landscape: The server performs precise first-order updates (i.e., BP), whereas clients conduct memory-efficient zeroth-order optimization. This hybrid design not only eliminates the need for client-side BP but also enables dimension-free model aggregation, drastically lowering communication costs. Crucially, we provide a theoretical convergence analysis, demonstrating that HO-SFL mitigates the dimension-dependent convergence slowdown of zeroth-order optimization, achieving a convergence rate comparable to first-order methods. Extensive experiments on tasks across vision and language modalities validate that HO-SFL achieves convergence speeds comparable to first-order baselines while significantly reducing communication costs and client memory footprints.
Figures & tables
Figure 1 : Overview of the HO-SFL training loop. Each selected client m computes an activation zmt=fc(xm;θct) and sends (zmt,ym) to the server. The server performs BP to update θs and returns the activation-gradient feedback λmt=∇zmℓ . In parallel, the client runs P ZO perturbation forward passes z~m,pt=fc(xm;θct+μupt) and computes scalar projections vm,pt=λmt⊤(z~m,pt−zmt) . The server then aggregates these scalars into vˉpt=K1∑m∈Stvm,pt , which are broadcast back for clients to construct g^ct=Pμ1∑p=1Pvˉptupt and update θc .
Figure 2 : Comparison of system efficiency. (a) Standard SFL incurs significant idle time on the client side waiting for gradients. (b) HO-SFL effectively masks the computational cost of multiple client-side zeroth-order perturbations by overlapping them with the server’s backpropagation and communication processes.
Figure 3 : Validation accuracy convergence on CIFAR-10 under IID (left) and Non-IID (right) settings. Solid lines denote the mean performance, and shaded regions represent the standard deviation across 10 independent random seeds.
Model
Task
SplitLoRA
ZO-SFL
HO-SFL
OPT (125M)
SST2
87.5
52.8
87.6
WSC
64.4
36.5
60.6
RTE
57.8
52.0
59.2
Gemma-3 (270M)
SST2
90.3
51.8
90.8
WSC
61.5
36.5
62.5
RTE
59.6
54.2
65.0
Table 1 : LLM Fine-tuning Accuracy (%) on GLUE tasks. Values in boldface indicate the highest accuracy.
Model
SplitLoRA
HO-SFL
OPT (125M)
0.5985
0.5744
LLaMA-3.2 (1B)
0.8804
0.8687
LLaMA-3.2 (3B)
0.9271
0.9238
Qwen3 (8B)
0.9413
0.9389
Table 2 : LLM Fine-tuning F1 Scores on the SQuAD task.
Figure 4 : Communication and memory profiling.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : Feasibility analysis of latency hiding under constrained edge resources.
First Order
Zeroth Order
Hybrid
Task (Dataset)
SFL
FSL SAGE
ZO-SFL
Mu-SplitFed
HO-SFL (Ours)
CIFAR-10 (IID)
77.5(0.5)
62.1(1.5)
12.3(1.0)
14.0(0.9)
75.0(0.3)
CIFAR-10 (Non-IID)
69.3(2.6)
37.9(2.9)
11.8(1.2)
13.6(1.0)
69.6(0.8)
CIFAR-100 (IID)
43.2(0.4)
8.3(4.2)
1.2(0.2)
1.3(0.1)
44.2(0.3)
CIFAR-100 (Non-IID)
39.5(0.6)
4.0(3.0)
1.2(0.1)
1.3(0.1)
42.5(0.9)
Appendix
Table 3 : Performance comparison across different vision tasks. The notation Acc(Std) denotes the test accuracy (%) and its standard deviation. Bold indicates the best performance among comparative methods for each task.
Figure 6 : Validation accuracy convergence on CIFAR-100 under IID (left) and Non-IID (right) settings.
Figure 7 : Impact of client-side model depth on convergence. We evaluate LLaMA-3.2-1B on the SST-2 task with varying numbers of transformer layers (2, 4, 6, 8) allocated to the client.
Figure 8 : Ablation studies on CIFAR-10 with ResNet-18, focusing on the perturbation number P and the smoothing parameter μ . Solid lines denote the mean performance, and shaded regions represent the standard deviation across 10 random seeds.
Component
HO-SFL
SFL
FSL-SAGE
ZO-SFL
MU-SplitFed
Uplink (Act)
1.22
1.26
0.29
2.52
3.67
Uplink (Model)
0.00
0.84
0.84
0.84
12.97
Uplink (Scalar)
0.0002
0.0
0.0
0.0
0.0
Downlink (Grad)
1.22
1.26
0.00
0.00
0.00
Downlink (Model)
0.00
0.84
2.58
0.84
12.97
Downlink (Scalar/Seed)
0.004
0.0
0.0
0.00002
0.00002
Appendix
Table 4 : Breakdown of total communication cost (GB).