Organizations: 6G Research Center, Khalifa University, 127788 Abu Dhabi, UAE · Department of Computer Science, Central South University, 410083, Changsha, China · Université de Lorraine, CNRS, CRAN, F-54000 Nancy, France
Federated Learning (FL) enables privacy-preserving fine-tuning of Large Language Models (LLMs), yet the massive communication overhead remains a critical bottleneck. Furthermore, applying Low-Rank Adaptation (LoRA) in FL faces a fundamental "aggregation dilemma" between the accurate Sum-of-Products (SoP) and the communication-efficient Product-of-Sums (PoS) implementations. To tackle these challenges, we propose FedFit. First, to significantly reduce communication overhead, we introduce a disjoint shared vector-bank parameterization that reconstructs high-dimensional adapter matrices from two compact and disjoint global vector banks. Second, to address the aggregation dilemma, we devise an alternating optimization schedule. By cycling between decoupled single-bank updates (which allow for accurate aggregation) and joint updates corrected by a Residual Spectral Aggregation mechanism, we resolve the conflict between SoP and PoS. Additionally, we integrate blockwise quantization with client-side error feedback to further compress the transmitted vectors. Furthermore, we establish theoretical convergence guarantees for the proposed algorithm. Extensive experiments on Qwen2.5 models demonstrate that FedFit achieves perplexity performance comparable to standard federated LoRA methods, while providing compression ratios up to 100x higher.
Figures & tables
Fig. 1 : Illustration of client k in FedFit framework
Model
Method
Communication Cost
Compression Ratio
Eval Perplexity
Average Score
Qwen2.5-0.5B-IT
FedIT [ 50 ]
4.40M
1.00
7.42
50.48
Flex-LoRA [ 2 ]
4.40M
1.00
7.41
50.18
FFA-LoRA [ 37 ]
2.20M
2.00
7.65
50.49
RoLoRA [ 9 ]
2.22M
2.25
7.79
50.78
Quantized-LoRA [ 22 ]
1.16M
4.29
7.64
50.63
LoRA-FAIR [ 5 ]
4.43M
1.13
7.56
50.57
TABLE I: Comparison of FL methods with LoRA rank r=4 with 4 clients and 20 FL rounds for Qwen2.5-0.5B/1.5B/3B/7B-Instruct models. Communication cost denotes the number of bytes per client per direction (downlink/uplink) assuming BF16 representation (2 bytes/weight). Compression ratio is based on FedIT with full precision. RTN4 stands for 4-bit round-to-nearest quantization. Proposed method achieves slightly better performance with a large compression ratio of 20-100 × depending on the model size compared to SOTA methods.
Model
Method
BoolQ
PIQA
HellaS.
WinoG.
ARC-e
ARC-c
OBQA
Average Score
Qwen2.5-0.5B-IT
FedIT [ 50 ]
67.92
69.97
40.00
56.43
65.15
29.10
24.80
50.48
FFA-LoRA [ 37 ]
67.89
70.35
40.03
55.96
65.28
29.35
24.60
50.49
Flex-LoRA [ 2 ]
68.50
69.97
39.82
55.88
64.60
28.50
24.00
50.18
RoLoRA [ 9 ]
68.01
70.89
40.25
56.83
65.66
30.03
23.80
50.78
Quantized-LoRA [ 22 ]
67.83
70.40
40.06
56.51
65.40
29.61
24.60
50.63
LoRA-FAIR [ 5 ]
68.20
70.46
40.00
56.04
65.61
29.10
24.60
50.57
TABLE II: Accuracy comparison of FL methods applied to Qwen2.5-0.5B/1.5B/3B/7B-Instruct models across 7 common-sense reasoning benchmarks. Proposed method outperforms all baselines with full precision and achieves comparable performances when 4-bit RTN quantization is applied during weight transmission.
Fig. 2 : Evaluation perplexity of different FL methods for Qwen2.5-0.5B-Instruct. The proposed method consistently outperforms all baselines and continues to improve after the others saturate.
Model
Method
2 clients
4 clients
8 clients
Qwen2.5-0.5B-IT
FedIT
7.42
7.41
7.40
FedFit
7.35
7.30
7.25
Qwen2.5-1.5B-IT
FedIT
5.67
5.67
5.66
FedFit
5.61
5.58
5.54
Qwen2.5-3B-IT
FedIT
5.10
5.10
5.09
FedFit
5.04
5.04
5.03
TABLE III: Evaluation perplexity of the proposed method compared with the FedIT baseline as a function of the number of clients for different model sizes.
Fig. 3 : Evaluation perplexity under different quantization settings for Qwen2.5-0.5B-Instruct. Using 4-bit RTN quantization for both uplink and downlink, which achieves a compression ratio of approximately 4× , results in nearly negligible performance degradation.
Scheduling [Both/A(B)]
Communication Cost
Eval Perplexity
Average Score
10
1.05M
7.25
49.73
11
0.70M
7.30
50.69
12
0.63M
7.32
50.44
TABLE IV: Non-alternating and our alternating scheduling. “ xy ” denotes the number of rounds for the A+B phase and A/B solo phases respectively.
Fig. 4 : Comparison of scheduler designs for Qwen2.5-0.5B-Instruct under different training-phase schedules.
Fig. 5 : Per-round bank norms (left) and round-to-round changes (right).
Phase
FedIT
FedFit
Local SFT (per client)
24.5 s
23.8 s
Model build + adapter init
<0.05 s
<0.05 s
Central eval (averaged)
0.5 s
0.5 s
Server aggregation
<1 ms
2.2 ms
RTN encode/decode
—
<10 ms
Other (RPC, transport)
7.4 s
7.4 s
TABLE V: Median per-round wall-clock breakdown for FedIT and FedFit over the run.
Method
0.5B
1.5B
3B
7B
FedIT [ 50 ]
7.04
14.77
23.95
32.35
Flex-LoRA [ 2 ]
7.04
23.95
23.95
32.35
FFA-LoRA [ 37 ]
3.52
11.97
11.97
17.82
RoLoRA [ 9 ]
3.55
7.41
12.02
16.18
Quantized-LoRA [ 22 ]
1.86
3.81
6.14
8.19
LoRA-FAIR [ 5 ]
7.09
14.83
24.02
32.35
TABLE VI: Uplink time, in seconds, at a representative 5 Mbps cellular link, derived from the communication cost in Tab. I .
Configuration
L
r
Uplink/round
Eval ppl
Avg Score
FedFit (smaller bank)
262 K
4
0.35 MB
5.67
60.50
FedFit (default)
524 K
4
0.70 MB
5.62
60.61
FedFit (larger bank)
1.05 M
4
1.40 MB
5.61
60.68
FedFit (rank 2 )
524 K
2
0.70 MB
5.63
60.84
FedFit (rank 8 )
524 K
8
0.70 MB
5.65
60.73
FedFit (rank 16 )
524 K
16
0.70 MB
5.67
60.53
TABLE VII: Sweep over vector-bank dimension and LoRA rank.
Aggregator (joint round)
Eval ppl
Average Score
Vector bank (FedFit)
no RSC (alt-phases + FedAvg )
5.62
60.76
pure RSC (no FedAvg term, γ=1.0 )
5.72
60.65
RSC applied to vA only
5.63
60.57
RSC applied to vB only
5.64
60.63
RSC applied to both banks (proposed)
5.64
60.87
TABLE VIII: Ablation of the residual spectral correction, and of one-sided correction on the raw LoRA factors.
Model
Method
GSM8K
MMLU
Qwen2.5-1.5B-IT
FedIT
64.67
60.12
FedFit
64.67
59.61
FedFit (RTN4)
63.38
59.63
Qwen2.5-3B-IT
FedIT
74.98
65.46
FedFit
74.75
65.30
FedFit (RTN4)
73.62
65.28
TABLE IX: MetaMathQA fine-tuning on Qwen2.5-1.5B/3B-Instruct; GSM8K and MMLU scores.
Dirichlet α
Eval ppl
Average Score
Uplink (MB/rd)
0.05
5.63
60.93
0.70
0.1
5.63
60.60
0.70
0.3
5.65
60.80
0.70
1.0
5.64
60.93
0.70
3.0
5.63
60.64
0.70
IID baseline
5.62
60.61
0.70
TABLE X: Dirichlet non-IID sweep over Dolly’s category field.
Clients K
Rounds T
Eval ppl
Average Score
4
50
5.63
60.54
16
50
5.60
60.96
32
50
5.61
60.47
TABLE XI: Scalability of FedFit across client counts K∈{4,16,32} .
Clients K
Joint rounds
mean σ1
mean ρ1
Correction
4
17
4.09
0.364
applied
16
17
1.07
0.079
withheld
32
17
0.56
0.041
withheld
TABLE XII: Spectrum of the joint-round core matrix and the resulting gate decision, averaged over the joint rounds of each run.