Organizations: 6G Research Center, Khalifa University, 127788 Abu Dhabi, UAE · Department of Computer Science, Central South University, 410083, Changsha, China · Université de Lorraine, CNRS, CRAN, F-54000 Nancy, France
Federated Learning (FL) enables privacy-preserving fine-tuning of Large Language Models (LLMs), yet the massive communication overhead remains a critical bottleneck. Furthermore, applying Low-Rank Adaptation (LoRA) in FL faces a fundamental "aggregation dilemma" between the accurate Sum-of-Products (SoP) and the communication-efficient Product-of-Sums (PoS) implementations. To tackle these challenges, we propose FedFit. First, to significantly reduce communication overhead, we introduce a disjoint shared vector-bank parameterization that reconstructs high-dimensional adapter matrices from two compact and disjoint global vector banks. Second, to address the aggregation dilemma, we devise an alternating optimization schedule. By cycling between decoupled single-bank updates (which allow for accurate aggregation) and joint updates corrected by a Residual Spectral Aggregation mechanism, we resolve the conflict between SoP and PoS. Additionally, we integrate blockwise quantization with client-side error feedback to further compress the transmitted vectors. Furthermore, we establish theoretical convergence guarantees for the proposed algorithm. Extensive experiments on Qwen2.5 models demonstrate that FedFit achieves perplexity performance comparable to standard federated LoRA methods, while providing compression ratios up to 100x higher.
Figures & tables
Fig. 1 : Illustration of client k in FedFit framework
Model
Method
Communication Cost
Compression Ratio
Eval Perplexity
Average Score
Qwen2.5-0.5B-IT
FedIT [ 50 ]
4.40M
1.00
7.42
50.48
Flex-LoRA [ 2 ]
4.40M
1.00
7.41
50.18
FFA-LoRA [ 37 ]
2.20M
2.00
7.65
50.49
RoLoRA [ 9 ]
2.22M
2.25
7.79
50.78
Quantized-LoRA [ 22 ]
1.16M
4.29
7.64
50.63
LoRA-FAIR [ 5 ]
4.43M
1.13
7.56
50.57
TABLE I: Comparison of FL methods with LoRA rank r=4 with 4 clients and 20 FL rounds for Qwen2.5-0.5B/1.5B/3B/7B-Instruct models. Communication cost denotes the number of bytes per client per direction (downlink/uplink) assuming BF16 representation (2 bytes/weight). Compression ratio is based on FedIT with full precision. RTN4 stands for 4-bit round-to-nearest quantization. Proposed method achieves slightly better performance with a large compression ratio of 20-100 × depending on the model size compared to SOTA methods.
Model
Method
BoolQ
PIQA
HellaS.
WinoG.
ARC-e
ARC-c
OBQA
Average Score
Qwen2.5-0.5B-IT
FedIT [ 50 ]
67.92
69.97
40.00
56.43
65.15
29.10
24.80
50.48
FFA-LoRA [ 37 ]
67.89
70.35
40.03
55.96
65.28
29.35
24.60
50.49
Flex-LoRA [ 2 ]
68.50
69.97
39.82
55.88
64.60
28.50
24.00
50.18
RoLoRA [ 9 ]
68.01
70.89
40.25
56.83
65.66
30.03
23.80
50.78
Quantized-LoRA [ 22 ]
67.83
70.40
40.06
56.51
65.40
29.61
24.60
50.63
LoRA-FAIR [ 5 ]
68.20
70.46
40.00
56.04
65.61
29.10
24.60
50.57
TABLE II: Accuracy comparison of FL methods applied to Qwen2.5-0.5B/1.5B/3B/7B-Instruct models across 7 common-sense reasoning benchmarks. Proposed method outperforms all baselines with full precision and achieves comparable performances when 4-bit RTN quantization is applied during weight transmission.
Fig. 2 : Evaluation perplexity of different FL methods for Qwen2.5-0.5B-Instruct. The proposed method consistently outperforms all baselines and continues to improve after the others saturate.
Model
Method
2 clients
4 clients
8 clients
Qwen2.5-0.5B-IT
FedIT
7.42
7.41
7.40
FedFit
7.35
7.30
7.25
Qwen2.5-1.5B-IT
FedIT
5.67
5.67
5.66
FedFit
5.61
5.58
5.54
Qwen2.5-3B-IT
FedIT
5.10
5.10
5.09
FedFit
5.04
5.04
5.03
TABLE III: Evaluation perplexity of the proposed method compared with the FedIT baseline as a function of the number of clients for different model sizes.
Fig. 3 : Evaluation perplexity under different quantization settings for Qwen2.5-0.5B-Instruct. Using 4-bit RTN quantization for both uplink and downlink, which achieves a compression ratio of approximately 4× , results in nearly negligible performance degradation.
Scheduling [Both/A(B)]
Communication Cost
Eval Perplexity
Average Score
10
1.05M
7.25
49.73
11
0.70M
7.30
50.69
12
0.63M
7.32
50.44
TABLE IV: Non-alternating and our alternating scheduling. “ xy ” denotes the number of rounds for the A+B phase and A/B solo phases respectively.
Fig. 4 : Comparison of scheduler designs for Qwen2.5-0.5B-Instruct under different training-phase schedules.
Fig. 5 : Per-round bank norms (left) and round-to-round changes (right).
Phase
FedIT
FedFit
Local SFT (per client)
24.5 s
23.8 s
Model build + adapter init
<0.05 s
<0.05 s
Central eval (averaged)
0.5 s
0.5 s
Server aggregation
<1 ms
2.2 ms
RTN encode/decode
—
<10 ms
Other (RPC, transport)
7.4 s
7.4 s
TABLE V: Median per-round wall-clock breakdown for FedIT and FedFit over the run.
Method
0.5B
1.5B
3B
7B
FedIT [ 50 ]
7.04
14.77
23.95
32.35
Flex-LoRA [ 2 ]
7.04
23.95
23.95
32.35
FFA-LoRA [ 37 ]
3.52
11.97
11.97
17.82
RoLoRA [ 9 ]
3.55
7.41
12.02
16.18
Quantized-LoRA [ 22 ]
1.86
3.81
6.14
8.19
LoRA-FAIR [ 5 ]
7.09
14.83
24.02
32.35
TABLE VI: Uplink time, in seconds, at a representative 5 Mbps cellular link, derived from the communication cost in Tab. I .
Configuration
L
r
Uplink/round
Eval ppl
Avg Score
FedFit (smaller bank)
262 K
4
0.35 MB
5.67
60.50
FedFit (default)
524 K
4
0.70 MB
5.62
60.61
FedFit (larger bank)
1.05 M
4
1.40 MB
5.61
60.68
FedFit (rank 2 )
524 K
2
0.70 MB
5.63
60.84
FedFit (rank 8 )
524 K
8
0.70 MB
5.65
60.73
FedFit (rank 16 )
524 K
16
0.70 MB
5.67
60.53
TABLE VII: Sweep over vector-bank dimension and LoRA rank.
Aggregator (joint round)
Eval ppl
Average Score
Vector bank (FedFit)
no RSC (alt-phases + FedAvg )
5.62
60.76
pure RSC (no FedAvg term, γ=1.0 )
5.72
60.65
RSC applied to vA only
5.63
60.57
RSC applied to vB only
5.64
60.63
RSC applied to both banks (proposed)
5.64
60.87
TABLE VIII: Ablation of the residual spectral correction, and of one-sided correction on the raw LoRA factors.
Model
Method
GSM8K
MMLU
Qwen2.5-1.5B-IT
FedIT
64.67
60.12
FedFit
64.67
59.61
FedFit (RTN4)
63.38
59.63
Qwen2.5-3B-IT
FedIT
74.98
65.46
FedFit
74.75
65.30
FedFit (RTN4)
73.62
65.28
TABLE IX: MetaMathQA fine-tuning on Qwen2.5-1.5B/3B-Instruct; GSM8K and MMLU scores.
Dirichlet α
Eval ppl
Average Score
Uplink (MB/rd)
0.05
5.63
60.93
0.70
0.1
5.63
60.60
0.70
0.3
5.65
60.80
0.70
1.0
5.64
60.93
0.70
3.0
5.63
60.64
0.70
IID baseline
5.62
60.61
0.70
TABLE X: Dirichlet non-IID sweep over Dolly’s category field.
Clients K
Rounds T
Eval ppl
Average Score
4
50
5.63
60.54
16
50
5.60
60.96
32
50
5.61
60.47
TABLE XI: Scalability of FedFit across client counts K∈{4,16,32} .
Clients K
Joint rounds
mean σ1
mean ρ1
Correction
4
17
4.09
0.364
applied
16
17
1.07
0.079
withheld
32
17
0.56
0.041
withheld
TABLE XII: Spectrum of the joint-round core matrix and the resulting gate decision, averaged over the joint rounds of each run.
Federated fine-tuning with Low-Rank Adaptation (LoRA) enables efficient collaborative adaptation of Large Language Models (LLMs) without centralizing private data. However, LoRA's two-factor parameterization creates an aggregation mismatch across clients: naively averaging the factors does not recover the average of their induced updates. This mismatch can be avoided by forming the exact aggregate in the full weight space and then recompressing it, but decomposing the resulting dense matrix is computationally expensive and memory-intensive. We propose FraQ, an efficient coordinate-space recompression method for federated LoRA. Starting from stacked factors that exactly represent the aggregate, FraQ factorizes it into an orthonormal basis and a compact coordinate matrix. It then recovers the singular spectrum from a small Gram matrix, selects the smallest rank satisfying a prescribed energy threshold, and maps the selected coordinate subspace back through the basis to construct the global adapter. Experiments on text classification and commonsense reasoning benchmarks show that FraQ achieves accuracy close to uncompressed baselines while substantially reducing downlink communication with low server-side recompression overhead.
Shenghui Li, Thiemo Voigt
Uppsala University, Uppsala, Sweden · Research Institutes of Sweden, Stockholm, Sweden
Federated fine-tuning of foundation models with Low-Rank Adaptation (LoRA) provides an efficient solution for reducing communication and computation costs while preserving data locality. However, the direct combination of FedAvg and LoRA suffers from three key issues: limited update space, which restricts the model's effective learning capacity; inter-round state mismatch, which disrupts cross-round local optimization continuity; and a client-agnostic starting state, which slows local convergence on clients. Although recent methods mitigate the limited update space issue by merging LoRA updates into the backbone across communication rounds, inter-round state mismatch and the client-agnostic starting state remain insufficiently addressed. To address these issues, we propose FedSmoothLoRA, a federated LoRA tuning framework that preserves the enlarged update space, improves cross-round local optimization continuity, and provides a client-aware starting state for local training. At each communication round, FedSmoothLoRA constructs the local LoRA initialization using two matrices: a Round-Matching matrix that preserves cross-round local state continuity, and a Gradient-Aligned matrix that provides client-specific optimization guidance from gradient signals estimated on local data. Together, these designs enable smoother and faster convergence. Extensive experiments on image classification and natural language generation tasks demonstrate that FedSmoothLoRA consistently outperforms existing federated LoRA tuning methods. Code: https://github.com/wangzehao0704/FedSmoothLoRA
Zehao Wang, Guanglei Yang, Yihan Zeng +4
1Harbin Institute of Technology · 2Huawei Noah’s Ark Lab · University College Dublin
Low-rank adaptation (LoRA) has emerged as a powerful tool for parameter-efficient fine-tuning of large language models (LLMs). This paper studies LoRA under a federated learning setting, enabling collaborative fine-tuning across clients while preserving parameter efficiency. We focus on a highly heterogeneous regime in which clients share only partial structure and a substantial subset may be contaminated. We propose Collaborative Low-rank Alignment and Identifiable Recovery (CLAIR), a contamination-aware framework that relies only on preliminary local estimators. Its formulation applies broadly, from linear regression to neural network and LLM modules, whenever local adaptation can be represented by matrix-valued updates. CLAIR recovers the shared LoRA subspace and detects contaminated clients via a structured low-rank plus block-sparse decomposition. We prove exact recovery of the shared LoRA subspace in the noiseless case, stable recovery under preliminary estimation error, and consistent collaborative-set recovery under mild separation conditions. We further quantify the gain from CLAIR refinement: it reduces off-subspace estimation error through cross-client averaging while preserving client-specific variation within the shared LoRA subspace, thus improves over local fine-tuning whenever this oracle gain outweighs the costs of subspace estimation and benign-client heterogeneity. Empirically, we demonstrate the benefits of CLAIR by fine-tuning a Transformer architecture on a text-copying task. The results show accurate contamination detection and improved benign-client performance compared with local fine-tuning and non-robust federated averaging.
Shuaida He, Liwen Chen, Long Feng
School of Computing & Data Science, The University of Hong Kong