Organizations: State Key Laboratory for Novel Software Technology · School of Artificial Intelligence School of Electronic Science and Engineering Nanjing University
Federated LoRA fine-tuning enables parameter-efficient adaptation of pre-trained models without sharing private data, but suffers from two fundamental mismatches under heterogeneous client data: a structural aggregation mismatch caused by independently averaging LoRA factors, and a statistical collaboration mismatch caused by enforcing a single global adapter across divergent clients. To address these issues, we propose CF-LoRA, a clustered federated LoRA fine-tuning framework that combines decoupled factor aggregation with adaptation-aware client clustering. CF-LoRA first learns a globally shared A factor while retaining personalized Bi factors, then identifies clients with similar adaptation patterns based on the cosine similarity of their learned Bi factors, and finally performs intra-cluster B-factor aggregation with a frozen A factor. By decoupling LoRA factor aggregation, CF-LoRA preserves the low-rank structure and mitigates the structural aggregation mismatch, while adaptation-aware clustering promotes collaboration among clients with similar adaptation patterns and reduces negative transfer caused by statistical heterogeneity. Experiments on four language tasks and four vision datasets with RoBERTa and ViT show that CF-LoRA achieves the highest average accuracy in both modalities while communicating only one LoRA factor per optimization round.
Figures & tables
Figure 1: Two mismatches in federated LoRA fine-tuning. Right : the structural aggregation mismatch arises because independently averaging the LoRA factors introduces spurious cross-client products, making the resulting update inconsistent with the ideal aggregated update. Left : the statistical collaboration mismatch arises because enforcing global collaboration among clients with heterogeneous data distributions can lead to negative transfer.
Figure 2: Cross-client cosine distances of learned LoRA factors on SST-2. Points show the mean pairwise cosine distance between flattened A or B factors under Dirichlet data partitioning; smaller α indicates greater heterogeneity. The A factors remain comparatively similar across clients, whereas the B factors diverge as heterogeneity increases.
Figure 3: Overview of the CF-LoRA framework. CF-LoRA consists of three stages: (I) global aggregation of the LoRA A matrices with local personalization of B , (II) client clustering based on the similarity of the learned B matrices, and (III) intra-cluster federated fine-tuning by aggregating B matrices with a frozen shared A . By aggregating only one low-rank matrix in each optimization stage, CF-LoRA avoids forming the product of independently averaged LoRA factors, while clustering similar clients effectively mitigates the negative effects of data heterogeneity.
Method
MNLI-m
MNLI-mm
MRPC
QQP
SST-2
Average
FL-LoRA (ICASSP’24)
86.83±0.06
86.56±0.08
87.25±0.25
89.61±0.05
94.76±0.07
89.00
FFA-LoRA (ICLR’24)
85.08±0.03
85.24±0.08
85.62±0.14
88.06±0.01
94.11±0.18
87.62
FedEx-LoRA (ACL’25)
86.67±0.14
86.53±0.03
86.93±0.37
89.50±0.04
94.57±0.26
88.84
FedSA-LoRA (ICLR’25)
87.16±0.05
86.65±0.10
85.13±0.38
90.60±0.10
94.53±0.07
88.81
FedRot-LoRA (ICML’26)
86.99±0.25
86.75±0.13
87.91±0.75
89.67±0.10
94.91±0.29
89.25
PACFL + FedIT (AAAI’23)
86.60±0.11
86.26±0.03
87.27±0.25
90.73±0.03
95.04±0.13
89.18
Table 1: Performance of different methods on five GLUE evaluation sets under Dirichlet-based data partitioning with α=1 . MNLI-m and MNLI-mm denote the matched and mismatched evaluation sets of MNLI, respectively. We report accuracy ( % ) where higher values indicate better performance. For all evaluation sets, we report accuracy evaluated across 3 runs with mean and standard deviation.
Method
Upload
Download
FL-LoRA
A+B
A+B
FFA-LoRA
B
B
FedEx-LoRA
A+B
A+B+ΔWres
FedSA-LoRA
A
A
FedRot-LoRA
A+B
A+B
PACFL + FedIT
A+B
A+B
Table 2: Comparison of uplink and downlink communication payloads across federated LoRA methods in each optimization round.
Figure 4: Mean accuracy across the five GLUE evaluation sets under different Dirichlet heterogeneity levels. Smaller α indicates stronger heterogeneity, and values above the bars report the corresponding mean accuracies.
Num
Method
M-m
M-mm
MRPC
QQP
SST-2
Avg.
12
FL-LoRA
86.83
86.56
87.25
89.61
94.76
89.00
Ours
87.72
87.11
91.17
91.35
95.20
90.51
50
FL-LoRA
86.55
86.58
79.17
89.58
94.38
87.25
Ours
86.58
86.14
88.93
89.80
94.69
89.23
Table 3: Accuracy (%) with 12 and 50 clients on the five GLUE evaluation sets.
Clustering
M-m
M-mm
MRPC
QQP
SST-2
Avg.
Random
86.24
86.46
84.06
89.13
94.05
87.99
K-means
86.10
86.54
89.70
89.00
94.05
89.08
Spectral
85.91
87.58
86.27
88.69
94.28
88.55
Hierarchical
87.72
87.11
91.17
91.35
95.20
90.51
Table 4: Ablation of client grouping strategies using the same LoRA B -based representations. We report accuracy (%).
Method
Flowers
DTD
UCF
Caltech
Average
FL-LoRA
97.53
61.76
74.43
93.96
81.92
FFA-LoRA
92.97
51.83
58.82
80.12
70.94
FedEx-LoRA
97.82
63.18
77.42
94.12
83.14
FedSA-LoRA
93.79
66.19
79.18
90.47
82.41
FedRot-LoRA
98.23
62.77
74.81
94.12
82.48
PACFL + FedIT
97.50
63.48
77.29
93.65
82.98
Table 5: Performance comparison on vision datasets using ViT-B/16 as the pre-trained backbone. We report classification accuracy (%).