Organizations: State Key Laboratory for Novel Software Technology · School of Artificial Intelligence School of Electronic Science and Engineering Nanjing University
Federated LoRA fine-tuning enables parameter-efficient adaptation of pre-trained models without sharing private data, but suffers from two fundamental mismatches under heterogeneous client data: a structural aggregation mismatch caused by independently averaging LoRA factors, and a statistical collaboration mismatch caused by enforcing a single global adapter across divergent clients. To address these issues, we propose CF-LoRA, a clustered federated LoRA fine-tuning framework that combines decoupled factor aggregation with adaptation-aware client clustering. CF-LoRA first learns a globally shared A factor while retaining personalized Bi factors, then identifies clients with similar adaptation patterns based on the cosine similarity of their learned Bi factors, and finally performs intra-cluster B-factor aggregation with a frozen A factor. By decoupling LoRA factor aggregation, CF-LoRA preserves the low-rank structure and mitigates the structural aggregation mismatch, while adaptation-aware clustering promotes collaboration among clients with similar adaptation patterns and reduces negative transfer caused by statistical heterogeneity. Experiments on four language tasks and four vision datasets with RoBERTa and ViT show that CF-LoRA achieves the highest average accuracy in both modalities while communicating only one LoRA factor per optimization round.
Figures & tables
Figure 1: Two mismatches in federated LoRA fine-tuning. Right : the structural aggregation mismatch arises because independently averaging the LoRA factors introduces spurious cross-client products, making the resulting update inconsistent with the ideal aggregated update. Left : the statistical collaboration mismatch arises because enforcing global collaboration among clients with heterogeneous data distributions can lead to negative transfer.
Figure 2: Cross-client cosine distances of learned LoRA factors on SST-2. Points show the mean pairwise cosine distance between flattened A or B factors under Dirichlet data partitioning; smaller α indicates greater heterogeneity. The A factors remain comparatively similar across clients, whereas the B factors diverge as heterogeneity increases.
Figure 3: Overview of the CF-LoRA framework. CF-LoRA consists of three stages: (I) global aggregation of the LoRA A matrices with local personalization of B , (II) client clustering based on the similarity of the learned B matrices, and (III) intra-cluster federated fine-tuning by aggregating B matrices with a frozen shared A . By aggregating only one low-rank matrix in each optimization stage, CF-LoRA avoids forming the product of independently averaged LoRA factors, while clustering similar clients effectively mitigates the negative effects of data heterogeneity.
Method
MNLI-m
MNLI-mm
MRPC
QQP
SST-2
Average
FL-LoRA (ICASSP’24)
86.83±0.06
86.56±0.08
87.25±0.25
89.61±0.05
94.76±0.07
89.00
FFA-LoRA (ICLR’24)
85.08±0.03
85.24±0.08
85.62±0.14
88.06±0.01
94.11±0.18
87.62
FedEx-LoRA (ACL’25)
86.67±0.14
86.53±0.03
86.93±0.37
89.50±0.04
94.57±0.26
88.84
FedSA-LoRA (ICLR’25)
87.16±0.05
86.65±0.10
85.13±0.38
90.60±0.10
94.53±0.07
88.81
FedRot-LoRA (ICML’26)
86.99±0.25
86.75±0.13
87.91±0.75
89.67±0.10
94.91±0.29
89.25
PACFL + FedIT (AAAI’23)
86.60±0.11
86.26±0.03
87.27±0.25
90.73±0.03
95.04±0.13
89.18
Table 1: Performance of different methods on five GLUE evaluation sets under Dirichlet-based data partitioning with α=1 . MNLI-m and MNLI-mm denote the matched and mismatched evaluation sets of MNLI, respectively. We report accuracy ( % ) where higher values indicate better performance. For all evaluation sets, we report accuracy evaluated across 3 runs with mean and standard deviation.
Method
Upload
Download
FL-LoRA
A+B
A+B
FFA-LoRA
B
B
FedEx-LoRA
A+B
A+B+ΔWres
FedSA-LoRA
A
A
FedRot-LoRA
A+B
A+B
PACFL + FedIT
A+B
A+B
Table 2: Comparison of uplink and downlink communication payloads across federated LoRA methods in each optimization round.
Figure 4: Mean accuracy across the five GLUE evaluation sets under different Dirichlet heterogeneity levels. Smaller α indicates stronger heterogeneity, and values above the bars report the corresponding mean accuracies.
Num
Method
M-m
M-mm
MRPC
QQP
SST-2
Avg.
12
FL-LoRA
86.83
86.56
87.25
89.61
94.76
89.00
Ours
87.72
87.11
91.17
91.35
95.20
90.51
50
FL-LoRA
86.55
86.58
79.17
89.58
94.38
87.25
Ours
86.58
86.14
88.93
89.80
94.69
89.23
Table 3: Accuracy (%) with 12 and 50 clients on the five GLUE evaluation sets.
Clustering
M-m
M-mm
MRPC
QQP
SST-2
Avg.
Random
86.24
86.46
84.06
89.13
94.05
87.99
K-means
86.10
86.54
89.70
89.00
94.05
89.08
Spectral
85.91
87.58
86.27
88.69
94.28
88.55
Hierarchical
87.72
87.11
91.17
91.35
95.20
90.51
Table 4: Ablation of client grouping strategies using the same LoRA B -based representations. We report accuracy (%).
Method
Flowers
DTD
UCF
Caltech
Average
FL-LoRA
97.53
61.76
74.43
93.96
81.92
FFA-LoRA
92.97
51.83
58.82
80.12
70.94
FedEx-LoRA
97.82
63.18
77.42
94.12
83.14
FedSA-LoRA
93.79
66.19
79.18
90.47
82.41
FedRot-LoRA
98.23
62.77
74.81
94.12
82.48
PACFL + FedIT
97.50
63.48
77.29
93.65
82.98
Table 5: Performance comparison on vision datasets using ViT-B/16 as the pre-trained backbone. We report classification accuracy (%).
Low-rank adaptation (LoRA) has emerged as a powerful tool for parameter-efficient fine-tuning of large language models (LLMs). This paper studies LoRA under a federated learning setting, enabling collaborative fine-tuning across clients while preserving parameter efficiency. We focus on a highly heterogeneous regime in which clients share only partial structure and a substantial subset may be contaminated. We propose Collaborative Low-rank Alignment and Identifiable Recovery (CLAIR), a contamination-aware framework that relies only on preliminary local estimators. Its formulation applies broadly, from linear regression to neural network and LLM modules, whenever local adaptation can be represented by matrix-valued updates. CLAIR recovers the shared LoRA subspace and detects contaminated clients via a structured low-rank plus block-sparse decomposition. We prove exact recovery of the shared LoRA subspace in the noiseless case, stable recovery under preliminary estimation error, and consistent collaborative-set recovery under mild separation conditions. We further quantify the gain from CLAIR refinement: it reduces off-subspace estimation error through cross-client averaging while preserving client-specific variation within the shared LoRA subspace, thus improves over local fine-tuning whenever this oracle gain outweighs the costs of subspace estimation and benign-client heterogeneity. Empirically, we demonstrate the benefits of CLAIR by fine-tuning a Transformer architecture on a text-copying task. The results show accurate contamination detection and improved benign-client performance compared with local fine-tuning and non-robust federated averaging.
Shuaida He, Liwen Chen, Long Feng
School of Computing & Data Science, The University of Hong Kong
Federated parameter-efficient fine-tuning enables clients to adapt pre-trained models without sharing raw data or communicating the full model, but statistical heterogeneity makes a single global adapter insufficient for personalized prediction. Existing personalized methods typically use the same low-rank structure for both shared and private adaptation, overlooking their distinct requirements for aggregation and personalization. We propose FedLAFP, a role-aware framework that couples a compact, globally aggregated LoRA branch with a client-private, full-rank-capable RandLoRA branch. The shared branch provides an efficient interface for transferring common knowledge, whereas the private branch combines fixed random low-rank bases with learned scaling coefficients to provide expressive client-specific adaptation without additional communication. Client- and layer-specific mixing coefficients jointly fuse the two branches, and only the shared LoRA parameters are exchanged. A controlled linear study supports this role assignment: LoRA yields more aligned client updates and lower aggregation error, while RandLoRA more accurately recovers client-specific residuals. Experiments across four visual recognition benchmarks show that FedLAFP consistently outperforms local-only and federated LoRA baselines, achieving an average personalized accuracy of 86.93% and exceeding the best baseline average by 1.30 percentage points.
Mengjun Yi, Huaian Gu, Yinghao Ai +2
State Key Laboratory for Novel Software Technology · School of Artificial Intelligence · School of Computer Science +1
Federated fine-tuning of foundation models using Low-Rank Adaptation (LoRA) offers a communication efficient solution for distributed learning. However, existing federated LoRA methods suffer from two fundamental limitations: (1) structural aggregation bias, where independently averaging low rank factors fails to approximate the true combined update, and (2) client side initialization lag, as clients repeatedly reinitialize LoRA parameters across communication rounds, slowing convergence. We propose HyperLoRA, a unified framework that addresses both issues through amortized federated adaptation through hypernetwork-driven LoRA generation and product space aggregation. Instead of iterative per-client optimization, HyperLoRA employs a learned generator that maps client distribution signatures to LoRA initializations, effectively amortizing per client adaptation. On the server side, we introduce a learned aggregation module that directly synthesizes updates in the low-rank product space, eliminating the inconsistencies of factor-wise averaging. A lightweight residual correction module further improves stability under heterogenous (non-IID) client distributions.By replacing iterative optimization and heuristic averaging with learned operators, HyperLoRA jointly enables efficient personalization, unbiased aggregation, and faster convergence. Experiments on federated vision and vision-language benchmarks show that HyperLoRA achieves improved convergence speed, greater robustness to distribution shift, and stronger personalization performance compared to prior federated LoRA methods.