Organizations: State Key Laboratory for Novel Software Technology · School of Artificial Intelligence · School of Computer Science · School of Electronic Science and Engineering
Federated parameter-efficient fine-tuning enables clients to adapt pre-trained models without sharing raw data or communicating the full model, but statistical heterogeneity makes a single global adapter insufficient for personalized prediction. Existing personalized methods typically use the same low-rank structure for both shared and private adaptation, overlooking their distinct requirements for aggregation and personalization. We propose FedLAFP, a role-aware framework that couples a compact, globally aggregated LoRA branch with a client-private, full-rank-capable RandLoRA branch. The shared branch provides an efficient interface for transferring common knowledge, whereas the private branch combines fixed random low-rank bases with learned scaling coefficients to provide expressive client-specific adaptation without additional communication. Client- and layer-specific mixing coefficients jointly fuse the two branches, and only the shared LoRA parameters are exchanged. A controlled linear study supports this role assignment: LoRA yields more aligned client updates and lower aggregation error, while RandLoRA more accurately recovers client-specific residuals. Experiments across four visual recognition benchmarks show that FedLAFP consistently outperforms local-only and federated LoRA baselines, achieving an average personalized accuracy of 86.93% and exceeding the best baseline average by 1.30 percentage points.
Figures & tables
Figure 1: Comparison of global-only adaptation, symmetric shared–private LoRA, and the proposed role-aware FedLAFP.
Figure 2: Controlled comparison of LoRA and RandLoRA under increasing client heterogeneity. Panels report (a) personalization error, (b) aggregation error, and (c) mean cross-client update similarity.
Figure 3: Overview of the FedLAFP framework. (a) The server broadcasts the shared LoRA parameters, and clients upload only their locally updated shared parameters for sample-weighted aggregation, while the private RandLoRA parameters remain local. (b) Within target layer ℓ of client i , the output of the frozen linear transformation is augmented by the weighted outputs of the shared LoRA and private RandLoRA branches, using πℓ,ish and πℓ,ipr , respectively. These client- and layer-specific fusion weights are derived from private logits and are never communicated.
Method
DTD
Oxford Pets
SUN397
UCF101
Average
Local-only LoRA
67.42±0.30
90.81±0.17
78.97±0.03
83.57±0.22
80.19
Local-only RandLoRA
67.99±0.36
91.11±0.22
78.99±0.07
83.70±0.20
80.45
FedLoRA [ICASSP’24]
74.59±0.41
95.47±0.07
83.02±0.04
87.81±0.11
85.22
FedRandLoRA
75.00±0.31
95.43±0.07
83.89±0.07
88.20±0.09
85.63
FFA-LoRA [ICLR’24]
74.35±0.36
95.48±0.08
82.86±0.02
87.86±0.23
85.14
FedEx-LoRA [ACL’25]
74.47±0.12
95.45±0.09
82.95±0.01
87.73±0.03
85.15
Table 1: Comparison with local-only and federated LoRA baselines on four visual recognition benchmarks. Accuracy (%) is reported as mean ± standard deviation over three runs. The best result in each column is shown in bold.
Method
MRPC
SST-2
Average
FedLoRA [ICASSP’24]
72.30
93.81
83.06
FedALT [AAAI’26]
88.24
95.41
91.83
FedLAFP (Ours)
88.73
95.76
92.25
Table 2: Accuracy (%) on GLUE language tasks using RoBERTa-Large.
Figure 4: Role-assignment matrices on DTD and UCF101. Rows specify the shared adapter and columns specify the private adapter. Each cell reports accuracy (%); the smaller annotation and cell color encode the gain over LoRA/LoRA. Bold values denote the best assignment per dataset.
Scheme
DTD
Pets
SUN
UCF
Private (0,1)
67.99
91.11
78.99
83.70
Shared (1,0)
74.59
95.47
83.02
87.81
Equal (0.5,0.5)
77.48
95.78
83.96
88.82
Instance gate
78.07
95.86
82.49
88.76
Client–layer (Ours)
78.62
95.96
84.00
89.12
Table 3: Comparison of strategies for fusing the shared and private branches.
Figure 5: Hyperparameter sensitivity on DTD and UCF101. Each panel varies one hyperparameter while holding the remaining settings fixed. The horizontal axis shows the accuracy change in percentage points relative to the default configuration (rL,rR,K,E)=(4,128,12,3) ; marker labels show absolute accuracy (%). The shaded row denotes the default value.
Figure 6: Accuracy gain of FedLAFP over FedLoRA under different Dirichlet concentration parameters. Positive values indicate an improvement over FedLoRA under the same client partition. The black curve shows the mean gain across the four datasets; larger α corresponds to less heterogeneous client data.
Large Language Models (LLMs) have achieved remarkable success across diverse domains, but their adaptation to privacy-sensitive, distributed datasets remains a challenge. While Federated Learning (FL) combined with Low-Rank Adaptation (LoRA) provides a resource-efficient paradigm for collaborative fine-tuning, practical deployments are hindered by the dual challenges of resource heterogeneity and data heterogeneity. Existing rank-heterogeneous methods primarily focus on bridging dimension mismatches for aggregation but typically provide a unified global model for all clients sharing the same rank, failing to capture client-specific features in non-IID scenarios. In this paper, we propose FedRoRA (Federated Rank-wise Personalized LoRA), a novel framework that enables fine-grained personalization within rank-heterogeneous federations. FedRoRA decouples adaptation into shared global directions and personalized rank-wise magnitudes governed by learnable diagonal scales. On the server side, it extracts a global subspace via singular value decomposition (SVD) and redistributes client-specific initializations through a personalized projection and top-k selection mechanism. Extensive experiments on NLU and NLG benchmarks demonstrate that FedRoRA consistently outperforms state-of-the-art methods.
Lei Wang, Jieming Bian, Letian Zhang +1
University of Florida · Gainesville, FL 32611 · Middle Tennessee State University +1
Federated fine-tuning of foundation models using Low-Rank Adaptation (LoRA) offers a communication efficient solution for distributed learning. However, existing federated LoRA methods suffer from two fundamental limitations: (1) structural aggregation bias, where independently averaging low rank factors fails to approximate the true combined update, and (2) client side initialization lag, as clients repeatedly reinitialize LoRA parameters across communication rounds, slowing convergence. We propose HyperLoRA, a unified framework that addresses both issues through amortized federated adaptation through hypernetwork-driven LoRA generation and product space aggregation. Instead of iterative per-client optimization, HyperLoRA employs a learned generator that maps client distribution signatures to LoRA initializations, effectively amortizing per client adaptation. On the server side, we introduce a learned aggregation module that directly synthesizes updates in the low-rank product space, eliminating the inconsistencies of factor-wise averaging. A lightweight residual correction module further improves stability under heterogenous (non-IID) client distributions.By replacing iterative optimization and heuristic averaging with learned operators, HyperLoRA jointly enables efficient personalization, unbiased aggregation, and faster convergence. Experiments on federated vision and vision-language benchmarks show that HyperLoRA achieves improved convergence speed, greater robustness to distribution shift, and stronger personalization performance compared to prior federated LoRA methods.
Federated LoRA fine-tuning enables parameter-efficient adaptation of pre-trained models without sharing private data, but suffers from two fundamental mismatches under heterogeneous client data: a structural aggregation mismatch caused by independently averaging LoRA factors, and a statistical collaboration mismatch caused by enforcing a single global adapter across divergent clients. To address these issues, we propose CF-LoRA, a clustered federated LoRA fine-tuning framework that combines decoupled factor aggregation with adaptation-aware client clustering. CF-LoRA first learns a globally shared A factor while retaining personalized Bi factors, then identifies clients with similar adaptation patterns based on the cosine similarity of their learned Bi factors, and finally performs intra-cluster B-factor aggregation with a frozen A factor. By decoupling LoRA factor aggregation, CF-LoRA preserves the low-rank structure and mitigates the structural aggregation mismatch, while adaptation-aware clustering promotes collaboration among clients with similar adaptation patterns and reduces negative transfer caused by statistical heterogeneity. Experiments on four language tasks and four vision datasets with RoBERTa and ViT show that CF-LoRA achieves the highest average accuracy in both modalities while communicating only one LoRA factor per optimization round.
Mengjun Yi, Langxing Yang, Suhan Guo +2
State Key Laboratory for Novel Software Technology · School of Artificial Intelligence School of Electronic Science and Engineering Nanjing University