FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection
Organizations: Tianjin University
Abstract
Federated training of foundation models is constrained by client memory and communication costs. LoRA-based methods reduce these costs through low-rank adapters, but their fixed rank budget can limit adaptation. Gradient low-rank optimization offers greater flexibility, yet independently chosen client subspaces create a problem we term \emph{subspace fragmentation}: local projections interact with data heterogeneity to bias aggregated directions, while aggregation can increase update rank and communication cost. Thus, accurate local gradient compression need not preserve global descent. We propose \texttt{FedLore}, which shares a low-rank optimization basis within each round and refreshes it across rounds. The shared basis enables exact aggregation in low-rank coordinates and eliminates the identified projection bias. Subspace refresh allows the accumulated model update to exceed the per-round rank budget. We characterize the aggregation bias and establish an stationarity bound for the projected-SGD variant under a global-gradient coverage condition and standard smoothness and variance assumptions, with bounded gradient heterogeneity. Experiments on vision and language tasks, including federated pre-training, show that \texttt{FedLore} outperforms the evaluated low-rank adapter baselines and matches or exceeds full-parameter training, while reducing communication and optimizer-state memory.
Figures & tables
| Method | Venue | CIFAR-100 | Tiny-ImageNet | Food-101 | Avg. | |||
| Swin-Base | ViT-Base | Swin-Base | ViT-Base | Swin-Base | ViT-Base | |||
| FedIT | ICASSP’24 | |||||||
| FFA-LoRA | ICLR’24 | |||||||
| FlexLoRA | ICLR’25 | |||||||
| LoRA-FAIR | ICCV’25 | |||||||
| RoLoRA | NeurIPS’25 | |||||||
| Method | Venue | SNLI | AG News | QQP | DBPedia 14 | QNLI | MNLI | Avg. |
| FedIT | ICASSP’24 | |||||||
| FFA-LoRA | ICLR’24 | |||||||
| FlexLoRA | ICLR’25 | |||||||
| LoRA-FAIR | ICCV’25 | |||||||
| RoLoRA | NeurIPS’25 | |||||||
| FLoRA | NeurIPS’24 |
| Method | 60M | 130M | 350M | 1B | Comm. | Time/Round | ||||
| Loss | PPL | Loss | PPL | Loss | PPL | Loss | PPL | on LLaMA 350M (s) | ||
| FedIT | 8.811 | 6707.62 | 8.544 | 5135.85 | 8.072 | 3203.50 | 7.561 | 1921.77 | 118.6 | |
| Local GaLore | 4.326 | 75.64 | 4.215 | 67.69 | 3.962 | 52.56 | 3.781 | 43.86 | 136.6 | |
| FedFull | 4.134 | 62.43 | 4.056 | 57.74 | 3.798 | 44.61 | 3.561 | 35.20 | 132.1 | |
| FedLore-Random | 4.112 | 61.07 | 4.036 | 56.60 | 3.781 | 43.86 | 3.523 | 33.89 | 119.8 | |
| FedLore | 3.833 | 46.20 | 3.611 | 37.00 | 3.554 | 34.95 | 3.345 | 28.36 | 122.4 | |
| Method | Comm. | Opt. Memory | Ablation Setting | Val Loss / PPL | Comm. |
| FedIT | Local GaLore | 3.962 / 52.56 | |||
| FedFull | FedLore | 3.554 / 34.95 | |||
| Local GaLore | w/o Shared Projector | 3.890 / 48.91 | |||
| FedLore | w/o Dynamic Subspace | 5.952 / 384.7 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Hidden | Intermediate | Heads | Layers |
| LLaMA-60M | 512 | 1376 | 8 | 8 |
| LLaMA-130M | 768 | 2048 | 12 | 12 |
| LLaMA-350M | 1024 | 2736 | 16 | 24 |
| LLaMA-1B | 2048 | 5461 | 32 | 24 |
| Configuration | Value |
| Dataset | C4 |
| Number of clients | 20 |
| Client participation | 100% |
| Communication rounds | 100 |
| Local steps per round | 50 |
| Batch size | 16 |
| Method | Communication Cost | Optimizer Memory |
| FedIT | ||
| FedFull | ||
| Local GaLore | ||
| FedLore |
| Method | LLaMA-60M | LLaMA-130M | LLaMA-350M | LLaMA-1B |
| FedIT | ||||
| Local GaLore | ||||
| FedFull | ||||
| FedLore-Random | ||||
| FedLore |
| Configuration | Value |
| Datasets | CIFAR-100, Tiny-ImageNet, Food-101 |
| Backbone | Swin-Base, ViT-Base |
| Number of clients | 50 |
| Data partition | Dirichlet ( ) |
| Communication rounds | 100 |
| Local steps | 50 |
| Configuration | Value |
| Datasets | SNLI, AG News, QQP, DBPedia-14, QNLI, MNLI |
| Backbone | RoBERTa-Base |
| Number of clients | 20 |
| Data partition | Dirichlet ( ) |
| Communication rounds | 100 |
| Local steps | 50 |
| Method | CIFAR-100 | Tiny-ImageNet | Food-101 | |||
| Swin-Base | ViT-Base | Swin-Base | ViT-Base | Swin-Base | ViT-Base | |
| FedIT | ||||||
| FFA-LoRA | ||||||
| FlexLoRA | ||||||
| LoRA-FAIR | ||||||
| RoLoRA | ||||||
| Method | SNLI | AG News | QQP | DBPedia-14 | QNLI | MNLI |
| FedIT | ||||||
| FFA-LoRA | ||||||
| FlexLoRA | ||||||
| LoRA-FAIR | ||||||
| RoLoRA | ||||||
| FLoRA |
| Method | Communication Cost | Optimizer Memory |
| FedIT | ||
| FedFull | ||
| Local GaLore | ||
| FedLore |
| Method | 60M | 130M | 350M | 1B |
| Full-Rank | 0.23G | 0.51G | 1.37G | 5.20G |
| GaLore | 0.13G | 0.28G | 0.54G | 1.78G |
| Low-Rank | 0.17G | 0.37G | 0.72G | 2.38G |
| LoRA | 0.17G | 0.37G | 0.72G | 2.38G |
| FedLore | 0.13G | 0.28G | 0.54G | 1.78G |