Organizations: Hong Kong JC STEM Lab of Smart City and Department of Computer Science, City University of Hong Kong, Kowloon, Hong Kong SAR, China · School of Computer Science and Technology, Xidian University, Xi’an, China · Department of Electrical and Computer Engineering, The University of Hong Kong, Pok Fu Lam, Hong Kong, China · School of Data Science, Lingnan University, Tuen Mun, Hong Kong, China
The growing demand for on-device large language model (LLM) services on mobile edge devices has driven the adoption of Mixture-of-Experts (MoE) architectures, which scale model capacity with limited computation. Since fine-tuning MoE-based LLMs relies on privacy-sensitive local data, federated learning (FL) offers a natural paradigm for collaborative training without exposing raw data. However, integrating MoE-based LLM fine-tuning into FL faces two critical challenges caused by data heterogeneity across clients: (i) divergent local data distributions drive clients to develop distinct gating preferences, so direct parameter aggregation yields a one-size-fits-none global gating network; and (ii) same-indexed experts develop disparate semantic roles across devices, leading to expert semantic blurring and degraded specialization. To address these challenges, we propose FedAlign-MoE, a federated aggregation alignment framework for edge computing systems that jointly enforces routing consistency and expert semantic alignment. Specifically, FedAlign-MoE aggregates gating behaviors by aligning routing distributions through consistency weighting and optimizes local gating networks through distribution regularization, maintaining cross-client stability while preserving discriminative local gating preferences. Meanwhile, FedAlign-MoE quantifies the semantic consistency of same-indexed experts across devices and selectively aggregates semantically aligned experts, ensuring stable and specialized global experts. Extensive experiments demonstrate that FedAlign-MoE outperforms state-of-the-art benchmarks, achieving faster convergence and higher accuracy in non-IID federated environments with lightweight computation and efficient communication.
Fig. 2: The expert selection preferences across edge devices before aggregation under non-IID data distributions.
Fig. 3: The t-SNE visualization of gating parameters and the performance of expert selection from direct parameter aggregation across mobile clients on AGNews dataset.
Fig. 4: The t-SNE representation visualization and local performance for expert specialization on AGNews dataset.
Fig. 5: An overview of FedAlign-MoE framework, where the server aligns routing behaviors across clients via consistency-aware alignment of routing distributions and performs semantic-aware expert aggregation to stabilize the semantics of the same-indexed experts in the global MoE model.
Fig. 6: The gating distribution alignment with routing consistency weighting during global aggregation and adaptive distribution regularization during local fine-tuning.
Fig. 7: The semantic-aware expert aggregation mechanism, which selectively aggregates experts based on semantic consistency, with adaptive thresholding to dynamically calibrate alignment sensitivity.
Model
Method
IID distribution
non-IID distribution
AGNews
PIQA
HellaSwag
MMLU
AGNews
PIQA
HellaSwag
MMLU
Switch-base-16
FedAvg
0.9263
0.7421
0.7172
0.4641
0.7740
0.6792
0.5206
0.3008
FedProx
0.9290
0.7473
0.7204
0.4707
0.7872
0.6878
0.5288
0.3176
PFL-MoE
0.9318
0.7596
0.7314
0.4837
0.7994
0.7008
0.5476
0.3325
FedMoE
0.9339
0.7681
0.7446
0.4812
0.8137
0.7126
0.5642
0.3506
FedAlign-MoE
0.9424
0.7822
0.7531
0.5131
0.8522
0.7325
0.5847
0.3845
Table I: The test accuracy across benchmarks on four datasets with Switch-base-16 and DeepSeek-MoE-16B models under IID and non-IID data distribution ( α=0.1 ).
Fig. 8: The testbed for our FedAlign-MoE framework.
Fig. 9: The average accuracy for local fine-tuning on four datasets under non-IID data distributions.
Fig. 10: The convergence time across benchmarks on AGNews and MMLU datasets under non-IID data distributions.
Fig. 11: The comparison of average communication overhead on AGNews and MMLU datasets.
Model
Method
AGNews
MMLU
α=0.1
α=0.5
α=1.0
α=0.1
α=0.5
α=1.0
Switch-base-16
FedAvg
0.7740
0.8866
0.8978
0.3008
0.3482
0.4012
FedProx
0.7872
0.9006
0.9048
0.3176
0.3615
0.4148
PFL-MoE
0.7994
0.9041
0.9109
0.3325
0.3789
0.4305
FedMoE
0.8137
0.9103
0.9164
0.3506
0.3921
0.4463
FedAlign-MoE
0.8522
0.9312
0.9355
0.3845
0.4250
0.4766
Table II: The test accuracy on the AGNews and MMLU datasets under non-IID data distributions using Switch-base-16 and DeepSeek-MoE-16B models.
Fig. 12: The comparison of average GPU memory footprint across clients on AGNews and MMLU datasets.
Fig. 13: The scalability of FedAlign-MoE and other baselines on AGNews and MMLU datasets under non-IID local data.
Fig. 14: Performance of FedAlign-MoE with varying expert overlap threshold η on AGNews and MMLU datasets.
Fig. 15: Performance of FedAlign-MoE with different balancing coefficient λ on AGNews and MMLU datasets.
Fig. 16: Consistency-based gating distribution alignment on AGNews and MMLU dataset using Switch-base-16 and DeepSeek-MoE-16B models.
Fig. 17: Semantic-aware expert aggregation on AGNews (Fig. 17 (a)-(b)) and MMLU (Fig. 17 (c)-(d)) datasets.
The continuous scaling of large language models (LLMs) incurs prohibitive computational costs, making Mixture-of-Experts (MoE) a scalable alternative for efficient fine-tuning via sparse activation. While federated learning (FL) emerges as the paradigm for privacy-preserving collaborative optimization, integrating MoE into FL under data heterogeneity may trigger conflicting expert optimizations. Client-specific data distributions force same-indexed experts to optimize under inconsistent or even conflicting feature-label correlations. This mismatch induces destructive interference during aggregation, thus destabilizing the optimization trajectory and degrading model performance. To address this issue, we propose FC-MoE, a federated conflict-aware framework for MoE fine-tuning. It employs an importance aware weighting scheme to prioritize reliable local updates and utilizes gradient consensus projection to suppress conflicting updates, ensuring a stable global optimization path. Moreover, a local knowledge retention mechanism further preserves specialized client expertise by re-anchoring domain-specific residuals. Extensive experiments demonstrate that FC-MoE accelerates convergence and enhances both global and local model performance in non-IID federated environments.
Yijun Lu, Zihan Fang, Pengpeng Qiao +6
Waseda University, Tokyo, Japan · City University of Hong Kong, China · Institute of Science Tokyo, Tokyo, Japan +4
Large Language Models (LLMs) have significantly propelled the advancement of edge intelligence and have been widely deployed across various scenarios, including autonomous driving, industrial inspection, and personalized IoT services. However, the collaborative adaptation of LLMs on edge devices continues to face formidable challenges due to strict data privacy constraints, highly heterogeneous computing and communication resources, and the non-independent and identically distributed (non-IID) nature of local data. Federated Fine-Tuning (FFT) enables the collaborative optimization of distributed models without exposing raw data. Yet, traditional synchronous aggregation suffers from a severe straggler effect, resulting in high system latency and low resource utilization. Existing asynchronous federated learning methods are predominantly designed for small-to-medium-scale models and struggle to address the specific challenges inherent in LLM fine-tuning namely, model drift caused by stale updates, aggravated client drift stemming from data heterogeneity, and aggregation fairness imbalance resulting from the dominance of fast clients. To address these issues, this paper proposes AlignFed, an asynchronous federated fine-tuning framework for LLMs tailored to heterogeneous edge environments. AlignFed employs a lightweight multi-stage semantic alignment mechanism comprising three core modules: version-aware update grouping, cross-version semantic alignment based on a mini-batch calibration set, and fairness-aware aggregation that integrates both update freshness and client participation frequency. This framework effectively mitigates cross-version model drift and client drift while enhancing aggregation fairness, thereby achieving stable and efficient asynchronous federated optimization in scenarios characterized by high heterogeneity and significant update staleness.
Yan Wang, Ziyi Gao, Rui Wang
Department of Computer and Communication Engineering, University of Science and Technology Beijing, Beijing 100083, China
Federated PEFT enables LLMs to collaboratively adapt to decentralized private data without sharing raw examples. However, task heterogeneity across clients can cause cross-task interference and gradient conflicts during aggregation. Federated MoE-LoRA addresses this challenge through specialized LoRA experts and conditional routing. Yet existing methods typically specialize at client granularity, implicitly assuming task-coherent clients. Our core insight is that experts need purity, namely pattern-coherent updates that preserve specialization, whereas routers need contrast, namely mixed-task observations that support expert comparison. We propose FedWeave, a framework that adopts asymmetric aggregation, separating expert aggregation from router optimization to meet these two requirements. FedWeave uses unsupervised prototype discovery to form local buckets and align them across clients, enabling prototype-level expert aggregation while retaining mixed-task client trajectories for router training. At inference, FedWeave performs sparse inference with one active expert while preserving nearly all soft-routing performance. Our theoretical analysis explains why asymmetric aggregation is advantageous: it controls expert convergence in stationarity through off-pattern contamination, identifies the consensus error induced by fragmented router trajectories, and bounds sparse-inference risk. On a heterogeneous multi-task benchmark with mainstream LLM backbones, FedWeave consistently outperforms strong baselines, while ablations verify the effectiveness of our design.
Donghang Duan, Xu Zheng, Lizong Zhang +2
University of Electronic Science and Technology of China · 2Zhejiang University