Federated Mixture-of-Experts Alignment on Mobile Edge Networks under Data Heterogeneity
Organizations: Hong Kong JC STEM Lab of Smart City and Department of Computer Science, City University of Hong Kong, Kowloon, Hong Kong SAR, China · School of Computer Science and Technology, Xidian University, Xi’an, China · Department of Electrical and Computer Engineering, The University of Hong Kong, Pok Fu Lam, Hong Kong, China · School of Data Science, Lingnan University, Tuen Mun, Hong Kong, China
Abstract
The growing demand for on-device large language model (LLM) services on mobile edge devices has driven the adoption of Mixture-of-Experts (MoE) architectures, which scale model capacity with limited computation. Since fine-tuning MoE-based LLMs relies on privacy-sensitive local data, federated learning (FL) offers a natural paradigm for collaborative training without exposing raw data. However, integrating MoE-based LLM fine-tuning into FL faces two critical challenges caused by data heterogeneity across clients: (i) divergent local data distributions drive clients to develop distinct gating preferences, so direct parameter aggregation yields a one-size-fits-none global gating network; and (ii) same-indexed experts develop disparate semantic roles across devices, leading to expert semantic blurring and degraded specialization. To address these challenges, we propose FedAlign-MoE, a federated aggregation alignment framework for edge computing systems that jointly enforces routing consistency and expert semantic alignment. Specifically, FedAlign-MoE aggregates gating behaviors by aligning routing distributions through consistency weighting and optimizes local gating networks through distribution regularization, maintaining cross-client stability while preserving discriminative local gating preferences. Meanwhile, FedAlign-MoE quantifies the semantic consistency of same-indexed experts across devices and selectively aggregates semantically aligned experts, ensuring stable and specialized global experts. Extensive experiments demonstrate that FedAlign-MoE outperforms state-of-the-art benchmarks, achieving faster convergence and higher accuracy in non-IID federated environments with lightweight computation and efficient communication.
Figures & tables
| Model | Method | IID distribution | non-IID distribution | ||||||
|---|---|---|---|---|---|---|---|---|---|
| AGNews | PIQA | HellaSwag | MMLU | AGNews | PIQA | HellaSwag | MMLU | ||
| Switch-base-16 | FedAvg | 0.9263 | 0.7421 | 0.7172 | 0.4641 | 0.7740 | 0.6792 | 0.5206 | 0.3008 |
| FedProx | 0.9290 | 0.7473 | 0.7204 | 0.4707 | 0.7872 | 0.6878 | 0.5288 | 0.3176 | |
| PFL-MoE | 0.9318 | 0.7596 | 0.7314 | 0.4837 | 0.7994 | 0.7008 | 0.5476 | 0.3325 | |
| FedMoE | 0.9339 | 0.7681 | 0.7446 | 0.4812 | 0.8137 | 0.7126 | 0.5642 | 0.3506 | |
| FedAlign-MoE | 0.9424 | 0.7822 | 0.7531 | 0.5131 | 0.8522 | 0.7325 | 0.5847 | 0.3845 | |
| Model | Method | AGNews | MMLU | ||||
|---|---|---|---|---|---|---|---|
| Switch-base-16 | FedAvg | 0.7740 | 0.8866 | 0.8978 | 0.3008 | 0.3482 | 0.4012 |
| FedProx | 0.7872 | 0.9006 | 0.9048 | 0.3176 | 0.3615 | 0.4148 | |
| PFL-MoE | 0.7994 | 0.9041 | 0.9109 | 0.3325 | 0.3789 | 0.4305 | |
| FedMoE | 0.8137 | 0.9103 | 0.9164 | 0.3506 | 0.3921 | 0.4463 | |
| FedAlign-MoE | 0.8522 | 0.9312 | 0.9355 | 0.3845 | 0.4250 | 0.4766 | |