Redundancy Meets Synergy: Dependency-aware Expert Selection for MoE via Submodular Optimization
Organizations: Interdisciplinary Centre for Security, Reliability and Trust, University of Luxembourg, Luxembourg · Department of Computer Science, Peking University, China · Institute of Space Internet, Fudan University, China · Department of Electrical and Computer Engineering, The University of British Columbia, Canada · Department of Computer Science, City University of Hong Kong, Hong Kong, China · School of Engineering, Edith Cowan University, Australia · College of Computing and Data Science, Nanyang Technological University, Singapore
Abstract
While Mixture-of-Experts (MoE) models effectively scale model capacity through sparse activation, their deployment is often bottlenecked by prohibitive memory requirements. Extracting a compact subset of experts presents a promising solution. However, existing expert selection heuristics predominantly rely on Top-k ranking, which isolates the evaluation of individual experts and ignores the intricate inter-expert dependencies introduced by the MoE gating network. In this paper, we propose DS-MoE, a theoretically grounded framework that redefines expert selection via difference-of-submodular (DS) optimization. By analyzing the second-order Taylor expansion of the loss degradation, we reveal functional duality within expert combinations: redundancy (where experts encode overlapping representations) and synergy (where experts provide complementary error cancellation). To navigate this duality, we mathematically decouple redundancy reduction from synergy maximization by formulating the selection objective as a DS function. Furthermore, we devise a tailored majorization-minimization (MM) algorithm with provable monotonicity guarantees to efficiently identify the optimal expert subset. Extensive experiments demonstrate that DS-MoE effectively preserves indispensable expert combinations, achieving superior performance compared to the state-of-the-art baselines.
Figures & tables
| Model | Method | Test Accuracy (%) | Convergence Rounds | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| AGNews | PIQA | HellaSwag | MMLU | Avg. | AGNews | PIQA | HellaSwag | MMLU | Avg. | ||
| Switch-Base-32 | Rand-MoE | 87.3 0.3 | 68.1 0.8 | 65.9 0.7 | 39.8 1.1 | 65.3 0.7 | 84 6 | 26 3 | 28 4 | 34 5 | 43.0 4.5 |
| S-MoE | 92.6 0.2 | 74.3 0.5 | 71.8 0.4 | 46.4 0.6 | 71.3 0.4 | 59 3 | 15 2 | 17 2 | 19 3 | 27.5 2.5 | |
| SEER-MoE | 93.1 0.3 | 75.2 0.4 | 72.7 0.5 | 46.9 0.5 | 72.0 0.4 | 49 3 | 12 1 | 13 2 | 17 2 | 22.8 2.0 | |
| DiEP | 93.3 0.2 | 75.8 0.3 | 73.2 0.4 | 47.5 0.3 | 72.5 0.3 | 48 2 | 12 2 | 12 1 | 15 2 | 21.8 1.8 | |
| DS-MoE | 94.7 0.1 | 78.6 0.2 | 75.6 0.2 | 51.8 0.3 | 75.2 0.2 | 32 1 | 8 1 | 10 1 | 11 2 | 15.3 1.3 | |
| Model | Method | Test Accuracy (%) | Convergence Rounds | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AGNews | PIQA | HellaSwag | MMLU | Avg. | AGNews | PIQA | HellaSwag | MMLU | Avg. | |||
| DeepSeek-16B | 1 | Rand-MoE | 83.2 0.6 | 66.8 1.2 | 65.9 0.9 | 28.9 1.5 | 61.2 1.1 | 162 14 | 59 7 | 56 9 | 57 11 | 83.5 10.3 |
| S-MoE | 90.1 0.5 | 73.4 0.9 | 72.1 0.7 | 37.7 1.2 | 68.3 0.8 | 118 9 | 26 4 | 23 5 | 29 6 | 49.0 6.0 | ||
| SEER-MoE | 90.9 0.4 | 75.7 0.7 | 73.2 0.8 | 39.2 0.9 | 69.8 0.7 | 112 6 | 19 4 | 19 3 | 21 4 | 42.8 4.3 | ||
| DiEP | 91.6 0.3 | 75.9 0.6 | 73.8 0.7 | 40.8 0.8 | 70.5 0.6 | 104 5 | 18 3 | 17 4 | 18 3 | 39.3 3.8 | ||
| DS-MoE | 93.2 0.3 | 79.8 0.4 | 76.3 0.3 | 43.9 0.5 | 73.3 0.4 | 66 3 | 12 2 | 13 2 | 15 3 | 26.5 2.5 | ||