Redundancy Meets Synergy: Dependency-aware Expert Selection for MoE via Submodular Optimization
Authors: Zheng Lin, Shaoke Fang, Yuxin Zhang, Jinfeng Xu, Zihan Fang, Zhe Chen, Wei Ni, Jun Luo, +1 more
Organizations: Interdisciplinary Centre for Security, Reliability and Trust, University of Luxembourg, Luxembourg · Department of Computer Science, Peking University, China · Institute of Space Internet, Fudan University, China · Department of Electrical and Computer Engineering, The University of British Columbia, Canada · Department of Computer Science, City University of Hong Kong, Hong Kong, China · School of Engineering, Edith Cowan University, Australia · College of Computing and Data Science, Nanyang Technological University, Singapore
While Mixture-of-Experts (MoE) models effectively scale model capacity through sparse activation, their deployment is often bottlenecked by prohibitive memory requirements. Extracting a compact subset of experts presents a promising solution. However, existing expert selection heuristics predominantly rely on Top-k ranking, which isolates the evaluation of individual experts and ignores the intricate inter-expert dependencies introduced by the MoE gating network. In this paper, we propose DS-MoE, a theoretically grounded framework that redefines expert selection via difference-of-submodular (DS) optimization. By analyzing the second-order Taylor expansion of the loss degradation, we reveal functional duality within expert combinations: redundancy (where experts encode overlapping representations) and synergy (where experts provide complementary error cancellation). To navigate this duality, we mathematically decouple redundancy reduction from synergy maximization by formulating the selection objective as a DS function. Furthermore, we devise a tailored majorization-minimization (MM) algorithm with provable monotonicity guarantees to efficiently identify the optimal expert subset. Extensive experiments demonstrate that DS-MoE effectively preserves indispensable expert combinations, achieving superior performance compared to the state-of-the-art baselines.
Figures & tables
Figure 1: Overview of the DS-MoE framework.
Figure 2: First-order gradients versus second-order Hessian. (a) First-order proxy fits a flat tangent plane at local stationary points, failing to distinguish between experts. (b) Hessian captures steep curvature, revealing the massive penalty of removing expert 1 ( ΔL(m)=100 ) versus expert 2 ( ΔL(m)=1 ).
Model
Method
Test Accuracy (%) ↑
Convergence Rounds ↓
AGNews
PIQA
HellaSwag
MMLU
Avg.
AGNews
PIQA
HellaSwag
MMLU
Avg.
Switch-Base-32
Rand-MoE
87.3 ± 0.3
68.1 ± 0.8
65.9 ± 0.7
39.8 ± 1.1
65.3 ± 0.7
84 ± 6
26 ± 3
28 ± 4
34 ± 5
43.0 ± 4.5
S-MoE
92.6 ± 0.2
74.3 ± 0.5
71.8 ± 0.4
46.4 ± 0.6
71.3 ± 0.4
59 ± 3
15 ± 2
17 ± 2
19 ± 3
27.5 ± 2.5
SEER-MoE
93.1 ± 0.3
75.2 ± 0.4
72.7 ± 0.5
46.9 ± 0.5
72.0 ± 0.4
49 ± 3
12 ± 1
13 ± 2
17 ± 2
22.8 ± 2.0
DiEP
93.3 ± 0.2
75.8 ± 0.3
73.2 ± 0.4
47.5 ± 0.3
72.5 ± 0.3
48 ± 2
12 ± 2
12 ± 1
15 ± 2
21.8 ± 1.8
DS-MoE
94.7 ± 0.1
78.6 ± 0.2
75.6 ± 0.2
51.8 ± 0.3
75.2 ± 0.2
32 ± 1
8 ± 1
10 ± 1
11 ± 2
15.3 ± 1.3
Table 1: Experimental results on test accuracy (%) ( ↑ ) and convergence rounds ( ↓ ). We evaluate DS-MoE and four baselines on Switch-Base-32, Switch-Base-64, and DeepSeek-MoE-16B.
Model
k
Method
Test Accuracy (%) ↑
Convergence Rounds ↓
AGNews
PIQA
HellaSwag
MMLU
Avg.
AGNews
PIQA
HellaSwag
MMLU
Avg.
DeepSeek-16B
1
Rand-MoE
83.2 ± 0.6
66.8 ± 1.2
65.9 ± 0.9
28.9 ± 1.5
61.2 ± 1.1
162 ± 14
59 ± 7
56 ± 9
57 ± 11
83.5 ± 10.3
S-MoE
90.1 ± 0.5
73.4 ± 0.9
72.1 ± 0.7
37.7 ± 1.2
68.3 ± 0.8
118 ± 9
26 ± 4
23 ± 5
29 ± 6
49.0 ± 6.0
SEER-MoE
90.9 ± 0.4
75.7 ± 0.7
73.2 ± 0.8
39.2 ± 0.9
69.8 ± 0.7
112 ± 6
19 ± 4
19 ± 3
21 ± 4
42.8 ± 4.3
DiEP
91.6 ± 0.3
75.9 ± 0.6
73.8 ± 0.7
40.8 ± 0.8
70.5 ± 0.6
104 ± 5
18 ± 3
17 ± 4
18 ± 3
39.3 ± 3.8
DS-MoE
93.2 ± 0.3
79.8 ± 0.4
76.3 ± 0.3
43.9 ± 0.5
73.3 ± 0.4
66 ± 3
12 ± 2
13 ± 2
15 ± 3
26.5 ± 2.5
Table 2: Robustness and Scalability Analysis on DeepSeek-MoE-16B. The training performance versus (a) expert sparsity k with N=64 and (b) expert pool size N with k=2 .
Figure 3: Comprehensive ablation study of the DS-MoE framework on DeepSeek-MoE-16B with k=2 and N=64 . (a)-(d): Effectiveness of second-order Hessian over first-order proxies. (e)-(h): Necessity of DS decomposition. (i)-(l): Optimization superiority of the proposed MM algorithm.
Mixture-of-Experts (MoE) models have become a leading approach for decoupling parameter count from computational cost in large language models, yet effectively scaling MoE performance remains a challenge. Prior work shows that fine-grained experts enlarge the space of expert combinations and improve flexibility, but they also impose substantial routing overhead, creating a new scalability bottleneck. In this paper, we explore a complementary axis for scaling -- how expert outputs are aggregated. We theoretically show that replacing the standard weighted-summation aggregation with structural aggregation expands the expert-combination space without altering the experts or router, and enables possible multi-step reasoning within a single MoE layer. To this end, we propose DAG-MoE, a sparse MoE framework that employs a lightweight module to automatically learn the optimal aggregation structure among the selected experts. Extensive experiments under standard language modeling settings show that DAG-MoE consistently improves performance in both pretraining and fine-tuning, surpassing traditional MoE baselines.
Jiarui Feng, Hanqing Zeng, Karish Grover +11
Meta MRS · Washington University in St. Louis · Carnegie Mellon University +1
Mixture-of-Experts (MoE) models scale by activating only a small subset of experts per token. However, training such models remains challenging because top-k routing is discrete and non-differentiable, requiring gradient estimators for expert selection whose design remains a central open problem. We introduce ProbMoE, a probabilistic routing framework that models expert selection as a distribution over cardinality-constrained expert subsets and formulates routing as probabilistic inference in this discrete subset space. We first propose ProbMoE Exact-k routing, which samples k-expert subsets in the forward pass, and the backward pass uses gradients through each expert's exact marginal probability as a tractable surrogate for the true gradient. ProbMoE naturally generalizes to a dynamic-k routing setting, where both training and inference constrain the routing cardinality to the same predefined range, allowing adaptive expert allocation per token. Across benchmarks and model backbones, ProbMoE Exact-k achieves strong performance compared to competitive baselines, with improved expert utilization and routing diversity; ProbMoE Dynamic-k achieves comparable performance with fewer activated experts.
Heng Zhao, Zilei Shao, Guy Van den Broeck +1
Department of Computer Science, University of Virginia, Charlottesville, USA · Department of Computer Science, University of California, Los Angeles, Los Angeles, USA
Mixture-of-experts (MoE) models have recently moved beyond routing a fixed number of complete experts. Shared-expert designs preserve reusable knowledge, fine-grained methods vary computation within experts, and dynamic routers adapt the number of active experts. Yet these decisions are usually made independently, overlooking a basic dependency: extracting reusable computation changes both what remains and how much expert capacity the remainder needs. We study this dependency by decomposing sparsely upcycled feed-forward experts into key-value channels. Co-activated experts align at a subset of value positions; removing these positions changes expert preference; and greater shared coverage is associated with lower residual expert demand. These observations lead to one principle: share first, then route what remains. We instantiate it in UniF-MoE, a unified framework for token-adaptive MoE computation. Each expert is partitioned into aligned blocks. A shared-demand score sets the shared block count and pathway weight, key prototypes select the shared content, and the complementary demand determines the residual expert count through cumulative routing mass. A Gram regularizer separates and normalizes router embeddings, promoting diverse routing directions, sparse expert overlap, and a simple routing geometry. Experiments on DomainBed and GLUE show that this unified design improves predictive performance over representative static and dynamic MoEs while reducing activated computation, inference latency, and memory. Code is available at https://github.com/existence0420/UniF-MoE.
Gongli Zhang, Zhulin Liu, C. L. Philip Chen
Guangdong Provincial Key Laboratory of Computational AI Models and Cognitive Intelligence, School of Computer Science and Engineering, South China University of Technology, Guangzhou 510006, China · Pazhou Lab, Guangzhou 510335, China · Engineering Research Center of the Ministry of Education on Health Intelligent Perception and Paralleled Digital-Human, Guangzhou 510641, China