EAT: Expert Account Tracker for Efficient MoE Inference
Authors: Yuexian Li, Yifei Yang, Zouying Cao, Hai Zhao
Organizations: Paris Elite Institute of Technology, Shanghai Jiao Tong University · AGI Institute, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China · Key Laboratory of Shanghai Education Commission for Intelligent Interaction and Cognitive Engineering, Shanghai Jiao Tong University · Ant Group
Mixture-of-Experts (MoE) models have emerged as a revolutionary method to scale Transformer models. However, traditional MoE architecture still suffers from inefficiency since a large number of experts are unnecessarily activated. Existing approaches for reducing the number of activated experts often overlook the historical performance of each expert. In this paper, we propose EAT, a novel method called Expert Account Tracker (EAT), which utilizes history-awareness metrics and adaptive thresholding to dynamically select the most important experts, thereby reducing the activated expert number while effectively maintaining the model performance. Experiments show that EAT outperforms the existing baseline Top-P method across multiple models and datasets, achieving over 25% an average reduction compared to the vanilla method in the number of activated experts and performing better token generation speed compared to the baseline. Furthermore, the performance of pruned models can be efficiently recovered via OPD using only 9K data. Additionally, through ablation studies, we find that excessively reducing the number of activated experts can significantly harm model performance, and the importance of experts varies across layers, with higher-level experts being generally more critical.
Figures & tables
Figure 1: Overview of MoE architecture and our proposed EAT method. EAT selects the most useful experts by continuously tracking their historical contributions and incorporating this information into the routing decision.
Hyperparameter
Value
Description
α
0.95
Smoothing factor for Cs(t) .
w1,w2,w3
0.3, 0.3, 0.4
Weights to compute Sh,i .
β
0.6
Balance factor for final importance Ii .
α1,α2
0.5, 1.0
Scaling coefficients for threshold τ .
κ1,κ2
0.7, 0.3
Weights for threshold adjustment.
K (Mixtral/Phi)
2
Max activated experts per token.
Table 1: Key hyperparameters of the proposed EAT routing strategy.
LLM
Method
Reasoning
Language
Know.
Examination
Und.
HeSw
PIQA
CHID
WSC
BoolQ
MMLU
CMMLU
XSum
Mixtral -8x7B -v0.1
Vanilla
77.11
81.07
37.51
61.54
69.11
71.67
53.11
9.19
Top-P
75.72
79.71
32.85
60.58
66.15
65.38
46.74
8.70
EAT
76.79
80.30
33.97
63.46
68.40
70.37
51.04
9.08
EAT + OPD
77.05
80.96
35.42
64.83
69.25
71.59
52.66
9.10
Phi -3.5-MoE -instruct
Vanilla
75.17
80.20
66.75
68.27
75.32
76.64
61.03
14.68
Table 2: Main experimental results evaluated on the OpenCompass Platform. Know. denotes Knowledge benchmarks and Und. denotes Understanding benchmark.
LLM
Method
Reasoning
Language
Know.
Examination
Und.
HeSw
PIQA
CHID
WSC
BoolQ
MMLU
CMMLU
XSum
Mixtral -8x7B -v0.1
Vanilla
2.00
2.00
2.00
2.00
2.00
2.00
2.00
2.00
Top-P
1.50
1.44
1.75
1.52
1.46
1.51
1.74
1.51
EAT
1.47
1.46
1.44
1.50
1.45
1.47
1.48
1.47
EAT + OPD
1.50
1.47
1.46
1.50
1.45
1.69
1.57
1.71
Phi -3.5-MoE -instruct
Vanilla
2.00
2.00
2.00
2.00
2.00
2.00
2.00
2.00
Table 3: Comparison of the average number of activated experts under different expert routing strategies.
Length
Mixtral-8x7B
Phi-3.5-MoE
Vanilla
Top-P
EAT
Vanilla
Top-P
EAT
2048+256
0.5214
0.5100
0.5196
1.1319
0.9479
1.1600
1024+128
0.3931
0.3739
0.3864
0.9295
0.9073
0.9158
512+32
0.3814
0.1596
0.3671
0.8060
0.8719
1.1421
Table 4: Average token generation speed (token/sec).
Figure 2: Curve of the Relationship Between PPL and Number of Activated Experts.
Figure 3: Curve of the Relationship Between layer number and Number of Activated Experts.
Figure 4: Curve of PPL on OPD with different subset size on Qwen30B on PIQA dataset.
Figure 5: Curve of PPL on OPD with different steps on Qwen30B on BOOLQ dataset.
Sparsely-activated Mixture-of-Experts (MoE) language models achieve remarkable inference efficiency by activating only a small fraction of parameters per token, yet their full expert banks reside in memory at all times, creating a prohibitive deployment bottleneck. Existing structured pruning methods, largely designed for dense transformers, assess expert importance using locally derived heuristics that are blind to the interdependent nature of MoE routing. We introduce MAESTRO (Markov-chain Approximated Expert Sparsification via Transition-based ROuting), a structured pruning framework designed for MoE architectures that models autoregressive expert activation trajectories as Ergodic Markov chains whose stationary distributions encode cross-layer dependencies, yielding a globally aware importance heuristic. Evaluated across five diverse domains including Safety, Bias, and Ethics, MAESTRO outperforms state-of-the-art baselines by up to 10.61% in average performance retention under a strict 50% compression regime, while exhibiting substantially lower cross-task variance, indicating that global, routing-congruent pruning produces models that generalize more consistently across heterogeneous tasks.
Mixture-of-Experts (MoE) scales language models efficiently through sparse expert activation, and its dynamic variant further reduces computation by adjusting the activated experts in an input-dependent manner. Existing dynamic MoE methods usually rely on pre-training from scratch or task-specific adaptation, leaving the practical conversion of fully trained MoE underexplored. Enabling such adaptation would directly alleviate the inference costs by allowing easy tokens to bypass unnecessary expert during serving. This paper introduces Zero-Expert Self-Distillation Adaptation (ZEDA), a low-cost framework that transforms post-trained static MoE models into efficient dynamic ones. To stabilize this architectural conversion, ZEDA injects parameter-free zero-output experts into each MoE layer and adapts the augmented model through two-stage self-distillation, utilizing the original MoE as a frozen teacher and applying a group-level balancing loss. On Qwen3-30B-A3B and GLM-4.7-Flash across 11 benchmarks spanning math, code, and instruction following, ZEDA eliminates over 50% of expert FLOPs at marginal accuracy loss. It outperforms the strongest dynamic MoE baseline by 6.1 and 4.0 points on the two models, and delivers ~1.20× end-to-end inference speedup.
Xingtai Lv, Li Sheng, Kaiyan Zhang +12
Tsinghua University · Shanghai AI Lab · WeChat AI +1
Mixture-of-Experts (MoE) models enable efficient scaling of large language models but face critical deployment challenges due to massive memory requirements. Existing pruning methods either incur prohibitive search costs or neglect the dynamic interdependencies between experts. To address these challenges, we present OMP-MoE, a novel training-free compression framework for reducing expert redundancy in MoE-based LLMs. Based on observations of expert contribution patterns, we reformulate the pruning problem as a sparse signal reconstruction task solved through Orthogonal Matching Pursuit. Specifically, our method first treats individual expert contributions as dictionary atoms and selects experts that greedily minimize reconstruction error with linear computational complexity. Then, we optimize cross-layer expert allocation through a water-filling strategy that accounts for both reconstruction quality and routing stability. Finally, we introduce OMP-MoE†, an adaptive inference mechanism that dynamically adjusts expert activation based on energy prediction. Comprehensive experiments on Qwen, DeepSeek-V2, GPT-OSS, and Mixtral MoE demonstrate consistent improvements over existing methods at 25-50% pruning ratios. For Qwen3-30B-A3B at 50% compression, we retain 93.3% of original performance, achieving 33× faster search and 1.55× inference speedup. Codes will be available after acceptance.
Dezhi Li, Lujun Li, Qiyuan Zhu +4
The Hong Kong University of Science and Technology