Mixture-of-Experts (MoE) models enable parameter scaling with limited per-token computation by activating only a small subset of experts for each token, but deploying them still requires loading the complete expert pool into memory. Structured expert pruning can effectively reduce the memory usage by removing experts. However, existing pruning methods either use expert ranking criteria that are not well aligned with model performance or rely on effective expert subset searching that is computationally expensive. Moreover, these methods typically overlook the routing-behavior redundancy among the retained experts. In this paper, we propose MoE Pruning via Router Bias Learning and Expert Approximation (MoRA), a framework for structured MoE expert pruning. We introduce a learnable router bias for each expert and optimize these biases by minimizing the language-modeling loss and a routing-diversity regularizer. The learned router biases sharpen the routing probability distributions to identify experts critical to model performance while encouraging the selection of experts with diverse routing preferences. In addition, we introduce an expert approximation mechanism as a post-pruning enhancement. It leverages the remaining experts to approximate the outputs of pruned experts by affine transformation, further improving the performance of the pruned model. We evaluate MoRA on Qwen3-30B-A3B, DeepSeek-V2-Lite, and Moonlight-16B-A3B, removing 25% and 50% of the routed experts in each MoE layer. Extensive experiments on nine zero-shot benchmarks show that MoRA outperforms state-of-the-art pruning algorithms. Our code will be released.
Figures & tables
Figure 1: Comparison of three ways to select retained MoE experts. (a) Statistic-based pruning assigns each expert a fixed score based on routing statistics and retains the highest-scoring experts. (b) Reconstruction-based pruning evaluates candidate expert subsets using layer-output reconstruction error and selects the subset with the lowest reconstruction error. (c) MoRA adds a router bias to each expert to sharpen the routing probability distribution. These biases are optimized to establish an expert ranking for subsequent selection.
Pruning ratio
Method
ARC-c
ARC-e
BoolQ
HellaSwag
MMLU
OBQA
PIQA
RTE
WinoGrande
Avg.
0%
Qwen3-30B-A3B
56.31
78.87
88.47
77.66
77.81
45.00
80.58
82.67
70.56
73.10
25%
Activation Count
55.48
76.83
87.35
76.44
72.83
43.00
79.76
79.06
68.98
71.08
Average Gate Mass
55.74
77.12
87.47
76.56
72.71
43.60
80.12
79.12
69.61
71.34
MoNE
54.61
76.68
86.79
74.69
72.87
44.80
78.84
80.87
69.85
71.11
MC-SMoE
52.52
80.86
87.38
77.42
70.62
43.80
79.82
79.23
70.59
71.36
HC-SMoE
47.18
71.84
83.55
61.98
64.35
38.00
72.14
80.50
65.27
64.98
Table 1: Zero-shot results (%) on Qwen3-30B-A3B with C4 as the calibration dataset. Bold marks the best result among MoRA and the baselines for each column at each pruning ratio.
Pruning ratio
Method
ARC-c
ARC-e
BoolQ
HellaSwag
MMLU
OBQA
PIQA
RTE
WinoGrande
Avg.
0%
DeepSeek-V2-Lite
49.23
76.60
79.94
77.92
54.97
43.80
80.03
60.65
71.27
66.05
25%
Activation Count
45.14
73.53
71.10
73.38
44.58
42.80
79.60
59.57
70.25
62.22
Average Gate Mass
45.82
74.41
70.52
75.26
46.39
41.80
79.71
60.29
69.22
62.60
MoNE
45.59
72.81
73.68
76.28
47.74
43.20
79.71
60.37
69.88
63.25
MC-SMoE
39.08
66.16
69.20
67.59
36.20
37.40
76.99
57.04
69.30
57.66
HC-SMoE
44.88
72.10
73.37
74.49
47.29
40.60
78.78
61.37
71.03
62.66
Table 2: Zero-shot results (%) on DeepSeek-V2-Lite with C4 as the calibration dataset. Bold marks the best result among MoRA and the baselines for each column at each pruning ratio.
Pruning ratio
Method
ARC-c
ARC-e
BoolQ
HellaSwag
MMLU
OBQA
PIQA
RTE
WinoGrande
Avg.
0%
Moonlight-16B-A3B
58.19
82.58
80.15
78.37
67.31
45.60
80.96
66.06
71.43
70.07
25%
Activation Count
55.12
80.01
78.17
76.96
41.45
45.40
81.18
60.73
70.48
65.50
Average Gate Mass
55.24
80.51
78.32
77.11
48.29
46.80
81.50
60.01
70.43
66.47
MoNE
53.63
78.67
77.56
77.00
53.35
45.20
80.23
59.01
71.43
66.23
MC-SMoE
48.55
76.73
78.10
75.27
39.74
44.80
80.31
57.40
70.25
63.46
HC-SMoE
41.72
70.45
70.34
53.93
53.80
36.40
68.99
59.57
56.59
56.87
Table 3: Zero-shot results (%) on Moonlight-16B-A3B with C4 as the calibration dataset. Bold marks the best result among MoRA and the baselines for each column at each pruning ratio.
Model
Pruning ratio
Retained experts
Model params. (B)
Model storage (GiB)
Affine params. (B)
Affine storage (GiB)
Qwen3-30B-A3B
0%
128
30.532
56.871
0
0
25%
96
23.284
43.371
0.0063
0.012
50%
64
16.037
29.871
0.0126
0.023
DeepSeek-V2-Lite
0%
64
15.706
29.256
0
0
25%
48
12.108
22.552
0.0017
0.0032
50%
32
8.509
15.849
0.0034
0.0063
Table 4: Compression efficiency of MoRA. We report the model parameter counts and storage, and the additional overhead of the affine transformations.
Components
Avg.
LLM
LDiv
Expert Approximation
25%
50%
✓
–
–
67.33
58.08
–
✓
–
67.48
54.57
✓
✓
–
72.17
65.30
✓
✓
✓
72.50
65.83
Table 5: Component ablation on Qwen3-30B-A3B. Avg. denotes the mean score across nine benchmarks.
Model
Pruning
Average
r′
MoRA w/o
ratio
Gate Mass
Gate Mass
Expert Approximation
Qwen3-30B-A3B
25%
71.34
71.55
72.17
50%
58.09
58.81
65.30
DeepSeek-V2-Lite
25%
62.60
62.79
63.25
50%
50.98
51.38
54.02
Moonlight-16B-A3B
25%
66.47
67.07
67.19
Table 6: Mean scores (%) across nine zero-shot benchmarks. Average Gate Mass uses the pretrained router, whereas r′ Gate Mass uses the bias-adjusted routing probabilities. We report the average score of MoRA after pruning without expert approximation.
Figure 2: Relationship between expert scores and deletion-induced loss increases on Qwen3-30B-A3B. Each analysis uses the top 50% of experts within a layer ranked by the corresponding score. Spearman correlations for learned biases and activation frequency are denoted by ρs and ρa , respectively.
Figure 3: Heatmap visualization of token-category routing probabilities in Qwen3 at Layers 0, 24, and 47. Rows correspond to word categories, and darker cells indicate larger routing probabilities. Experts are ordered from left to right by decreasing activation frequency (a) and learned biases (b).
Sparsely-activated Mixture-of-Experts (MoE) language models achieve remarkable inference efficiency by activating only a small fraction of parameters per token, yet their full expert banks reside in memory at all times, creating a prohibitive deployment bottleneck. Existing structured pruning methods, largely designed for dense transformers, assess expert importance using locally derived heuristics that are blind to the interdependent nature of MoE routing. We introduce MAESTRO (Markov-chain Approximated Expert Sparsification via Transition-based ROuting), a structured pruning framework designed for MoE architectures that models autoregressive expert activation trajectories as Ergodic Markov chains whose stationary distributions encode cross-layer dependencies, yielding a globally aware importance heuristic. Evaluated across five diverse domains including Safety, Bias, and Ethics, MAESTRO outperforms state-of-the-art baselines by up to 10.61% in average performance retention under a strict 50% compression regime, while exhibiting substantially lower cross-task variance, indicating that global, routing-congruent pruning produces models that generalize more consistently across heterogeneous tasks.
Mixture-of-Experts (MoE) language models reduce per-token computation through sparse expert activation, yet deployment still requires storing the full expert pool, making one-shot expert pruning a practical approach for reducing memory usage. Although effective, existing criteria are largely heuristic, and no single criterion is universally optimal. Thus, establishing a principle for selecting pruning criteria suited to different deployment objectives remains an important yet largely underexplored problem in one-shot expert pruning. To this end, we introduce a unified formulation for one-shot MoE expert pruning organized around three factors: routing frequency, gate weighting, and activation strength. The formulation yields a criteria selection principle: task-agnostic pruning should favor routed-token-averaged, gate-free activation-based criteria, whereas task-specific pruning can benefit from retaining routing-frequency and gate-weight information. Beyond this principle, the formulation also provides a systematic view of existing heuristic criteria and gives rise to two new task-agnostic criteria, Mean Activation Norm (MAN) and Mean Squared Activation Norm (MSAN). Across four representative MoE models and 16 diverse benchmarks, MAN and MSAN are consistently strong in the task-agnostic setting, obtain the top-two average ranks, and improve average performance by up to 8.8 points over the strongest baseline.
Zongfang Liu, Jinghui Zhang, Zijian Ma +2
1Zhejiang University · 2Westlake University · 3Mohamed bin Zayed University of Artificial Intelligence +1
Mixture-of-Experts (MoE) models decouple parameter count from per-token compute, but deployment still requires hosting every expert in memory. Recent theory shows that experts whose router weights change least during fine-tuning can be pruned with provable accuracy preservation, yet the guarantee assumes full fine-tuning. We show that the signal can be elicited through a brief parameter-efficient adaptation. We fine-tune with a lightweight adapter, rank experts by the induced router change, and prune the least-changed experts in one shot. On Mixtral-8×7B-Instruct, router-only LoRA trains 0.002% of parameters and retains 27.54% MMLU-Pro accuracy with half the experts removed, against roughly 16% for magnitude and random pruning. Signal quality improves monotonically with adapter size, reaching 28.76%, and declines as adaptation spreads beyond the router. Under their shared budget, IA3 reaches 28.04% while Houlsby reaches 25.39%. The criterion transfers to Qwen1.5-MoE fine-tuned for mathematical reasoning, retaining 49.7% mean accuracy over eleven benchmarks with half the experts removed. Structural pruning reduces memory by 49% and per-token latency by 37%. Lightweight router sensitivity therefore makes provably motivated, task-conditioned expert pruning practical at scale.
Ali Janati, Kaoutar El Maghraoui, Chengke Zou +2
Data Science Institute, Columbia University, New York, NY, USA · Department of Computer Science, Columbia University, New York, NY, USA