Mixture-of-Experts (MoE) models enable parameter scaling with limited per-token computation by activating only a small subset of experts for each token, but deploying them still requires loading the complete expert pool into memory. Structured expert pruning can effectively reduce the memory usage by removing experts. However, existing pruning methods either use expert ranking criteria that are not well aligned with model performance or rely on effective expert subset searching that is computationally expensive. Moreover, these methods typically overlook the routing-behavior redundancy among the retained experts. In this paper, we propose MoE Pruning via Router Bias Learning and Expert Approximation (MoRA), a framework for structured MoE expert pruning. We introduce a learnable router bias for each expert and optimize these biases by minimizing the language-modeling loss and a routing-diversity regularizer. The learned router biases sharpen the routing probability distributions to identify experts critical to model performance while encouraging the selection of experts with diverse routing preferences. In addition, we introduce an expert approximation mechanism as a post-pruning enhancement. It leverages the remaining experts to approximate the outputs of pruned experts by affine transformation, further improving the performance of the pruned model. We evaluate MoRA on Qwen3-30B-A3B, DeepSeek-V2-Lite, and Moonlight-16B-A3B, removing 25% and 50% of the routed experts in each MoE layer. Extensive experiments on nine zero-shot benchmarks show that MoRA outperforms state-of-the-art pruning algorithms. Our code will be released.
Figures & tables
Figure 1: Comparison of three ways to select retained MoE experts. (a) Statistic-based pruning assigns each expert a fixed score based on routing statistics and retains the highest-scoring experts. (b) Reconstruction-based pruning evaluates candidate expert subsets using layer-output reconstruction error and selects the subset with the lowest reconstruction error. (c) MoRA adds a router bias to each expert to sharpen the routing probability distribution. These biases are optimized to establish an expert ranking for subsequent selection.
Pruning ratio
Method
ARC-c
ARC-e
BoolQ
HellaSwag
MMLU
OBQA
PIQA
RTE
WinoGrande
Avg.
0%
Qwen3-30B-A3B
56.31
78.87
88.47
77.66
77.81
45.00
80.58
82.67
70.56
73.10
25%
Activation Count
55.48
76.83
87.35
76.44
72.83
43.00
79.76
79.06
68.98
71.08
Average Gate Mass
55.74
77.12
87.47
76.56
72.71
43.60
80.12
79.12
69.61
71.34
MoNE
54.61
76.68
86.79
74.69
72.87
44.80
78.84
80.87
69.85
71.11
MC-SMoE
52.52
80.86
87.38
77.42
70.62
43.80
79.82
79.23
70.59
71.36
HC-SMoE
47.18
71.84
83.55
61.98
64.35
38.00
72.14
80.50
65.27
64.98
Table 1: Zero-shot results (%) on Qwen3-30B-A3B with C4 as the calibration dataset. Bold marks the best result among MoRA and the baselines for each column at each pruning ratio.
Pruning ratio
Method
ARC-c
ARC-e
BoolQ
HellaSwag
MMLU
OBQA
PIQA
RTE
WinoGrande
Avg.
0%
DeepSeek-V2-Lite
49.23
76.60
79.94
77.92
54.97
43.80
80.03
60.65
71.27
66.05
25%
Activation Count
45.14
73.53
71.10
73.38
44.58
42.80
79.60
59.57
70.25
62.22
Average Gate Mass
45.82
74.41
70.52
75.26
46.39
41.80
79.71
60.29
69.22
62.60
MoNE
45.59
72.81
73.68
76.28
47.74
43.20
79.71
60.37
69.88
63.25
MC-SMoE
39.08
66.16
69.20
67.59
36.20
37.40
76.99
57.04
69.30
57.66
HC-SMoE
44.88
72.10
73.37
74.49
47.29
40.60
78.78
61.37
71.03
62.66
Table 2: Zero-shot results (%) on DeepSeek-V2-Lite with C4 as the calibration dataset. Bold marks the best result among MoRA and the baselines for each column at each pruning ratio.
Pruning ratio
Method
ARC-c
ARC-e
BoolQ
HellaSwag
MMLU
OBQA
PIQA
RTE
WinoGrande
Avg.
0%
Moonlight-16B-A3B
58.19
82.58
80.15
78.37
67.31
45.60
80.96
66.06
71.43
70.07
25%
Activation Count
55.12
80.01
78.17
76.96
41.45
45.40
81.18
60.73
70.48
65.50
Average Gate Mass
55.24
80.51
78.32
77.11
48.29
46.80
81.50
60.01
70.43
66.47
MoNE
53.63
78.67
77.56
77.00
53.35
45.20
80.23
59.01
71.43
66.23
MC-SMoE
48.55
76.73
78.10
75.27
39.74
44.80
80.31
57.40
70.25
63.46
HC-SMoE
41.72
70.45
70.34
53.93
53.80
36.40
68.99
59.57
56.59
56.87
Table 3: Zero-shot results (%) on Moonlight-16B-A3B with C4 as the calibration dataset. Bold marks the best result among MoRA and the baselines for each column at each pruning ratio.
Model
Pruning ratio
Retained experts
Model params. (B)
Model storage (GiB)
Affine params. (B)
Affine storage (GiB)
Qwen3-30B-A3B
0%
128
30.532
56.871
0
0
25%
96
23.284
43.371
0.0063
0.012
50%
64
16.037
29.871
0.0126
0.023
DeepSeek-V2-Lite
0%
64
15.706
29.256
0
0
25%
48
12.108
22.552
0.0017
0.0032
50%
32
8.509
15.849
0.0034
0.0063
Table 4: Compression efficiency of MoRA. We report the model parameter counts and storage, and the additional overhead of the affine transformations.
Components
Avg.
LLM
LDiv
Expert Approximation
25%
50%
✓
–
–
67.33
58.08
–
✓
–
67.48
54.57
✓
✓
–
72.17
65.30
✓
✓
✓
72.50
65.83
Table 5: Component ablation on Qwen3-30B-A3B. Avg. denotes the mean score across nine benchmarks.
Model
Pruning
Average
r′
MoRA w/o
ratio
Gate Mass
Gate Mass
Expert Approximation
Qwen3-30B-A3B
25%
71.34
71.55
72.17
50%
58.09
58.81
65.30
DeepSeek-V2-Lite
25%
62.60
62.79
63.25
50%
50.98
51.38
54.02
Moonlight-16B-A3B
25%
66.47
67.07
67.19
Table 6: Mean scores (%) across nine zero-shot benchmarks. Average Gate Mass uses the pretrained router, whereas r′ Gate Mass uses the bias-adjusted routing probabilities. We report the average score of MoRA after pruning without expert approximation.
Figure 2: Relationship between expert scores and deletion-induced loss increases on Qwen3-30B-A3B. Each analysis uses the top 50% of experts within a layer ranked by the corresponding score. Spearman correlations for learned biases and activation frequency are denoted by ρs and ρa , respectively.
Figure 3: Heatmap visualization of token-category routing probabilities in Qwen3 at Layers 0, 24, and 47. Rows correspond to word categories, and darker cells indicate larger routing probabilities. Experts are ordered from left to right by decreasing activation frequency (a) and learned biases (b).