Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughput benefit. We present SlimWise, a serving framework that tailors the expert pool to each inference phase. SlimWise performs prefill with the full model and decode with a pruned model that directly reuses the prefill-generated KV cache without conversion. Across two MoE backbones and three pruning criteria, this training-free KV cache handoff substantially narrows accuracy gaps relative to the full model in many settings. We also show that benchmark accuracy can conceal substantial pruning-induced changes in generation length. To address these distortions and residual accuracy loss, SlimWise introduces a low-cost distillation stage that trains the decoder to continue from full-model KV caches while updating only a small subset of parameters. Implemented in vLLM, SlimWise supports both prefill-decode (PD) disaggregation and PD-colocated serving. On Qwen3.6-35B-A3B, SlimWise improves decode throughput by up to 1.81x at 50% expert pruning with minimal accuracy loss.
Figures & tables
Math
Coding
Tool-Use (BFCLv4)
Model
Method
GSM8K
MATH-500
HE+
MBPP+
Non-Live
Live
Qwen3.6 (35B-A3B)
Baseline
0.962 (1.7k)
0.983 (2.4k)
0.917 (2.3k)
0.969 (1.9k)
0.884 (663)
0.811 (567)
Mass
0.925 (1.7k)
0.966 (2.6k)
0.888 (3.7k)
0.825 (5.3k)
0.793 (484)
0.680 (526)
SlimWise
0.946 (1.6k)
0.971 (2.6k)
0.907 (2.5k)
0.883 (2.9k)
0.803 (532)
0.740 (563)
+ Distill.
0.946 (1.6k)
0.967 (2.4k)
0.921 (2.4k)
0.919 (2.1k)
0.872 (611)
0.792 (542)
EAN
0.941 (2.6k)
0.961 (8.5k)
0.900 (1.1k)
0.930 (1.0k)
0.829 (387)
0.730 (473)
Table 1: Accuracy and output length at 50% expert pruning across two MoE backbones and three pruning criteria. All results use sampled decoding ( T =1.0, top- p =0.95; top- k =20 for Qwen3.6, 64 for Gemma 4) with thinking enabled. Accuracy is averaged over three runs; the parenthesized value is the median output length over all responses pooled across runs. Bold marks the highest accuracy for each model, pruning criterion, and benchmark; underlining marks the highest accuracy across all pruned configurations for each model and benchmark.
Math
Coding
GSM8K
MATH-500
HE+
MBPP+
m′
Method
Acc
p50
p90
Acc
p50
p90
Acc
p50
p90
Acc
p50
p90
256
Baseline
0.962
1.7k
2.5k
0.983
2.4k
6.1k
0.917
2.3k
4.5k
0.969
1.9k
4.4k
192
REAP
0.956
1.8k
2.7k
0.983
2.5k
6.3k
0.915
2.3k
4.6k
0.969
2.5k
8.1k
SlimWise
0.959
1.7k
2.7k
0.978
2.5k
6.2k
0.919
2.4k
4.5k
0.959
2.2k
6.7k
+ Distill.
0.954
1.6k
2.5k
0.981
2.5k
6.1k
0.915
2.3k
4.4k
0.974
1.9k
4.2k
Table 2: Accuracy and output length across retained expert counts m′ on Qwen3.6-35B-A3B with REAP pruning, using sampled decoding with thinking enabled. Accuracy (Acc) is averaged over three runs; p50 and p90 denote the median and 90th-percentile output lengths, respectively, in thousands of tokens, pooled across runs. † denotes students distilled using their own prefill KV caches ( s=0 in Section 3.3 ). Bold marks the highest accuracy for each pruning level and benchmark.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Math
Coding
GSM8K
MATH-500
HE+
MBPP+
m′
Method
Acc
p50
p90
Acc
p50
p90
Acc
p50
p90
Acc
p50
p90
256
Baseline
0.962
1.7k
2.5k
0.983
2.4k
6.1k
0.917
2.3k
4.5k
0.969
1.9k
4.4k
128
REAP
0.949
1.7k
3.2k
0.964
6.6k
15.5k
0.902
0.3k
1.5k
0.932
2.8k
8.9k
Reverse
0.962
1.7k
2.6k
0.976
2.4k
9.9k
0.890
0.3k
2.6k
0.966
2.0k
4.1k
SlimWise
0.959
1.7k
2.9k
0.971
5.8k
13.5k
0.929
2.3k
4.3k
0.967
3.4k
8.6k
Appendix
Table 3: Accuracy and output length across prefill–decode model configurations on Qwen3.6-35B-A3B with REAP-selected expert sets. REAP uses the pruned model for both phases; ‘Reverse’ uses pruned prefill followed by full-model decode; and SlimWise uses full-model prefill followed by pruned decode. All results use sampled decoding ( T=1.0 , top- p=0.95 , top- k=20 ) with thinking enabled and a 32,768-token generation budget.
Math
Coding
Tool-Use (BFCLv4)
Model
Method
GSM8K
MATH-500
HE+
MBPP+
Non-Live
Live
Qwen3.6 (35B-A3B)
Baseline
0.962 (0.006)
0.983 (0.003)
0.917 (0.006)
0.969 (0.003)
0.884 (0.003)
0.811 (0.008)
Mass
0.925 (0.009)
0.966 (0.006)
0.888 (0.010)
0.825 (0.022)
0.793 (0.011)
0.680 (0.004)
SlimWise
0.946 (0.003)
0.971 (0.003)
0.907 (0.008)
0.883 (0.020)
0.803 (0.003)
0.740 (0.001)
+ Distill.
0.946 (0.007)
0.967 (0.003)
0.921 (0.005)
0.919 (0.012)
0.872 (0.005)
0.792 (0.007)
EAN
0.941 (0.008)
0.961 (0.003)
0.900 (0.012)
0.930 (0.018)
0.829 (0.004)
0.730 (0.004)
Appendix
Table 4: Accuracy at 50% expert pruning across two MoE backbones and three pruning criteria. All evaluations use sampled decoding ( T=1.0 , top- p=0.95 ; top- k=20 for Qwen3.6 and 64 for Gemma 4) with thinking enabled. Each cell reports mean accuracy over three runs, with the standard deviation in parentheses. BFCL uses the same decoding settings with a 32,768-token budget per call; we report its Non-Live and Live function-calling aggregates. Bold indicates the highest accuracy within each model, pruning criterion, and benchmark; underlining indicates the highest accuracy across all pruned configurations for each model and benchmark.