Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughput benefit. We present SlimWise, a serving framework that tailors the expert pool to each inference phase. SlimWise performs prefill with the full model and decode with a pruned model that directly reuses the prefill-generated KV cache without conversion. Across two MoE backbones and three pruning criteria, this training-free KV cache handoff substantially narrows accuracy gaps relative to the full model in many settings. We also show that benchmark accuracy can conceal substantial pruning-induced changes in generation length. To address these distortions and residual accuracy loss, SlimWise introduces a low-cost distillation stage that trains the decoder to continue from full-model KV caches while updating only a small subset of parameters. Implemented in vLLM, SlimWise supports both prefill-decode (PD) disaggregation and PD-colocated serving. On Qwen3.6-35B-A3B, SlimWise improves decode throughput by up to 1.81x at 50% expert pruning with minimal accuracy loss.
Figures & tables
Math
Coding
Tool-Use (BFCLv4)
Model
Method
GSM8K
MATH-500
HE+
MBPP+
Non-Live
Live
Qwen3.6 (35B-A3B)
Baseline
0.962 (1.7k)
0.983 (2.4k)
0.917 (2.3k)
0.969 (1.9k)
0.884 (663)
0.811 (567)
Mass
0.925 (1.7k)
0.966 (2.6k)
0.888 (3.7k)
0.825 (5.3k)
0.793 (484)
0.680 (526)
SlimWise
0.946 (1.6k)
0.971 (2.6k)
0.907 (2.5k)
0.883 (2.9k)
0.803 (532)
0.740 (563)
+ Distill.
0.946 (1.6k)
0.967 (2.4k)
0.921 (2.4k)
0.919 (2.1k)
0.872 (611)
0.792 (542)
EAN
0.941 (2.6k)
0.961 (8.5k)
0.900 (1.1k)
0.930 (1.0k)
0.829 (387)
0.730 (473)
Table 1: Accuracy and output length at 50% expert pruning across two MoE backbones and three pruning criteria. All results use sampled decoding ( T =1.0, top- p =0.95; top- k =20 for Qwen3.6, 64 for Gemma 4) with thinking enabled. Accuracy is averaged over three runs; the parenthesized value is the median output length over all responses pooled across runs. Bold marks the highest accuracy for each model, pruning criterion, and benchmark; underlining marks the highest accuracy across all pruned configurations for each model and benchmark.
Math
Coding
GSM8K
MATH-500
HE+
MBPP+
m′
Method
Acc
p50
p90
Acc
p50
p90
Acc
p50
p90
Acc
p50
p90
256
Baseline
0.962
1.7k
2.5k
0.983
2.4k
6.1k
0.917
2.3k
4.5k
0.969
1.9k
4.4k
192
REAP
0.956
1.8k
2.7k
0.983
2.5k
6.3k
0.915
2.3k
4.6k
0.969
2.5k
8.1k
SlimWise
0.959
1.7k
2.7k
0.978
2.5k
6.2k
0.919
2.4k
4.5k
0.959
2.2k
6.7k
+ Distill.
0.954
1.6k
2.5k
0.981
2.5k
6.1k
0.915
2.3k
4.4k
0.974
1.9k
4.2k
Table 2: Accuracy and output length across retained expert counts m′ on Qwen3.6-35B-A3B with REAP pruning, using sampled decoding with thinking enabled. Accuracy (Acc) is averaged over three runs; p50 and p90 denote the median and 90th-percentile output lengths, respectively, in thousands of tokens, pooled across runs. † denotes students distilled using their own prefill KV caches ( s=0 in Section 3.3 ). Bold marks the highest accuracy for each pruning level and benchmark.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Math
Coding
GSM8K
MATH-500
HE+
MBPP+
m′
Method
Acc
p50
p90
Acc
p50
p90
Acc
p50
p90
Acc
p50
p90
256
Baseline
0.962
1.7k
2.5k
0.983
2.4k
6.1k
0.917
2.3k
4.5k
0.969
1.9k
4.4k
128
REAP
0.949
1.7k
3.2k
0.964
6.6k
15.5k
0.902
0.3k
1.5k
0.932
2.8k
8.9k
Reverse
0.962
1.7k
2.6k
0.976
2.4k
9.9k
0.890
0.3k
2.6k
0.966
2.0k
4.1k
SlimWise
0.959
1.7k
2.9k
0.971
5.8k
13.5k
0.929
2.3k
4.3k
0.967
3.4k
8.6k
Appendix
Table 3: Accuracy and output length across prefill–decode model configurations on Qwen3.6-35B-A3B with REAP-selected expert sets. REAP uses the pruned model for both phases; ‘Reverse’ uses pruned prefill followed by full-model decode; and SlimWise uses full-model prefill followed by pruned decode. All results use sampled decoding ( T=1.0 , top- p=0.95 , top- k=20 ) with thinking enabled and a 32,768-token generation budget.
Math
Coding
Tool-Use (BFCLv4)
Model
Method
GSM8K
MATH-500
HE+
MBPP+
Non-Live
Live
Qwen3.6 (35B-A3B)
Baseline
0.962 (0.006)
0.983 (0.003)
0.917 (0.006)
0.969 (0.003)
0.884 (0.003)
0.811 (0.008)
Mass
0.925 (0.009)
0.966 (0.006)
0.888 (0.010)
0.825 (0.022)
0.793 (0.011)
0.680 (0.004)
SlimWise
0.946 (0.003)
0.971 (0.003)
0.907 (0.008)
0.883 (0.020)
0.803 (0.003)
0.740 (0.001)
+ Distill.
0.946 (0.007)
0.967 (0.003)
0.921 (0.005)
0.919 (0.012)
0.872 (0.005)
0.792 (0.007)
EAN
0.941 (0.008)
0.961 (0.003)
0.900 (0.012)
0.930 (0.018)
0.829 (0.004)
0.730 (0.004)
Appendix
Table 4: Accuracy at 50% expert pruning across two MoE backbones and three pruning criteria. All evaluations use sampled decoding ( T=1.0 , top- p=0.95 ; top- k=20 for Qwen3.6 and 64 for Gemma 4) with thinking enabled. Each cell reports mean accuracy over three runs, with the standard deviation in parentheses. BFCL uses the same decoding settings with a 32,768-token budget per call; we report its Non-Live and Live function-calling aggregates. Bold indicates the highest accuracy within each model, pruning criterion, and benchmark; underlining indicates the highest accuracy across all pruned configurations for each model and benchmark.
Mixture-of-Experts (MoE) models enable efficient scaling of large language models but face critical deployment challenges due to massive memory requirements. Existing pruning methods either incur prohibitive search costs or neglect the dynamic interdependencies between experts. To address these challenges, we present OMP-MoE, a novel training-free compression framework for reducing expert redundancy in MoE-based LLMs. Based on observations of expert contribution patterns, we reformulate the pruning problem as a sparse signal reconstruction task solved through Orthogonal Matching Pursuit. Specifically, our method first treats individual expert contributions as dictionary atoms and selects experts that greedily minimize reconstruction error with linear computational complexity. Then, we optimize cross-layer expert allocation through a water-filling strategy that accounts for both reconstruction quality and routing stability. Finally, we introduce OMP-MoE†, an adaptive inference mechanism that dynamically adjusts expert activation based on energy prediction. Comprehensive experiments on Qwen, DeepSeek-V2, GPT-OSS, and Mixtral MoE demonstrate consistent improvements over existing methods at 25-50% pruning ratios. For Qwen3-30B-A3B at 50% compression, we retain 93.3% of original performance, achieving 33× faster search and 1.55× inference speedup. Codes will be available after acceptance.
Dezhi Li, Lujun Li, Qiyuan Zhu +4
The Hong Kong University of Science and Technology
Fine-grained Mixture-of-Experts (MoE) models sparsely activate only a subset of experts per token, reducing activated computation while maintaining high model capacity. However, in memory-constrained inference scenarios, only a small set of experts can be cached. Experts not in the cache must be fetched from slow external storage (e.g., UFS), leading to frequent evictions and substantial I/O overhead. We propose ReMoE, a router fine-tuning framework designed to boost token-wise expert reuse. ReMoE biases the router toward recently selected experts, producing temporally stable routing that better matches cache locality constraints. By increasing short-horizon expert reuse, ReMoE reduces expert fetches from storage without adding inference-time computation. Experiments on DeepSeek and Qwen models show that ReMoE improves expert reuse by 26% while maintaining downstream task performance. Real-system evaluations further confirm these benefits, improving output throughput by 8.4% under vLLM GPU-CPU expert offloading and reducing TPOT by 43.6-49.8% under llama.cpp on Jetson Orin NX, corresponding to a 1.77-1.99× decode speedup across diverse workloads. Checkpoints and usage instructions are available at https://github.com/BUAA-OSCAR/ReMoE.
Xiongwei Zhu, Xiaojian Liao, Tianyang Jiang +3
School of Computer Science and Engineering, Beihang University, Beijing 100191, China · Huawei Technologies Ltd.
Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, but their routed expert weights often exceed accelerator memory. We target latency-critical single-user settings where routed experts are staged on demand from CPU memory to a GPU or from Flash to a mobile NPU. In this setting, self-speculative decoding faces a new bottleneck: increasing the draft expert set improves accuracy but triggers extra expert loading, while cheap small-footprint drafts have low acceptance; moreover, verifying a multi-token block activates the union of target experts and is no longer close to one target step. We propose DraftExpert, an expansion-aware self-speculative decoding framework for expert-offloaded MoE inference. DraftExpert trains one lightweight accelerator-resident draft expert per layer by self-distilling residual, logit/token, and router-agreement signals from the frozen target MoE. At inference time, it uses a fixed-footprint shared+top-1+draft-expert drafter together with confidence--expansion truncation and target-expert prefetching, while final tokens are still exactly verified by the target model. On DeepSeek-V2-Lite and Moonlight-16B-A3B across CPU-GPU and Flash-NPU offload, DraftExpert improves decode throughput by 1.45x on average, raises draft acceptance to 8487%, and achieves 8688% prefetch hit rates.