cs.LGSep 24, 2026

RAZOR: Pruning Replaceable Experts in LLMs

Authors: Mingyang Song, Mao Zheng

Organizations: Foundation Model Department, Tencent, China

Abstract

Mixture-of-experts (MoE) models activate only a few experts per token but store the entire expert pool. Pruning this pool requires identifying experts whose removal preserves model behavior. Routing frequency and output magnitude do not fully describe deletion damage, which also depends on how the surviving and replacement experts compensate for the removed output. We introduce RAZOR, a training-free pruning method based on consensus residuals, the deviations of expert outputs from their original weighted mixture. At a fixed layer input, these residuals give the exact output change for a single deletion under survivor renormalization and router refill. RAZOR aggregates this damage by conditional root mean square and selects experts under a layerwise budget using forward computation alone, without gradients, subset search, or recovery training. Against frequency, activation-norm, and REAP baselines on GLM-4.7-Flash and Qwen3.6-35B-A3B at 25% and 50% expert removal, it attains the highest macro average over nine reasoning-intensive tasks in all four model-budget settings, gaining 2.12-5.59 points over REAP and lowering reverse KL in all four. On DeepSeek-V4-Flash-0731 and Hy3, it also achieves the highest macro average among the three residual criteria. Local exactness does not guarantee better joint pruning. Generation analyses show changes in diversity, formatting, and termination despite higher task scores.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Aug 7, 2026cs.LG

Shape Mutating Expert Compression:LorExperts and BTExperts

Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing their many expert weight matrices. Expert pruning (e.g., REAP) and merging reduce cost but sacrifice accuracy and require retraining the router; low-rank delta decomposition of experts (e.g., D^2-MoE) preserves all experts and the router, but degrades sharply as the expert count grows because a single shared component cannot approximate many near-orthogonal experts. Because MoE expert weights are near-orthogonal, a single shared component (as in prior delta decomposition) scales poorly with the expert count; we show that experts nonetheless organize into functional co-activation communities that are decoupled from weight similarity. Building on this, we introduce LorExperts, a router-preserving compression method that clusters experts, keeps one full-precision dominant per cluster, and represents the remaining members as low-rank corrections to their local dominant. LorExperts retains all experts and the original router (no router retraining). At ~50% expert compression on Qwen3-30B-A3B and Gemma-4-26B-A4B, LorExperts preserves downstream accuracy and perplexity better than the baselines on most of the tasks; the margin over D^2-MoE grows with expert count E. We further give a reconstruction fine-tuning procedure for LorExperts, and BTExperts, a tree organization of dominants and corrections that enables inference-time amortization of shared computation.
Jul 2, 2026cs.AI

Generic Expert Coverage for Pruning SparseMixture-of-Experts Language Models

Sparsely activated Mixture-of-Experts (MoE) language models contain substantial structured redundancy among routed experts, but pruning them without downstream calibration data remains challenging. Existing expert-pruning methods typically rely on a single aggregated importance score, which can bias the retained set toward experts favored by dominant calibration patterns. We propose \textbf{Generic TB-Coverage}, a coverage-aware expert pruning method that uses only generic text corpora (WikiText2 and C4) for calibration. Instead of collapsing expert utility into one score, our method profiles per-expert utility separately on each corpus and enforces a fixed-budget coverage rule that preserves high-utility experts from each corpus before constructing the final pruning mask. Across Qwen1.5-MoE-A2.7B and DeepSeek-MoE-16B-Base at 25%, 50%, and 75% retention budgets, our method improves average accuracy on six common zero-shot benchmarks over random pruning, REAP, and ExpertSparsity, while also reducing perplexity degradation on WikiText2 and C4. The gains are largest under aggressive pruning (25% and 50% retain), suggesting that preserving cross-corpus expert coverage is an effective generic-data prior for MoE pruning. Our improvements hold with fixed pruning budgets and no downstream calibration data.
Aug 8, 2026cs.LG

Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models

Mixture-of-Experts (MoE) models decouple parameter count from per-token compute, but deployment still requires hosting every expert in memory. Recent theory shows that experts whose router weights change least during fine-tuning can be pruned with provable accuracy preservation, yet the guarantee assumes full fine-tuning. We show that the signal can be elicited through a brief parameter-efficient adaptation. We fine-tune with a lightweight adapter, rank experts by the induced router change, and prune the least-changed experts in one shot. On Mixtral-8×\times7B-Instruct, router-only LoRA trains 0.002% of parameters and retains 27.54% MMLU-Pro accuracy with half the experts removed, against roughly 16% for magnitude and random pruning. Signal quality improves monotonically with adapter size, reaching 28.76%, and declines as adaptation spreads beyond the router. Under their shared budget, IA3 reaches 28.04% while Houlsby reaches 25.39%. The criterion transfers to Qwen1.5-MoE fine-tuned for mathematical reasoning, retaining 49.7% mean accuracy over eleven benchmarks with half the experts removed. Structural pruning reduces memory by 49% and per-token latency by 37%. Lightweight router sensitivity therefore makes provably motivated, task-conditioned expert pruning practical at scale.