Mixture-of-experts (MoE) models activate only a few experts per token but store the entire expert pool. Pruning this pool requires identifying experts whose removal preserves model behavior. Routing frequency and output magnitude do not fully describe deletion damage, which also depends on how the surviving and replacement experts compensate for the removed output. We introduce RAZOR, a training-free pruning method based on consensus residuals, the deviations of expert outputs from their original weighted mixture. At a fixed layer input, these residuals give the exact output change for a single deletion under survivor renormalization and router refill. RAZOR aggregates this damage by conditional root mean square and selects experts under a layerwise budget using forward computation alone, without gradients, subset search, or recovery training. Against frequency, activation-norm, and REAP baselines on GLM-4.7-Flash and Qwen3.6-35B-A3B at 25% and 50% expert removal, it attains the highest macro average over nine reasoning-intensive tasks in all four model-budget settings, gaining 2.12-5.59 points over REAP and lowering reverse KL in all four. On DeepSeek-V4-Flash-0731 and Hy3, it also achieves the highest macro average among the three residual criteria. Local exactness does not guarantee better joint pruning. Generation analyses show changes in diversity, formatting, and termination despite higher task scores.
Figures & tables
Figure 1: Replaceability depends on output geometry and router refill. A constructed single-token example without (a) and with (b) refill. Damage is L2 distance to the original mixture. Bold marks row minima. Appendix A gives the construction and plotting details.
Figure 2: Domain-wise reverse KL at 25% (top) and 50% (bottom) removal. Values are in nats, with lower (inward) better. Compare methods along each spoke, not by polygon area. Whiskers show pointwise 95% paired-batch intervals. Shading distinguishes ID/OOD domains.
Model
Method
Math
Instruction Following
Know.
Tool
LC
Coding
Overall Avg
AIME’26
IFEval
IFBench
Avg
SuperGPQA
BFCL v4
LongB. v2
HE+
LCB’26
SWE
Avg
[][]
0%
—
31.7
79.7
40.0
59.9
38.5
65.3
25.3
79.9
37.4
5.2
40.8
44.8
[][]
Frequency
26.3
77.3
37.3
57.3
29.5
62.8
21.9
74.4
29.3
5.2
36.3
40.4
[][]
EAN
30.4
76.3
39.3
57.8
31.5
64.7
27.2
75.0
33.8
4.8
37.9
42.6
[][]
REAP ⋆
30.0
74.9
36.7
55.8
32.6
64.9
27.2
75.0
30.6
3.6
36.4
41.7
[][]
REAP
30.8
74.9
37.3
56.1
33.7
65.0
27.0
75.0
34.3
3.6
37.6
42.4
Table 1: Nine-task benchmark results on GLM-4.7-Flash and Qwen3.6-35B-A3B at 25% and 50% removal. Scores are scaled by 100. Bold marks the best pruned score in each column within a model–budget block.
Model
Method
Math
Instruction Following
Know.
Tool
LC
Coding
Overall Avg
AIME’26
IFEval
IFBench
Avg
SuperGPQA
BFCL v4
LongB. v2
HE+
LCB’26
SWE
Avg
[][]
0%
—
67.9
86.1
37.7
61.9
60.3
72.6
33.8
83.5
69.3
21.2
58.0
59.2
[][]
RCS
73.8
82.6
37.0
59.8
55.4
70.2
35.4
91.5
68.3
12.4
57.4
58.5
[][]
RCS-LOO
76.7
83.7
38.7
61.2
55.6
70.4
35.2
89.6
68.8
18.8
59.1
59.7
[][]
25%
Razor
81.3
84.7
36.3
60.5
56.0
72.8
37.4
93.3
69.1
18.2
60.2
61.0
[][]
RCS
75.8
78.4
36.7
57.6
46.7
66.5
30.4
84.2
58.0
8.4
50.2
53.9
Table 2: Nine-task benchmark results on DeepSeek-V4-Flash-0731 and Hy3 at 25% and 50% removal. These two backbones compare the residual criteria against the unpruned reference, without the Frequency, EAN, and REAP baselines of Table 1 . Other conventions follow Table 1 .
Figure 3: Qwen3.6-35B-A3B response diversity and response form. The left panel shows Distinct- n changes from Original. The right panels show completions containing </think> per 10,000 LiveCodeBench completions, median token length, and loop share among length-limited completions. Dashed lines mark Original, with 95% bands in the form panels. Whiskers show pointwise 95% bootstrap intervals.
Figure 4: RMS aggregation and expert selection for Razor . Left to right, panels show relative KL change from mean (negative favors RMS), expert CV, RMS-only retained experts per layer, and deletion-set Gram ratios. Shading marks 50% removal, used in the last three panels.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Domain
Samples
Source
Coding
512
Dolci-Instruct-SFT and Nemotron competitive programming, SWE, and OpenCode
Mathematics
256
Nemotron-Math-v2, high-reasoning AoPS subset
Science
256
Nemotron-Science-v1
Chinese-STEM
256
Nemotron-SFT-Multilingual-v1
Instruction following
256
Dolci constraint-following mixture
Tool calling
256
Nemotron-Agentic-v1
Appendix
Table 3: RazorCal domain allocations and calibration sources. Counts denote source examples, not tokens.
Figure 5: Overview of the RazorCal calibration pool. (a) Cosine t-SNE projection of Qwen3-Embedding-8B representations of the rendered conversations, with color and marker shape encoding domain and labels marking domain medians. (b) Per-domain character-length coverage over the full observed range, counting message content only. The curve is a kernel-density estimate in log-character space, the bar is the interquartile range with the median marked, the thin line spans minimum to maximum, and each dot is one of the 2,048 examples. Domain counts are given in Table 3 .
Qwen3.6-
GLM-4.7-
DeepSeek-V4-
Hy3
35B-A3B
Flash
Flash-0731
Architecture
Total parameters
35B
30B
284B
295B
Active parameters per token
∼ 3B
∼ 3B
∼ 13B
∼ 21B
Decoder layers
40
47
43
80
MoE layers
40
46
43
79
Appendix
Table 4: Architectures and pruning settings of the four main backbones. Parameter counts are nominal backbone sizes, and expert counts are per MoE layer.
Instance
Token quantity
Evidence in this paper
RCS-Refill ( Razor )
1−wi+wr∥wiri−wrrr∥2
Nine-task benchmarks on all four backbones, with baselines on two (Tables 1 , 2 ). Reverse KL (Figure 2 , Table 6 ), score distributions (Table 7 ), and refill mechanism (Table 8 ). Mean–RMS reduction and factorial comparisons (Figure 6 , Table 9 ). Per-axis, routing-stratified and PPL fidelity, selection probes, and Qwen3.6-35B-A3B response behavior (Table 10 ).
RCS-LOO
1−wiwi∥ri∥2
Nine-task benchmarks on all four backbones, with baselines on two (Tables 1 , 2 ). Output-fidelity and diversity figures, response behavior, reverse KL, score distributions, and calibration-budget and corpus diagnostics.
RCS
wi∥ri∥2
Nine-task benchmarks on all four backbones, with baselines on two (Tables 1 , 2 ). Reverse KL (Figure 2 , Table 6 ) and the reduction axis (Figure 6 ).
REAP-RMS
wi∥fi∥2
Reverse KL in the separate factorial collection (Appendix C.5 ). No benchmarks.
Appendix
Table 5: Scoring criteria and experimental coverage.
Removal
REAP
RCS
RCS-LOO
Razor
Δ vs REAP
Refill gain
Qwen3.6-35B-A3B
25%
0.07461
0.06669
0.07141
0.06872
−4.3%
+3.8%
50%
0.45654
0.40337
0.45108
0.44397
−1.2%
+1.6%
GLM-4.7-Flash
25%
0.22850
0.20481
0.20351
0.20459
−10.9%
−0.5%
50%
1.01690
0.92310
0.87103
0.85457
−14.3%
+1.9%
Appendix
Table 6: Reverse KL across scoring rules in nats (lower is better).
Qwen3.6-35B-A3B
GLM-4.7-Flash
RCS-LOO
Razor
RCS-LOO
Razor
Sample
Scored tokens
785,703
785,632
Coefficient of variation
CV median
0.643
0.599
0.577
0.574
CVp5 / p95
0.370 / 1.255
0.373 / 1.184
0.350 / 1.002
0.362 / 0.978
Appendix
Table 7: Damage variability and mean–RMS selection differences for RCS-LOO and Razor .
Figure 6: Reverse KL under different calibration reductions (lower is better). The RCS quantity uses conditional RMS, conditional mean, and corpus sum. The RCS-LOO and RCS-Refill quantities each use the two conditional reductions. Selection frequency is excluded because it discards the token quantity rather than reducing it.
Quantity
Qwen3.6-35B-A3B
GLM-4.7-Flash
Routing geometry (token-weighted)
Selected weight wi ( =1/k )
0.1250
0.2500
Promoted pseudo-weight wr
0.0740
0.1325
cos(ri,rr)
0.1080
0.1034
Score ratio sirf/siloo (per expert)
Median
1.0281
0.9744
Appendix
Table 8: Calibration statistics of routing weights and refill-induced score changes.
Token quantity
Aggregate
Qwen3.6-35B-A3B
GLM-4.7-Flash
25%
50%
25%
50%
wi∥fi∥2
Sum
0.14712
0.67725
0.22309
0.89425
wi∥fi−c∥2
Sum
0.16005
0.67173
0.23513
0.96880
1−wiwi∥fi−c∥2
Sum
0.17659
0.69294
0.23339
0.91110
1−wi+wr∥wiri−wrrr∥2
Sum
0.18798
0.70264
0.23176
0.88086
wi∥fi∥2
Mean
0.07469
0.45691
0.22852
1.01692
Appendix
Table 9: Fourteen scoring configurations varying reference, routing factor, and aggregation. Reverse KL is in nats at 25% and 50% removal, with lower values better. Thirteen rows share four evaluation shards. The conditional-mean refill row comes from the main sweep in Figure 6 , which differs in one shard. Comparisons involving this row are not fully matched. Sum denotes ∑tgi(t)2 . No corpus-sum result is available for 1−wiwi∥fi∥2 . Refill damage with RMS is Razor .
Figure 7: Calibration budget and selection agreement on Qwen3.6-35B-A3B. Subset agreement (left pair) compares two disjoint calibration subsets, while full-set agreement (right pair) compares each against the full 32-shard selection. Tokens per shard count are given in the text.
Figure 8: Routing-stratified fidelity at 50% removal. The three residual criteria use conditional RMS and are evaluated on the same batches, reference prefixes, and token positions. All panels compare against REAP at the same decile, with a dashed zero line and higher values indicating better fidelity. The top two rows show reverse-KL reduction in percent. The bottom two show REAP’s ∣PPLq−PPLp∣ minus the criterion’s, in PPL points (Equation 24 ), so positive values mean closer to the original. Differences avoid unstable ratios where REAP’s deviation approaches zero ( 0.0016 on Hy3 ID). Bands are pointwise 95% paired batch-bootstrap intervals, and panels use independent vertical scales.
Criterion
Cut
Stray </think> /10k
No fence
Prefix
Len-lim.
Med. kchar
Qwen3.6-35B-A3B, 44,800 completions per checkpoint
Original
–
3.6 [1.6, 6.0]
7.1
61.4
16.0
8.3
EAN
25%
0.4 [0.0, 1.1]
8.7
52.3
19.5
9.7
Frequency
25%
2.0 [0.7, 3.6]
5.2
60.5
15.1
9.2
REAP
25%
1.6 [0.4, 3.1]
8.4
55.1
19.1
10.3
RCS-LOO
25%
0.2 [0.0, 0.7]
8.4
52.3
17.0
8.6
Appendix
Table 10: Response form on LiveCodeBench by criterion and budget, Qwen3.6-35B-A3B. Rates are percentages except stray </think> per 10,000. Lengths are in kchar. Brackets give pointwise 95% question-bootstrap intervals. GLM-4.7-Flash is measured but not printed here because that collection covers only REAP and RCS-LOO, with no Razor row for comparison. Its RCS-LOO ladder appears in Table 11 .
Quantity
0%
25%
50%
75%
ρ
p
Qwen3.6-35B-A3B
Stray </think> /10k
3.6
0.2
68.5
42.0
+0.75
0.020
No fence (%)
7.1
8.4
10.2
25.8
+0.82
0.007
Prefix (%)
61.4
52.3
16.6
43.6
−0.28
0.460
Length-limited (%)
16.0
17.0
25.1
42.0
+0.82
0.007
Median kchar
8.3
8.6
7.3
8.4
+0.43
0.244
Appendix
Table 11: The RCS-LOO budget ladder on LiveCodeBench. Spearman correlations use the nine pruned checkpoints, not only the four shown and not the 0% reference. Units follow Table 10 .
Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing their many expert weight matrices. Expert pruning (e.g., REAP) and merging reduce cost but sacrifice accuracy and require retraining the router; low-rank delta decomposition of experts (e.g., D^2-MoE) preserves all experts and the router, but degrades sharply as the expert count grows because a single shared component cannot approximate many near-orthogonal experts. Because MoE expert weights are near-orthogonal, a single shared component (as in prior delta decomposition) scales poorly with the expert count; we show that experts nonetheless organize into functional co-activation communities that are decoupled from weight similarity. Building on this, we introduce LorExperts, a router-preserving compression method that clusters experts, keeps one full-precision dominant per cluster, and represents the remaining members as low-rank corrections to their local dominant. LorExperts retains all experts and the original router (no router retraining). At ~50% expert compression on Qwen3-30B-A3B and Gemma-4-26B-A4B, LorExperts preserves downstream accuracy and perplexity better than the baselines on most of the tasks; the margin over D^2-MoE grows with expert count E. We further give a reconstruction fine-tuning procedure for LorExperts, and BTExperts, a tree organization of dominants and corrections that enables inference-time amortization of shared computation.
Sparsely activated Mixture-of-Experts (MoE) language models contain substantial structured redundancy among routed experts, but pruning them without downstream calibration data remains challenging. Existing expert-pruning methods typically rely on a single aggregated importance score, which can bias the retained set toward experts favored by dominant calibration patterns. We propose \textbf{Generic TB-Coverage}, a coverage-aware expert pruning method that uses only generic text corpora (WikiText2 and C4) for calibration. Instead of collapsing expert utility into one score, our method profiles per-expert utility separately on each corpus and enforces a fixed-budget coverage rule that preserves high-utility experts from each corpus before constructing the final pruning mask. Across Qwen1.5-MoE-A2.7B and DeepSeek-MoE-16B-Base at 25%, 50%, and 75% retention budgets, our method improves average accuracy on six common zero-shot benchmarks over random pruning, REAP, and ExpertSparsity, while also reducing perplexity degradation on WikiText2 and C4. The gains are largest under aggressive pruning (25% and 50% retain), suggesting that preserving cross-corpus expert coverage is an effective generic-data prior for MoE pruning. Our improvements hold with fixed pruning budgets and no downstream calibration data.
Mixture-of-Experts (MoE) models decouple parameter count from per-token compute, but deployment still requires hosting every expert in memory. Recent theory shows that experts whose router weights change least during fine-tuning can be pruned with provable accuracy preservation, yet the guarantee assumes full fine-tuning. We show that the signal can be elicited through a brief parameter-efficient adaptation. We fine-tune with a lightweight adapter, rank experts by the induced router change, and prune the least-changed experts in one shot. On Mixtral-8×7B-Instruct, router-only LoRA trains 0.002% of parameters and retains 27.54% MMLU-Pro accuracy with half the experts removed, against roughly 16% for magnitude and random pruning. Signal quality improves monotonically with adapter size, reaching 28.76%, and declines as adaptation spreads beyond the router. Under their shared budget, IA3 reaches 28.04% while Houlsby reaches 25.39%. The criterion transfers to Qwen1.5-MoE fine-tuned for mathematical reasoning, retaining 49.7% mean accuracy over eleven benchmarks with half the experts removed. Structural pruning reduces memory by 49% and per-token latency by 37%. Lightweight router sensitivity therefore makes provably motivated, task-conditioned expert pruning practical at scale.
Ali Janati, Kaoutar El Maghraoui, Chengke Zou +2
Data Science Institute, Columbia University, New York, NY, USA · Department of Computer Science, Columbia University, New York, NY, USA