OMP-MoE: Efficient Expert Pruning for Mixture-of-Experts LLMs via Orthogonal Matching Pursuit
Organizations: The Hong Kong University of Science and Technology
Abstract
Mixture-of-Experts (MoE) models enable efficient scaling of large language models but face critical deployment challenges due to massive memory requirements. Existing pruning methods either incur prohibitive search costs or neglect the dynamic interdependencies between experts. To address these challenges, we present OMP-MoE, a novel training-free compression framework for reducing expert redundancy in MoE-based LLMs. Based on observations of expert contribution patterns, we reformulate the pruning problem as a sparse signal reconstruction task solved through Orthogonal Matching Pursuit. Specifically, our method first treats individual expert contributions as dictionary atoms and selects experts that greedily minimize reconstruction error with linear computational complexity. Then, we optimize cross-layer expert allocation through a water-filling strategy that accounts for both reconstruction quality and routing stability. Finally, we introduce OMP-MoE†, an adaptive inference mechanism that dynamically adjusts expert activation based on energy prediction. Comprehensive experiments on Qwen, DeepSeek-V2, GPT-OSS, and Mixtral MoE demonstrate consistent improvements over existing methods at 25-50% pruning ratios. For Qwen3-30B-A3B at 50% compression, we retain 93.3% of original performance, achieving 33 faster search and 1.55 inference speedup. Codes will be available after acceptance.
Figures & tables
| Method | ARC-c | BoolQ | HellaS. | MMLU | OBQA | WinoG. | Avg. | ARC-c | BoolQ | HellaS. | MMLU | OBQA | WinoG. | Avg. |
| Qwen3-30B-A3B | DeepSeek-V2-Lite | |||||||||||||
| Original | 0.528 | 0.887 | 0.596 | 0.778 | 0.346 | 0.703 | 0.640 | 0.465 | 0.799 | 0.587 | 0.551 | 0.348 | 0.710 | 0.577 |
| Pruning Ratio 25% | Pruning Ratio 25% | |||||||||||||
| MC-SMoE | 0.392 | 0.777 | 0.415 | 0.540 | 0.318 | 0.588 | 0.505 | 0.367 | 0.713 | 0.531 | 0.422 | 0.366 | 0.687 | 0.514 |
| HC-SMoE | 0.458 | 0.865 | 0.515 | 0.669 | 0.410 | 0.704 | 0.603 | 0.420 | 0.722 | 0.560 | 0.458 | 0.280 | 0.695 | 0.523 |
| NAEE | 0.481 | 0.870 | 0.555 | 0.701 | 0.290 | 0.693 | 0.598 | 0.375 | 0.669 | 0.531 | 0.365 | 0.290 | 0.669 | 0.483 |
| Method | Qwen3-30B-A3B | DeepSeek-V2-Lite | GPT-OSS-20B | Mixtral-8 7B | ||||
| Time (s) | Mem (GB) | Time (s) | Mem (GB) | Time (s) | Mem (GB) | Time (s) | Mem (GB) | |
| NAEE | 21301 | 79.4 | 8017 | 42.5 | 12787 | 66.3 | 7168 | 102.4 |
| DiEP | 8972 | 267.0 | 2745 | 98.8 | 4189 | 122.3 | 5957 | 242.3 |
| OMP-MoE (Ours) | 641 | 75.0 | 358 | 40.6 | 275 | 56.8 | 188 | 97.5 |
| Prun. Ratio | OMP-MoE | OMP-MoE † | Avg. Acc | Cost | Speedup |
| 0 | - | - | 0.609 | 2459s | 1.00 |
| 25% | ✓ | - | 0.576 | 2222s | 1.11 |
| 25% | ✓ | ✓ | 0.581 | 2019s | 1.22 |
| 50% | ✓ | - | 0.489 | 1825s | 1.35 |
| 50% | ✓ | ✓ | 0.476 | 1690s | 1.46 |
| Method | Qwen3-30B-A3B | DeepSeek-V2-Lite | ||
| 25% | 50% | 25% | 50% | |
| Sub-MoE | 0.601 | 0.574 | - | - |
| HEAPr | 0.590 | 0.480 | - | - |
| HC-SMoE | 0.644 | 0.536 | 0.552 | 0.456 |
| DiEP | 0.640 | 0.511 | 0.566 | 0.435 |
| OMP-MoE | ||||
| Risk Function | Routing-Risk Weight | |||
| 0 | 1 | 2 | 3 | |
| 0.618 | 0.622 | 0.625 | 0.643 | |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Search Complexity | Forward Passes | Optimization Type |
| NAEE | Massive | Combinatorial | |
| DiEP | Extensive | Gradient-based | |
| OMP-MoE | Singular | Greedy Pursuit |
| Architecture | Active Experts ( ) | Normalized FLOPs | Efficiency Gain ( ) | |
| Mixtral-8 7B | 2 | 1.22 | 1.000 0.610 | 39.0% |
| DeepSeek-V2-Lite | 6 | 3.90 | 1.000 0.650 | 35.0% |
| Qwen3-30B-A3B | 8 | 5.14 | 1.000 0.643 | 35.8% |
| Model | Total Params | Active Params | Layers ( ) | Experts ( ) | Active ( ) | Hidden Size ( ) | Expert Intermediate | Dense Intermediate |
| Qwen/Qwen3-30B-A3B | 30.5B | 3.3B | 48 | 128 | 8 | 2048 | 768 | 6144 |
| deepseek-ai/DeepSeek-V2-Lite | 15.7B | 2.4B | 27 | 64 | 6 | 2048 | 1408 | 10944 |
| openai/gpt-oss-20b | 21.0B | 3.6B | 24 | 32 | 4 | 2880 | 2880 | – |
| mistralai/Mixtral-8x7B-v0.1 | 46.7B | 12.9B | 32 | 8 | 2 | 4096 | 14336 | – |
| Model | Top-K | Shared Experts |
| Qwen3-30B-A3B | 8 | No |
| DeepSeek-V2-Lite | 6 | Yes |
| GPT-OSS-20B | 4 | No |
| Mixtral-8 7B-v0.1 | 2 | No |
| Dataset | Domain | Processing Method | Block Length | Blocks |
| C4 | Web Corpus | Non-overlapping fixed-length chunking | 4096 | 64 |
| WikiText-2 | Wikipedia | Non-overlapping fixed-length chunking | 4096 | 64 |
| Tulu-3 SFT Personas Math | Mathematics | Raw-token fixed-length construction | 4096 | 64 |
| Evol-CodeAlpaca-v1 | Code | Raw-token fixed-length construction | 4096 | 64 |
| Category | Task | Evaluation | Primary Metric |
| Perplexity | WikiText-2, C4 | 0-shot | Perplexity |
| 8 General Tasks | ARC-Challenge, ARC-Easy, BoolQ, HellaSwag, MMLU, OpenBookQA, RTE, WinoGrande | 0-shot | Accuracy |
| Compositional Reasoning | BBH | 0-shot | Exact Match |
| Mathematics | GSM8K | 5-shot | Exact Match |
| Long-context | LongBench | 0-shot | Score |
| Code Generation | HumanEval | 0-shot | pass@1 |
| Subgroup | Parameter | Value |
| OMP Search | risk penalty coefficient | 3.0 |
| Numerical stability | 1e-6 | |
| Global Allocation | Min expert ratio | 0.375 |
| Max expert ratio | 0.875 | |
| OMP-MoE † | Adaptive threshold epsilon | 0.1 |
| Numerical stability | 1e-6 |
| Parameter | Configuration |
| Calibration Dataset | C4 and WikiText-2 |
| Calibration Samples per Layer | 64 samples |
| Search Batch Size | 1 |
| Vectorization | PyTorch torch.matmul |
| Layer Offloading | Enabled |
| Category | Component | Specification |
| Hardware | GPU | 8 NVIDIA H20 (96GB VRAM) |
| CPU | Intel(R) Xeon(R) Platinum 8358 @ 2.60GHz | |
| RAM | 1TB DDR4 | |
| Software | OS | Ubuntu 22.04.3 LTS |
| Python | 3.10.12 | |
| PyTorch | 2.9.1+cu124 |
| Expert | Method | ARC-c | ARC-e | BoolQ | HellaS. | MMLU | OBQA | RTE | WinoG. | Avg. |
| Qwen3-30B-A3B | ||||||||||
| Num=128 | Original | 0.528 | 0.792 | 0.887 | 0.596 | 0.778 | 0.346 | 0.823 | 0.703 | 0.682 |
| Num=96 | MC-SMoE | 0.392 | 0.647 | 0.777 | 0.415 | 0.540 | 0.318 | 0.823 | 0.588 | 0.562 |
| HC-SMoE | 0.458 | 0.760 | 0.865 | 0.515 | 0.669 | 0.410 | 0.769 | 0.704 | 0.644 | |
| NAEE | 0.481 | 0.769 | 0.870 | 0.555 | 0.701 | 0.290 | 0.773 | 0.693 | 0.642 | |
| DiEP | 0.505 | 0.774 | 0.871 | 0.567 | 0.646 | 0.334 | 0.718 | 0.702 | 0.640 | |
| Expert | Method | ARC-c | ARC-e | BoolQ | HellaS. | MMLU | OBQA | RTE | WinoG. | Avg. |
| GPT-OSS-20B | ||||||||||
| Num=32 | Original | 0.451 | 0.775 | 0.757 | 0.415 | 0.566 | 0.270 | 0.700 | 0.657 | 0.574 |
| Num=24 | MC-SMoE | 0.399 | 0.739 | 0.725 | 0.399 | 0.488 | 0.258 | 0.617 | 0.665 | 0.536 |
| HC-SMoE | 0.285 | 0.592 | 0.622 | 0.336 | 0.425 | 0.182 | 0.610 | 0.617 | 0.459 | |
| NAEE | 0.422 | 0.729 | 0.731 | 0.400 | 0.547 | 0.232 | 0.664 | 0.624 | 0.544 | |
| DiEP | 0.379 | 0.712 | 0.751 | 0.393 | 0.527 | 0.224 | 0.736 | 0.593 | 0.539 | |
| Pruned | Method | ARC-c | ARC-e | BoolQ | HellaS. | MMLU | OBQA | RTE | WinoG. | Avg. |
| 0% | Original | 0.528 | 0.792 | 0.887 | 0.596 | 0.778 | 0.346 | 0.823 | 0.703 | 0.682 |
| 25% | REAP | 0.555 | 0.797 | 0.867 | 0.579 | 0.733 | 0.330 | 0.791 | 0.690 | 0.668 |
| EASY-EP | 0.511 | 0.791 | 0.887 | 0.593 | 0.734 | 0.344 | 0.809 | 0.695 | 0.670 | |
| OMP-MoE | 0.534 | 0.803 | 0.890 | 0.592 | 0.747 | 0.338 | 0.805 | 0.699 | 0.676 | |
| 50% | REAP | 0.455 | 0.741 | 0.821 | 0.464 | 0.546 | 0.316 | 0.737 | 0.651 | 0.591 |
| EASY-EP | 0.506 | 0.780 | 0.865 | 0.527 | 0.613 | 0.320 | 0.737 | 0.672 | 0.627 |
| Methods | C4 | WikiText-2 | ARC-c | ARC-e | BoolQ | HellaS. | MMLU | OBQA | RTE | WinoG. | Avg. |
| OMP-MoE 25% | 10.618 | 13.824 | 0.456 | 0.774 | 0.713 | 0.583 | 0.494 | 0.320 | 0.574 | 0.697 | 0.576 |
| + GPTQ | 12.438 | 14.236 | 0.404 | 0.739 | 0.743 | 0.405 | 0.508 | 0.254 | 0.621 | 0.657 | 0.541 |
| + SparseGPT (4:8) | 13.709 | 15.673 | 0.369 | 0.711 | 0.720 | 0.501 | 0.368 | 0.302 | 0.578 | 0.655 | 0.525 |
| + Wanda | 13.321 | 14.430 | 0.394 | 0.728 | 0.720 | 0.501 | 0.390 | 0.308 | 0.585 | 0.649 | 0.534 |
| Calibration Dataset | Number of Samples | ARC-c | ARC-e | BoolQ | HellaS. | MMLU | OBQA | RTE | WinoG. | Avg. |
| C4 | 16 | 0.428 | 0.708 | 0.867 | 0.575 | 0.544 | 0.312 | 0.758 | 0.702 | 0.612 |
| 32 | 0.460 | 0.742 | 0.870 | 0.572 | 0.554 | 0.308 | 0.791 | 0.692 | 0.624 | |
| 64 | 0.457 | 0.737 | 0.872 | 0.579 | 0.574 | 0.332 | 0.765 | 0.703 | 0.627 | |
| 128 | 0.482 | 0.777 | 0.874 | 0.575 | 0.578 | 0.320 | 0.736 | 0.695 | 0.630 | |
| 256 | 0.497 | 0.778 | 0.880 | 0.575 | 0.583 | 0.338 | 0.675 | 0.702 | 0.628 | |
| WikiText-2 | 16 | 0.517 | 0.778 | 0.873 | 0.548 | 0.606 | 0.316 | 0.762 | 0.681 | 0.635 |
| Threshold | Average Experts | Skip Ratio | WikiText-2 | C4 | Average Accuracy |
| 0 (Static) | 6 | 0.00% | 13.824 | 10.618 | 0.576 |
| 0.001 | 5.97 | 0.50% | 13.180 | 10.405 | 0.575 |
| 0.01 | 5.73 | 4.50% | 13.061 | 10.333 | 0.576 |
| 0.05 | 4.68 | 22.00% | 13.062 | 10.341 | 0.578 |
| 0.1 (Default) | 3.90 | 35.00% | 13.576 | 10.522 | 0.581 |
| 0.2 | 2.90 | 51.70% | 13.580 | 10.785 | 0.570 |
| Pruning Ratio | OMP-MoE | OMP-MoE † | Avg. Acc | Cost | Speedup |
| 0 | - | - | 0.574 | 3188s | 1.00 |
| 0.25 | ✓ | - | 0.565 | 2140s | 1.49 |
| 0.25 | ✓ | ✓ | 0.562 | 2129s | 1.50 |
| 0.50 | ✓ | - | 0.515 | 1705s | 1.87 |
| 0.50 | ✓ | ✓ | 0.510 | 1596s | 2.00 |
| Pruning Ratio | OMP-MoE | OMP-MoE † | Avg. Acc | Cost | Speedup |
| 0 | - | - | 0.675 | 2954s | 1.00 |
| 0.25 | ✓ | - | 0.649 | 2685s | 1.10 |
| 0.25 | ✓ | ✓ | 0.644 | 2517s | 1.17 |
| 0.50 | ✓ | - | 0.607 | 2382s | 1.24 |
| 0.50 | ✓ | ✓ | 0.600 | 2315s | 1.28 |
| Pruning Ratio | OMP-MoE | OMP-MoE † | Avg. Acc | Cost | Speedup |
| 0 | - | - | 0.682 | 9980 | 1.00 |
| 0.25 | ✓ | - | 0.676 | 8911s | 1.12 |
| 0.25 | ✓ | ✓ | 0.674 | 8382s | 1.19 |
| 0.50 | ✓ | - | 0.643 | 6698s | 1.49 |
| 0.50 | ✓ | ✓ | 0.641 | 6447s | 1.55 |
| Pruning Ratio | OMP-MoE | OMP-MoE † | Avg. Acc | Cost | Speedup |
| 0 | - | - | 0.609 | 2459s | 1.00 |
| 0.25 | ✓ | - | 0.576 | 2222s | 1.11 |
| 0.25 | ✓ | ✓ | 0.581 | 2019s | 1.22 |
| 0.50 | ✓ | - | 0.489 | 1825s | 1.35 |
| 0.50 | ✓ | ✓ | 0.476 | 1690s | 1.46 |
| Retained | Method | ARC-c | ARC-e | BoolQ | HellaS. | MMLU | OBQA | RTE | WinoG. | Avg. |
| 100% | Original | 0.528 | 0.792 | 0.887 | 0.596 | 0.778 | 0.346 | 0.823 | 0.703 | 0.682 |
| 40% | NAEE | 0.276 | 0.457 | 0.623 | 0.401 | 0.239 | 0.222 | 0.610 | 0.617 | 0.431 |
| 40% | DiEP | 0.299 | 0.562 | 0.775 | 0.425 | 0.375 | 0.224 | 0.570 | 0.630 | 0.482 |
| 40% | OMP-MoE | 0.372 | 0.636 | 0.848 | 0.553 | 0.235 | 0.284 | 0.736 | 0.705 | 0.546 |
| 25% | NAEE | 0.195 | 0.283 | 0.419 | 0.262 | 0.233 | 0.132 | 0.588 | 0.507 | 0.327 |
| 25% | DiEP | 0.212 | 0.423 | 0.622 | 0.316 | 0.229 | 0.142 | 0.545 | 0.556 | 0.381 |
| Num | Method | ARC-c | ARC-e | BoolQ | HellaS. | MMLU | OBQA | RTE | WinoG. | Avg. | Time (s) |
| 128 | Original | 0.618 | 0.854 | 0.896 | 0.678 | 0.849 | 0.372 | 0.812 | 0.787 | 0.733 | – |
| 64 | NAEE | 0.486 | 0.737 | 0.841 | 0.591 | 0.680 | 0.296 | 0.747 | 0.740 | 0.640 | 20491 |
| 64 | DiEP | 0.513 | 0.789 | 0.854 | 0.611 | 0.698 | 0.344 | 0.711 | 0.757 | 0.660 | 10473 |
| 64 | OMP-MoE | 0.559 | 0.822 | 0.898 | 0.667 | 0.724 | 0.358 | 0.682 | 0.774 | 0.685 | 846 |
| Model | ARC-c | ARC-e | BoolQ | HellaS. | MMLU | OBQA | RTE | WinoG. | Avg. |
| Qwen3-1.7B dense | 0.399 | 0.725 | 0.777 | 0.461 | 0.524 | 0.278 | 0.708 | 0.615 | 0.561 |
| Qwen3-4B dense | 0.503 | 0.752 | 0.850 | 0.522 | 0.530 | 0.296 | 0.747 | 0.654 | 0.607 |
| Qwen3-30B-A3B, 50% pruned | 0.533 | 0.795 | 0.879 | 0.546 | 0.605 | 0.324 | 0.765 | 0.695 | 0.643 |
| Task | Method | Full (0%) | 25% | 50% |
| GSM8K | Full model | 0.891 | – | – |
| OMP-MoE | – | 0.897 | 0.898 | |
| REAP | – | 0.895 | 0.895 | |
| EASY-EP | – | 0.895 | 0.888 | |
| HumanEval | Full model | 0.933 | – | – |
| OMP-MoE | – | 0.939 | 0.896 |
| Pruning Ratio | Reconstruction Error | Average Accuracy |
| 0% | 0.000 | 0.682 |
| 25% | 0.193 | 0.676 |
| 50% | 2.633 | 0.643 |
| 60% | 5.388 | 0.546 |
| 75% | 17.047 | 0.429 |
| Model | Layer Index | Residual | Expert ID |
| DeepSeek-V2-Lite | 3 | 0.0041 | 54 |
| GPT-OSS-20B | 6 | 0.0407 | 5 |
| Mixtral-8 7B | 1 | 0.0001 | 3 |
| Qwen3-30B-A3B | 2 | 0.0118 | 92 |
| Qwen3-30B-A3B | 3 | 0.0322 | 82 |