Mixture-of-Experts (MoE) models enable efficient scaling of large language models but face critical deployment challenges due to massive memory requirements. Existing pruning methods either incur prohibitive search costs or neglect the dynamic interdependencies between experts. To address these challenges, we present OMP-MoE, a novel training-free compression framework for reducing expert redundancy in MoE-based LLMs. Based on observations of expert contribution patterns, we reformulate the pruning problem as a sparse signal reconstruction task solved through Orthogonal Matching Pursuit. Specifically, our method first treats individual expert contributions as dictionary atoms and selects experts that greedily minimize reconstruction error with linear computational complexity. Then, we optimize cross-layer expert allocation through a water-filling strategy that accounts for both reconstruction quality and routing stability. Finally, we introduce OMP-MoE†, an adaptive inference mechanism that dynamically adjusts expert activation based on energy prediction. Comprehensive experiments on Qwen, DeepSeek-V2, GPT-OSS, and Mixtral MoE demonstrate consistent improvements over existing methods at 25-50% pruning ratios. For Qwen3-30B-A3B at 50% compression, we retain 93.3% of original performance, achieving 33× faster search and 1.55× inference speedup. Codes will be available after acceptance.
Figures & tables
Figure 1 : Analysis of Dynamic Expert Selection. We visualize the selection process across three representative layers in Mixtral-8 × 7B: ( a ) Decay of normalized residual rate curve as a function of the number of selected experts. ( b ) Dynamic importance of candidate experts sorted by selection order. We provide more details and visualizations across various MoE LLMs in Appendix E.1 and E.2 .
Figure 2 : Overview of our OMP-MoE framework: (a) Greedy pursuit of expert contribution atoms through iterative selection to minimize layer wise reconstruction error. (b) Cross-layer resource optimization via water-filling to align expert counts with marginal benefits. (c) Energy-based dynamic pruning during runtime to achieve streamlined inference without performance degradation.
Method
ARC-c
BoolQ
HellaS.
MMLU
OBQA
WinoG.
Avg.
ARC-c
BoolQ
HellaS.
MMLU
OBQA
WinoG.
Avg.
Qwen3-30B-A3B
DeepSeek-V2-Lite
Original
0.528
0.887
0.596
0.778
0.346
0.703
0.640
0.465
0.799
0.587
0.551
0.348
0.710
0.577
Pruning Ratio 25%
Pruning Ratio 25%
MC-SMoE
0.392
0.777
0.415
0.540
0.318
0.588
0.505
0.367
0.713
0.531
0.422
0.366
0.687
0.514
HC-SMoE
0.458
0.865
0.515
0.669
0.410
0.704
0.603
0.420
0.722
0.560
0.458
0.280
0.695
0.523
NAEE
0.481
0.870
0.555
0.701
0.290
0.693
0.598
0.375
0.669
0.531
0.365
0.290
0.669
0.483
Table 1 : Zero-shot performance comparison across 6 tasks on Qwen, DeepSeek, GPT-OSS and Mixtral MoE models at 25% and 50% expert pruning ratios. Bold indicates the best performance. ‘Avg.’ represents the average accuracy across 6 tasks.
Method
Qwen3-30B-A3B
DeepSeek-V2-Lite
GPT-OSS-20B
Mixtral-8 × 7B
Time (s)
Mem (GB)
Time (s)
Mem (GB)
Time (s)
Mem (GB)
Time (s)
Mem (GB)
NAEE
21301
79.4
8017
42.5
12787
66.3
7168
102.4
DiEP
8972
267.0
2745
98.8
4189
122.3
5957
242.3
OMP-MoE (Ours)
641
75.0
358
40.6
275
56.8
188
97.5
Table 2: Comparison of search efficiency and peak memory usage of NAEE, DiEP and OMP-MoE at a 50% pruning ratio. Measurements are taken on 8 × NVIDIA H20 GPUs.
Prun. Ratio
OMP-MoE
OMP-MoE †
Avg. Acc
Cost ↓
Speedup ↑
0
-
-
0.609
2459s
1.00 ×
25%
✓
-
0.576
2222s
1.11 ×
25%
✓
✓
0.581
2019s
1.22 ×
50%
✓
-
0.489
1825s
1.35 ×
50%
✓
✓
0.476
1690s
1.46 ×
Table 3: Inference performance and efficiency trade offs on DeepSeek-V2-Lite model. We report the average zero-shot accuracy across 8 benchmarks with the total inference time, and the speedup at 25% and 50% pruning ratios. The analysis compares the static pruning of OMP-MoE against the OMP-MoE † algorithm with cross layer allocation enabled.
Table 6
Method
Qwen3-30B-A3B
DeepSeek-V2-Lite
25%
50%
25%
50%
Sub-MoE
0.601
0.574
-
-
HEAPr
0.590
0.480
-
-
HC-SMoE
0.644
0.536
0.552
0.456
DiEP
0.640
0.511
0.566
0.435
OMP-MoE
Figure 3 : Performance of different cross-layer allocation methods on 8 tasks zero-shot average accuracy.
Risk Function ψl(k)
Routing-Risk Weight λ
0
1
2
3
−log(Cl(k)+ζ)
0.618
0.622
0.625
0.643
Figure 4 : Performance of different routing-risk weights on Qwen3-30B-A3B at a 50% pruning ratio.
Figure 5 : Ablations across eight benchmarks. (a) Accuracy and expert skipping under different adaptive thresholds on 25%-pruned DeepSeek-V2-Lite. (b) Accuracy under different calibration sizes and domains on 50%-pruned Qwen3-30B-A3B.
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Search Complexity
Forward Passes
Optimization Type
NAEE
O(L⋅Niter⋅(nNe))
Massive
Combinatorial
DiEP
O(Tft⋅L⋅B⋅d2)
Extensive
Gradient-based
OMP-MoE
O(L⋅n⋅Ne⋅Bd)
Singular
Greedy Pursuit
Appendix
Table 6: Asymptotic complexity comparison between OMP MoE and state of the art pruning methodologies. We denote Niter as the number of iterations required for combinatorial search and Tft as the duration of the fine-tuning stage.
Architecture
Active Experts ( K )
E[m∗∣ϵ]
Normalized FLOPs ↓
Efficiency Gain ( η )
Mixtral-8 × 7B
2
1.22
1.000 → 0.610
39.0%
DeepSeek-V2-Lite
6
3.90
1.000 → 0.650
35.0%
Qwen3-30B-A3B
8
5.14
1.000 → 0.643
35.8%
Appendix
Table 7: Normalized routed-expert FLOPs per token under OMP-MoE † with ϵ=0.1 . The baseline cost for each architecture is normalized to 1.000.
Model
Total Params
Active Params
Layers ( L )
Experts ( Ne )
Active ( K )
Hidden Size ( d )
Expert Intermediate
Dense Intermediate
Qwen/Qwen3-30B-A3B
30.5B
3.3B
48
128
8
2048
768
6144
deepseek-ai/DeepSeek-V2-Lite
15.7B
2.4B
27
64
6
2048
1408
10944
openai/gpt-oss-20b
21.0B
3.6B
24
32
4
2880
2880
–
mistralai/Mixtral-8x7B-v0.1
46.7B
12.9B
32
8
2
4096
14336
–
Appendix
Table 8: Detailed architectural specifications of the evaluated MoE models.
Model
Top-K
Shared Experts
Qwen3-30B-A3B
8
No
DeepSeek-V2-Lite
6
Yes
GPT-OSS-20B
4
No
Mixtral-8 × 7B-v0.1
2
No
Appendix
Table 9: Detailed routing mechanisms and gating configurations.
Dataset
Domain
Processing Method
Block Length
Blocks
C4
Web Corpus
Non-overlapping fixed-length chunking
4096
64
WikiText-2
Wikipedia
Non-overlapping fixed-length chunking
4096
64
Tulu-3 SFT Personas Math
Mathematics
Raw-token fixed-length construction
4096
64
Evol-CodeAlpaca-v1
Code
Raw-token fixed-length construction
4096
64
Appendix
Table 10: Characteristics and processing configurations of the calibration datasets. C4 and WikiText-2 are also used for perplexity evaluation.
Table 11: Benchmark evaluation tasks and their corresponding performance metrics.
Subgroup
Parameter
Value
OMP Search
risk penalty coefficient λ
3.0
Numerical stability ζ
1e-6
Global Allocation
Min expert ratio
0.375
Max expert ratio
0.875
OMP-MoE †
Adaptive threshold epsilon
0.1
Numerical stability ζ
1e-6
Appendix
Table 12: Hyperparameter settings for OMP-MoE pruning and OMP-MoE † inference.
Parameter
Configuration
Calibration Dataset
C4 and WikiText-2
Calibration Samples per Layer
64 samples
Search Batch Size
1
Vectorization
PyTorch torch.matmul
Layer Offloading
Enabled
Appendix
Table 13: Search and statistics collection configurations.
Category
Component
Specification
Hardware
GPU
8 × NVIDIA H20 (96GB VRAM)
CPU
Intel(R) Xeon(R) Platinum 8358 @ 2.60GHz
RAM
1TB DDR4
Software
OS
Ubuntu 22.04.3 LTS
Python
3.10.12
PyTorch
2.9.1+cu124
Appendix
Table 14: Hardware and software environment configurations.
Expert
Method
ARC-c
ARC-e
BoolQ
HellaS.
MMLU
OBQA
RTE
WinoG.
Avg. ↑
Qwen3-30B-A3B
Num=128
Original
0.528
0.792
0.887
0.596
0.778
0.346
0.823
0.703
0.682
Num=96
MC-SMoE
0.392
0.647
0.777
0.415
0.540
0.318
0.823
0.588
0.562
HC-SMoE
0.458
0.760
0.865
0.515
0.669
0.410
0.769
0.704
0.644
NAEE
0.481
0.769
0.870
0.555
0.701
0.290
0.773
0.693
0.642
DiEP
0.505
0.774
0.871
0.567
0.646
0.334
0.718
0.702
0.640
Appendix
Table 15: Comprehensive zero-shot performance comparison on Qwen3-30B-A3B and DeepSeek-V2-Lite model. We report accuracy across eight benchmarks at 25% and 50% expert pruning ratios, corresponding to retained expert counts of Num=96 and Num=64 for Qwen3, with analogous settings for DeepSeek. Bold indicates the best performance among pruning methods.
Expert
Method
ARC-c
ARC-e
BoolQ
HellaS.
MMLU
OBQA
RTE
WinoG.
Avg. ↑
GPT-OSS-20B
Num=32
Original
0.451
0.775
0.757
0.415
0.566
0.270
0.700
0.657
0.574
Num=24
MC-SMoE
0.399
0.739
0.725
0.399
0.488
0.258
0.617
0.665
0.536
HC-SMoE
0.285
0.592
0.622
0.336
0.425
0.182
0.610
0.617
0.459
NAEE
0.422
0.729
0.731
0.400
0.547
0.232
0.664
0.624
0.544
DiEP
0.379
0.712
0.751
0.393
0.527
0.224
0.736
0.593
0.539
Appendix
Table 16: Comprehensive zero-shot performance comparison on GPT-OSS-20B and Mixtral-8 × 7B model. We report accuracy across eight benchmarks at 25% and 50% expert pruning ratios, corresponding to retained expert counts of Num=24 and Num=16 for GPT-OSS, with analogous settings for Mixtral. Bold indicates the best performance among pruning methods.
Pruned
Method
ARC-c
ARC-e
BoolQ
HellaS.
MMLU
OBQA
RTE
WinoG.
Avg.
0%
Original
0.528
0.792
0.887
0.596
0.778
0.346
0.823
0.703
0.682
25%
REAP
0.555
0.797
0.867
0.579
0.733
0.330
0.791
0.690
0.668
EASY-EP
0.511
0.791
0.887
0.593
0.734
0.344
0.809
0.695
0.670
OMP-MoE
0.534
0.803
0.890
0.592
0.747
0.338
0.805
0.699
0.676
50%
REAP
0.455
0.741
0.821
0.464
0.546
0.316
0.737
0.651
0.591
EASY-EP
0.506
0.780
0.865
0.527
0.613
0.320
0.737
0.672
0.627
Appendix
Table 17: Comparison with REAP and EASY-EP on Qwen3-30B-A3B. Bold indicates the best result among the three pruning methods at the same pruning ratio.
Methods
C4 ↓
WikiText-2 ↓
ARC-c
ARC-e
BoolQ
HellaS.
MMLU
OBQA
RTE
WinoG.
Avg. ↑
OMP-MoE 25%
10.618
13.824
0.456
0.774
0.713
0.583
0.494
0.320
0.574
0.697
0.576
+ GPTQ
12.438
14.236
0.404
0.739
0.743
0.405
0.508
0.254
0.621
0.657
0.541
+ SparseGPT (4:8)
13.709
15.673
0.369
0.711
0.720
0.501
0.368
0.302
0.578
0.655
0.525
+ Wanda
13.321
14.430
0.394
0.728
0.720
0.501
0.390
0.308
0.585
0.649
0.534
Appendix
Table 18: Performance comparison of OMP-MoE combined with post-training quantization and weight pruning methods on DeepSeek-V2-Lite at a 25% pruning ratio.
Calibration Dataset
Number of Samples
ARC-c
ARC-e
BoolQ
HellaS.
MMLU
OBQA
RTE
WinoG.
Avg. ↑
C4
16
0.428
0.708
0.867
0.575
0.544
0.312
0.758
0.702
0.612
32
0.460
0.742
0.870
0.572
0.554
0.308
0.791
0.692
0.624
64
0.457
0.737
0.872
0.579
0.574
0.332
0.765
0.703
0.627
128
0.482
0.777
0.874
0.575
0.578
0.320
0.736
0.695
0.630
256
0.497
0.778
0.880
0.575
0.583
0.338
0.675
0.702
0.628
WikiText-2
16
0.517
0.778
0.873
0.548
0.606
0.316
0.762
0.681
0.635
Appendix
Table 19: Impact of calibration sample size and dataset source on Qwen3-30B-A3B performance at a 50% pruning ratio.
Threshold ϵ
Average Experts
Skip Ratio
WikiText-2 ↓
C4 ↓
Average Accuracy ↑
0 (Static)
6
0.00%
13.824
10.618
0.576
0.001
5.97
0.50%
13.180
10.405
0.575
0.01
5.73
4.50%
13.061
10.333
0.576
0.05
4.68
22.00%
13.062
10.341
0.578
0.1 (Default)
3.90
35.00%
13.576
10.522
0.581
0.2
2.90
51.70%
13.580
10.785
0.570
Appendix
Table 20: Impact of the OMP-MoE † threshold ϵ on the inference efficiency and accuracy of DeepSeek-V2-Lite at a 25% pruning ratio.
Pruning Ratio
OMP-MoE
OMP-MoE †
Avg. Acc
Cost ↓
Speedup ↑
0
-
-
0.574
3188s
1.00 ×
0.25
✓
-
0.565
2140s
1.49 ×
0.25
✓
✓
0.562
2129s
1.50 ×
0.50
✓
-
0.515
1705s
1.87 ×
0.50
✓
✓
0.510
1596s
2.00 ×
Appendix
Table 21: Inference performance and efficiency trade offs for the GPT-OSS-20B model.
Pruning Ratio
OMP-MoE
OMP-MoE †
Avg. Acc
Cost ↓
Speedup ↑
0
-
-
0.675
2954s
1.00 ×
0.25
✓
-
0.649
2685s
1.10 ×
0.25
✓
✓
0.644
2517s
1.17 ×
0.50
✓
-
0.607
2382s
1.24 ×
0.50
✓
✓
0.600
2315s
1.28 ×
Appendix
Table 22: Inference performance and efficiency trade offs for the Mixtral-8 × 7B model.
Pruning Ratio
OMP-MoE
OMP-MoE †
Avg. Acc
Cost ↓
Speedup ↑
0
-
-
0.682
9980
1.00 ×
0.25
✓
-
0.676
8911s
1.12 ×
0.25
✓
✓
0.674
8382s
1.19 ×
0.50
✓
-
0.643
6698s
1.49 ×
0.50
✓
✓
0.641
6447s
1.55 ×
Appendix
Table 23: Inference performance and efficiency trade offs for the Qwen3-30B-A3B model.
Pruning Ratio
OMP-MoE
OMP-MoE †
Avg. Acc
Cost ↓
Speedup ↑
0
-
-
0.609
2459s
1.00 ×
0.25
✓
-
0.576
2222s
1.11 ×
0.25
✓
✓
0.581
2019s
1.22 ×
0.50
✓
-
0.489
1825s
1.35 ×
0.50
✓
✓
0.476
1690s
1.46 ×
Appendix
Table 24: Inference performance and efficiency trade offs for the DeepSeek-V2-Lite model.
Retained
Method
ARC-c
ARC-e
BoolQ
HellaS.
MMLU
OBQA
RTE
WinoG.
Avg.
100%
Original
0.528
0.792
0.887
0.596
0.778
0.346
0.823
0.703
0.682
40%
NAEE
0.276
0.457
0.623
0.401
0.239
0.222
0.610
0.617
0.431
40%
DiEP
0.299
0.562
0.775
0.425
0.375
0.224
0.570
0.630
0.482
40%
OMP-MoE
0.372
0.636
0.848
0.553
0.235
0.284
0.736
0.705
0.546
25%
NAEE
0.195
0.283
0.419
0.262
0.233
0.132
0.588
0.507
0.327
25%
DiEP
0.212
0.423
0.622
0.316
0.229
0.142
0.545
0.556
0.381
Appendix
Table 25: Performance under stronger retained-expert constraints on Qwen3-30B-A3B. “Retained” denotes the fraction of experts kept in each MoE layer.
Num
Method
ARC-c
ARC-e
BoolQ
HellaS.
MMLU
OBQA
RTE
WinoG.
Avg.
Time (s)
128
Original
0.618
0.854
0.896
0.678
0.849
0.372
0.812
0.787
0.733
–
64
NAEE
0.486
0.737
0.841
0.591
0.680
0.296
0.747
0.740
0.640
20491
64
DiEP
0.513
0.789
0.854
0.611
0.698
0.344
0.711
0.757
0.660
10473
64
OMP-MoE
0.559
0.822
0.898
0.667
0.724
0.358
0.682
0.774
0.685
846
Appendix
Table 26: Performance on Qwen3-235B-A22B at a 50% pruning ratio. We report eight-task accuracy and search time.
Model
ARC-c
ARC-e
BoolQ
HellaS.
MMLU
OBQA
RTE
WinoG.
Avg.
Qwen3-1.7B dense
0.399
0.725
0.777
0.461
0.524
0.278
0.708
0.615
0.561
Qwen3-4B dense
0.503
0.752
0.850
0.522
0.530
0.296
0.747
0.654
0.607
Qwen3-30B-A3B, 50% pruned
0.533
0.795
0.879
0.546
0.605
0.324
0.765
0.695
0.643
Appendix
Table 27: Comparison between the pruned Qwen3-30B-A3B MoE model and smaller dense Qwen3 models across eight tasks.
Task
Method
Full (0%)
25%
50%
GSM8K
Full model
0.891
–
–
OMP-MoE
–
0.897
0.898
REAP
–
0.895
0.895
EASY-EP
–
0.895
0.888
HumanEval
Full model
0.933
–
–
OMP-MoE
–
0.939
0.896
Appendix
Table 28: Domain-specific calibration on Qwen3-30B-A3B. GSM8K uses 5-shot exact match with mathematical calibration, while HumanEval uses 0-shot pass@1 with code calibration and the humaneval_instruct evaluation protocol. Bold indicates the best pruned result at each pruning ratio.
Pruning Ratio
Reconstruction Error E
Average Accuracy A
0%
0.000
0.682
25%
0.193
0.676
50%
2.633
0.643
60%
5.388
0.546
75%
17.047
0.429
Appendix
Table 29: Reconstruction error and average downstream performance on Qwen3-30B-A3B.
Figure 6 : Layer-wise reconstruction gain heatmaps for the Qwen3-30B-A3B and DeepSeek-V2-Lite models. Each row represents a specific MoE layer and each column represents a selection iteration. Bright cells indicate experts that provide the maximum reconstruction gain at that step.
Figure 7 : Layer-wise reconstruction gain heatmaps for the GPT-OSS-20B and Mixtral-8 × 7B models. Each row represents a specific MoE layer and each column represents a selection iteration. Bright cells indicate experts that provide the maximum reconstruction gain at that step.
Figure 8 : Normalized residual and risk curves for all layers on Qwen3-30B-A3B and DeepSeek-V2-Lite models. The horizontal axis represents the number of retained experts. The vertical axis represents the normalized residual rate and risk rate. Curves with rapid decay indicate layers with high expert redundancy.
Figure 9 : Normalized residual and risk curves for all layers on GPT-OSS-20B and Mixtral-8 × 7B models. The horizontal axis represents the number of retained experts. The vertical axis represents the normalized residual rate and risk rate. Curves with rapid decay indicate layers with high expert redundancy.
Model
Layer Index
Residual rl(1)
Expert ID
DeepSeek-V2-Lite
3
0.0041
54
GPT-OSS-20B
6
0.0407
5
Mixtral-8 × 7B
1
0.0001
3
Qwen3-30B-A3B
2
0.0118
92
Qwen3-30B-A3B
3
0.0322
82
Appendix
Table 30: High-contribution experts across four MoE architectures, identified by the initial OMP reconstruction residual.
Mixture-of-Experts (MoE) language models scale model ability with sparsely activated experts, making this architecture a standard recipe for modern large models. However, sparse activation does not remove the deployment burden of storing and serving all experts, and the available deployment budget can vary substantially across devices, users, and workloads. Existing MoE compression methods are still largely fixed-budget, typically optimizing one compressed endpoint at each chosen target budget. We study a different setting: converting a large pretrained MoE LLM into a nested family of deployable subnetworks across budgets. Our method first ranks expert FFN channels by their importance, then lets each expert learn a discrete action to prune its channels. By gradually increasing cost pressure, a single action-training run exports a series of action masks from high to low budgets, each of which identifies a reliable smaller subnetwork nested in the ranked base model. Moreover, we use a single recovery fine-tune at a mid pruning budget (40%) to recover degraded model quality and transfer the recovered model to other unseen budgets. Overall, our framework surpasses recent MoE compression baselines. Specifically, on Qwen2-57B-A14B, our method retains ~99.8% of base performance while pruning 50% of routed expert parameters even without fine-tuning. For deployment, our pruned subnetworks deliver real memory reduction and throughput gains, and further support realtime online budget switching with kernel-level co-design.
Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck. Expert pruning is the most direct approach for reducing this parameter count, yet existing methods make pruning decisions for each expert independently, and assume experts' contributions are purely additive. In reality, expert usage in MoEs is inherently cooperative. We derive HOPE (Higher-Order Pruning of Experts), a second-order pruning objective which provably minimizes an upper bound on the error resulting from pruning. We show that REAP (a state-of-the-art first-order pruning method) is a special case of HOPE where interaction terms are ignored. Across three frontier MoE models (up to 122B parameters), two distinct calibration sets, and multiple benchmarks (including math, instruction following, coding, and an agentic suite), we demonstrate that HOPE produces better pruning decisions than existing methods, and its advantage is most pronounced at high pruning rates and on challenging agentic workloads. At 50% pruning, HOPE outperforms all baselines and achieves an average rank of 1.58 out of 5 methods (versus 2.42 for the next-best method, REAP), with gains of up to +6.1% on agentic coding. Over all conditions, HOPE again achieves the best average rank and surpasses every other method in the majority of head-to-head comparisons. By preserving cooperative expert structure that first-order methods ignore, HOPE enables aggressive compression with minimal degradation, particularly on complex tasks where diverse expert combinations are invoked over long sequences.
Mixture-of-Experts (MoE) models have shown strong potential in scaling language models efficiently by activating only a small subset of experts per input. However, their widespread deployment remains limited due to the high memory overhead associated with storing all expert parameters, particularly as the number of experts increases. To address this challenge, prior works have explored expert dropping and merging strategies, yet they often suffer from performance drop at high compression ratios. In this paper, we introduce PuzzleMoE, a training-free MoE compression method that achieves both high accuracy and efficient inference through two key innovations: First, PuzzleMoE performs sparse expert merging by identifying element-wise weight redundancy and specialization. It uses a dual-mask to capture both shared and expert-specific parameters. Second, to avoid the overhead of storing binary masks and signs, PuzzleMoE introduces a bit-packed encoding scheme that reuses underutilized exponent bits, enabling efficient MoE inference on GPUs. Extensive experiments demonstrate that PuzzleMoE can compress MoE models by up to 50% while maintaining accuracy across various tasks. Specifically, it outperforms prior MoE compression methods by up to 16.7% on MMLU at 50% compression ratio, and achieves up to 1.28\times inference speedup.
Yushu Zhao, Zheng Wang, Minjia Zhang
1Tsinghua University · †Work done while intern at UIUC. · University of Illinois Urbana-Champaign.