Organizations: Department of Electrical and Computer Engineering, Northeastern University · Khoury College of Computer Science, Northeastern University
Sparsely activated Mixture-of-Experts (MoE) models increase model capacity without a proportional increase in per-token computation. Dense-to-MoE upcycling reuses pretrained dense models to construct such systems, commonly by copying feed-forward networks (FFNs) into independently trained experts. We introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN. Each expert is defined by a learned binary mask, and a token-level router selects which masked FFNs to execute and combine. The router and mask scores are optimized jointly, while the underlying FFN weight values remain unchanged. This formulation supports neuron-structured, semi-structured, and unstructured experts within the same routing architecture. Our main configuration uses four 2:4 experts with top-2 routing, where two half-dense expert passes have the nominal FFN arithmetic of one dense pass, without requiring independent expert weight matrices. On five vision-language benchmarks with Qwen and Gemma backbones, this configuration achieves the highest performance among the compared baselines. Comparisons across mask granularities, routing interventions, and compute-matched controls distinguish the effects of learned connectivity from expert activation count. These results establish mask learning over frozen weights as a practical alternative for constructing token-routed MoE experts.
Figures & tables
Figure 1: Conventional MoE and MASKerade. Tokens select independent FFNs (left) or learned masks over one frozen FFN (right).
Method
Act. Params.
GQA ↑
MME ↑
POPE ↑
SQA ↑
TextVQA ↑
Qwen3-1.7B
Zero-shot ( Yang et al., 2025 )
1.72B
25.97
962.0
83.59
60.93
32.38
Dense MLP
1.72B
58.41 ± 0.35
1300.4 ± 18.4
85.01 ± 0.24
61.16 ± 0.32
49.19 ± 0.32
MoEfication ( Zhang et al., 2022 )
1.46B
53.60 ± 0.20
1280.5 ± 15.8
84.75 ± 0.33
60.11 ± 0.96
49.94 ± 0.55
LLaVA-MoLE ( Chen et al., 2024a )
1.91B
58.05 ± 0.38
1295.0 ± 18.3
84.97 ± 0.36
60.24 ± 0.09
50.25 ± 0.06
PESC ( Wu et al., 2024 )
1.73B
56.81 ± 0.36
1309.1 ± 11.3
84.31 ± 0.29
61.60 ± 0.57
49.95 ± 0.25
LLaMA-MoE ( Zhu et al., 2024 )
1.46B
57.85 ± 0.16
1233.3 ± 17.5
85.45 ± 0.39
60.56 ± 0.58
50.90 ± 0.58
Table 1: Main comparison with existing methods on small models. All methods start from the same Zero-shot checkpoint. We learn masks as experts over frozen FFN weights. Best and second-best results in each metric are shown in bold and underlined , respectively.
Figure 2: Initialization loss for 2:4 masks on Qwen3-1.7B.
Method
Act. Params.
GQA ↑
MME ↑
POPE ↑
SQA ↑
TextVQA ↑
Qwen3.6-27B
Zero-shot ( Yang et al., 2025 )
26.90B
61.39
1692.1
86.23
92.74
71.65
Dense MLP
26.90B
64.13
1716.0
86.90
93.38
73.41
MoEfication ( Zhang et al., 2022 )
22.62B
61.05
1691.2
86.15
91.55
72.00
LLaVA-MoLE ( Chen et al., 2024a )
29.82B
65.71
1702.5
86.47
92.17
74.08
PESC ( Wu et al., 2024 )
26.94B
65.61
1717.4
87.08
91.18
77.33
LLaMA-MoE ( Zhu et al., 2024 )
22.62B
64.33
1715.5
86.78
93.46
78.63
Table 2: Main comparison with existing methods on large models. All methods start from the same Zero-shot checkpoint. We learn masks as experts over frozen FFN weights. Best and second-best results in each metric are shown in bold and underlined , respectively.
Method
GQA
MME
POPE
SQA
TextVQA
Qwen3-1.7B
Zero-shot
25.97
962.0
83.59
60.93
32.38
Dense MLP
58.41 ± 0.35
1300.4 ± 18.4
85.01 ± 0.24
61.16 ± 0.32
49.19 ± 0.32
Random 2:4
43.36 ± 0.30
871.8 ± 21.3
82.15 ± 0.16
40.75 ± 0.42
28.22 ± 0.48
Magnitude 2:4
51.54 ± 0.26
1127.6 ± 14.7
85.38 ± 0.53
58.34 ± 0.80
43.15 ± 0.71
Wanda 2:4 ( Sun et al., 2024 )
51.97 ± 0.53
1148.2 ± 23.8
83.96 ± 0.44
60.81 ± 0.64
44.01 ± 0.27
Wanda Unstructured ( Sun et al., 2024 )
52.24 ± 0.48
1187.4 ± 37.4
86.06 ± 0.27
63.01 ± 0.52
45.21 ± 0.22
Table 3: Mask-based MoE comparison. Best and second-best results in each metric are shown in bold and underlined , respectively.
Setting
GQA ↑
MME ↑
POPE ↑
SQA ↑
TextVQA ↑
Qwen3-1.7B
Trained
62.00 ± 0.28
1332.4 ± 11.9
86.49 ± 0.16
63.48 ± 0.57
52.50 ± 0.37
Random (frozen during training)
60.02 ± 0.07
1307.7 ± 21.0
85.96 ± 0.45
61.27 ± 1.05
50.97 ± 0.28
Random (after training)
59.30 ± 0.39
1297.8 ± 12.6
84.60 ± 0.11
60.32 ± 0.66
49.71 ± 0.58
Label-permuted
56.86 ± 0.52
1270.0 ± 12.2
83.82 ± 0.66
60.82 ± 0.84
49.63 ± 0.61
Gemma3-1B
Table 4: Router learning and inference-time interventions with four 2:4 experts and top-2 routing. Frozen-random routers remain fixed during mask learning.
Table 7
Configuration
Masks
Calculation
Act. params
GQA ↑
MME ↑
POPE ↑
SQA ↑
TextVQA ↑
MoEfication ( Zhang et al., 2022 )
-
-
1.46B
53.60 ± 0.20
1280.5 ± 15.8
84.75 ± 0.33
60.11 ± 0.96
49.94 ± 0.55
LLaMA-MoE ( Zhu et al., 2024 )
-
-
1.46B
57.85 ± 0.16
1233.3 ± 17.5
85.45 ± 0.39
60.56 ± 0.58
50.90 ± 0.58
DIVE ( Feng et al., 2025 )
-
-
1.46B
57.87 ± 0.16
1300.6 ± 16.7
83.80 ± 0.34
60.84 ± 0.76
50.55 ± 0.62
ToMoE ( Gao et al., 2026 )
-
-
1.46B
58.54 ± 0.22
1302.2 ± 16.2
85.66 ± 0.40
61.61 ± 0.50
50.77 ± 0.22
1 mask only
2:4
50%×1
1.46B
56.51 ± 0.32
1258.1 ± 15.5
84.50 ± 0.25
59.67 ± 0.56
49.13 ± 0.68
[Ours] 4e/top-2
1:4
25%×2
1.46B
61.40 ± 0.32
1322.2 ± 11.8
85.96 ± 0.46
61.79 ± 0.53
50.94 ± 0.34
Table 7: Matched nominal activation budgets on Qwen3-1.7B. Calculation denotes retained density per expert multiplied by the number of activated experts.
Figure 3: Learned mask structure on Qwen3-1.7B: (a) 2:4 mask Jaccard overlap; (b)-(d) layer-wise sharing of retained connections.
Figure 4: Layer-local expert selection by training source on Qwen3-1.7B. Panels show layers 0, 14, 26, and an equal index-wise mean over all 14 converted layers. Expert IDs are not aligned across layers. 50% indicates balanced top-2 usage.
Figure 5: Random versus importance initialization across mask granularities on Qwen3-1.7B.
Figure 6: Random versus importance initialization across mask granularities on Gemma3-1B.
Table 14
Configuration
Masks
Calculation
Act. params
GQA ↑
MME ↑
POPE ↑
SQA ↑
TextVQA ↑
MoEfication ( Zhang et al., 2022 )
-
-
0.84B
55.36 ± 0.28
1212.9 ± 13.5
80.92 ± 0.35
50.33 ± 0.55
40.68 ± 0.58
LLaMA-MoE ( Zhu et al., 2024 )
-
-
0.84B
57.67 ± 0.33
1234.7 ± 11.2
85.00 ± 0.42
52.71 ± 0.63
40.96 ± 0.67
DIVE ( Feng et al., 2025 )
-
-
0.85B
51.66 ± 0.27
1129.3 ± 13.2
82.80 ± 0.28
50.84 ± 0.63
41.79 ± 0.58
ToMoE ( Gao et al., 2026 )
-
-
0.84B
53.96 ± 0.15
1123.0 ± 17.9
81.37 ± 0.43
51.98 ± 0.52
40.63 ± 0.32
1 mask only
2:4
50%×1
0.84B
58.30 ± 0.30
1239.9 ± 13.7
84.87 ± 0.38
52.59 ± 0.69
40.07 ± 0.68
[Ours] 4e/top-2
1:4
25%×2
0.84B
59.09 ± 0.33
1248.7 ± 15.5
85.11 ± 0.36
52.99 ± 0.61
42.47 ± 0.64
Appendix
Table 12: Matched nominal activation budgets on Gemma3-1B. Calculation denotes retained density per expert times the number of activated experts. Only protocol-matched runs are included.
Figure 7: Learned mask structure on Gemma3-1B, averaged over three runs per family: (a) 2:4 Jaccard overlap; (b)-(d) layer-wise sharing of retained connections, normalized by dense weight positions.
Figure 8: Layer-local expert selection by training source on Gemma3-1B. Panels show layers 0, 12, 24, and the unaligned index-wise mean over all 13 converted layers. 50% denotes balanced top-2 usage.
Table 18Table 19
Layer
Token type
Cosine
Distance
e0
e1
e2
e3
Qwen3-1.7B
0
Vision
0.700
0.390
13.60
13.70
13.50
13.80
0
Prompt
0.780
0.330
9.20
9.30
9.10
9.40
0
Response
0.720
0.370
11.60
11.40
11.80
12.00
14
Vision
0.650
0.420
52.00
53.00
51.00
54.00
14
Prompt
0.730
0.370
44.00
45.00
46.00
44.50
Appendix
Table 19: Identical-input expert outputs by token type. Cosine and Distance are pairwise averages. Distance normalizes L2 differences by summed output norms, and ei denotes the mean output norm.
Layer
Slot
Alt.
Mean Δ NLL
Median
Δ NLL >0 (%)
Δ Acc
C → W
W → C
Qwen3-1.7B
0
0
0
+0.025
+0.008
62.0
-1.00
1.50
0.50
0
0
1
+0.040
+0.012
67.0
-1.50
2.00
0.50
0
1
0
+0.015
+0.004
60.0
-0.50
1.00
0.50
0
1
1
+0.020
+0.006
63.0
-1.00
1.50
0.50
14
0
0
+0.080
+0.035
73.0
-3.00
4.00
1.00
Appendix
Table 20: Single-layer expert replacement with routing weights fixed. Slot and Alt. give zero-based router ranks among selected and unselected experts. Δ compares with original routing. C/W denote correct/incorrect answers. Accuracy changes and flips are in percentage points and percent, respectively.
Mixture-of-experts (MoE) language models often exceed the memory of a single GPU. Expert offloading keeps most experts in host memory and loads them on demand, so decoding speed depends on how many experts each token must fetch. Caching and prefetching reduce this cost only as far as the routing allows. Router-only fine-tuning can reshape the routing to reuse experts, but it keeps the experts frozen, so they cannot adapt to the tokens the new routing sends them. We propose MaskCoFT, a masked co-adaptive fine-tuning method that trains routers and experts together with the cross-entropy loss alone. During fine-tuning, a learnable binary mask restricts the Top-K routing of each layer to a subset of experts, and the experts adapt to the tokens redirected to them. At inference, the learned mask becomes a soft prior that re-ranks experts, so every expert remains selectable. We simulate a GPU cache of 4 experts per layer for Mixtral-8x7B and 12 for DeepSeek-V2-Lite. MaskCoFT cuts expert fetches per token by 23.7% and 10.1% relative to the base model. In real offloading system serving, it lowers the time per output token by up to 16.4% and 5.5%, respectively. Its average accuracy over nine benchmarks stays above the base model by 0.92 and 0.53 points.
Junfeng Wu, Zehao Fan, Hadjer Benmeziane +3
Department of Industrial and Systems Engineering, Rensselaer Polytechnic Institute · Department of Electrical, Computer, and Systems Engineering, Rensselaer Polytechnic Institute · IBM Research
Mixture-of-Experts (MoE) architectures enhance the efficiency of large language models by activating only a subset of experts per token. However, standard MoE employs a fixed Top-K routing strategy, leading to redundant computation and suboptimal inference latency. Existing acceleration methods either require costly retraining with architectural changes or suffer from severe performance drop at high sparsity due to train-inference mismatch. To address these limitations, we propose BEAM (Binary Expert Activation Masking), a novel method that learns token-adaptive expert selection via trainable binary masks. With a straight-through estimator and an auxiliary regularization loss, BEAM induces dynamic expert sparsity through end-to-end training while maintaining model capability. We further implement an efficient custom CUDA kernel for BEAM, ensuring seamless integration with the vLLM inference framework. Experiments show that BEAM retains over 98% of the original model's performance while reducing MoE layer FLOPs by up to 85%, achieving up to 2.5× faster decoding and 1.4× higher throughput, demonstrating its effectiveness as a practical, plug-and-play solution for efficient MoE inference.
Juntong Wu, Jialiang Cheng, Qishen Yin +5
Taobao & Tmall Group of Alibaba · Shenzhen Graduate School, Peking University
The scaling of Large Language Models (LLMs) has driven significant performance gains but created substantial challenges in inference efficiency. While Mixture of Experts (MoEs) architectures address this by decoupling model size from inference cost, training MoEs from scratch is often unstable and compute intensive. Conversion of pre-trained dense models into sparse MoEs has emerged as an alternative solution; however, existing methods typically rely on heuristic neuron clustering or random splitting to partition the Feed-Forward Network (FFN) into experts. In this work, we propose DOT-MoE, a novel framework that formulates the decomposition of dense layers as a Differentiable Optimal Transport (DOT) problem. Instead of static heuristics, we model neuron assignment as a balanced transport problem, utilizing differentiable Sinkhorn-Knopp iterations to enforce strict expert capacity constraints. Furthermore, we utilize Straight-Through Estimators (STE) to jointly learn the discrete neuron-to-expert assignment and the token-to-expert routing policy end-to-end. Extensive experiments across multiple architectures and benchmarks demonstrate that DOT-MoE significantly outperforms structured pruning, heuristic clustering, and random-split baselines, retaining 90% of the original dense model's performance while reducing active parameters by 50%.