MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference
Authors: Junfeng Wu, Zehao Fan, Hadjer Benmeziane, Kaoutar El Maghraoui, Liu Liu, Yinan Wang
Organizations: Department of Industrial and Systems Engineering, Rensselaer Polytechnic Institute · Department of Electrical, Computer, and Systems Engineering, Rensselaer Polytechnic Institute · IBM Research
Mixture-of-experts (MoE) language models often exceed the memory of a single GPU. Expert offloading keeps most experts in host memory and loads them on demand, so decoding speed depends on how many experts each token must fetch. Caching and prefetching reduce this cost only as far as the routing allows. Router-only fine-tuning can reshape the routing to reuse experts, but it keeps the experts frozen, so they cannot adapt to the tokens the new routing sends them. We propose MaskCoFT, a masked co-adaptive fine-tuning method that trains routers and experts together with the cross-entropy loss alone. During fine-tuning, a learnable binary mask restricts the Top-K routing of each layer to a subset of experts, and the experts adapt to the tokens redirected to them. At inference, the learned mask becomes a soft prior that re-ranks experts, so every expert remains selectable. We simulate a GPU cache of 4 experts per layer for Mixtral-8x7B and 12 for DeepSeek-V2-Lite. MaskCoFT cuts expert fetches per token by 23.7% and 10.1% relative to the base model. In real offloading system serving, it lowers the time per output token by up to 16.4% and 5.5%, respectively. Its average accuracy over nine benchmarks stays above the base model by 0.92 and 0.53 points.
Figures & tables
Figure 1: Routed expert weights dominate memory, and loading them dominates offloaded decoding latency. MaskCoFT reduces expert fetches while preserving accuracy. (a) Routed experts hold 96.6% of Mixtral-8×7B’s parameters and 91.7% of DeepSeek-V2-Lite’s. (b) Decode latency breakdown of Mixtral-8×7B in MoE-Offloading ( Eliseev and Mazur, 2023 ) , measured with NVIDIA Nsight Systems ( NVIDIA Corporation, 2026 ) . (c) Expert fetches per token and average accuracy over nine benchmarks under GPU cache holding 4 experts per layer for Mixtral-8×7B and 12 for DeepSeek-V2-Lite, both relative to the base model. The GPU cache holds 4 experts per layer for Mixtral-8×7B and 12 for DeepSeek-V2-Lite. MaskCoFT cuts fetches by 23.7% on Mixtral-8×7B and by 10.1% on DeepSeek-V2-Lite. It keeps average accuracy above the base model.
Figure 2: MaskCoFT learns a concentration set of experts per layer and co-adapts routers and experts. (a) At every step, Stage 1 samples a set of E′ experts from πℓ and runs Top- K routing only inside it. A straight-through gradient, accumulated over T steps, updates m~ℓ . (b) Stage 2 fixes the sets of all layers, then the routers and experts adapt to the rerouted tokens. (c) At inference, the learned prior π re-ranks experts toward the concentration set, so decoding fetches fewer experts.
Generative
Multiple choice
Method
GSM8K
HumanEval
MMLU
ARC-E
ARC-C
HellaSwag
WinoGrande
PIQA
BoolQ
Avg.
DeepSeek-V2-Lite
Baseline
38.36
26.83
55.53
74.37
45.99
77.73
71.27
80.25
80.34
61.19
CE-only
37.38
28.05
54.73
77.69
50.51
79.07
72.22
80.69
79.30
62.18
MaskCoFT (Ours)
38.06
28.05
53.36
75.72
49.57
76.06
71.82
80.36
82.45
61.72
ΔBaseline
−0.30
+1.22
−2.17
+1.35
+3.58
−1.67
+0.55
+0.11
+2.11
+0.53
Table 1: Accuracy (%) on nine benchmarks. GSM8K is 5-shot with flexible-extract scoring. HumanEval is 0-shot pass@1. The other seven tasks are 0-shot multiple choice. Avg. is the mean over the nine benchmarks. Bold marks the best of the three models in each column. The Δ rows give MaskCoFT minus each reference, in percentage points.
LRU
LFU
FIFO
Method
HR (%) ↑
Fetches/tok ↓
HR (%) ↑
Fetches/tok ↓
HR (%) ↑
Fetches/tok ↓
DeepSeek-V2-Lite B=12
Baseline
44.45
86.66
43.45
88.21
43.34
88.39
CE-only
43.08
88.80
41.48
91.29
41.59
91.12
MaskCoFT (Ours)
47.85
81.36
49.18
79.28
45.75
84.63
ΔBaseline
+3.40
−6.1%
+5.73
−10.1%
+2.41
−4.3%
Table 2: Trace-driven cache simulation. HR is the hit rate, and Fetches/tok is the number of expert fetches per token (Eq. ( 5 )). The GPU cache holds B=12 experts per layer for DeepSeek-V2-Lite and B=4 for Mixtral-8×7B. The Δ rows give the change of MaskCoFT against each reference, in percentage points for HR and as a relative change for Fetches/tok.
Figure 3: Hit rate and Fetches/tok under LRU, LFU, and FIFO as the cache budget B varies. Left : DeepSeek-V2-Lite with B∈{12,24,36} . Right : Mixtral-8x7B with B∈{2,4,6} .
Figure 4: System performance for Mixtral-8×7B (Top) DeepSeek-V2-Lite (Bottom) and under different output token lengths.
Accuracy (%)
HR (%), B=12
Method
GSM8K (flex)
GSM8K (strict)
HumanEval
MMLU
LRU
LFU
FIFO
Reported by ReMoE
Baseline
39.04
38.89
26.83
57.72
45.19
45.97
44.32
CE-only
37.23
36.92
28.05
57.44
N/A
N/A
N/A
ReMoE
38.36
38.13
29.27
57.81
50.35
51.51
49.30
ΔBaseline
−0.68
−0.76
+2.44
+0.09
+5.16
+5.54
+4.98
Table 3: Comparison with the numbers reported by ReMoE ( Zhu et al., 2026 ) on DeepSeek-V2-Lite. Our runs use 5-shot GSM8K, 0-shot HumanEval pass@1 and 0-shot MMLU. GSM8K reports both the flexible-extract and the strict-match score. ReMoE does not report its GSM8K and MMLU shot counts. HR is the hit rate with a cache of B=12 routed experts per layer and batch size 1. The Δ rows give each method minus its own Baseline or CE-only run, in percentage points. ReMoE reports no CE-only hit rates (N/A).
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Dense
MoE
Shared
Routed
Active
Total
Active
Context
BF16 size
Model
layers
layers
experts
experts
experts
params (B)
params (B)
length
(GB)
DeepSeek-V2-Lite
1
26
2
64
6 + 2 shared
15.7
2.4
32K
≈ 31.4
Mixtral-8x7B
0
32
0
8
2
46.7
12.9
32K
≈ 93.4
Appendix
Table 4: Specifications of the MoE models. Expert counts are per MoE layer.
Figure 5: The convergence of the concentration expert set during Stage 1 of MaskCoFT. It shows the overlap ratio between the concentration expert set and the final converged concentration set in Stage 2 every 100 steps. The mask converges early for both models. Mixtral-8×7B reaches 100 % by step 1.7k, and DeepSeek-V2-Lite reaches 99 % by step 4.5k.
Figure 6: Decoding latency breakdown for Mixtral-8×7B under different output token lengths.
Figure 7: Decoding latency breakdown for DeepSeek-V2-Lite under different output token lengths.