MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference
Authors: Junfeng Wu, Zehao Fan, Hadjer Benmeziane, Kaoutar El Maghraoui, Liu Liu, Yinan Wang
Organizations: Department of Industrial and Systems Engineering, Rensselaer Polytechnic Institute · Department of Electrical, Computer, and Systems Engineering, Rensselaer Polytechnic Institute · IBM Research
Mixture-of-experts (MoE) language models often exceed the memory of a single GPU. Expert offloading keeps most experts in host memory and loads them on demand, so decoding speed depends on how many experts each token must fetch. Caching and prefetching reduce this cost only as far as the routing allows. Router-only fine-tuning can reshape the routing to reuse experts, but it keeps the experts frozen, so they cannot adapt to the tokens the new routing sends them. We propose MaskCoFT, a masked co-adaptive fine-tuning method that trains routers and experts together with the cross-entropy loss alone. During fine-tuning, a learnable binary mask restricts the Top-K routing of each layer to a subset of experts, and the experts adapt to the tokens redirected to them. At inference, the learned mask becomes a soft prior that re-ranks experts, so every expert remains selectable. We simulate a GPU cache of 4 experts per layer for Mixtral-8x7B and 12 for DeepSeek-V2-Lite. MaskCoFT cuts expert fetches per token by 23.7% and 10.1% relative to the base model. In real offloading system serving, it lowers the time per output token by up to 16.4% and 5.5%, respectively. Its average accuracy over nine benchmarks stays above the base model by 0.92 and 0.53 points.
Figures & tables
Figure 1: Routed expert weights dominate memory, and loading them dominates offloaded decoding latency. MaskCoFT reduces expert fetches while preserving accuracy. (a) Routed experts hold 96.6% of Mixtral-8×7B’s parameters and 91.7% of DeepSeek-V2-Lite’s. (b) Decode latency breakdown of Mixtral-8×7B in MoE-Offloading ( Eliseev and Mazur, 2023 ) , measured with NVIDIA Nsight Systems ( NVIDIA Corporation, 2026 ) . (c) Expert fetches per token and average accuracy over nine benchmarks under GPU cache holding 4 experts per layer for Mixtral-8×7B and 12 for DeepSeek-V2-Lite, both relative to the base model. The GPU cache holds 4 experts per layer for Mixtral-8×7B and 12 for DeepSeek-V2-Lite. MaskCoFT cuts fetches by 23.7% on Mixtral-8×7B and by 10.1% on DeepSeek-V2-Lite. It keeps average accuracy above the base model.
Figure 2: MaskCoFT learns a concentration set of experts per layer and co-adapts routers and experts. (a) At every step, Stage 1 samples a set of E′ experts from πℓ and runs Top- K routing only inside it. A straight-through gradient, accumulated over T steps, updates m~ℓ . (b) Stage 2 fixes the sets of all layers, then the routers and experts adapt to the rerouted tokens. (c) At inference, the learned prior π re-ranks experts toward the concentration set, so decoding fetches fewer experts.
Generative
Multiple choice
Method
GSM8K
HumanEval
MMLU
ARC-E
ARC-C
HellaSwag
WinoGrande
PIQA
BoolQ
Avg.
DeepSeek-V2-Lite
Baseline
38.36
26.83
55.53
74.37
45.99
77.73
71.27
80.25
80.34
61.19
CE-only
37.38
28.05
54.73
77.69
50.51
79.07
72.22
80.69
79.30
62.18
MaskCoFT (Ours)
38.06
28.05
53.36
75.72
49.57
76.06
71.82
80.36
82.45
61.72
ΔBaseline
−0.30
+1.22
−2.17
+1.35
+3.58
−1.67
+0.55
+0.11
+2.11
+0.53
Table 1: Accuracy (%) on nine benchmarks. GSM8K is 5-shot with flexible-extract scoring. HumanEval is 0-shot pass@1. The other seven tasks are 0-shot multiple choice. Avg. is the mean over the nine benchmarks. Bold marks the best of the three models in each column. The Δ rows give MaskCoFT minus each reference, in percentage points.
LRU
LFU
FIFO
Method
HR (%) ↑
Fetches/tok ↓
HR (%) ↑
Fetches/tok ↓
HR (%) ↑
Fetches/tok ↓
DeepSeek-V2-Lite B=12
Baseline
44.45
86.66
43.45
88.21
43.34
88.39
CE-only
43.08
88.80
41.48
91.29
41.59
91.12
MaskCoFT (Ours)
47.85
81.36
49.18
79.28
45.75
84.63
ΔBaseline
+3.40
−6.1%
+5.73
−10.1%
+2.41
−4.3%
Table 2: Trace-driven cache simulation. HR is the hit rate, and Fetches/tok is the number of expert fetches per token (Eq. ( 5 )). The GPU cache holds B=12 experts per layer for DeepSeek-V2-Lite and B=4 for Mixtral-8×7B. The Δ rows give the change of MaskCoFT against each reference, in percentage points for HR and as a relative change for Fetches/tok.
Figure 3: Hit rate and Fetches/tok under LRU, LFU, and FIFO as the cache budget B varies. Left : DeepSeek-V2-Lite with B∈{12,24,36} . Right : Mixtral-8x7B with B∈{2,4,6} .
Figure 4: System performance for Mixtral-8×7B (Top) DeepSeek-V2-Lite (Bottom) and under different output token lengths.
Accuracy (%)
HR (%), B=12
Method
GSM8K (flex)
GSM8K (strict)
HumanEval
MMLU
LRU
LFU
FIFO
Reported by ReMoE
Baseline
39.04
38.89
26.83
57.72
45.19
45.97
44.32
CE-only
37.23
36.92
28.05
57.44
N/A
N/A
N/A
ReMoE
38.36
38.13
29.27
57.81
50.35
51.51
49.30
ΔBaseline
−0.68
−0.76
+2.44
+0.09
+5.16
+5.54
+4.98
Table 3: Comparison with the numbers reported by ReMoE ( Zhu et al., 2026 ) on DeepSeek-V2-Lite. Our runs use 5-shot GSM8K, 0-shot HumanEval pass@1 and 0-shot MMLU. GSM8K reports both the flexible-extract and the strict-match score. ReMoE does not report its GSM8K and MMLU shot counts. HR is the hit rate with a cache of B=12 routed experts per layer and batch size 1. The Δ rows give each method minus its own Baseline or CE-only run, in percentage points. ReMoE reports no CE-only hit rates (N/A).
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Dense
MoE
Shared
Routed
Active
Total
Active
Context
BF16 size
Model
layers
layers
experts
experts
experts
params (B)
params (B)
length
(GB)
DeepSeek-V2-Lite
1
26
2
64
6 + 2 shared
15.7
2.4
32K
≈ 31.4
Mixtral-8x7B
0
32
0
8
2
46.7
12.9
32K
≈ 93.4
Appendix
Table 4: Specifications of the MoE models. Expert counts are per MoE layer.
Figure 5: The convergence of the concentration expert set during Stage 1 of MaskCoFT. It shows the overlap ratio between the concentration expert set and the final converged concentration set in Stage 2 every 100 steps. The mask converges early for both models. Mixtral-8×7B reaches 100 % by step 1.7k, and DeepSeek-V2-Lite reaches 99 % by step 4.5k.
Figure 6: Decoding latency breakdown for Mixtral-8×7B under different output token lengths.
Figure 7: Decoding latency breakdown for DeepSeek-V2-Lite under different output token lengths.
Mixture-of-Experts (MoE) architectures enhance the efficiency of large language models by activating only a subset of experts per token. However, standard MoE employs a fixed Top-K routing strategy, leading to redundant computation and suboptimal inference latency. Existing acceleration methods either require costly retraining with architectural changes or suffer from severe performance drop at high sparsity due to train-inference mismatch. To address these limitations, we propose BEAM (Binary Expert Activation Masking), a novel method that learns token-adaptive expert selection via trainable binary masks. With a straight-through estimator and an auxiliary regularization loss, BEAM induces dynamic expert sparsity through end-to-end training while maintaining model capability. We further implement an efficient custom CUDA kernel for BEAM, ensuring seamless integration with the vLLM inference framework. Experiments show that BEAM retains over 98% of the original model's performance while reducing MoE layer FLOPs by up to 85%, achieving up to 2.5× faster decoding and 1.4× higher throughput, demonstrating its effectiveness as a practical, plug-and-play solution for efficient MoE inference.
Juntong Wu, Jialiang Cheng, Qishen Yin +5
Taobao & Tmall Group of Alibaba · Shenzhen Graduate School, Peking University
Fine-grained Mixture-of-Experts (MoE) models sparsely activate only a subset of experts per token, reducing activated computation while maintaining high model capacity. However, in memory-constrained inference scenarios, only a small set of experts can be cached. Experts not in the cache must be fetched from slow external storage (e.g., UFS), leading to frequent evictions and substantial I/O overhead. We propose ReMoE, a router fine-tuning framework designed to boost token-wise expert reuse. ReMoE biases the router toward recently selected experts, producing temporally stable routing that better matches cache locality constraints. By increasing short-horizon expert reuse, ReMoE reduces expert fetches from storage without adding inference-time computation. Experiments on DeepSeek and Qwen models show that ReMoE improves expert reuse by 26% while maintaining downstream task performance. Real-system evaluations further confirm these benefits, improving output throughput by 8.4% under vLLM GPU-CPU expert offloading and reducing TPOT by 43.6-49.8% under llama.cpp on Jetson Orin NX, corresponding to a 1.77-1.99× decode speedup across diverse workloads. Checkpoints and usage instructions are available at https://github.com/BUAA-OSCAR/ReMoE.
Xiongwei Zhu, Xiaojian Liao, Tianyang Jiang +3
School of Computer Science and Engineering, Beihang University, Beijing 100191, China · Huawei Technologies Ltd.
Mixture-of-Experts (MoE) models decouple parameter count from per-token compute, but deployment still requires hosting every expert in memory. Recent theory shows that experts whose router weights change least during fine-tuning can be pruned with provable accuracy preservation, yet the guarantee assumes full fine-tuning. We show that the signal can be elicited through a brief parameter-efficient adaptation. We fine-tune with a lightweight adapter, rank experts by the induced router change, and prune the least-changed experts in one shot. On Mixtral-8×7B-Instruct, router-only LoRA trains 0.002% of parameters and retains 27.54% MMLU-Pro accuracy with half the experts removed, against roughly 16% for magnitude and random pruning. Signal quality improves monotonically with adapter size, reaching 28.76%, and declines as adaptation spreads beyond the router. Under their shared budget, IA3 reaches 28.04% while Houlsby reaches 25.39%. The criterion transfers to Qwen1.5-MoE fine-tuned for mathematical reasoning, retaining 49.7% mean accuracy over eleven benchmarks with half the experts removed. Structural pruning reduces memory by 49% and per-token latency by 37%. Lightweight router sensitivity therefore makes provably motivated, task-conditioned expert pruning practical at scale.
Ali Janati, Kaoutar El Maghraoui, Chengke Zou +2
Data Science Institute, Columbia University, New York, NY, USA · Department of Computer Science, Columbia University, New York, NY, USA