cs.LGOct 6, 2026

MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling

Authors: Mingyuan Zhang, Yue Bai, Zhongruo Wang, Yupin Huang, Yiyang Huang, Hailing Wang, Huimin Zeng, Yun Fu

Organizations: Department of Electrical and Computer Engineering, Northeastern University · Khoury College of Computer Science, Northeastern University

Abstract

Sparsely activated Mixture-of-Experts (MoE) models increase model capacity without a proportional increase in per-token computation. Dense-to-MoE upcycling reuses pretrained dense models to construct such systems, commonly by copying feed-forward networks (FFNs) into independently trained experts. We introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN. Each expert is defined by a learned binary mask, and a token-level router selects which masked FFNs to execute and combine. The router and mask scores are optimized jointly, while the underlying FFN weight values remain unchanged. This formulation supports neuron-structured, semi-structured, and unstructured experts within the same routing architecture. Our main configuration uses four 2:4 experts with top-2 routing, where two half-dense expert passes have the nominal FFN arithmetic of one dense pass, without requiring independent expert weight matrices. On five vision-language benchmarks with Qwen and Gemma backbones, this configuration achieves the highest performance among the compared baselines. Comparisons across mask granularities, routing interventions, and compute-matched controls distinguish the effects of learned connectivity from expert activation count. These results establish mask learning over frozen weights as a practical alternative for constructing token-routed MoE experts.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference

    Sep 28, 2026Junfeng Wu, Zehao Fan, Hadjer Benmeziane +3Mixture-Of-ExpertsParameter-Efficient Fine-Tuning Methods

  2. BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE

    May 14, 2026Juntong Wu, Jialiang Cheng, Qishen Yin +5Mixture-Of-Expert InferenceMixture-Of-Experts

  3. DOT-MoE: Differentiable Optimal Transport for MoEfication

    Jun 1, 2026Udbhav Bamba, Arnav Chavan, Aryamaan Thakur +2Localized Lora-MoeInference Cost