cs.LGSep 28, 2026

MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference

Authors: Junfeng Wu, Zehao Fan, Hadjer Benmeziane, Kaoutar El Maghraoui, Liu Liu, Yinan Wang

Organizations: Department of Industrial and Systems Engineering, Rensselaer Polytechnic Institute · Department of Electrical, Computer, and Systems Engineering, Rensselaer Polytechnic Institute · IBM Research

Abstract

Mixture-of-experts (MoE) language models often exceed the memory of a single GPU. Expert offloading keeps most experts in host memory and loads them on demand, so decoding speed depends on how many experts each token must fetch. Caching and prefetching reduce this cost only as far as the routing allows. Router-only fine-tuning can reshape the routing to reuse experts, but it keeps the experts frozen, so they cannot adapt to the tokens the new routing sends them. We propose MaskCoFT, a masked co-adaptive fine-tuning method that trains routers and experts together with the cross-entropy loss alone. During fine-tuning, a learnable binary mask restricts the Top-K routing of each layer to a subset of experts, and the experts adapt to the tokens redirected to them. At inference, the learned mask becomes a soft prior that re-ranks experts, so every expert remains selectable. We simulate a GPU cache of 4 experts per layer for Mixtral-8x7B and 12 for DeepSeek-V2-Lite. MaskCoFT cuts expert fetches per token by 23.7% and 10.1% relative to the base model. In real offloading system serving, it lowers the time per output token by up to 16.4% and 5.5%, respectively. Its average accuracy over nine benchmarks stays above the base model by 0.92 and 0.53 points.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE

    May 14, 2026Juntong Wu, Jialiang Cheng, Qishen Yin +5Mixture-Of-Expert InferenceMixture-Of-Experts

  2. ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference

    May 26, 2026Xiongwei Zhu, Xiaojian Liao, Tianyang Jiang +3Mixture-Of-Expert InferenceMixture-Of-Experts

  3. Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models

    Aug 8, 2026Ali Janati, Kaoutar El Maghraoui, Chengke Zou +2ExpertsLow-Rank Adaptation Framework