cs.CLOct 8, 2026

Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts

Authors: Yunkai Chai, Tong Zhu, Xiaoye Qu, Xuyang Hu, Guanjie Chen, Qipeng Guo, Yu Cheng

Organizations: Shanghai AI Laboratory, Shanghai, China · Shanghai Jiao Tong University, Shanghai, China · Nanyang Technological University, Singapore

Abstract

Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token. However, traditional training paradigms enforce a static choice of top-kk experts, which converts a continuous routing distribution into a rigid step function. This constraint introduces a brittle boundary where highly competitive experts are arbitrarily separated into full-supervision and zero-feedback zones based on minor score fluctuations. To address this issue, we propose Elastic Expert Routing, which stochastically samples the active expert budget from a localized discrete distribution centered at kk. Over multiple training iterations, this mechanism softens the sharp threshold into a gradual probability distribution. Because the sampling neighborhood remains symmetric, this approach matches the expected computational cost of deterministic training, while preserving the inference budget. Extensive experiments demonstrate the efficacy of our method on both supervised fine-tuning and from-scratch pretraining settings. During supervised fine-tuning, elastic routing improves downstream macro-averages on OLMoE-1B-7B and Qwen3-30B-A3B by +0.84+0.84 and +2.02+2.02 points, respectively. In addition, in from-scratch pretraining, it outperforms the static top-kk baseline by 1.61.6 points on average across downstream tasks.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models

    Sep 29, 2026Yury Nahshan, Nati Daniel, Jacob Goldberger +1Mixture-of-Experts Language ModelsParameter-Free Mixture-of-Experts Routing

  2. How Sparse Probability Maps Shape Mixture-of-Experts Routing

    Oct 5, 2026Tomás Brogueira, Marcos Treviso, Miguel CouceiroMixture-of-Experts Language ModelsMixture-of-Experts Inference