cs.LGOct 5, 2026

Distributionally Robust Mixture-of-Experts Training

Authors: Xin Teng, Muxiao Li, Hongyi Wen

Organizations: New York University Center for Data Science, NYU Shanghai

Abstract

Mixture-of-Experts (MoE) transformers scale capacity by activating only a few experts per token, but this sparsity creates a hidden reliability problem: when routing is imperfect, load-balanced models may send tokens to experts that are insufficiently trained for the assigned inputs. We propose Distributionally Robust MoE Training (DRMoET), a drop-in objective that treats layer-wise experts as endogenous robustness groups and optimizes high-loss routing outcomes rather than merely equalizing traffic. DRMoET updates a per-layer expert distribution by an entropy-regularized softmax rule on EMA-smoothed, activation-weighted expert losses, strengthening plausible non-top routing paths while preserving standard MoE computation. Under the FLAME-MoE recipe at 746M-total and 10.3B-total scales, DRMoET improves downstream averages over both standard FLAME-MoE and auxiliary-loss-free balancing. At 10.3B total parameters and 67B training tokens, DRMoET improves the seven-task average from 0.6625 to 0.6767, while the auxiliary-loss-free baseline achieves 0.6431. Mechanistic analyses show lower expert-loss variance with nearly unchanged mean loss, 4.3% lower excess loss under forced mid-kk misrouting, and improved domain-expert specialization. These results position routing robustness-not only utilization balance-as a practical objective for reliable sparse MoE scaling. Project page and code are available at: https://drmoet.github.io/.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ProbMoE: Differentiable Probabilistic Routing for Mixture-of-Experts

    Jun 1, 2026Heng Zhao, Zilei Shao, Guy Van den Broeck +1Large Language Model RoutingProbabilistic Inference

  2. Hierarchical Mixture-of-Experts with Two-Stage Optimization

    May 8, 2026Gleb Molodtsov, Alexander Miasnikov, Aleksandr BeznosikovMixture-Of-ExpertsLoad Balancing