cs.AISep 29, 2026

Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models

Authors: Yury Nahshan, Nati Daniel, Jacob Goldberger, Yoli Shavit

Organizations: Bar-Ilan University, Ramat-Gan, Israel · NVIDIA, Israel

Abstract

Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing affinities and token-level error. We introduce token-error supervision for sparse routing in two forms. The first form predicts an error score per expert. The affinity-weighted aggregate of these scores is aligned to the next-token cross-entropy loss, while the individual scores attenuate affinity before top-KK selection. The second directly aligns the router's affinities to the model's objective without requiring an additional head or inference-time modification. Both formulations use the Itakura--Saito divergence or an exponential negative log-likelihood for aligning affinities and token errors. Across two sparse MoE backbones and four multiple-choice question-answering benchmarks, we evaluate both supervision mechanisms. On Granite, our method improves accuracy by approximately 2.3 percentage points on average over a parameter-matched routing baseline. With stronger supervision, the gain on ARC-Challenge reaches 2.94 points. Both mechanisms preserve the native sparse execution budget and aggregation policy. Our code is available in the supplementary materials.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models

    May 8, 2026Youngsik Yoon, Siwei Wang, Wei Chen +1Mixture-Of-ExpertsExperts

  2. Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing

    Jul 30, 2026Huiyuan Tian, Bonan Xu, Shijian LiMixture-Of-ExpertsComplementarity

  3. TEXAS: Task-Expert-Aware Supervision for Downstream Mixture-of-Experts LLM Adaptation

    Jul 31, 2026Guanzhi Deng, Haibo Wang, Kuan Wu +5Token-Level Supervision