cs.AIOct 8, 2026

RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing

Authors: Ilya Lasy, Nora Yinuo Cai, Kola Ayonrinde

Organizations: Faculty of Informatics, TU Wien · Independent · UK AI Security Institute

Abstract

Sparse Mixture of Experts (MoE) models scale more efficiently than dense models by routing tokens to modular expert networks that are only active for processing a fraction of tokens. A leading hypothesis for the performance of MoE models is that each expert specialises in a single, coherent domain. However, interpretability efforts that assume this hypothesis have generally been unsuccessful. We propose and present evidence for an alternative account that we call the Superposed Specialisation Hypothesis (SSH): experts specialise in a disjoint union of fine-grained features rather than one broad domain. Leveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most predictive of routing decisions and produces unified natural language explanations. On gpt-oss-20b, RouterInterp explains expert routing with ∼65%{\sim}65\% higher detection accuracy than prior token statistics based methods. This work provides a scalable method for generating more accurate explanations of expert routing and increases our understanding of a previously uninterpretable component of foundation models.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models

    May 8, 2026Youngsik Yoon, Siwei Wang, Wei Chen +1Mixture-of-Experts Language ModelsLLM Routing

  2. Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing

    Jul 30, 2026Huiyuan Tian, Bonan Xu, Shijian LiSparse Mixture-of-ExpertsExpert Routing

  3. Self-Routing: Parameter-Free Expert Routing from Hidden States

    Apr 1, 2026Jama Hussein Mohamud, Drew Wagner, Mirco RavanelliParameter-Free Mixture-of-Experts RoutingLLM Routing