RouterInterp: Understanding Superposed Specialisation in Mixture of Experts Routing
Organizations: Faculty of Informatics, TU Wien · Independent · UK AI Security Institute
Abstract
Sparse Mixture of Experts (MoE) models scale more efficiently than dense models by routing tokens to modular expert networks that are only active for processing a fraction of tokens. A leading hypothesis for the performance of MoE models is that each expert specialises in a single, coherent domain. However, interpretability efforts that assume this hypothesis have generally been unsuccessful. We propose and present evidence for an alternative account that we call the Superposed Specialisation Hypothesis (SSH): experts specialise in a disjoint union of fine-grained features rather than one broad domain. Leveraging the SSH, we introduce RouterInterp, a method for interpreting expert routing that identifies Sparse Autoencoder features most predictive of routing decisions and produces unified natural language explanations. On gpt-oss-20b, RouterInterp explains expert routing with higher detection accuracy than prior token statistics based methods. This work provides a scalable method for generating more accurate explanations of expert routing and increases our understanding of a previously uninterpretable component of foundation models.
Figures & tables
| Layer 4 | Layer 8 | Layer 12 | Layer 16 | Layer 20 | |
| Unigram Lookup | 0.295 | 0.300 | 0.278 | 0.305 | 0.315 |
| Bigram Lookup | 0.309 | 0.320 | 0.283 | 0.320 | 0.331 |
| Unigram AutoInterp | 0.332 | 0.406 | 0.320 | 0.356 | 0.348 |
| Bigram AutoInterp | 0.429 | 0.453 | 0.385 | 0.407 | 0.411 |
| Expert Activations AutoInterp | 0.415 | 0.475 | 0.345 | 0.425 | 0.425 |
| RouterInterp ( ) | 0.490 | 0.542 | 0.465 | 0.509 | 0.470 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| OLMoE-1B-7B | gpt-oss-20b | |
|---|---|---|
| Unigram Baseline | 0.564 | 0.296 |
| Bigram Baseline | 0.633 | 0.356 |
| Neuron Basis Probe | 0.666 | 0.586 |
| PCA Basis Probe | 0.680 | 0.536 |
| SAE Predictor (Ours) | 0.740 | 0.730 |
| Layer 4 | Layer 8 | Layer 12 | Layer 16 | Layer 20 | |
|---|---|---|---|---|---|
| SAE Predictor ( ) | 0.569 | 0.702 | 0.730 | 0.439 | 0.789 |
| SAE Predictor ( ) | 0.552 | 0.725 | 0.685 | 0.297 | 0.773 |
| Neuron Probe ( ) | 0.519 | 0.638 | 0.586 | 0.559 | 0.630 |
| Neuron Probe ( ) | 0.476 | 0.603 | 0.509 | 0.524 | 0.570 |
| PCA Probe ( ) | 0.544 | 0.623 | 0.536 | 0.494 | 0.366 |
| PCA Probe ( ) | 0.517 | 0.603 | 0.523 | 0.482 | 0.384 |
| Layer 4 | Layer 8 | Layer 12 | Layer 16 | Layer 20 | |
| Unigram Baseline | 0.472 | 0.442 | 0.296 | 0.343 | 0.369 |
| Bigram Baseline | 0.535 | 0.511 | 0.356 | 0.406 | 0.411 |
| SAE Predictor ( ) | 0.183 | 0.260 | 0.185 | 0.089 | 0.280 |
| SAE Predictor ( ) | 0.235 | 0.380 | 0.267 | 0.249 | 0.617 |
| SAE Predictor ( ) | 0.349 | 0.408 | 0.434 | 0.387 | 0.694 |
| SAE Predictor ( ) | 0.490 | 0.640 | 0.632 | 0.414 | 0.749 |
| Layer | FVU | Dead features (%) |
|---|---|---|
| 3 | 0.007 | 4.11 |
| 7 | 0.024 | 0.51 |
| 11 | 0.064 | 0.59 |
| 15 | 0.218 | 0.51 |
| FVU | Dead features (%) | |||
|---|---|---|---|---|
| Layer | ||||
| 4 | 0.050 | 0.037 | 0.04 | 0.90 |
| 8 | 0.101 | 0.077 | 0.02 | 0.29 |
| 12 | 0.172 | 0.137 | 0.00 | 0.01 |
| 16 | 0.232 | 0.191 | 0.00 | 0.00 |
| 20 | 0.146 | 0.117 | 0.02 | 0.00 |
| Metric | Human–Human | Human–LLM |
|---|---|---|
| Pooled agreement | 0.807 0.070 | 0.857 0.050 |
| Cohen’s | 0.487 0.165 | 0.602 0.150 |
| Layer | Macro-F1 | Spearman | MSE |
|---|---|---|---|
| 4 | 0.569 | 0.258 | 0.0278 |
| 8 | 0.702 | 0.809 | 0.0294 |
| 12 | 0.730 | 0.741 | 0.0295 |
| 16 | 0.439 | 0.413 | 0.0285 |
| 20 | 0.789 | 0.817 | 0.0296 |
| Level | Explanation |
|---|---|
| Expert 2 | This expert activates across following distinct contexts: (1) Physical trauma and medical conditions—high activation on “bruis” and variants (“bruised”, “bruises”), plus fragments like “blister”, “abras”, “haemat”, “infarct”, and “stunned”. (2) Technical and structured writing—code, mathematical expressions, scientific notation, legal citations, and medical terminology from programming, mathematics, science, law, and academia, including punctuation and formatting elements. (3) Brand and term suffixes—final segments of brand names, company names, or technical terms before spaces or punctuation. (4) Electronic components—activation on “diode”, “resistor”, and “rectifier” in patents, diagrams, and technical specifications. |
| Feature 41223 | Physical trauma terms—“bruis” and variants (“bruised”, “bruises”), plus fragments of injury-related terms (“blister”, “abras”, “haemat”, “infarct”, “stunned”) in clinical or descriptive contexts. |
| Activating Examples: | |
| [1] suffered scr apes , bruis es , bleeding | |
| [2] . Her face was bruis ed and there | |
| Feature 71261 | Programming and scientific notation—code snippets, structured data, and formal keywords with high activation on punctuation, symbols, and domain-specific terms. |