Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models
Organizations: Bar-Ilan University, Ramat-Gan, Israel · NVIDIA, Israel
Abstract
Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing affinities and token-level error. We introduce token-error supervision for sparse routing in two forms. The first form predicts an error score per expert. The affinity-weighted aggregate of these scores is aligned to the next-token cross-entropy loss, while the individual scores attenuate affinity before top- selection. The second directly aligns the router's affinities to the model's objective without requiring an additional head or inference-time modification. Both formulations use the Itakura--Saito divergence or an exponential negative log-likelihood for aligning affinities and token errors. Across two sparse MoE backbones and four multiple-choice question-answering benchmarks, we evaluate both supervision mechanisms. On Granite, our method improves accuracy by approximately 2.3 percentage points on average over a parameter-matched routing baseline. With stronger supervision, the gain on ARC-Challenge reaches 2.94 points. Both mechanisms preserve the native sparse execution budget and aggregation policy. Our code is available in the supplementary materials.
Figures & tables
| Model | Method | ARC-Challenge | OpenBookQA | SciQ † | MedMCQA |
|---|---|---|---|---|---|
| Granite 3.1 | CE | ||||
| ACS-ENLL (Ours) | |||||
| ACS-IS (Ours) | |||||
| CE + affinity tuning | |||||
| ACS-ENLL + affinity tuning (Ours) | |||||
| ACS-IS + affinity tuning (Ours) |
| Model | Mechanism | Objective | ARC-C | OpenBookQA | SciQ | MedMCQA |
|---|---|---|---|---|---|---|
| Granite | TES, frozen | IS | ||||
| ENLL | ||||||
| Granite | ACS, trainable | IS | ||||
| ENLL | ||||||
| OLMoE | ACS, frozen | IS | ||||
| ENLL |
| Objective | Supervised MoE layers | ARC-Challenge | OpenBookQA | SciQ | MedMCQA |
|---|---|---|---|---|---|
| ACS-ENLL | Final layer | ||||
| ACS-ENLL | All 16 layers | ||||
| ACS-IS | Final layer | ||||
| ACS-IS | All 16 layers |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Routing transform | Pointwise | Group | Group | Test NLL | Accuracy (%) |
|---|---|---|---|---|---|
| 0.1586 | 0.3975 | 0.4977 | 0.9723 | 67.27 | |
| 0.2486 | 0.4521 | 0.5637 | 0.9673 | 66.87 |
| Parameter | Value | Accuracy (%) | Choice NLL |
|---|---|---|---|
| Error-head learning rate | |||
| Attenuation scale | |||
| Setting | Granite 3.1 3B-A800M | OLMoE-1B-7B-SFT |
| Training examples | 2,000 (ARC); 4,992 (others) | 2,000 (ARC); 5,000 (others) |
| Validation / test examples | 50 / 500 | 50 / 500 |
| LoRA rank / / dropout | 8 / 8 / 0.05 | 8 / 8 / 0.05 |
| LoRA targets | Q/K/V and routed-expert projections | Q/K/V/O and routed-expert projections |
| LoRA learning rate | ||
| Batch / accumulation | 8 / 2 | 8 / 1 |
| Model | Affinity | Objective | ARC-C | OpenBook QA | SciQ | MedMCQA | |
|---|---|---|---|---|---|---|---|
| Granite | Trainable | IS | |||||
| ENLL | |||||||
| OLMoE | Frozen | IS | |||||
| Objective | ||||
|---|---|---|---|---|
| ACS–IS | ||||
| ACS–ENLL |
| Model | Method | ACS and affinity- training scope | ARC-Challenge | OpenBookQA | SciQ | MedMCQA |
|---|---|---|---|---|---|---|
| Granite 3.1 | Router CE | Final 1 | ||||
| First 8 | ||||||
| Last 8 | ||||||
| All 32 | ||||||
| ACS-ENLL | Final 1 | |||||
| First 8 |
| Method | Split | |||
|---|---|---|---|---|
| Router CE | Train | |||
| Test | ||||
| ACS–IS | Train | |||
| Test | ||||
| ACS–ENLL | Train | |||
| Test |
| Scope | Objective | Depth- scaled | scaled | Change |
|---|---|---|---|---|
| Last 8 | IS | |||
| ENLL | ||||
| All 32 | IS | |||
| ENLL |
| Dataset | Scope | Objective | Fixed | Depth- scaled | Change |
|---|---|---|---|---|---|
| ARC-C | Last 8 | IS | |||
| ENLL | |||||
| All 32 | IS | ||||
| ENLL | |||||
| OpenBookQA | Last 8 | IS | |||
| ENLL |