Distributionally Robust Mixture-of-Experts Training
Organizations: New York University Center for Data Science, NYU Shanghai
Abstract
Mixture-of-Experts (MoE) transformers scale capacity by activating only a few experts per token, but this sparsity creates a hidden reliability problem: when routing is imperfect, load-balanced models may send tokens to experts that are insufficiently trained for the assigned inputs. We propose Distributionally Robust MoE Training (DRMoET), a drop-in objective that treats layer-wise experts as endogenous robustness groups and optimizes high-loss routing outcomes rather than merely equalizing traffic. DRMoET updates a per-layer expert distribution by an entropy-regularized softmax rule on EMA-smoothed, activation-weighted expert losses, strengthening plausible non-top routing paths while preserving standard MoE computation. Under the FLAME-MoE recipe at 746M-total and 10.3B-total scales, DRMoET improves downstream averages over both standard FLAME-MoE and auxiliary-loss-free balancing. At 10.3B total parameters and 67B training tokens, DRMoET improves the seven-task average from 0.6625 to 0.6767, while the auxiliary-loss-free baseline achieves 0.6431. Mechanistic analyses show lower expert-loss variance with nearly unchanged mean loss, 4.3% lower excess loss under forced mid- misrouting, and improved domain-expert specialization. These results position routing robustness-not only utilization balance-as a practical objective for reliable sparse MoE scaling. Project page and code are available at: https://drmoet.github.io/.
Figures & tables
| Scale | Method | ARC-C | ARC-E | HellaSwag | PIQA | WinoGrande | SciQ | ReCoRD* | Average |
| 290M-746M | Training tokens: 33.5B | ||||||||
| FLAME-MoE | 0.2329 | 0.5295 | 0.3447 | 0.6697 | 0.5012 | 0.8020 | 0.7023 | 0.5403 | |
| Aux-free ( ) | 0.2261 | 0.4878 | 0.3283 | 0.6670 | 0.4949 | 0.7100 | 0.6445 | 0.5084 | |
| Aux-free ( ) | 0.2184 | 0.4895 | 0.3301 | 0.6627 | 0.4949 | 0.6990 | 0.6440 | 0.5055 | |
| DRMoET ( ) | 0.2338 | 0.5522 | 0.3472 | 0.6844 | 0.5209 | 0.7950 | 0.7115 | 0.5493 | |
| DRMoET ( ) | 0.2270 | 0.5400 | 0.3470 | 0.6839 | 0.4941 | 0.8180 | 0.7007 | 0.5444 | |
| Metric | Baseline | DRMoET | (%) |
| Worst loss | 2.6930 | 2.6379 | |
| Best loss | 2.1149 | 2.0722 | |
| Mean loss | 2.3712 | 2.3699 | |
| Range | 0.5781 | 0.5656 | |
| Std. | 0.1516 | 0.1386 | |
| CV | 0.0639 | 0.0585 |
| Model | Normal | Misrouting | Deg. |
| Baseline | 3.187 | 7.249 | |
| DRMoET | 3.182 | 7.071 |
| Tier | FLAME-MoE | DRMoET | Improvement(%) |
| Top | 3.176 | 3.188 | |
| Mid | 3.100 | 3.077 | |
| Bottom | 3.072 | 3.058 |
| Metric | Improvement (%) |
| Avg Max Selectivity | |
| Avg (specialization margin) | |
| Avg (competence advantage) | |
| Avg MI |
| Ablation | Setting | ARC-E | HellaSwag | PIQA | SciQ | ReCoRD* | Average |
| Credit | FLAME-MoE | 0.5295 | 0.3447 | 0.6697 | 0.8020 | 0.7023 | 0.6096 |
| Activation weighted, | 0.5400 | 0.3470 | 0.6839 | 0.8180 | 0.7007 | 0.6179 | |
| Activation weighted, | 0.5682 | 0.3452 | 0.6888 | 0.8190 | 0.7105 | 0.6263 | |
| Raw probability, | 0.5455 | 0.3384 | 0.6839 | 0.7920 | 0.6938 | 0.6107 | |
| Raw probability, | 0.5421 | 0.3435 | 0.6861 | 0.7910 | 0.7065 | 0.6138 | |
| Top-12 routing | FLAME-MoE | 0.5459 | 0.3506 | 0.6991 | 0.7950 | 0.7033 | 0.6188 |
| Metric | Baseline | DRMoET | |
| TFLOP/s/GPU | |||
| Time/iter (ms) | |||
| Throughput std |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Step | Baseline CV | DRMoET CV | Diff |
| 540 | 2.75% | 3.32% | +0.57% |
| 2160 | 2.46% | 3.04% | +0.58% |
| 3780 | 1.78% | 2.73% | +0.95% |
| 5400 | 2.00% | 2.64% | +0.64% |
| 7020 | 6.47% | 2.21% | -4.26% |
| 8640 | 1.60% | 2.07% | +0.46% |
| Layer | Variance | Gini | Top-3 Conc. | ||||||
| Baseline | DRMoET | Baseline | DRMoET | Baseline | DRMoET | ||||
| 2 | 0.00518 | 0.00483 | -6.6% | 0.248 | 0.238 | -4.2% | 17.7% | 17.5% | -1.2% |
| 3 | 0.00800 | 0.00678 | -15.3% | 0.301 | 0.275 | -8.4% | 19.7% | 19.3% | -2.4% |
| 4 | 0.01534 | 0.01395 | -9.0% | 0.413 | 0.393 | -4.9% | 24.3% | 23.8% | -2.4% |
| 5 | 0.02487 | 0.02188 | -12.0% | 0.500 | 0.475 | -5.0% | 30.7% | 28.7% | -6.4% |
| 6 | 0.01477 | 0.01549 | +4.9% | 0.422 | 0.427 | +1.1% | 24.0% | 25.1% | +4.3% |
| Method | ARC-C | ARC-E | HellaSwag | PIQA | WinoGrande | SciQ | ReCoRD* | Average |
| FLAME-MoE (with balancing) | 0.2329 | 0.5295 | 0.3447 | 0.6697 | 0.5012 | 0.8020 | 0.7023 | 0.5403 |
| DRMoET without balancing | 0.2389 | 0.5539 | 0.3430 | 0.6877 | 0.5067 | 0.7770 | 0.7066 | 0.5448 |
| DRMoET with balancing | 0.2295 | 0.5682 | 0.3452 | 0.6888 | 0.5051 | 0.8190 | 0.7105 | 0.5523 |