ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control
Organizations: Qwen Team, Alibaba Token Hub, Alibaba Group
Abstract
Scaling Large Language Models (LLMs) via Mixture-of-Experts (MoE) enables massive parameter growth with nearly constant per-token computation. However, further scaling the parameter count requires increasingly sparse routing, where expert load imbalance becomes more severe. This imbalance reduces parameter utilization and training efficiency, and can undermine training stability, becoming a bottleneck to reliable scaling. In this work, we unify two representative auxiliary-loss-free methods as incomplete Proportional-Integral-Derivative (PID) controllers: DeepSeek's loss-free method acts as a fixed-step integral controller, while Kimi K3's Quantile Balancing functions as a generalized proportional controller. Building on this control perspective, we propose ID Balancing, an Integral-Derivative controller. It scales its integral term with load error and activates its derivative term only when imbalance worsens, enabling stronger corrections for large or worsening errors and smaller updates near balance. Evaluated across Top-, Top-, and Top- routing over experts, ID Balancing reduces worst-case backbone MaxVio and training-average backbone MinVio by over and , respectively, relative to the best baselines in the Top- setting. When the total parameter count increases from B to B (Top--of-), ID Balancing's worst-case backbone MaxVio remains nearly unchanged and is approximately lower than that of the auxiliary-loss baseline. ID Balancing also maintains competitive language-modeling and downstream performance. The advantages of ID Balancing grow as sparsity increases, making it a promising solution for scaling larger, sparser MoE models.
Figures & tables
| Backbone | MTP Module | |||||||||||
| Method | Act./Total | LM | MaxVio | MinVio | MaxVio | MinVio | ||||||
| loss | Last1k | Avg. | Worst | Last1k | Avg. | Last1k | Avg. | Worst | Last1k | Avg. | ||
| Top-3 of 768 experts, 120B tokens | ||||||||||||
| Auxiliary loss | 0.87B / 18.9B | 1.7527 | 1.4809 | 1.7239 | 70.76 | 0.5462 | 0.6232 | 3.8952 | 4.4639 | 43.38 | 0.6580 | 0.7502 |
| DeepSeek loss-free | 1.7426 | 0.7314 | 1.9356 | 211.19 | 0.7487 | 0.7441 | 0.7477 | 0.8930 | 17.32 | 0.6868 | 0.6533 | |
| Quantile Balancing | 1.7432 | 0.8093 | 0.7985 | 31.11 | 0.6504 | 0.6373 | 1.8777 | 1.2071 | 30.89 | 0.7106 | 0.7083 | |
| Method | Knowledge | STEM | Reasoning | Multilingual | Code | Avg. | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| MMLU | MMLU-Pro | SuperGPQA | MATH | GSM8K | BBH | MMMLU | EvalPlus | MultiPL-E | ||
| Top-8 of 256 experts, 3.0B / 24.8B, 560B tokens | ||||||||||
| Auxiliary loss | 68.00 | 47.10 | 27.36 | 47.12 | 74.94 | 69.57 | 59.82 | 51.51 | 46.22 | 54.63 |
| DeepSeek loss-free | 70.08 | 47.23 | 27.43 | 45.84 | 76.54 | 71.09 | 60.61 | 53.96 | 46.66 | 55.49 |
| Quantile Balancing | 70.49 | 48.06 | 27.61 | 48.54 | 74.37 | 69.94 | 60.17 | 52.27 | 42.46 | 54.88 |
| ID Balancing | 69.58 | 48.19 | 27.52 | 47.92 | 74.49 | 72.91 | 60.78 | 56.48 | 44.40 | 55.81 |
| Backbone | MTP Module | ||||||||||
| Method | LM | MaxVio | MinVio | MaxVio | MinVio | ||||||
| loss | Last1k | Avg. | Worst | Last1k | Avg. | Last1k | Avg. | Worst | Last1k | Avg. | |
| Auxiliary loss | 2.0099 | 3.1383 | 2.8379 | 20.89 | 0.8423 | 0.8458 | 3.5640 | 3.0533 | 4.83 | 0.6569 | 0.7130 |
| DeepSeek loss-free | 2.0031 | 1.3755 | 1.7811 | 72.54 | 0.8730 | 0.8269 | 0.5033 | 0.7124 | 17.98 | 0.6105 | 0.5514 |
| Quantile Balancing | 2.0009 | 0.5021 | 0.4979 | 2.96 | 0.4083 | 0.3938 | 0.6243 | 0.5836 | 10.77 | 0.4346 | 0.4308 |
| ID Balancing | 2.0007 | 0.6437 | 0.7274 | 11.97 | 0.4828 | 0.4950 | 0.4642 | 0.5068 | 7.20 | 0.3940 | 0.3927 |
| Backbone | MTP module | ||||||||||
| Setting | LM | MaxVio | MinVio | MaxVio | MinVio | ||||||
| loss | Last1k | Avg. | Worst | Last1k | Avg. | Last1k | Avg. | Worst | Last1k | Avg. | |
| Integral gain , integral term only, | |||||||||||
| 1.7422 | 0.6759 | 0.9459 | 23.26 | 0.5190 | 0.5777 | 0.6375 | 0.7991 | 9.14 | 0.5603 | 0.5775 | |
| 1.7431 | 0.6924 | 0.8003 | 14.96 | 0.5233 | 0.5527 | 0.8557 | 0.8862 | 15.18 | 0.5848 | 0.5714 | |
| 1.7429 | 0.6831 | 0.7872 | 20.13 | 0.5263 | 0.5436 | 0.9222 | 1.0388 | 35.56 | 0.6279 | 0.5800 | |
| Backbone | MTP module | ||||||||||
| Setting | LM | – k steps | – k steps | Avg. | – k steps | Avg. | |||||
| loss | MaxVio | MinVio | MaxVio | MinVio | MaxVio | MinVio | MaxVio | MinVio | MaxVio | MinVio | |
| Derivative gain , fixed | |||||||||||
| (I only) | 1.7431 | 2.0020 | 0.7955 | 0.9059 | 0.5989 | 0.8003 | 0.5527 | 1.4965 | 0.6222 | 0.8862 | 0.5714 |
| 1.7428 | 1.8900 | 0.7853 | 0.8869 | 0.5992 | 0.7910 | 0.5485 | 1.7454 | 0.6201 | 0.9459 | 0.5799 | |
| 1.7426 | 1.9330 | 0.7791 | 0.8612 | 0.5899 | 0.7931 | 0.5480 | 1.7724 | 0.6336 | 0.9698 | 0.5782 | |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Specification | Small MoE | Medium MoE | Large MoE |
|---|---|---|---|
| Transformer layers | 20 | 28 | 28 |
| Hidden size | 1024 | 2048 | 2048 |
| Vocabulary size | 248320 | 248320 | 248320 |
| Softmax-attention heads | 16 | 16 | 16 |
| Query groups | 2 | 2 | 2 |
| Softmax-attention head dimension | 256 | 256 | 256 |