Scaling Large Language Models (LLMs) via Mixture-of-Experts (MoE) enables massive parameter growth with nearly constant per-token computation. However, further scaling the parameter count requires increasingly sparse routing, where expert load imbalance becomes more severe. This imbalance reduces parameter utilization and training efficiency, and can undermine training stability, becoming a bottleneck to reliable scaling. In this work, we unify two representative auxiliary-loss-free methods as incomplete Proportional-Integral-Derivative (PID) controllers: DeepSeek's loss-free method acts as a fixed-step integral controller, while Kimi K3's Quantile Balancing functions as a generalized proportional controller. Building on this control perspective, we propose ID Balancing, an Integral-Derivative controller. It scales its integral term with load error and activates its derivative term only when imbalance worsens, enabling stronger corrections for large or worsening errors and smaller updates near balance. Evaluated across Top-10, Top-5, and Top-3 routing over 768 experts, ID Balancing reduces worst-case backbone MaxVio and training-average backbone MinVio by over 50% and 12%, respectively, relative to the best baselines in the Top-3 setting. When the total parameter count increases from 18.9B to 69.9B (Top-10-of-768), ID Balancing's worst-case backbone MaxVio remains nearly unchanged and is approximately 89.6% lower than that of the auxiliary-loss baseline. ID Balancing also maintains competitive language-modeling and downstream performance. The advantages of ID Balancing grow as sparsity increases, making it a promising solution for scaling larger, sparser MoE models.
Figures & tables
Figure 1: ID Balancing maintains load control as routing becomes sparser and model capacity grows. (a) Worst backbone MaxVio for Top- 10 , Top- 5 , and Top- 3 routing over 768 experts. (b) Training-average backbone MinVio, expressed as a decimal ratio, across the same settings. Models in (a) and (b) have 18.9 B total parameters and are trained on 120 B tokens. (c) Worst backbone MaxVio as total parameters increase from 18.9 B to 69.9 B ( 1.03 B to 3.2 B active) under fixed Top- 10 -of- 768 routing, using the first 30 k and 50 k steps of the small and large models, respectively.
Figure 2: ID Balancing limits transient overload and sustained underload in highly sparse MoE training. Results use Top- 3 -of- 768 routing in an 18.9 B-parameter model trained on 120 B tokens, with metrics averaged over 20 backbone layers. (a) Mean MaxVio on a logarithmic scale ( ↓ ). (b) Mean MinVio ( ↓ ). ID Balancing achieves the lowest training-average underload.
Figure 3: Magnitude-aware integral control reduces early overload and sustained underload. We compare the integral-only update with DeepSeek’s loss-free method in an 18.9 B model using Top- 3 -of- 768 routing and 120 B training tokens. (a) Maximum MaxVio across the 20 backbone layers, shown on a logarithmic scale ( ↓ ). (b) Mean MinVio across the same layers ( ↓ ).
Figure 4: The worsening-gated derivative term reduces early overload and expert concentration. We compare magnitude-aware integral control with and without the derivative term in an 18.9 B model using Top- 3 -of- 768 routing and 120 B training tokens. Metrics are averaged over 20 backbone layers. (a) gˉ(t) is the fraction of active gates averaged over the 20 backbone layers. (b) Mean MaxVio ( ↓ ). (c) Share of token assignments routed to the busiest 10% of experts ( →0.10 ), with a balanced target of 0.10 .
Figure 5: EMA smoothing reduces bias drift but slows load correction in Quantile Balancing. We compare the original update ( ρ=0 ), its EMA variants, DeepSeek’s loss-free method, and ID Balancing in an 18.9 B model using Top- 3 -of- 768 routing and 120 B training tokens. All other comparisons use Quantile Balancing without EMA. (a) Mean absolute bias change ∣Δb∣ between checkpoints 1 k steps apart during the final 5 k steps ( ↓ ). (b) Mean MaxVio on a logarithmic scale ( ↓ ). (c) Mean MinVio ( ↓ ).
Backbone
MTP Module
Method
Act./Total
LM
MaxVio ↓
MinVio ↓
MaxVio ↓
MinVio ↓
loss
Last1k
Avg.
Worst
Last1k
Avg.
Last1k
Avg.
Worst
Last1k
Avg.
Top-3 of 768 experts, 120B tokens
Auxiliary loss
0.87B / 18.9B
1.7527
1.4809
1.7239
70.76
0.5462
0.6232
3.8952
4.4639
43.38
0.6580
0.7502
DeepSeek loss-free
1.7426
0.7314
1.9356
211.19
0.7487
0.7441
0.7477
0.8930
17.32
0.6868
0.6533
Quantile Balancing
1.7432
0.8093
0.7985
31.11
0.6504
0.6373
1.8777
1.2071
30.89
0.7106
0.7083
Table 1: ID Balancing improves backbone load control while maintaining competitive LM loss. We compare Top- 3 , Top- 5 , and Top- 10 routing over 768 experts in 18.9 B models trained on 120 B tokens. All summaries use the same 30 k-step window. Backbone Last1k and Avg. metrics are averaged over 20 layers, while Worst MaxVio is the maximum across layers and steps. MTP metrics are reported separately.
Figure 6: Reduced ID Balancing gains limit load drift while retaining competitive LM loss during continued pretraining. We compare full gains ( Ki=Kd=6×10−3 ), gains reduced by 1000× ( 6×10−6 ), and a frozen bias in 28 -layer models at a constant learning rate of 3×10−5 . The frozen setting matches the continued-pretraining baselines for DeepSeek’s loss-free method and Quantile Balancing. The top and bottom rows use Top- 8 -of- 256 and Top- 10 -of- 768 routing, respectively. (a, d) LM loss with 201 -step smoothing ( ↓ ). (b, e) Mean MaxVio ( ↓ ). (c, f) Mean MinVio ( ↓ ).
Method
Knowledge
STEM
Reasoning
Multilingual
Code
Avg.
MMLU
MMLU-Pro
SuperGPQA
MATH
GSM8K
BBH
MMMLU
EvalPlus
MultiPL-E
Top-8 of 256 experts, 3.0B / 24.8B, 560B tokens
Auxiliary loss
68.00
47.10
27.36
47.12
74.94
69.57
59.82
51.51
46.22
54.63
DeepSeek loss-free
70.08
47.23
27.43
45.84
76.54
71.09
60.61
53.96
46.66
55.49
Quantile Balancing
70.49
48.06
27.61
48.54
74.37
69.94
60.17
52.27
42.46
54.88
ID Balancing
69.58
48.19
27.52
47.92
74.49
72.91
60.78
56.48
44.40
55.81
Table 2: ID Balancing maintains competitive downstream performance. Results use Top- 8 -of- 256 and Top- 10 -of- 768 routing. Avg. is the mean across nine benchmarks covering knowledge, STEM, reasoning, multilingual understanding, and code generation. Block headers list active and total parameter counts. Higher scores are better, and bold marks the best result in each column within a block.
Backbone
MTP Module
Method
LM
MaxVio ↓
MinVio ↓
MaxVio ↓
MinVio ↓
loss
Last1k
Avg.
Worst
Last1k
Avg.
Last1k
Avg.
Worst
Last1k
Avg.
Auxiliary loss
2.0099
3.1383
2.8379
20.89
0.8423
0.8458
3.5640
3.0533
4.83
0.6569
0.7130
DeepSeek loss-free
2.0031
1.3755
1.7811
72.54
0.8730
0.8269
0.5033
0.7124
17.98
0.6105
0.5514
Quantile Balancing
2.0009
0.5021
0.4979
2.96
0.4083
0.3938
0.6243
0.5836
10.77
0.4346
0.4308
ID Balancing
2.0007
0.6437
0.7274
11.97
0.4828
0.4950
0.4642
0.5068
7.20
0.3940
0.3927
Table 3: ID Balancing and Quantile Balancing maintain load control at a higher learning rate. Results use a 1.03 B-active/ 18.9 B-total model with Top- 10 -of- 768 routing, trained on 120 B tokens at a constant learning rate of 5.86×10−3 ( 2.3× the standard peak). Both methods yield lower backbone load violations than Auxiliary loss and DeepSeek’s loss-free method.
Figure 7: ID Balancing and Quantile Balancing maintain lower backbone overload and underload at a higher learning rate. Results use Top- 10 -of- 768 routing and 120 B training tokens at a constant learning rate of 5.86×10−3 ( 2.3× the standard peak). Metrics are averaged over 20 backbone layers. (a) Mean MaxVio on a logarithmic scale ( ↓ ), showing a large early spike for DeepSeek’s loss-free method. (b) Mean MinVio ( ↓ ), showing more severe underload for Auxiliary loss and DeepSeek’s loss-free method.
Figure 8: Larger auxiliary-loss coefficients worsen LM loss and load balance at a higher learning rate. We compare α∈{0.05,0.10,0.50} in a 1.03 B-active/ 18.9 B-total model with Top- 10 -of- 768 routing at a constant learning rate of 5.86×10−3 ( 2.3× the standard peak). Training targets 120 B tokens, but both larger- α runs terminate before 30 k steps. Load metrics are averaged over 20 backbone layers. (a) LM loss ( ↓ ), with higher values at larger coefficients. (b) Mean MaxVio on a logarithmic scale ( ↓ ), exceeding 50 for α=0.50 . (c) Mean MinVio ( ↓ ), reaching 1.0 for both larger coefficients.
Figure 9: Load-balancing methods show distinct gradient and activation dynamics. Results use a 3.0 B-active/ 24.8 B-total model with Top- 8 -of- 256 routing over 50 k steps. (a) Global gradient norm. ID Balancing and Quantile Balancing have lower, smoother trajectories, while DeepSeek’s loss-free method shows stronger late-stage fluctuations. (b) MoE-output magnitude averaged over the 28 backbone layers.
Figure 10: ID Balancing maintains consistent overload and underload control across layers. Results use Top- 3 -of- 768 routing over 30 k steps and cover all 20 backbone layers. The top and bottom rows show the training average and the final 1 k-step average, respectively. (a, c) MaxVio ( ↓ ). (b, d) MinVio ( ↓ ). ID Balancing achieves lower or comparable MaxVio and lower MinVio than Quantile Balancing across all layers. MinVio generally increases with depth for all four methods.
Figure 11: ID Balancing reduces expert inactivity relative to Auxiliary loss. Results are averaged by layer over nine downstream benchmarks. The top and bottom rows use Top- 8 -of- 256 and Top- 10 -of- 768 routing, respectively. (a, d) Inactive-expert ratio, the fraction of experts receiving no evaluation token ( ↓ ). (b, e) Top- K gating-score entropy, where a larger value means a more even mixture over selected experts. (c, f) Full-pool selection-score entropy, where a larger value means a flatter selection-score distribution.
Backbone
MTP module
Setting
LM
MaxVio ↓
MinVio ↓
MaxVio ↓
MinVio ↓
loss
Last1k
Avg.
Worst
Last1k
Avg.
Last1k
Avg.
Worst
Last1k
Avg.
Integral gain Ki , integral term only, Kd=0
Ki=3×10−3
1.7422
0.6759
0.9459
23.26
0.5190
0.5777
0.6375
0.7991
9.14
0.5603
0.5775
Ki=6×10−3
1.7431
0.6924
0.8003
14.96
0.5233
0.5527
0.8557
0.8862
15.18
0.5848
0.5714
Ki=9×10−3
1.7429
0.6831
0.7872
20.13
0.5263
0.5436
0.9222
1.0388
35.56
0.6279
0.5800
Table 4: An integral gain of Ki=6×10−3 yields the lowest worst-case backbone MaxVio among the tested gains. Results use a 0.87 B-active/ 18.9 B-total model with Top- 3 -of- 768 routing over 30 k steps. We fix Kd=0 and vary Ki∈{3,6,9}×10−3 . Increasing Ki to 9×10−3 further reduces training-average backbone violations but worsens MTP balance. Lower values are better for all metrics.
Backbone
MTP module
Setting
LM
0 – 1 k steps
1 – 5 k steps
Avg.
0 – 5 k steps
Avg.
loss
MaxVio ↓
MinVio ↓
MaxVio ↓
MinVio ↓
MaxVio ↓
MinVio ↓
MaxVio ↓
MinVio ↓
MaxVio ↓
MinVio ↓
Derivative gain Kd , fixed Ki=6×10−3
Kd=0 (I only)
1.7431
2.0020
0.7955
0.9059
0.5989
0.8003
0.5527
1.4965
0.6222
0.8862
0.5714
Kd=3×10−3
1.7428
1.8900
0.7853
0.8869
0.5992
0.7910
0.5485
1.7454
0.6201
0.9459
0.5799
Kd=6×10−3
1.7426
1.9330
0.7791
0.8612
0.5899
0.7931
0.5480
1.7724
0.6336
0.9698
0.5782
Table 5: Larger derivative gains reduce backbone overload over 1 – 5 k steps but increase MTP overload. Results use a 0.87 B-active/ 18.9 B-total model with Top- 3 -of- 768 routing over 30 k steps. We fix Ki=6×10−3 and vary Kd∈{0,3,6,12}×10−3 , with Kd=0 as the integral-only baseline. Load metrics summarize the indicated early intervals and the full training window (Avg.).
Figure 12: A small integral gain responds slowly to the early imbalance, while larger gains converge to similar backbone trajectories. Results use Top- 3 -of- 768 routing over 30 k steps with Kd=0 . We compare Ki∈{3,6,9}×10−3 . (a) Mean backbone MaxVio ( ↓ ). (b) Mean backbone MinVio ( ↓ ).
Figure 13: The derivative term reduces early backbone overload. Results use Top- 3 -of- 768 routing with Ki=6×10−3 and Kd∈{0,3,6,12}×10−3 , where Kd=0 is the integral-only baseline. Let Vt denote mean MaxVio over the 20 backbone layers. (a) Vt over the first 10 k steps, smoothed with a five-point moving average. (b) Cumulative excess above Vt=1 over 1 – 5 k steps.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Specification
Small MoE
Medium MoE
Large MoE
Transformer layers
20
28
28
Hidden size
1024
2048
2048
Vocabulary size
248320
248320
248320
Softmax-attention heads
16
16
16
Query groups
2
2
2
Softmax-attention head dimension
256
256
256
Appendix
Table A: Architectural specifications for the evaluated MoE models. Parameter counts include the input embedding and untied output head, and exclude the MTP module. Active parameters include all dense and shared components together with the routed experts selected by Top- K .