Organizations: Tongji University · Cornell University · Harbin Institute of Technology, Shenzhen · AI Training Platform Team, Shenzhen Loop Area Institute
Mixture-of-Experts (MoE) has been widely adopted in recent large language model (LLM) architectures. However, scaling up MoE in LLM training introduces system-level challenges on training, where non-uniform token routing can lead to highly imbalanced workloads across experts and devices, further destabilizing the training process. With trillion-scale LLMs, imbalanced expert workloads further amplify the resource cost of MoE training, resulting in degraded training efficiency and hardware utilization for underloaded experts, while hot experts require additional resources to accommodate excessive workloads. Recent studies address imbalanced MoE training through intricate parallelism strategies or resource reallocation. However, these system-level approaches often introduce additional resource requirements and considerable orchestration complexity, which become increasingly difficult to afford when training trillion-parameter LLMs under constrained computational resources. This work introduces CIPHER-MoE, which mitigates MoE workload imbalance while keeping the router's token-side Top-K selection unchanged. CIPHER-MoE applies affinity-aware Expert-to-Token filtering with explicit capacity control to reduce hotspot expert workloads without additional hardware resources or complex runtime design. The proposed method has been evaluated on large-scale MoE models, including DeepSeek-V4-Pro, showing up to 64.9 percentage points Top-1 expert workload reduction and 1.10×-1.94× training acceleration, while preserving the training quality. The source code will be released soon.
Figures & tables
Figure 1: Expert workload imbalance in DeepSeek-V4-Flash. (a) Illustration of the MoE architecture with data-dependent routing, which leads to the imbalanced workload. (b) Lopsided workload statistics across all layers of the DeepSeek-V4-Flash model. With the domain-specific dataset, 10% of the experts consume over 88% routing workload, leading to degraded training efficiency.
Models
Method
Latency / Iteration (sec)
NL4Opt
OptiBench
Bench4Opt Feasi.
Bench4Opt OR
DeepSeek-V4-Flash
Fixed Router
28.30 ( 1.00 × )
88.93
66.17
70.93
48.73
DeepSeek-V4-Flash
Vanilla MoE Training
38.69 (1.37 × )
93.08
68.00
71.51
51.02
DeepSeek-V4-Pro
Fixed Router
24.68 ( 1.00 × )
89.27
64.67
74.42
54.31
DeepSeek-V4-Pro
Vanilla MoE Training
OOM
-
-
-
-
Table 1: Post-training SFT of the DeepSeek-V4-Flash and DeepSeek-V4-Pro model on linear programming and operations research (OR) dataset. OR capabilities evaluated on NL4Opt ( Ramamonjison et al., 2023 ) , OptiBench ( Yang et al., 2025 ) , Bench4Opt-Feasi. ( Wang et al., 2025 ) , and Bench4Opt-OR ( Wang et al., 2025 ) . Fixed router exhibits suboptimal performance compared to the vanilla MoE training due to the homogenized expert training. Detailed training setup and resource configurations are summarized in the Appendix A.5 .
Figure 2: The overall design of the proposed CIPHER-MoE. The incoming tokens start with the normal Top-K expert selection, followed by the proposed expert-side token drop and reroute.
Model
Method
Operations Research
MedMCQA
SIQA
NL4Opt
OptiBench
B4O-Feas.
B4O-OR
W. Avg.
Strain
Acc.
Strain
Acc.
Strain
DeepSeek-V4-Flash
Vanilla MoE Training (Baseline)
93.08
68.00
71.51
51.02
69.09
-
76.60
-
51.84
-
CIPHER-strict
93.77
66.67
70.64
51.27
68.59
1.42 ×
77.36
1.50 ×
54.81
1.58 ×
CIPHER-reroute
90.66
66.33
69.77
51.27
67.73
1.38 ×
76.79
1.46 ×
57.83
1.52 ×
LocMoE ( τ=0.005 )
94.12
68.17
66.28
50.76
68.16
1.17 ×
–
–
–
–
LocMoE ( τ=0.01 )
94.81
67.67
70.06
53.05
69.46
1.12 ×
77.22
1.12 ×
55.94
1.19 ×
Table 2: Evaluation results and training speedup of domain-specific SFT. The proposed CIPHER-MoE achieves the optimal speedup and quality.
Figure 3: Training efficiency and memory footprint on DeepSeek-V4-Flash. (a) Per-step training time over the training progress, and CIPHER-MoE achieves up to 1.47 × speedup. (b) Step-time distributions, where CIPHER-strict and CIPHER-reroute achieve an average speedup of 1.38–1.42 × over the baseline. (c) Allocated memory throughout training.
Figure 4: Training stability and convergence. (a) LM loss over elapsed training time. (b) Ten-step trailing LM loss and time to reach the 0.10 target.
Figure 5: Router-update diversity and expert specialization. (a) Five consecutive training stages connected by arrows, with darker colors denoting later stages. Lower update-vector cosine similarity indicates clearer expert differentiation, while higher entropy effective rank indicates more dispersed, less homogeneous updates; the upper-left is better. (b) PCA of router-input hidden states. Compact same-color clusters indicate similar tokens routed to the same expert, while separation across colors indicates distinct token assignments and stronger specialization.
Table 3: DeepSeek-V4-Pro training efficiency and OR-domain evaluation.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Gradient stability of CIPHER-MoE over training progress for DeepSeek-V4-Flash, DeepSeek-V3, and GLM-5 on MedMCQA, SIQA, and OR, shown on logarithmic scales.
Figure 7: Expert routing behavior. Expert traffic across representative MoE layers (left), activated experts per token (middle), and load entropy during training (right).
Figure 8: Training efficiency and optimization stability on Qwen3.6-35B-A3B. (a) Step time over training progress, shown with a 10-step rolling mean and interquartile range. (b) Distribution of post-warm-up step times. (c) Raw gradient norms on a logarithmic scale. CIPHER maintains stable step times and convergent gradient norms, with CIPHER-strict exhibiting smoother late-stage optimization.
Model
Method
AIME-26
AMC
ARC
BBH
DROP
GPQA-D
GSM8K
HellaSwag
IFEval
MATH-500
MMLU-Pro
DeepSeek-V4-Flash
Baseline
33.33
71.60
98.10
84.00
86.00
64.14
95.90
85.40
87.80
91.20
80.02
CIPHER-strict
30.00
70.90
98.20
71.20
84.50
67.17
95.70
80.80
83.70
91.80
80.98
CIPHER-reroute
40.00
70.60
97.80
58.00
72.30
55.56
95.70
81.80
80.40
90.40
75.47
LocMoE ( τ =0.005)
33.33
70.90
98.10
71.00
86.30
66.67
96.80
82.00
86.50
92.20
80.59
LocMoE ( τ =0.01)
33.33
70.90
98.10
71.30
86.30
65.15
96.30
81.80
85.20
93.20
80.86
LocMoE ( τ =0.02)
46.67
71.60
98.00
71.20
86.40
65.15
96.40
81.70
85.40
91.80
81.44
Appendix
Table 4: Evaluation results of general benchmarks (%) for OR SFT.
Model
Method
AIME-26
AMC
ARC
BBH
DROP
GPQA-D
GSM8K
HellaSwag
IFEval
MATH-500
MMLU-Pro
DeepSeek-V4-Flash
Baseline
50.00
80.60
97.50
85.00
85.20
62.63
95.70
80.20
70.20
92.60
76.38
CIPHER-strict
56.67
75.56
98.03
85.16
81.62
66.67
96.44
76.35
78.74
94.00
80.31
CIPHER-reroute
50.00
75.56
97.74
86.13
81.40
60.10
96.13
80.33
74.12
92.80
78.77
LocMoE ( τ =0.01)
56.67
80.00
97.86
84.53
81.51
66.67
96.29
76.86
77.45
94.20
80.38
DeepSeek-V3
Baseline
46.67
77.61
96.65
87.70
82.65
60.61
96.66
74.38
80.10
92.20
74.28
CIPHER-strict
46.67
74.63
97.97
86.28
86.77
55.56
96.29
87.48
82.07
88.00
76.32
Appendix
Table 5: Evaluation results of general benchmarks (%) for MedMCQA SFT.
Model
Method
AIME-26
AMC
ARC
BBH
DROP
GPQA-D
GSM8K
HellaSwag
IFEval
MATH-500
MMLU-Pro
DeepSeek-V4-Flash
Baseline
53.33
75.40
96.30
87.90
84.70
60.10
96.10
87.10
79.30
90.80
77.68
CIPHER-strict
50.00
77.78
98.03
85.64
81.53
62.63
96.44
85.12
80.59
92.00
78.26
CIPHER-reroute
60.00
73.33
98.17
85.73
79.46
60.10
95.75
88.06
81.48
91.60
74.67
LocMoE ( τ =0.01)
53.33
73.33
97.91
85.73
81.91
56.57
96.51
84.87
81.33
91.20
78.27
DeepSeek-V3
Baseline
53.33
72.39
96.62
85.30
84.31
54.55
93.10
78.87
76.14
87.60
71.67
CIPHER-strict
46.67
65.67
97.94
86.35
88.39
57.58
95.98
87.61
92.81
87.40
73.99
Appendix
Table 6: Evaluation results of general benchmarks (%) for SIQA SFT.
Certified margin δ
Twofold reduction ( κ=2 )
Tenfold reduction ( κ=10 )
0.05
d≥6,586
d≥7,872
0.06
d≥4,571
d≥5,464
0.07
d≥3,357
d≥4,012
Appendix
Table 7: Sufficient representation dimensions for the illustrative capacity-separation regimes under A1–A3.
Most recent state-of-the-art (SOTA) large language models (LLMs) use Mixture-of-Experts (MoE) architectures to scale model capacity without proportional per-token compute, enabling higher-quality outputs at manageable serving costs. However, MoE inference at scale is fundamentally bottlenecked by expert load imbalance and inefficient token routing, especially in multi-node deployments where tokens are not guaranteed to be routed to local experts, resulting in significant inter-node all-to-all communication overhead. To systematically characterize these challenges, we profile SOTA open-source MoE models, including Llama 4 Maverick, DeepSeek V3-671B, and Qwen3-230B-A22B, on various datasets and collected over 100k real expert activation traces. Upon studying the expert activation patterns, we uncover various persistent properties across all the frontier MoE models: variable expert load imbalance, domain-specific expert activation where expert popularity shifts across task families (code, math, chat, general), and a strong correlation between prefill and decode expert activations. Motivated by these findings, we propose workload-aware micro-batch grouping and an expert placement strategy to maximize token locality to the destination expert, thereby reducing inter-node communication. Across models and datasets, these optimizations help reduce all2all communication data up to 20, resulting in lower MoE decode latency and better accelerator utilization.
Abhimanyu Bambhaniya, Geonhwa Jeong, Jason Park +6
1Georgia Institute of Technology, USA · 2Meta Platforms, Inc.
Mixture-of-experts (MoE) architectures enable trillion-parameter LLMs with sparsely activated experts. Expert parallelism (EP) is a widely adopted MoE training strategy, but it suffers from severe all-to-all communication bottlenecks, which is exaggerated by the limited inter-node network bandwidth as the growing model size requires distributing experts across GPU nodes. Prior work focused on overlapping these all-to-all communications with feed-forward network (FFN) and self-attention computations, which often leaves residual network-bound stalls due to inherent imbalance in attention and FFN layers' computation-communication ratios. We present DisagMoE, a disaggregated MoE training system that jointly optimizes model placement and scheduling for maximal efficiency. DisagMoE separates attention and FFN layers into disjoint GPU groups, introduces a multi-stage pipeline with uni-directional, many-to-many communications, and employs a computation-communication roofline model to balance GPU and network bandwidth allocation among the attention and FFN groups. DisagMoE is implemented on Megatron-LM, and evaluation shows that DisagMoE improves training efficiency across multiple MoE models with up to 1.8x speedup on 16-node 8xH800 clusters.
Zhichen Zeng, Chi-Chih Chang, Jiayi Wang +10
ByteDance Seed · University of Washington · Cornell University
Large Language Models (LLMs) increasingly adopt Mixture-of-Experts (MoE) architectures to scale efficiently under stringent resource constraints. However, MoE's sparse activation causes severe expert load imbalance, where a few experts become stragglers while others remain underutilized, leading to inflated inference latency and cost. Existing solutions assume static, serverful model deployments, limiting expert elasticity and often incurring costly expert swapping or degraded output quality. We present MoEless, an efficient serverless MoE serving framework that mitigates expert load imbalance via elastic expert execution. MoEless leverages lightweight, layer-aware predictors to estimate incoming expert load distributions and proactively identify stragglers. We design optimized scaling and placement strategies to improve function locality, GPU utilization, and cross-expert load balance. MoEless is prototyped on top of Megatron-LM and deployed on an eight-GPU testbed. Experiments with open-source MoE models and real-world workloads show that MoEless reduces inference latency by 43% and inference cost by 84% compared to state-of-the-art solutions.