CIPHER-MoE: Balancing Efficiency and Routing Fidelity in Trillion-Scale MoE Training
Organizations: Tongji University · Cornell University · Harbin Institute of Technology, Shenzhen · AI Training Platform Team, Shenzhen Loop Area Institute
Abstract
Mixture-of-Experts (MoE) has been widely adopted in recent large language model (LLM) architectures. However, scaling up MoE in LLM training introduces system-level challenges on training, where non-uniform token routing can lead to highly imbalanced workloads across experts and devices, further destabilizing the training process. With trillion-scale LLMs, imbalanced expert workloads further amplify the resource cost of MoE training, resulting in degraded training efficiency and hardware utilization for underloaded experts, while hot experts require additional resources to accommodate excessive workloads. Recent studies address imbalanced MoE training through intricate parallelism strategies or resource reallocation. However, these system-level approaches often introduce additional resource requirements and considerable orchestration complexity, which become increasingly difficult to afford when training trillion-parameter LLMs under constrained computational resources. This work introduces CIPHER-MoE, which mitigates MoE workload imbalance while keeping the router's token-side Top-K selection unchanged. CIPHER-MoE applies affinity-aware Expert-to-Token filtering with explicit capacity control to reduce hotspot expert workloads without additional hardware resources or complex runtime design. The proposed method has been evaluated on large-scale MoE models, including DeepSeek-V4-Pro, showing up to 64.9 percentage points Top-1 expert workload reduction and 1.10-1.94 training acceleration, while preserving the training quality. The source code will be released soon.
Figures & tables
| Models | Method | Latency / Iteration (sec) | NL4Opt | OptiBench | Bench4Opt Feasi. | Bench4Opt OR |
| DeepSeek-V4-Flash | Fixed Router | 28.30 ( 1.00 ) | 88.93 | 66.17 | 70.93 | 48.73 |
| DeepSeek-V4-Flash | Vanilla MoE Training | 38.69 (1.37 ) | 93.08 | 68.00 | 71.51 | 51.02 |
| DeepSeek-V4-Pro | Fixed Router | 24.68 ( 1.00 ) | 89.27 | 64.67 | 74.42 | 54.31 |
| DeepSeek-V4-Pro | Vanilla MoE Training | OOM | - | - | - | - |
| Model | Method | Operations Research | MedMCQA | SIQA | |||||||
| NL4Opt | OptiBench | B4O-Feas. | B4O-OR | W. Avg. | Acc. | Acc. | |||||
| DeepSeek-V4-Flash | Vanilla MoE Training (Baseline) | 93.08 | 68.00 | 71.51 | 51.02 | 69.09 | - | 76.60 | - | 51.84 | - |
| CIPHER-strict | 93.77 | 66.67 | 70.64 | 51.27 | 68.59 | 1.42 | 77.36 | 1.50 | 54.81 | 1.58 | |
| CIPHER-reroute | 90.66 | 66.33 | 69.77 | 51.27 | 67.73 | 1.38 | 76.79 | 1.46 | 57.83 | 1.52 | |
| LocMoE ( ) | 94.12 | 68.17 | 66.28 | 50.76 | 68.16 | 1.17 | – | – | – | – | |
| LocMoE ( ) | 94.81 | 67.67 | 70.06 | 53.05 | 69.46 | 1.12 | 77.22 | 1.12 | 55.94 | 1.19 | |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Method | AIME-26 | AMC | ARC | BBH | DROP | GPQA-D | GSM8K | HellaSwag | IFEval | MATH-500 | MMLU-Pro |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash | Baseline | 33.33 | 71.60 | 98.10 | 84.00 | 86.00 | 64.14 | 95.90 | 85.40 | 87.80 | 91.20 | 80.02 |
| CIPHER-strict | 30.00 | 70.90 | 98.20 | 71.20 | 84.50 | 67.17 | 95.70 | 80.80 | 83.70 | 91.80 | 80.98 | |
| CIPHER-reroute | 40.00 | 70.60 | 97.80 | 58.00 | 72.30 | 55.56 | 95.70 | 81.80 | 80.40 | 90.40 | 75.47 | |
| LocMoE ( =0.005) | 33.33 | 70.90 | 98.10 | 71.00 | 86.30 | 66.67 | 96.80 | 82.00 | 86.50 | 92.20 | 80.59 | |
| LocMoE ( =0.01) | 33.33 | 70.90 | 98.10 | 71.30 | 86.30 | 65.15 | 96.30 | 81.80 | 85.20 | 93.20 | 80.86 | |
| LocMoE ( =0.02) | 46.67 | 71.60 | 98.00 | 71.20 | 86.40 | 65.15 | 96.40 | 81.70 | 85.40 | 91.80 | 81.44 |
| Model | Method | AIME-26 | AMC | ARC | BBH | DROP | GPQA-D | GSM8K | HellaSwag | IFEval | MATH-500 | MMLU-Pro |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash | Baseline | 50.00 | 80.60 | 97.50 | 85.00 | 85.20 | 62.63 | 95.70 | 80.20 | 70.20 | 92.60 | 76.38 |
| CIPHER-strict | 56.67 | 75.56 | 98.03 | 85.16 | 81.62 | 66.67 | 96.44 | 76.35 | 78.74 | 94.00 | 80.31 | |
| CIPHER-reroute | 50.00 | 75.56 | 97.74 | 86.13 | 81.40 | 60.10 | 96.13 | 80.33 | 74.12 | 92.80 | 78.77 | |
| LocMoE ( =0.01) | 56.67 | 80.00 | 97.86 | 84.53 | 81.51 | 66.67 | 96.29 | 76.86 | 77.45 | 94.20 | 80.38 | |
| DeepSeek-V3 | Baseline | 46.67 | 77.61 | 96.65 | 87.70 | 82.65 | 60.61 | 96.66 | 74.38 | 80.10 | 92.20 | 74.28 |
| CIPHER-strict | 46.67 | 74.63 | 97.97 | 86.28 | 86.77 | 55.56 | 96.29 | 87.48 | 82.07 | 88.00 | 76.32 |
| Model | Method | AIME-26 | AMC | ARC | BBH | DROP | GPQA-D | GSM8K | HellaSwag | IFEval | MATH-500 | MMLU-Pro |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DeepSeek-V4-Flash | Baseline | 53.33 | 75.40 | 96.30 | 87.90 | 84.70 | 60.10 | 96.10 | 87.10 | 79.30 | 90.80 | 77.68 |
| CIPHER-strict | 50.00 | 77.78 | 98.03 | 85.64 | 81.53 | 62.63 | 96.44 | 85.12 | 80.59 | 92.00 | 78.26 | |
| CIPHER-reroute | 60.00 | 73.33 | 98.17 | 85.73 | 79.46 | 60.10 | 95.75 | 88.06 | 81.48 | 91.60 | 74.67 | |
| LocMoE ( =0.01) | 53.33 | 73.33 | 97.91 | 85.73 | 81.91 | 56.57 | 96.51 | 84.87 | 81.33 | 91.20 | 78.27 | |
| DeepSeek-V3 | Baseline | 53.33 | 72.39 | 96.62 | 85.30 | 84.31 | 54.55 | 93.10 | 78.87 | 76.14 | 87.60 | 71.67 |
| CIPHER-strict | 46.67 | 65.67 | 97.94 | 86.35 | 88.39 | 57.58 | 95.98 | 87.61 | 92.81 | 87.40 | 73.99 |
| Certified margin | Twofold reduction ( ) | Tenfold reduction ( ) |
|---|---|---|
| Model | Nodes | Parallel configuration | Seq. length | GBS | MoE configuration |
|---|---|---|---|---|---|
| DeepSeek-V4-Pro | 32 | TP2/PP8/EP64/CP1/ETP1/VPP4 + SP | 4K | 256/8192 | 384 experts; learned Top-6 |
| DeepSeek-V4-Flash | 8 | TP2/PP4/EP16/CP1/ETP1/DP16 + SP | 8K | 128 | 256 experts; learned Top-6; sqrt-softplus gate |
| DeepSeek-V3 | 16 | TP4/PP16/EP4/CP1/DP4 + SP | 8K | 128 | 256 experts; Top-8 routing; sigmoid gate |
| GLM-5 | 16 | TP2/PP8/EP32/CP1/ETP1/VPP2/DP16/EDP1 | 8K | 128 | 256 experts; Top-8 routing; sigmoid gate |