HAPMoE: Heterogeneity-Aware Automatic Parallelism Planning for Mixture-of-Experts Models Training
Authors: Mengyuan Fan, Peizhuang Cong, Zixiao Huang, Si Xu, Tong Qiao, Yanghao Li, Jing Yang, Tong Yang, +2 more
Organizations: State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · Infinigence AI · Tsinghua University
As model sizes continue to scale, distributed training has become inevitable. Automatic parallelization techniques can derive efficient training parallelism strategies at low cost while achieving superior performance. The difficulty of this problem is jointly determined by the complexity of the model and the underlying compute cluster. Meanwhile, mixture-of-experts (MoE) models are increasingly emerging as the dominant architecture and the rapid evolution of accelerator hardware has made cluster heterogeneity commonplace, posing substantial challenges to automatic parallelization. However, existing approaches typically target either MoE architectures or heterogeneous clusters, failing to generalize to scenarios where both challenges coexist. To this end, we present HAPMoE, a heterogeneity-aware automatic parallelism planner for MoE training. HAPMoE builds a lightweight MoE-aware cost model and efficiently searches a six-dimensional parallel space, producing parallel plans directly deployable on Megatron-LM. Experiments show that HAPMoE improves end-to-end training throughput by up to 3.2× over baselines across heterogeneous clusters. Its non-uniform pipeline partitioning yields an additional up to 78% gains, and its pruning-enhanced dynamic programming algorithm completes the search within 1 minute, demonstrating high efficiency and practical value in complex hardware environments.
Figures & tables
Figure 1: Overview of HAPMoE . HAPMoE follows a profile-model-search pipeline: it first performs lightweight profiling to build device/primitive lookup tables, then constructs MoE-aware latency/communication/memory models under heterogeneity, and finally searches the 6D parallel space with pruning to output a Megatron-LM-deployable plan.
Cluster Scale
Hardware Configuration
Symbol
16 (Homo.)
H800 (2 × 8)
16-1-h
910B (2 × 8)
16-2-h
MI300X (2 × 8)
16-3-h
16 (Heter.)
H800 (1 × 8)+910B (1 × 8)
16-1
H800 (1 × 8)+MI300X (1 × 8)
16-2
910B (1 × 8)+MI300X (1 × 8)
16-3
Table 1: Symbol definitions of clusters
Figure 2: Performance of dense models on heterogeneous clusters: Throughput and MFU.
Figure 3: Performance of MoE model on homogeneous clusters: Throughput and MFU.
Figure 4: Performance of MoE model on heterogeneous clusters: Throughput, MFU and Latency.
Model
Method
Thpt. ↑
MFU ↑
Lat. ↓
Mixtral-S
MI
1.00 ×
1.00 ×
1.00 ×
HeterMoE-style
1.40 ×
1.46 ×
0.72 ×
Metis-style
1.45 ×
1.50 ×
0.69 ×
HAPMoE
1.67 ×
1.72 ×
0.56 ×
Mixtral-L
MI
1.00 ×
1.00 ×
1.00 ×
HeterMoE-style
1.47 ×
1.50 ×
0.73 ×
Table 2: Geometric mean performance on heterogeneous MoE clusters, normalized to MI. Higher is better for throughput and MFU, lower is better for latency.
Figure 5: Search time and accuracy of HAPMoE
Cluster
16-1/2/3-h
16-2
32-1/2/3
Latency
1.12/1.18/1.10 ×
2.10 ×
2.30/1.35/2.80 ×
Thpt.
0.95/0.93/0.96 ×
0.69 ×
0.62/0.85/0.56 ×
Table 3: Performance change ratio by disabling non-uniform PP/DP
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A: Illustration of the 1F1B pipeline schedule. Forward and backward passes are interleaved across micro-batches, and steady-state throughput is governed by the bottleneck stage.
Cluster
Model
Latency (ms)
Error (%)
Profile
Training
Pred.
Meas.
H800 (4 × 8)
H800 (16 × 8)
M1
620
658
−5.78
MI300X (4 × 8)
MI300X (16 × 8)
M1
760
718
5.85
H800 (2 × 8)+MI300X (2 × 8)
H800 (8 × 8)+MI300X (8 × 8)
M1
780
847
−7.91
H800 (2 × 8)+910B (2 × 16)
H800 (4 × 8)+910B (2 × 16)
M1
1150
1336
−13.92
H800 (4 × 8)
H800 (16 × 8)
M2
700
637
9.89
Appendix
Table A: Predicted vs. measured latency of extrapolating 4-node profiling to larger-scale execution