HAPMoE: Heterogeneity-Aware Automatic Parallelism Planning for Mixture-of-Experts Models Training
Authors: Mengyuan Fan, Peizhuang Cong, Zixiao Huang, Si Xu, Tong Qiao, Yanghao Li, Jing Yang, Tong Yang, +2 more
Organizations: State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · Infinigence AI · Tsinghua University
As model sizes continue to scale, distributed training has become inevitable. Automatic parallelization techniques can derive efficient training parallelism strategies at low cost while achieving superior performance. The difficulty of this problem is jointly determined by the complexity of the model and the underlying compute cluster. Meanwhile, mixture-of-experts (MoE) models are increasingly emerging as the dominant architecture and the rapid evolution of accelerator hardware has made cluster heterogeneity commonplace, posing substantial challenges to automatic parallelization. However, existing approaches typically target either MoE architectures or heterogeneous clusters, failing to generalize to scenarios where both challenges coexist. To this end, we present HAPMoE, a heterogeneity-aware automatic parallelism planner for MoE training. HAPMoE builds a lightweight MoE-aware cost model and efficiently searches a six-dimensional parallel space, producing parallel plans directly deployable on Megatron-LM. Experiments show that HAPMoE improves end-to-end training throughput by up to 3.2× over baselines across heterogeneous clusters. Its non-uniform pipeline partitioning yields an additional up to 78% gains, and its pruning-enhanced dynamic programming algorithm completes the search within 1 minute, demonstrating high efficiency and practical value in complex hardware environments.
Figures & tables
Figure 1: Overview of HAPMoE . HAPMoE follows a profile-model-search pipeline: it first performs lightweight profiling to build device/primitive lookup tables, then constructs MoE-aware latency/communication/memory models under heterogeneity, and finally searches the 6D parallel space with pruning to output a Megatron-LM-deployable plan.
Cluster Scale
Hardware Configuration
Symbol
16 (Homo.)
H800 (2 × 8)
16-1-h
910B (2 × 8)
16-2-h
MI300X (2 × 8)
16-3-h
16 (Heter.)
H800 (1 × 8)+910B (1 × 8)
16-1
H800 (1 × 8)+MI300X (1 × 8)
16-2
910B (1 × 8)+MI300X (1 × 8)
16-3
Table 1: Symbol definitions of clusters
Figure 2: Performance of dense models on heterogeneous clusters: Throughput and MFU.
Figure 3: Performance of MoE model on homogeneous clusters: Throughput and MFU.
Figure 4: Performance of MoE model on heterogeneous clusters: Throughput, MFU and Latency.
Model
Method
Thpt. ↑
MFU ↑
Lat. ↓
Mixtral-S
MI
1.00 ×
1.00 ×
1.00 ×
HeterMoE-style
1.40 ×
1.46 ×
0.72 ×
Metis-style
1.45 ×
1.50 ×
0.69 ×
HAPMoE
1.67 ×
1.72 ×
0.56 ×
Mixtral-L
MI
1.00 ×
1.00 ×
1.00 ×
HeterMoE-style
1.47 ×
1.50 ×
0.73 ×
Table 2: Geometric mean performance on heterogeneous MoE clusters, normalized to MI. Higher is better for throughput and MFU, lower is better for latency.
Figure 5: Search time and accuracy of HAPMoE
Cluster
16-1/2/3-h
16-2
32-1/2/3
Latency
1.12/1.18/1.10 ×
2.10 ×
2.30/1.35/2.80 ×
Thpt.
0.95/0.93/0.96 ×
0.69 ×
0.62/0.85/0.56 ×
Table 3: Performance change ratio by disabling non-uniform PP/DP
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A: Illustration of the 1F1B pipeline schedule. Forward and backward passes are interleaved across micro-batches, and steady-state throughput is governed by the bottleneck stage.
Cluster
Model
Latency (ms)
Error (%)
Profile
Training
Pred.
Meas.
H800 (4 × 8)
H800 (16 × 8)
M1
620
658
−5.78
MI300X (4 × 8)
MI300X (16 × 8)
M1
760
718
5.85
H800 (2 × 8)+MI300X (2 × 8)
H800 (8 × 8)+MI300X (8 × 8)
M1
780
847
−7.91
H800 (2 × 8)+910B (2 × 16)
H800 (4 × 8)+910B (2 × 16)
M1
1150
1336
−13.92
H800 (4 × 8)
H800 (16 × 8)
M2
700
637
9.89
Appendix
Table A: Predicted vs. measured latency of extrapolating 4-node profiling to larger-scale execution
Frontier models increasingly adopt Mixture-of-Experts (MoE) architectures to achieve large-model performance at reduced cost. However, training MoE models on HPC platforms is hindered by large memory footprints, frequent large-scale communication across heterogeneous networks, and severe workload imbalance. To characterize these challenges, we develop a mathematical model that quantifies memory, compute, and communication requirements for MoE configurations under various parallelization schemes, verified through micro-benchmarking, code instrumentation, and hardware profiling. Our analysis identifies performance bottlenecks: all-to-all latency at scale from expert parallelism, insufficient compute-communication overlap, low GPU utilization from imbalanced skinny GEMMs, and the absence of platform-aware hybrid parallelization strategies. To address these, we introduce Piper, a framework that leverages resource modeling to identify efficient training strategies for MoE models on target HPC platforms, applying pipeline parallelism with optimized schedules. Piper achieves 2-3.5X higher MFU than state-of-the-art frameworks such as X-MoE, and a novel all-to-all algorithm delivers 1.2-9X bandwidth over vendor implementation.
This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models. It is a training paradigm that combines and specializes various existing and novel parallelism techniques at different layers and stages of the Mixture-of-Experts (MoE) model training pipeline. It leverages these techniques to achieve maximal efficiency given the physical constraints of CPU, CPU memory, GPU HBM memory, and the CPU-GPU, GPU-GPU, and node-node communication bandwidth of the GPU cluster. It also contains a novel strategy for the optimizer step to achieve high throughput and memory efficiency, enabling practitioners to conduct lossless pre-training/fine-tuning of trillion-parameter scale models, at a million context length, with just under 12 8x H200 GPU nodes, with state-of-the-art throughput and memory efficiency. In our experiments, MoP delivers 4.7x--8.2x higher per-GPU throughput than a strongly-tuned FSDP2 baseline (with the gap widening at larger scale) and sustains training at context lengths up to 1M tokens, where the baseline runs out of memory beyond 64--128K.
Mixture-of-experts (MoE) architectures enable trillion-parameter LLMs with sparsely activated experts. Expert parallelism (EP) is a widely adopted MoE training strategy, but it suffers from severe all-to-all communication bottlenecks, which is exaggerated by the limited inter-node network bandwidth as the growing model size requires distributing experts across GPU nodes. Prior work focused on overlapping these all-to-all communications with feed-forward network (FFN) and self-attention computations, which often leaves residual network-bound stalls due to inherent imbalance in attention and FFN layers' computation-communication ratios. We present DisagMoE, a disaggregated MoE training system that jointly optimizes model placement and scheduling for maximal efficiency. DisagMoE separates attention and FFN layers into disjoint GPU groups, introduces a multi-stage pipeline with uni-directional, many-to-many communications, and employs a computation-communication roofline model to balance GPU and network bandwidth allocation among the attention and FFN groups. DisagMoE is implemented on Megatron-LM, and evaluation shows that DisagMoE improves training efficiency across multiple MoE models with up to 1.8x speedup on 16-node 8xH800 clusters.
Zhichen Zeng, Chi-Chih Chang, Jiayi Wang +10
ByteDance Seed · University of Washington · Cornell University