Dynamic routing creates severe load imbalance in large-scale expert-parallel Mixture-of-Experts (MoE) training, turning GPUs that host hot experts into stragglers. As each MoE layer waits for its slowest rank, these stragglers prolong the expert-parallel stage and reduce overall training efficiency. Existing expert-parallelism load-balancing (EPLB) systems commonly compute load-balancing plans on the CPU, incurring device--host data transfers and cross-rank synchronization that make scheduling at every layer and microbatch expensive. Their planning formulations also overlook the hierarchical communication costs of modern scale-up and scale-out GPU clusters. We present \textit{TopoEP}, a GPU-native, topology-aware load-balancing system for large-scale MoE training. At each MoE layer and training microbatch, \textit{TopoEP} converts the current routing result into hot-expert replication and token-rerouting decisions and executes the resulting plan without data-dependent host synchronization, reducing critical-path overhead. To generate these decisions, \textit{TopoEP} uses a deterministic GPU solver that performs inter-node placement followed by intra-node refinement, allowing all ranks to independently produce bitwise-identical plans. On a 32-GPU NVIDIA H800 cluster, integrating \textit{TopoEP} with Megatron-LM improves end-to-end training throughput by 6.2%--11.4% across three representative MoE models.
Figures & tables
Figure 1: Execution flow of an expert-parallel MoE layer.
Figure 2: Expert-load skew across the 48 MoE layers of Qwen3-30B-A3B on DAPO-Math and StarCoderData. (a)–(b) Fraction of each layer’s token–expert assignments routed to each expert (horizontal axis: logical expert ID 0–127; vertical axis: MoE layer 1–48). (c) Per-layer maximum-to-mean expert-load ratio. (d) Fraction of the eight highest-load experts that changes between adjacent microbatches.
Figure 3: Inter-node communication for two replica placements. Each node contributes N=4096 tokens to expert E . Both placements use one replica and have the same maximum per-rank load, Lmax=N . (a) A replica on Node 1 incurs 64 MiB of token traffic (solid arrow). (b) A replica on Node 0 incurs 18 MiB of parameter traffic and 9 MiB of replica-gradient traffic (dashed arrow).
Figure 4: Critical-path execution of host-side and GPU-native EPLB planning. GPU-native planning removes the host round trip between routing and dependent dispatch operations.
Figure 5: Architecture and execution flow of TopoEP .
Figure 6: Device-initiated replica transfer using intra-node TMA and inter-node NCCL GIN: forward pulls expert parameters into replica slots, and backward returns replica gradients to their main ranks.
Figure 7: Two-chunk forward pipeline overlapping expert-parameter transfers and token communication with expert FFN computation.
Figure 8: Two-chunk backward pipeline overlapping expert-parameter re-pulls, replica-gradient pushes, and token communication with expert FFN backward computation.
Figure 9: Example of two-stage expert replication and token rerouting. Inter-node placement removes cross-domain token traffic, after which intra-node refinement balances rank loads through same-domain rerouting over NVLink.
Model
MoE layers
Experts
Top- k
Hidden
Interm. size
PP / EP
Qwen3-30B-A3B [ 31 ]
5
128
8
2,048
768
1/32
GLM-4.5-Air [ 36 ]
5
128
8
4,096
1,408
2/16
DeepSeek-V2 [ 3 ]
4
160
6
5,120
1,536
2/16
Table 1: MoE model configurations used in the evaluation.
Figure 10: Rank-load imbalance versus additional replica slots per rank for a 32-rank, 640-expert Top-8 workload. Lower is better, and the dashed line denotes ideal balance.
Figure 11: Mean solver-kernel latency with varying logical EP sizes and expert counts, measured over 200 executions after 20 warm-up iterations.
Figure 12: Average end-to-end training throughput under natural and synthetic routing. The synthetic settings vary router_skew from 0 to −4 . Stars mark interpolated throughput crossovers, and double-headed arrows show the TopoEP /Megatron-LM throughput ratio at router_skew =−4 .
Figure 13: Comparison with baselines at router_skew =−4 . (a) Average end-to-end throughput. (b) Average rank-load imbalance, defined as the maximum per-rank token-expert load divided by the average across ranks, where 1 denotes ideal balance.
Figure 14: Qwen3-30B-A3B training loss during 10,000-step runs under natural routing.
Figure 15: Per-layer MoE latency breakdown, with expert-computation time measured on straggler ranks.
Figure 16: Physical execution of token–expert assignments. (a) Fraction whose source and execution ranks are on different nodes. (b) Fraction executed by replicas.