cs.CLOct 7, 2026

Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling

Authors: Radha Gulhane, Quentin Anthony, Beren Millidge

Organizations: Zyphra

Abstract

Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts. Under expert parallelism (EP) the experts are distributed across GPUs, and every MoE layer runs all-to-all collectives in the forward and backward passes to dispatch tokens to their experts and then combine the results. On a cluster with 8 AMD Instinct MI300X GPUs per node, these collectives can take 45% of the training step at EP32 with top-2 routing and 60% with top-6 routing. We find that early in pretraining routers have already learned to assign tokens to experts in correlated patterns, both within a layer and across layers. At top-2, 0.8% of the expert pairs in a layer are selected together by 42% of tokens, and the experts a token selects at one layer predict the experts it selects at the next layer. We use these correlations to keep more token--expert assignments on the token's own GPU, which reduces communication across GPUs and across nodes. Correlated expert placement puts experts that are often selected together on the same GPU. Combined with a dispatcher that sends each token to each GPU once, it removes up to 58% of dispatched rows. Token shuffling applies when sequence parallelism shards tokens across the EP group. It moves each token to the GPU predicted to hold its next-layer experts during the reduce-scatter that follows attention. On one node this raises the share of token--expert assignments served on the token's GPU from 12.5% to 59%. In Megatron-LM, across EP degrees from 8 to 64 with top-2 and top-6 routing, the two methods reduce all-to-all time by 1.16-2.63X and end-to-end step time by up to 1.41X. Neither method changes the models' underlying routing decisions or expert parameters.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism

    May 10, 2026Zhichen Zeng, Chi-Chih Chang, Jiayi Wang +10Mixture-Of-Expert InferenceMixture-Of-Experts

  2. Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns

    Apr 25, 2026Abhimanyu Bambhaniya, Geonhwa Jeong, Jason Park +6Mixture-Of-Expert InferenceMixture-Of-Experts

  3. Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation

    Aug 11, 2026Gongli Zhang, Zhulin Liu, C. L. Philip ChenMixture-Of-ExpertsUnified Framework