Efficient Expert-Parallel Communication on PCIe-Connected Consumer GPUs
Authors: Jaehwan Lee, Sangmin Lee, Chaewon Kim, Junsik Shin, Jaejin Lee
Organizations: Dept. of Computer Science and Engineering Seoul National University · Dept. of Computer Science and Engineering Graduate School of Data Science Seoul National University
Expert parallelism (EP) enables inference of large Mixture-of-Experts (MoE) models by placing their experts across multiple GPUs, but requires substantial communication between GPUs at every MoE layer. As contemporary MoE models activate more experts per token, this communication accounts for a growing fraction of inference time. The cost becomes particularly pronounced on PCIe-based consumer GPU systems, where all inter-GPU transfers traverse CPU memory. However, existing MoE-specialized EP communication libraries assume that direct GPU-to-GPU access is available, largely overlooking consumer GPUs. Therefore, most LLM frameworks instead rely on NCCL, whose CPU-staged communication incurs redundant PCIe transfers and competes with expert computation for GPU resources, limiting their overlap. We present ThunderEP, a novel communication design for such systems that removes the relay hops of traditional ring algorithm, moves data through DMA engines to avoid compute resource contention, and minimizes synchronization latency by reducing the polling overhead of completion flags in CPU memory. We integrate the proposed design into vLLM and evaluate it on three widely used MoE models. Experiments on two PCIe systems equipped with RTX 4090 and RTX 5090 GPUs show that ThunderEP achieves average speedups of 2.00× and 1.53× over NCCL for dispatch and combine, respectively, and up to 1.66× end-to-end speedup over state-of-the-art MoE inference frameworks.
Figures & tables
Figure 1 . A single MoE layer (a) without and (b) with expert parallelism. I , E , O denote input tokens, experts, and outputs.
Figure 2 . Expert granularity of a MoE layer.
Figure 3 . An All-Gather and Reduce-Scatter communication pattern for token dispatch and combine in a MoE layer. Top-2 routing with 2 experts on each of 3 devices, where each token selects one expert on each remote device.
Figure 4 . Communication between GPUs in a node.
Figure 5 . All-Gather communication algorithm for dispatch in (a) NCCL and (b) ThunderEP.
Figure 6 . Proposed Reduce-Scatter algorithm for combine with (a) GPU reduction and (b) CPU reduction.
Figure 7 . DMA engine-based communication for token dispatch and combine in MoE with EP.
Figure 8 . (a) Serialized and (b) pipelined bidirectional chunked data transfer over a full-duplex PCIe link.
Figure 9 . Comparison of the synchronization protocols used by NCCL and the proposed method. (a) NCCL pairs every 4 bytes of data with a 4 byte completion flag. (b) ThunderEP stores completion flags separately from the data and uses a small number of threads to poll them. D denotes data, and F denotes a completion flag.
Figure 10 . Fine-grained synchronization policy for the Reduce-Scatter in combine, which mitigates the idle time when senders finish at different times because the router gives each device a different amount of expert work. (a) No receive begins before the last sender finishes. (b) Each begins on its own sender’s completion flag. GpSd is GPU p ’s contribution block for shard d , reduced on GPU d .
Figure 11 . All-Gather performance for token dispatch on a single node (6 GPUs).
Figure 12 . Reduce-Scatter performance for token combine on a single node (6 GPUs).
Systems
RTX 4090 system
RTX 5090 system
Mainboard
ASRock Rack GENOAD8X-2T
CPU
1 × AMD EPYC 9124 (16-Core)
Main memory
8 × DDR5-4800 32 GB
GPU
6 × RTX 4090
6 × RTX 5090
GPU memory
24 GB
32 GB
GPU interconnect
PCIe 4.0 × 16
PCIe 5.0 × 16
Table 1. System configurations used in our evaluation.
Figure 13 . Comparison of prefill performance across different input sequence lengths.
Figure 14 . Comparison of decode performance across different batch sizes.
Figure 15 . Latency breakdown of a Transformer layer.
Datasets
Synthetic
ShareGPT
LMSYS-Chat-1M
Prefill speedup
1.42 ×
1.38 ×
1.38 ×
Decode speedup
1.16 ×
1.13 ×
1.14 ×
Table 2. Speedup over vLLM with real-world prompts.
Figure 16 . Comparison of prefill performance across different input sequence lengths on the RTX 4090 system. We use a batch size of 24 and 256 output tokens, and vary the prompt length from 256 to 4K. We exclude GPT-OSS-120B since all baselines run GPU OOM.
Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication. Consequently, state-of-the-art inference systems require high-bandwidth, P2P interconnects (e.g., NVLink) in datacenter GPUs to handle massive token routing, making deployment prohibitively expensive. Consumer GPUs offer comparable compute power at significantly lower cost, promising to democratize MoE inference for individuals and enable privacy-preserving local deployments. However, their bandwidth-limited (only weak PCIe bus bandwidth) and host-mediated interconnects (no P2P support) introduce severe communication bottlenecks. We present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing. Our key insight is that the unique communication topology provides the opportunity to elevate the host to an active routing hub, which can fundamentally reduce communication volume and eliminate global synchronization-induced stalls. Specifically, for token dispatch, we introduce host-backed token multicast to write shared tokens to the host exactly once, eliminating outbound transmission redundancy. For token combine, we propose a fine-grained, token-level aggregation mechanism using host staging buffers, which replaces rigid global synchronization and mitigates straggler effects. Evaluation on RTX 5090 GPUs shows that CoMoE improves inference throughput by up to 1.46x, approaching the performance of NVLink-capable A800 GPUs at only 23.4% of the hardware cost.
Ruwen Fan, Yuezhi Zu, Junru Li +5
Jimmy · Alibaba Cloud Computing · Tsinghua University
Mixture-of-experts (MoE) architectures enable trillion-parameter LLMs with sparsely activated experts. Expert parallelism (EP) is a widely adopted MoE training strategy, but it suffers from severe all-to-all communication bottlenecks, which is exaggerated by the limited inter-node network bandwidth as the growing model size requires distributing experts across GPU nodes. Prior work focused on overlapping these all-to-all communications with feed-forward network (FFN) and self-attention computations, which often leaves residual network-bound stalls due to inherent imbalance in attention and FFN layers' computation-communication ratios. We present DisagMoE, a disaggregated MoE training system that jointly optimizes model placement and scheduling for maximal efficiency. DisagMoE separates attention and FFN layers into disjoint GPU groups, introduces a multi-stage pipeline with uni-directional, many-to-many communications, and employs a computation-communication roofline model to balance GPU and network bandwidth allocation among the attention and FFN groups. DisagMoE is implemented on Megatron-LM, and evaluation shows that DisagMoE improves training efficiency across multiple MoE models with up to 1.8x speedup on 16-node 8xH800 clusters.
Zhichen Zeng, Chi-Chih Chang, Jiayi Wang +10
ByteDance Seed · University of Washington · Cornell University
Mixture-of-Experts (MoE) architectures increase model capacity without proportionally increasing computation cost and have become a key building block for scaling large language models (LLMs) to trillion-parameter regimes. Efficient deployment of these MoE models relies on distributed execution across multiple GPUs, where each MoE layer involves two all-to-all communications: dispatching tokens to expert ranks and returning the expert outputs to their source ranks. Conventional MoE implementations launch this return all-to-all after expert compute completes, exposing communication latency on the critical path and reducing GPU utilization. We present a fine-grained approach that overlaps expert compute with the second all-to-all via tile-level signaling and scheduling. Our producer-consumer co-design combines: (1) a persistent per-rank computation kernel (producer) that covers all local experts on the rank to eliminate repeated kernel launch overhead and prioritizes remote-critical tiles, and (2) a persistent communication kernel (consumer) on a small dedicated partition of streaming multiprocessors (SMs) that issues segment-granular transfers as tiles become ready. Our co-design avoids intrusive changes to the underlying computation operators or communication primitives, making it practical for improving distributed MoE execution efficiency on multi-GPU systems. On a 4-A100 GPU platform, evaluated on three MoE models against four state-of-the-art MoE systems, our approach achieves up to 2.64x end-to-end speedup and 2.74x MoE-layer speedup. Compared with a conventional non-overlap baseline, our approach consistently improves both operator- and MoE-layer-level performance across varying GEMM shapes, router modes, and a broad range of producer/consumer SM partitions, while preserving correctness.