Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication. Consequently, state-of-the-art inference systems require high-bandwidth, P2P interconnects (e.g., NVLink) in datacenter GPUs to handle massive token routing, making deployment prohibitively expensive. Consumer GPUs offer comparable compute power at significantly lower cost, promising to democratize MoE inference for individuals and enable privacy-preserving local deployments. However, their bandwidth-limited (only weak PCIe bus bandwidth) and host-mediated interconnects (no P2P support) introduce severe communication bottlenecks. We present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing. Our key insight is that the unique communication topology provides the opportunity to elevate the host to an active routing hub, which can fundamentally reduce communication volume and eliminate global synchronization-induced stalls. Specifically, for token dispatch, we introduce host-backed token multicast to write shared tokens to the host exactly once, eliminating outbound transmission redundancy. For token combine, we propose a fine-grained, token-level aggregation mechanism using host staging buffers, which replaces rigid global synchronization and mitigates straggler effects. Evaluation on RTX 5090 GPUs shows that CoMoE improves inference throughput by up to 1.46x, approaching the performance of NVLink-capable A800 GPUs at only 23.4% of the hardware cost.
Figures & tables
Figure 1. Comparison of communication modes between datacenter and consumer-grade GPUs.
Figure 2. Example of token routing in Expert Parallelism.
Specification
A800-SXM
RTX 5090
Ratio
FP16 TFLOPs
312
419
1.34 ×
BF16 TFLOPs
312
209.5
0.67 ×
Memory (GB)
80
32
0.4 ×
Memory BW (TB/s)
2.0
1.8
0.9 ×
Interconnect
NVLink
PCIe 5.0
–
Inter-GPU BW (GB/s)
400
64 †
0.16 ×
Table 1. Comparison of datacenter GPUs and consumer-grade GPUs. Consumer GPUs offer comparable compute power but limited interconnect bandwidth.
Figure 3. Time to first token (TTFT) and time per output token (TPOT) comparison of SGLang on 8 NVIDIA A800 (datacenter) and RTX 5090 (consumer) GPUs.
Figure 4. Time breakdown of a prefill iteration. Comm: communication; Comp: computation.
Figure 5. (a) Average token-to-GPU dispatch fanout across models. See detailed model configuration in Table 2 . (b) Stall latency caused by coarse-grained synchronization.
Figure 6. Comparison of two dispatch schemes. (a) Existing system SGLang ( Zheng et al., 2024 ) exhibits transmission redundancy when dispatching tokens via All-to-All primitives. (b) CoMoE achieves zero-redundancy via host-backed token multicast.
Figure 7. Comparison of traditional All-to-All and CoMoE ’s decomposed combine.
Figure 8. Fused transfer kernel for the combine phase. Assume this kernel runs on GPU 1: sender blocks 1/3/5 send tokens to the corresponding staging buffers on GPUs 0/2/3, and receiver blocks 0/2/4 read from the staging buffers of GPUs 0/2/3.
Figure 9. Pipelined reduction. We assume that token T0/1/2/4 processed during inference on GPU 1 are routed to experts on GPU 0, while T1/2/3/4 are routed to experts on GPU 2.
Model
Size
Hid. Dim.
Exp.
Active Exp.
DeepSeek-V2-Lite ( DeepSeek-AI, 2024 )
16B
2048
64
2(S) + 6
GPT-OSS-20B ( OpenAI, 2025 )
20B
2880
32
4
Qwen3-30B-A3B ( Team, 2025 )
30B
2048
128
8
GLM-4.7-Flash ( GLM, 2025 )
31B
2048
64
1(S) + 4
Table 2. MoE models. Size: total parameters; Hid. Dim.: hidden size; Exp.: routed experts per MoE layer; Active Exp.: experts activated per token.
Figure 10. Time breakdown of a prefill iteration.
Figure 11. Dispatch time comparison across different models.
Figure 12. Performance comparison of the combine phase with and without global synchronization.
Figure 13. CoMoE combine latency across ablations.
Figure 14. TTFT across different request rates.
Figure 15. TPOT across different request rates.
Figure 16. Throughput across different models.
Metric
Qwen3 30B-A3B
GLM-4.7 Flash
DeepSeek V2-Lite
GPT-OSS 20B
CoCo Thpt.
5,506.00
3,367.13
7,386.28
7,304.18
A800 Thpt.
6,200.60
3,739.40
8,656.30
8,353.00
CoCo Tokens/$
0.344
0.211
0.462
0.457
A800 Tokens/$
0.078
0.047
0.108
0.104
Tokens/$ Ratio
4.44 ×
4.50 ×
4.27 ×
4.37 ×
Table 3. Cost-effectiveness of CoCo on RTX 5090 versus A800 with DeepEP.
Figure 17. PDF and histogram of latencies of DeepResearch tasks.
The Mixture-of-Experts (MoE) architecture improves computational efficiency via sparse expert activation, but throughput-oriented inference faces substantial GPU memory pressure due to a significant parameter size and intermediate data. Prior works attempt to mitigate this using expert offloading with micro-batching or by offloading computation to the CPU. However, the fragmented workload resulting from micro-batching degrades operational intensity, causing expert execution to become memory-bound. Meanwhile, CPU offloading is constrained by slow PCIe transfers and its limited applicability to attention computation in the decode stage. Consequently, these inefficiencies prevent effective system utilization, severely restricting the end-to-end throughput of MoE inference. To address these challenges, this paper proposes CoX-MoE, an Advanced Matrix Extensions (AMX)-enabled CPU-GPU collaborative system that comprehensively optimizes MoE inference by combining coalesced expert execution with strategic workload orchestration for higher throughput. CoX-MoE introduces (i) a coalescing-aware orchestration policy to jointly optimize resource allocation by adopting ordinary batch, instead of micro-batch, for expert computation and selective attention offloading, and (ii) a static expert-aware stratification scheme that pre-assigns frequently activated experts to the GPU, mitigating PCIe transfer overhead and balancing workload for the CPU and GPU during inference. Compared to state-of-the-art frameworks, CoX-MoE delivers significant gains, achieving up to 7.1x and 2.4x higher throughput than FlexGen and MoE-Lightning, respectively.
Expert parallelism (EP) enables inference of large Mixture-of-Experts (MoE) models by placing their experts across multiple GPUs, but requires substantial communication between GPUs at every MoE layer. As contemporary MoE models activate more experts per token, this communication accounts for a growing fraction of inference time. The cost becomes particularly pronounced on PCIe-based consumer GPU systems, where all inter-GPU transfers traverse CPU memory. However, existing MoE-specialized EP communication libraries assume that direct GPU-to-GPU access is available, largely overlooking consumer GPUs. Therefore, most LLM frameworks instead rely on NCCL, whose CPU-staged communication incurs redundant PCIe transfers and competes with expert computation for GPU resources, limiting their overlap. We present ThunderEP, a novel communication design for such systems that removes the relay hops of traditional ring algorithm, moves data through DMA engines to avoid compute resource contention, and minimizes synchronization latency by reducing the polling overhead of completion flags in CPU memory. We integrate the proposed design into vLLM and evaluate it on three widely used MoE models. Experiments on two PCIe systems equipped with RTX 4090 and RTX 5090 GPUs show that ThunderEP achieves average speedups of 2.00× and 1.53× over NCCL for dispatch and combine, respectively, and up to 1.66× end-to-end speedup over state-of-the-art MoE inference frameworks.
Jaehwan Lee, Sangmin Lee, Chaewon Kim +2
Dept. of Computer Science and Engineering Seoul National University · Dept. of Computer Science and Engineering Graduate School of Data Science Seoul National University
Mixture-of-Experts (MoE) architectures increase model capacity without proportionally increasing computation cost and have become a key building block for scaling large language models (LLMs) to trillion-parameter regimes. Efficient deployment of these MoE models relies on distributed execution across multiple GPUs, where each MoE layer involves two all-to-all communications: dispatching tokens to expert ranks and returning the expert outputs to their source ranks. Conventional MoE implementations launch this return all-to-all after expert compute completes, exposing communication latency on the critical path and reducing GPU utilization. We present a fine-grained approach that overlaps expert compute with the second all-to-all via tile-level signaling and scheduling. Our producer-consumer co-design combines: (1) a persistent per-rank computation kernel (producer) that covers all local experts on the rank to eliminate repeated kernel launch overhead and prioritizes remote-critical tiles, and (2) a persistent communication kernel (consumer) on a small dedicated partition of streaming multiprocessors (SMs) that issues segment-granular transfers as tiles become ready. Our co-design avoids intrusive changes to the underlying computation operators or communication primitives, making it practical for improving distributed MoE execution efficiency on multi-GPU systems. On a 4-A100 GPU platform, evaluated on three MoE models against four state-of-the-art MoE systems, our approach achieves up to 2.64x end-to-end speedup and 2.74x MoE-layer speedup. Compared with a conventional non-overlap baseline, our approach consistently improves both operator- and MoE-layer-level performance across varying GEMM shapes, router modes, and a broad range of producer/consumer SM partitions, while preserving correctness.