Expert Parallelism

Momentum

3 papers in the last four weeks, against 1 the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 15

All topics
CardsList
  1. Efficient Expert-Parallel Communication on PCIe-Connected Consumer GPUs

    Sep 30, 2026Jaehwan Lee, Sangmin Lee, Chaewon Kim +2Expert ParallelismMixture-of-Experts Inference

  2. HAPMoE: Heterogeneity-Aware Automatic Parallelism Planning for Mixture-of-Experts Models Training

    Sep 30, 2026Mengyuan Fan, Peizhuang Cong, Zixiao Huang +7Expert ParallelismHeterogeneous Computing

  3. TopoEP: Topology-Aware Load Balancing for Expert-Parallel MoE Training

    Sep 28, 2026Jiacheng Zhu, Xie Zhao, Gongming Zhao +3Expert ParallelismExpert Load Balancing

  4. Tevatron Meets Megatron: Expert-Parallel LLM Reranker Training on an Academic Budget

    Aug 2, 2026Zhichao Xu, Xueguang Ma, Shengyao Zhuang +5Expert ParallelismCross-Encoder Reranking

  5. UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods

    Jul 7, 2026Yipeng Liu, Chang Liu, Si Shen +16Expert ParallelismHigh-Performance Computing

  6. FoMoE: Breaking the Full-Replica Barrier with a Federation of MoEs

    Jun 17, 2026Lorenzo Sani, Zeyu Cao, Meghdad Kurmanji +5Expert ParallelismCommunication-Efficient Distributed Training

  7. Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement

    Jun 13, 2026Qianli Liu, Kaibin Guo, Zicong Hong +5Expert ParallelismLLM Serving

  8. Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design

    Jun 9, 2026Wenxin Wang, Yule Hou, Yu Ji +2Expert ParallelismOn-Device Language Model Inference

  9. UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing

    Jun 2, 2026Xinming Wei, Chao Jin, Tuo Dai +10Expert ParallelismExpert Load Balancing

  10. DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism

    May 10, 2026Zhichen Zeng, Chi-Chih Chang, Jiayi Wang +10Expert ParallelismCommunication-Efficient Distributed Training

  11. MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving

    May 3, 2026Zhaoyuan Su, Olatunji Ruwase, Karthik Ganesan +5Expert ParallelismLLM Serving