Expert Load Balancing

Momentum

3 papers in the last four weeks, against 2 the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 30

All topics
CardsList
  1. BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models

    Oct 5, 2026Gang Fu, Adel Javanmard, MohammadHossein Bateni +1Expert Load BalancingMixture-of-Experts Models

  2. CIPHER-MoE: Balancing Efficiency and Routing Fidelity in Trillion-Scale MoE Training

    Oct 5, 2026Jing Li, Jian Meng, Yingmeng Gao +8Expert Load BalancingExpert Routing

  3. ID Balancing: Stable Training of Extremely Sparse MoE via PID-Based Load Control

    Sep 30, 2026Peng Jin, Zihan Qiu, Zekun Wang +8Expert Load BalancingMixture-of-Experts Language Models

  4. TopoEP: Topology-Aware Load Balancing for Expert-Parallel MoE Training

    Sep 28, 2026Jiacheng Zhu, Xie Zhao, Gongming Zhao +3Expert ParallelismExpert Load Balancing

  5. EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference

    Aug 8, 2026Yize Wu, Ke Gao, Ling Li +1Expert Load BalancingMixture-of-Experts Inference

  6. REFLEX: Rethinking MoE Inference as Refinement-Aware Compute Allocation in Diffusion Language Models

    Aug 3, 2026Xiang Xia, Cheng Yan, Yiming Zhang +3Expert Load BalancingMixture-of-Experts Inference

  7. Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts

    Aug 1, 2026Ziang Wu, Peng Jin, Qishen Yin +4Expert Load BalancingVision-Language Models

  8. UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing

    Jun 2, 2026Xinming Wei, Chao Jin, Tuo Dai +10Expert ParallelismExpert Load Balancing

  9. ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving

    May 30, 2026Seokjin Go, Marko Scrbak, Ephrem Wu +2Expert Load BalancingMixture-of-Experts Inference

  10. GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems

    May 19, 2026Sourish Wawdhane, Avinash Kumar, Poulami DasExpert Load BalancingMixture-of-Experts Inference

  11. UB-SMoE: Universally Balanced Sparse Mixture-of-Experts for Resource-adaptive Federated Fine-tuning of Foundation Models

    May 15, 2026Van-Tuan Tran, Hong-Hanh Nguyen-Le, Marco Ruffini +1Expert Load BalancingSparse Mixture-of-Experts

  12. φφ-Balancing for Mixture-of-Experts Training

    May 14, 2026Lizhang Chen, Jonathan Li, Qi Wang +5Expert Load BalancingMirror Descent

  13. Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts

    May 12, 2026Sagi Ahrac, Noya Hochwald, Mor GevaExpert Load BalancingParameter-Free Mixture-of-Experts Routing

  14. Fast MoE Inference via Predictive Prefetching and Expert Replication

    May 12, 2026Ankit Jyothish, Ali Jannesari, Aishwarya Sarkar +1Expert Load BalancingGPU Acceleration

  15. ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning

    May 9, 2026Chao Jin, Xinming Wei, Yinmin Zhong +6RL for Language ModelsExpert Load Balancing

  16. Hierarchical Mixture-of-Experts with Two-Stage Optimization

    May 8, 2026Gleb Molodtsov, Alexander Miasnikov, Aleksandr BeznosikovExpert Load BalancingMixture of Experts

  17. E = T*H/(O+B): A Dimensionless Control Parameter for Mixture-of-Experts Ecology

    May 7, 2026Qingjun ZhangExpert Load BalancingMixture of Experts

  18. Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns

    Apr 25, 2026Abhimanyu Bambhaniya, Geonhwa Jeong, Jason Park +6Expert Load BalancingMixture-of-Experts Inference

  19. Mixture of Heterogeneous Grouped Experts for Language Modeling

    Apr 25, 2026Zhicheng Ma, Xiang Liu, Zhaoxiang Liu +5Expert Load BalancingMixture-of-Experts Language Models

  20. MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference

    Apr 19, 2026Bo Li, Chuan Wu, Shaolin ZhuEfficient Multimodal InferenceExpert Load Balancing

  21. MoEless: Efficient MoE LLM Serving with Serverless Experts

    Mar 6, 2026Hanfei Yu, Bei Ouyang, Shwai He +2LLM Inference EfficiencyExpert Load Balancing

  22. A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs

    Feb 23, 2026Zijie Liu, Jie Peng, Jinhao Duan +7Expert Load BalancingMixture of Experts