Mixture-of-Experts Inference

Latest papers 109

All topics
CardsList
  1. SpaceMoE: Realizing Distributed Mixture-of-Experts Inference over Space Networks

    May 1, 2026Zhanwei Wang, Huiling Yang, Min Sheng +2Mixture-of-Experts Language ModelsMixture-of-Experts Inference

  2. Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding

    May 1, 2026Lehan Pan, Ziyang Tao, Ruoyu Pang +3Mixture-of-Experts InferenceSpeculative Decoding

  3. Rethinking Network Topologies for Cost-Effective Mixture-of-Experts LLM Serving

    Apr 30, 2026Junsun Choi, Sam Son, Sunjin Choi +5LLM ServingCost-Aware Inference

  4. FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving

    Apr 29, 2026Minghe Wang, Trever Schirmer, Mohammadreza Malekabbasi +1Mixture-of-Experts Language ModelsLLM Serving

  5. Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns

    Apr 25, 2026Abhimanyu Bambhaniya, Geonhwa Jeong, Jason Park +6Expert Load BalancingMixture-of-Experts Inference

  6. Mixture of Heterogeneous Grouped Experts for Language Modeling

    Apr 25, 2026Zhicheng Ma, Xiang Liu, Zhaoxiang Liu +5Expert Load BalancingMixture-of-Experts Language Models

  7. Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs

    Apr 20, 2026Afsara Benazir, Felix Xiaozhu LinAI Accelerator InferenceEnergy-Efficient ML

  8. MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference

    Apr 19, 2026Bo Li, Chuan Wu, Shaolin ZhuEfficient Multimodal InferenceExpert Load Balancing

  9. MoEless: Efficient MoE LLM Serving with Serverless Experts

    Mar 6, 2026Hanfei Yu, Bei Ouyang, Shwai He +2LLM Inference EfficiencyExpert Load Balancing

  10. A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs

    Feb 23, 2026Zijie Liu, Jie Peng, Jinhao Duan +7Expert Load BalancingMixture of Experts

  11. OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale

    Feb 5, 2026Jingze Shi, Zhangyang Peng, Yizhang Zhu +3Mixture-of-Experts InferenceLLM Inference Acceleration

  12. 3D-MoE: Towards Spatial Intelligence with Mixture-of-Experts for 3D Reasoning and Action Generation

    Jan 28, 2025Yueen Ma, Zenglin Xu, Irwin King3D Spatial ReasoningMixture-of-Experts Inference

  13. Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference

    Date pendingKexin Chu, Dawei Xiang, Zixu Shen +3Mixed-Precision QuantizationGPU Acceleration

  14. FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving

    Date pendingQingxiu Liu, Yongchao He, Runhan Jiang +4Expert OffloadingGPU Acceleration

  15. Cache-Aware Joint Router Adaptation for Memory-Efficient MoE Inference

    Date pendingZhenhe Wu, Yaping Jin, Qinghua Xing +6Mixture of ExpertsMemory-Efficient Optimization