Mixture-of-Experts Inference

Latest papers 109

All topics
CardsList
  1. Democratizing MoE inference on commodity GPUs with CoMoE

    Oct 7, 2026Ruwen Fan, Yuezhi Zu, Junru Li +5Mixture-of-Experts Inference

  2. How Sparse Probability Maps Shape Mixture-of-Experts Routing

    Oct 5, 2026Tomás Brogueira, Marcos Treviso, Miguel CouceiroMixture-of-Experts Language ModelsMixture-of-Experts Inference

  3. ITC-MoE: Importance-guided Token-aware Compression for MoE Diffusion Language Models

    Oct 1, 2026Lianjun Liu, Shipeng Li, You Huang +5Mixture-of-Experts Language ModelsLLM Compression

  4. Efficient Expert-Parallel Communication on PCIe-Connected Consumer GPUs

    Sep 30, 2026Jaehwan Lee, Sangmin Lee, Chaewon Kim +2Expert ParallelismMixture-of-Experts Inference

  5. MoRE: Scaling mixture of experts with hardware-aware low-rank routing

    Sep 28, 2026Honam Wong, Surbhi Goel, Enric Boix-AdseràParameter-Free Mixture-of-Experts RoutingLow-Rank Matrix Decomposition

  6. MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference

    Sep 28, 2026Junfeng Wu, Zehao Fan, Hadjer Benmeziane +3Expert OffloadingMixture-of-Experts Language Models

  7. EEG-Fusion: Failure-Informed Source-Free Expert Routing for Robust Motor Imagery EEG Decoding

    Sep 27, 2026Abdul Basit, Saim Rehman, Muhammad ShafiqueSource-Free Domain AdaptationMI Classification

  8. EAT: Expert Account Tracker for Efficient MoE Inference

    Sep 27, 2026Yuexian Li, Yifei Yang, Zouying Cao +1Mixture-of-Experts PruningMixture-of-Experts Language Models

  9. OLED-MoE: Accelerating MoE-Based dLLM Inference via Inter-Iteration Locality-Aware Expert Offloading

    Sep 27, 2026Jingyuan Xiao, Jiayue Wang, Yitao Hu +7Expert OffloadingMixture of Experts

  10. You Only Need 2/3 of the Chosen Experts: An Empirical Study of Dynamic Expert Pruning in Fine-Grained MoE LLMs

    Sep 22, 2026Yuanteng Chen, Qiwei Lai, Chen Tianqi +7Mixture-of-Experts PruningMixture-of-Experts Language Models

  11. Evidence-Aligned Local Composition of Discrete Experts for Sequence Restoration

    Sep 14, 2026Mohammad Panahazari, Usman A. Khan, Shuchin AeronMixture-of-Experts InferenceMixture-of-Experts Models

  12. Quality-Constrained Routing over a Fixed Pool of Quantized Mixture-of-Experts Instances

    Sep 14, 2026Zhenghong Huang, Hongfan Wu, Jiheng ZhangMixture-of-Experts InferenceMixture-of-Experts Quantization

  13. SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading

    Sep 14, 2026Zihan Wang, Yuqi Wang, Lei Gong +5Expert OffloadingLLM Inference Scheduling

  14. PCoMoE: Shifting MoE Inference from Monolithic Expert Selection to Fine-Grained Path Composition

    Sep 1, 2026Ziyan Gan, Fangxin Liu, Chenyang Guan +10Mixture-of-Experts PruningMixture-of-Experts Inference

  15. Candidate-Expanding Routing with Permutation-Stabilized Experts for Mixed-Format Medical VQA

    Sep 1, 2026Hai-Dang Nguyen, Huy-Hieu PhamMemory-Augmented VLMsVisual Question Answering

  16. DynaNDE: Dynamic Near-Data Expert Scheduling for Batched MoE Inference

    Aug 31, 2026Xiaoyang Lu, Belthangady Akash Vi Narayana Pai, Xian-He SunEfficient Neural Network InferenceLLM Inference Scheduling

  17. APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference

    Aug 12, 2026Alish Kanani, Layan Badawi, Umit Y. OgrasEdge InferenceMixture-of-Experts Inference

  18. Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation

    Aug 11, 2026Gongli Zhang, Zhulin Liu, C. L. Philip ChenMixture of ExpertsMixture-of-Experts Inference

  19. UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models

    Aug 9, 2026Lei Xin, Bin Gu, Peize Li +8Mixture of ExpertsMixture-of-Experts Inference

  20. EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference

    Aug 8, 2026Yize Wu, Ke Gao, Ling Li +1Expert Load BalancingMixture-of-Experts Inference