Mixture-of-Experts Inference

Latest papers 109

All topics
CardsList
  1. UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing

    Jun 2, 2026Xinming Wei, Chao Jin, Tuo Dai +10Expert ParallelismExpert Load Balancing

  2. Beyond Task-Agnostic: Task-Aware Grouping for Communication-Efficient Multi-Task MoE Inference

    May 31, 2026Zhiyao Xu, Aoxue Liu, Zhanjie Ding +3Mixture-of-Experts InferenceMixture-of-Experts Models

  3. ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving

    May 30, 2026Seokjin Go, Marko Scrbak, Ephrem Wu +2Expert Load BalancingMixture-of-Experts Inference

  4. Composing Non-Conjugate Factor Graphs with Closed-Form Variational Inference

    May 28, 2026Mykola Lukashchuk, Kyrylo Yemets, Wouter M. Kouw +4Factor Graph OptimizationMixture-of-Experts Inference

  5. How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving

    May 27, 2026Hanjiang Wu, Abhimanyu Rajeshkumar Bambhaniya, Sarbartha Banerjee +9LLM ServingLLM Inference

  6. ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference

    May 26, 2026Xiongwei Zhu, Xiaojian Liao, Tianyang Jiang +3LLM Inference EfficiencyExpert Offloading

  7. Dense2MoE: Pushing the Pareto Frontier of On-Device LLMs via Unified Pruning and Upcycling

    May 26, 2026Fengfa Li, Hongjin Ji, Yifeng Ding +2Mixture-of-Experts PruningLLM Pruning

  8. DepthAgent: Towards Better Universal Depth Estimation via Sample-wise Expert Selection

    May 22, 2026Jie Zhu, Girish Chandar Ganesan, Xiaoming LiuMetric Depth EstimationMonocular Depth Estimation

  9. TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload

    May 19, 2026Zhiben Chen, Youpeng Zhao, Yang Sui +2Expert OffloadingLLM Inference Scheduling

  10. GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems

    May 19, 2026Sourish Wawdhane, Avinash Kumar, Poulami DasExpert Load BalancingMixture-of-Experts Inference

  11. FPED: A Functional-Network Prior-Guided Mixture-of-Experts Framework for Interpretable Brain Decoding

    May 19, 2026Yudan Ren, Pengcheng Shi, Zihan Ma +2Neural DecodingMixture-of-Experts Inference

  12. Post-Trained MoE Can Skip Half Experts via Self-Distillation

    May 18, 2026Xingtai Lv, Li Sheng, Kaiyan Zhang +12Mixture of ExpertsMixture-of-Experts Inference

  13. BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE

    May 14, 2026Juntong Wu, Jialiang Cheng, Qishen Yin +5Sparse Mixture-of-ExpertsMixture-of-Experts Inference

  14. ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems

    May 12, 2026Wenyong Zhou, Yuannuo Feng, Yizhe Chen +6Compute-in-MemoryMixture-of-Experts Language Models

  15. Fast MoE Inference via Predictive Prefetching and Expert Replication

    May 12, 2026Ankit Jyothish, Ali Jannesari, Aishwarya Sarkar +1Expert Load BalancingGPU Acceleration

  16. TRACE: Temporal Routing with Autoregressive Cross-channel Experts for EEG Representation Learning

    May 12, 2026Fan Ma, Qier An, Peng Chen +6Representation LearningSelf-Supervised Pre-Training

  17. DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices

    May 11, 2026Chenyang Song, Weilin Zhao, Xu Han +3Sparse Mixture-of-ExpertsMixture-of-Experts Inference

  18. TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D Tiling

    May 10, 2026Hongyaoxing Gu, Xinzhe Chen, Lijuan Hu +1Efficient InferenceMixture-of-Experts Inference

  19. Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution

    May 9, 2026Jongseok Park, Sunga Kim, Zhenyu Gu +2Activation SparsityMixture-of-Experts Inference

  20. EMO: Pretraining Mixture of Experts for Emergent Modularity

    May 7, 2026Ryan Wang, Akshita Bhagia, Sewon MinLanguage Model PretrainingMixture-of-Experts Language Models

  21. VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading

    May 7, 2026Cheng Xu, Xiaofeng Hou, Jiacheng Liu +1Expert OffloadingVision-Language Models

  22. Adaptive Inverted-Index Routing for Granular Mixtures-of-Experts

    May 6, 2026Klaus-Rudolf Kladny, Maximilian Mordig, Bernhard Schölkopf +1Mixture-of-Experts InferenceExpert Routing

  23. MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving

    May 3, 2026Zhaoyuan Su, Olatunji Ruwase, Karthik Ganesan +5Expert ParallelismLLM Serving