Mixture-of-Experts Inference

Latest papers 109

All topics
CardsList
  1. Elbow-Based MoE Routing: A Training-Free Inference Time Plugin for Expert Selection

    Aug 5, 2026Robin Pan, Raymond Liu, Daniel Fang +2Mixture-of-Experts InferenceLLM Inference Acceleration

  2. Uncertainty Is Not Enough: Value-of-Information Routing for Mixtures of LoRA Experts

    Aug 3, 2026Tom Saliencro, Rohan Desai, Priya Nair +2Adaptive Model RoutingSelective Prediction

  3. REFLEX: Rethinking MoE Inference as Refinement-Aware Compute Allocation in Diffusion Language Models

    Aug 3, 2026Xiang Xia, Cheng Yan, Yiming Zhang +3Expert Load BalancingMixture-of-Experts Inference

  4. TopoCompress: Topology Aware Token Compression Algorithm for Distributed Edge MoE Inference

    Aug 1, 2026Ning Li, Xinyu Wang, Xin Yuan +3Edge InferenceMixture-of-Experts Inference

  5. HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference

    Aug 1, 2026Xin Yuan, Ning Li, Wenchao Xu +2Edge InferenceCost-Aware Inference

  6. TrimMoE A communication aware and adaptive depth framework for distributed edge inference

    Aug 1, 2026Ning Li, Shuting Bai, Xin Yuan +3Adaptive InferenceEdge Inference

  7. Incast-Free MoE Rate-Based Scheduling

    Jul 28, 2026Evyatar Cohen, Jose Yallouz, Alexander Shpiner +3LLM Inference SchedulingMixture-of-Experts Inference

  8. OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis

    Jul 27, 2026Zihan Li, Feiyang Liu, Dandan Shan +2Test-Time AdaptationMedical Image Analysis

  9. Context-Adaptive Inference: A Unified Statistical and Foundation-Model View

    Jul 25, 2026Yue Yao, Caleb N. Ellington, Jingyun Jia +9Adaptive InferenceMeta-Learning

  10. OrderMoE: An expert similarity driven distributed edge MoE inference

    Jul 19, 2026Xin Yuan, Ning Li, Quan Chen +2Edge InferenceMixture-of-Experts Inference

  11. ThAME: 3D Memory-Enabled Heterogeneous Accelerator for LLM Mixture of Experts

    Jul 19, 2026Pratyush Dhingra, Pramit Kumar Pal, Janardhan Rao Doppa +1AI Accelerator InferenceMixture-of-Experts Inference

  12. Less Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Experts

    Jul 14, 2026Jincheng Xie, Runheng Liu, Heyan Huang +4Mixture-of-Experts InferenceSpeculative Decoding

  13. HCRMap: Pressure-Aware Hot-Expert Residency Mapping for 3.5D MoE Chiplet Inference

    Jul 13, 2026Yongqin ZhangAI Accelerator InferenceMixture-of-Experts Inference

  14. UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods

    Jul 7, 2026Yipeng Liu, Chang Liu, Si Shen +16Expert ParallelismHigh-Performance Computing

  15. Stable Global Weighting of Flow Mixtures using Simplex Exponential Moving Average

    Jul 4, 2026Benjamin Wiriyapong, Oktay Karakus, Can Eyupoglu +1Normalizing FlowsMixture-of-Experts Inference

  16. Separating Expert Retention from Autonomous Source Inference in Raw-ECG-Replay-Free Continual ECG Deployment

    Jul 2, 2026Yufan Lu, Xinhui Liu, Chenyang Xu +2Continual LearningDomain Generalization

  17. Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study

    Jun 19, 2026Alfarizy Alfarizy, Hung Truong Thanh Nguyen, René Richard +2LLM Inference EfficiencyOn-Device Language Model Inference

  18. CogniRoute: Learning to Route Social Evidence in Omni-Modal Models

    Jun 18, 2026Yifan Shen, Pei Tian, Xinzhuo Li +8RL for Language Model ReasoningMixture-of-Experts Inference

  19. A Spatio-Temporal Expert Prefetching Framework for Efficient MoE-based LLM Inference

    Jun 13, 2026Yingnan Zhao, Razvan Bunescu, Ahmed Louri +2AI Accelerator InferenceMixture-of-Experts Inference

  20. Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement

    Jun 13, 2026Qianli Liu, Kaibin Guo, Zicong Hong +5Expert ParallelismLLM Serving

  21. Decoupled Mixture-of-Experts for Parametric Knowledge Injection

    Jun 12, 2026Baoqing Yue, Weihang Su, Qingyao Ai +5Mixture-of-Experts Language ModelsKnowledge Augmentation for Language Models

  22. Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design

    Jun 9, 2026Wenxin Wang, Yule Hou, Yu Ji +2Expert ParallelismOn-Device Language Model Inference