cs.DCMar 6, 2026

MoEless: Efficient MoE LLM Serving with Serverless Experts

Authors: Hanfei Yu, Bei Ouyang, Shwai He, Ang Li, Hao Wang

Organizations: University of Maryland College Park

Abstract

Large Language Models (LLMs) increasingly adopt Mixture-of-Experts (MoE) architectures to scale efficiently under stringent resource constraints. However, MoE's sparse activation causes severe expert load imbalance, where a few experts become stragglers while others remain underutilized, leading to inflated inference latency and cost. Existing solutions assume static, serverful model deployments, limiting expert elasticity and often incurring costly expert swapping or degraded output quality. We present MoEless, an efficient serverless MoE serving framework that mitigates expert load imbalance via elastic expert execution. MoEless leverages lightweight, layer-aware predictors to estimate incoming expert load distributions and proactively identify stragglers. We design optimized scaling and placement strategies to improve function locality, GPU utilization, and cross-expert load balance. MoEless is prototyped on top of Megatron-LM and deployed on an eight-GPU testbed. Experiments with open-source MoE models and real-world workloads show that MoEless reduces inference latency by 43% and inference cost by 84% compared to state-of-the-art solutions.

Figures & tables

Explore similar work

CardsList
  1. Fast MoE Inference via Predictive Prefetching and Expert Replication

    May 12, 2026Ankit Jyothish, Ali Jannesari, Aishwarya Sarkar +1Mixture-Of-Expert InferenceLLM Inference Optimization

  2. Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns

    Apr 25, 2026Abhimanyu Bambhaniya, Geonhwa Jeong, Jason Park +6Mixture-Of-Expert InferenceMixture-Of-Experts

  3. A Spatio-Temporal Expert Prefetching Framework for Efficient MoE-based LLM Inference

    Jun 13, 2026Yingnan Zhao, Razvan Bunescu, Ahmed Louri +2Mixture-Of-Expert InferenceLLM Inference Optimization