Large Language Models (LLMs) increasingly adopt Mixture-of-Experts (MoE) architectures to scale efficiently under stringent resource constraints. However, MoE's sparse activation causes severe expert load imbalance, where a few experts become stragglers while others remain underutilized, leading to inflated inference latency and cost. Existing solutions assume static, serverful model deployments, limiting expert elasticity and often incurring costly expert swapping or degraded output quality. We present MoEless, an efficient serverless MoE serving framework that mitigates expert load imbalance via elastic expert execution. MoEless leverages lightweight, layer-aware predictors to estimate incoming expert load distributions and proactively identify stragglers. We design optimized scaling and placement strategies to improve function locality, GPU utilization, and cross-expert load balance. MoEless is prototyped on top of Megatron-LM and deployed on an eight-GPU testbed. Experiments with open-source MoE models and real-world workloads show that MoEless reduces inference latency by 43% and inference cost by 84% compared to state-of-the-art solutions.
Figures & tables
Figure 1 . Expert load imbalance across layers for different MoE models and datasets: (a) Mixtral-8 × 7B on ShareGPT and (b) Phi-3.5-MoE on LMSYS-Chat-1M.
Figure 2 . Illustration of serving Mixture-of-Experts (MoE) based Large Language Models under expert parallelism, where tokens are routed by per-layer gate networks to a sparse set of experts distributed across GPUs. Expert load imbalance triggers inefficient resource provisioning ( e.g. , over-scaling hot experts or under-utilizing cold ones), thereby increasing serving cost. Embed : embedding layer, TB : Transformer Block, Head : language modeling head, Attention : attention layer, Gate : gate networks, DP : data parallelism, MP : model parallelism.
Figure 3 . Serving Phi-3.5-MoE on LMSYS-Chat-1M using Azure LLM traces: (a) request arrivals, (b) aggregated token loads, and (c) total number of active experts.
Figure 4 . Per-layer parameter memory breakdown of three MoE models. Expert parameters dominate the per-layer memory footprint, accounting for 91.4%–97.1% of total layer memory.
Figure 5 . Inference performance of three approaches when serving Phi-3.5-MoE on ShareGPT.
Figure 6 . The architecture and workflow of MoEless .
Figure 7 . Characterizing Phi-3.5-MoE on LMSYS-Chat-1M: (a) cosine similarity of gate network inputs, and (b) expert load prediction accuracy across layers with different prediction distances.
Figure 8 . Expert load prediction accuracy on LMSYS-Chat-1M with and without fine-tuning at different prediction distances: (a) Mixtral-8 × 7B, and (b) Phi-3.5-MoE.
MoE Model
Parameters
Experts Per Layer
Num. of
(active / total)
(active / total)
Layers
Mixtral-8×7B
12.9B / 46.7B
2 / 8
32
Phi-3.5-MoE
6.6B / 42B
2 / 16
32
Llama-4-Scout
17B / 109B
1 / 16
48
Table 1 . Characterizations of MoE models used in the evaluation.
Figure 9 . Characterization of LMSYS-Chat-1M and ShareGPT datasets.
Figure 10 . Characterization of Azure LLM inference traces.
Figure 11 . MoE layer forward time of four approaches across three models on LMSYS-Chat-1M.
Figure 12 . MoE layer forward time of four approaches across three models on ShareGPT.
Figure 13 . Total inference cost of four approaches across three models on two datasets.
Figure 14 . Performance tradeoff of MoE layer forward latency speedup vs. expert memory cost.
Figure 15 . Expert load prediction accuracy for three prediction methods at different prediction distances.
Figure 16 . Correlations between predicted and actual expert load distributions across layers of two models. Heavier color means more correlation results fall in the slot.
Figure 17 . Sensitivity analysis of MoEless ’s expert prediction distance on LMSYS-Chat-1M.
Figure 18 . Sensitivity analysis of MoEless ’s expert prediction distance on ShareGPT.
Figure 19 . Sensitivity analysis of MoEless ’s expert load CV threshold on LMSYS-Chat-1M.
Figure 20 . Sensitivity analysis of MoEless ’s expert load CV threshold on ShareGPT.
Figure 21 . Ablation study of MoEless .
Model
Mixtral-offloading
ProMoE
Ours
Mixtral-8×7B
1.92 MB
128.32 MB
1.92 MB
Phi-3.5-MoE
4.16 MB
128.64 MB
4.16 MB
Llama-4-Scout
3.84 MB
120.48 MB
3.84 MB
Table 2 . Predictor memory footprints across different models and methods.