LLM Inference Efficiency

Recent momentum

-5%

61 papers in the last 28 days · 1.0% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

21 new papers

A weekly snapshot of new work published in LLM Inference Efficiency.

Period ending 2026-09-14

12 new papers

A weekly snapshot of new work published in LLM Inference Efficiency.

Period ending 2026-09-07

16 new papers

A weekly snapshot of new work published in LLM Inference Efficiency.

541 papers

Latest in LLM Inference Efficiency

Date pendingcs.LG

FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving

Mixture-of-Experts (MoE) models have become mainstream for scaling language models to hundreds of billions of expert parameters. Despite sparse expert activation, existing inference engines keep all experts GPU-resident, crowding out the key-value cache in large-batch, long-output offline workloads. We present FluxMoE, which decouples experts from physical GPU residency and adapts their footprint to available memory through a new \emph{expert paging} abstraction. FluxMoE combines PagedTensor for transparent remapping, a bandwidth-balanced hierarchy spanning losslessly compressed GPU memory and host DRAM, and a budget-aware residency planner. Unlike CPU-GPU co-inference and whole-layer offloading, FluxMoE streams weights on demand while keeping expert computation on GPUs. We implement FluxMoE atop vLLM and evaluate it on three MoE models. For GLM-4.5 on 8×\timesH20 GPUs, FluxMoE delivers up to 7.2×\times vLLM's throughput and 79.0% lower average Time-Per-Output-Token (TPOT), without measurable model-quality loss using lossless compression. For Mixtral-8×\times7B-Instruct on 2×\timesL40S GPUs, where weight-resident vLLM cannot fit, FluxMoE delivers 4.3×\times KTransformers's throughput and 29.1% lower average TPOT.
Qingxiu Liu, Yongchao He, Runhan Jiang +4