LLM Inference Efficiency

Recent momentum

-5%

61 papers in the last 28 days · 1.0% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

21 new papers

A weekly snapshot of new work published in LLM Inference Efficiency.

Period ending 2026-09-14

12 new papers

A weekly snapshot of new work published in LLM Inference Efficiency.

Period ending 2026-09-07

16 new papers

A weekly snapshot of new work published in LLM Inference Efficiency.

541 papers

Latest in LLM Inference Efficiency

  1. MiniMax Sparse Attention

    Jun 11, 2026Xunhao Lai, Weiqi Xu, Yufeng Yang +14Dynamic Sparse AttentionLLM Inference Efficiency

  2. UniSVQ: 2-bit Unified Scalar-Vector Quantization

    Jun 9, 2026Haoyu Wang, Haiyan Zhao, Xingyu Yu +4QuantizationLow-Bit

  3. LLM Explainability with Counterfactual Chains and Causal Graphs

    Jun 4, 2026Nirit Nussbaum-Hoffer, Nitay Calderon, Liat Ein-Dor +1Causal GraphExplainability