LLM Inference Efficiency

Recent momentum

-56%

33 papers in the last 28 days · 0.9% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-14

12 new papers

A weekly snapshot of new work published in LLM Inference Efficiency.

Period ending 2026-09-07

16 new papers

A weekly snapshot of new work published in LLM Inference Efficiency.

513 papers

Latest in LLM Inference Efficiency

Open your feed →
CardsList
  1. Learning to Evict from Key-Value Cache

    Feb 10, 2026Luca Moschella, Laura Manduchi, Ozan SenerKey-Value Cache EvictionKey-Value Cache

  2. Controllably Efficient Language Models

    Nov 7, 2025Jatin Prakash, Aahlad Puli, Rajesh RanganathLLM Inference EfficiencyProtein Language Model