LLM Inference Efficiency

Recent momentum

-5%

61 papers in the last 28 days · 1.0% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

21 new papers

A weekly snapshot of new work published in LLM Inference Efficiency.

Period ending 2026-09-14

12 new papers

A weekly snapshot of new work published in LLM Inference Efficiency.

Period ending 2026-09-07

16 new papers

A weekly snapshot of new work published in LLM Inference Efficiency.

541 papers

Latest in LLM Inference Efficiency

  1. Attention Drift: What Autoregressive Speculative Decoding Models Learn

    May 11, 2026Doğaç Eldenk, Payal Mohapatra, Yigitcan Comlek +3Speculative DecodingDrafter

  2. Test-Time Speculation

    May 10, 2026Avinash Kumar, Sujay Sanghavi, Poulami DasSpeculative DecodingSelf-Speculative