LLM Inference Efficiency

Recent momentum

-5%

61 papers in the last 28 days · 1.0% of indexed attention

Twelve weeks of publication activity for this topic as it is defined today.

Weekly history

Recent digests

What was published in this topic, kept on the site without email delivery.

Period ending 2026-09-21

21 new papers

A weekly snapshot of new work published in LLM Inference Efficiency.

Period ending 2026-09-14

12 new papers

A weekly snapshot of new work published in LLM Inference Efficiency.

Period ending 2026-09-07

16 new papers

A weekly snapshot of new work published in LLM Inference Efficiency.

541 papers

Latest in LLM Inference Efficiency

  1. ELiTeFormer: An Efficient Transformer for FPGAs

    Jul 4, 2026Victor Agostinelli, Nicolas Bohm Agostini, Antonino TumeoField-Programmable Gate ArraysLLM Inference Efficiency

  2. WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs

    Jul 2, 2026Mauricio Fadel Argerich, Jonathan Fürst, Marta Patiño-MartínezLLM Inference EfficiencyInference Workloads

  3. Closing the Operational Gap in Semantic Caching

    Jun 18, 2026Aditeya Baral, Radoslav Ralev, Iliya Sotirov Zhechev +2CacheSemantic Gap