LLM Inference Efficiency

LLM: Large Language Model

Latest papers 234

All topics
CardsList
  1. Attention via Black-Box Vector Search

    Oct 7, 2026Stepan Zharkov, Krish Singal, Ashwin Padaki +1Efficient AttentionSparse Attention

  2. ReFold: Training-Free Reversible Inter-Turn Context Folding for Long-Horizon Agents

    Oct 6, 2026Yupeng Su, Jiayi Tian, Zheng Zhang +1LLM Inference EfficiencyLong-Context Language Model Inference

  3. Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can Tell

    Oct 6, 2026Hochan Son, Kyungdoe Han, Jaehan Koh +3LLM Inference EfficiencyMulti-Agent LLM Systems

  4. Monte Carlo Estimation for KV Cache Eviction

    Oct 6, 2026Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer +5LLM Inference EfficiencyKV-Cache Eviction

  5. Evaluating Inference Compute for Generative AI: A Framework for Enterprise Workloads

    Oct 5, 2026Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona ZahidLLM Inference EfficiencyDisaggregated LLM Serving

  6. Expanding LLM Reasoning

    Oct 4, 2026Rian Atri, Evan LuoLLM Inference EfficiencyInference-Time Scaling

  7. Beyond State-of-the-Art: Standardising Environmental Impact Metrics for AI Research

    Oct 1, 2026Lachlan McGinness, Dan Pagendam, Robert OffnerLLM Inference EfficiencyEnergy-Efficient ML

  8. Beyond Leaderboards: Tokenomics of Agentic Small Language Model Ensembles

    Oct 1, 2026Alexei N. Skurikhin, Emily M. Taylor, Nathan A. DeBardelebenLLM Inference EfficiencyLanguage Model Generation Evaluation

  9. Characterizing High Bandwidth Flash for LLM Serving

    Sep 30, 2026Zack Yu, Chloe Wong, Coleman Hooper +7LLM Inference EfficiencyLLM Serving

  10. Adaptive Self-Consistency: From Black-Box Sampling to Distribution-Valued Feedback

    Sep 30, 2026Jingkai Huang, Yunfan Zhang, Will Ma +2LLM Inference EfficiencyAdaptive Inference

  11. Purlin: Separating Orchestration from the Datapath of Collectives

    Sep 29, 2026Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos KozyrakisLLM Inference EfficiencyHigh-Performance Computing

  12. Beyond Energy: When Sustainability Dimensions Reshape LLM Serving Decisions

    Sep 28, 2026Tianyao Shi, Xipeng Shen, Yi DingLLM Inference EfficiencyLLM Serving

  13. SCBO: Semantically Coherent Batching and Ordering for LLM-Based Social Surveys

    Sep 28, 2026Yuanzi Li, Lingjie Wang, Zihang Tian +2LLM Inference EfficiencyLLM Prompting

  14. AgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMs

    Sep 28, 2026Cheuk Hang Lau, Zeyu Cao, Kevin Wong Cheuk Yin +6LLM Inference EfficiencyLLM Inference

  15. Spexis: Speculative Lookahead Scheduling for LLM Inference

    Sep 28, 2026Hyungyu Jung, Jaehyeok Yu, Hoonseo Choi +3LLM Inference EfficiencyLLM Inference Scheduling

  16. Scaling of Capability and Efficiency at Inference Time in Large Reasoning Models

    Sep 22, 2026Moritz Laber, Zohair Shafi, Germans Savcisens +6LLM Inference EfficiencyLanguage Model Scaling Laws

  17. REFLEX with Jev for Efficient Selective Control in LLM Agents

    Sep 22, 2026Tiantong Wu, Wei Yang Bryan LimLLM Inference EfficiencyLLM Agent Evaluation

  18. Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference

    Sep 22, 2026Md Mostafizer Rahman, Md Faizul Ibne Amin, Md Shahajada Mia +2LLM Inference EfficiencyLong-Context Language Model Inference

  19. Rollout Efficiency in Reinforcement Learning for Reasoning Large Language Models: A Taxonomy and Future Directions

    Sep 21, 2026Niloofar Gholipour, Marcos Assuncao, Gursimran Singh +8LLM Inference EfficiencyReinforcement Learning

  20. ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs

    Sep 20, 2026Junyoung Park, Jungwook Choi, Mingu LeeLLM Inference EfficiencyKV-Cache Eviction

  21. Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics

    Sep 17, 2026Wonmi Choi, Minuk Park, Zhixiong Niu +3LLM Inference EfficiencyLLM Agent Workflow Optimization

  22. PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving

    Sep 17, 2026Omkar Shewale, Deepak Kumar, Divakar Kumar YadavLLM Inference EfficiencyLLM Inference

  23. Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling

    Sep 16, 2026Mobina Kashaniyan, Ali JannesariLLM Inference EfficiencyLLM Inference

  24. Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It

    Sep 16, 2026Yipeng Liu, Yingqiang Zhang, Feifei Li +1LLM Inference EfficiencyKV Caching

  25. Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning

    Sep 16, 2026Zhuo Chen, Zhen Zhang, Xinyu Wang +1LLM Inference EfficiencyLLM Agent Workflow Optimization

  26. Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum

    Sep 16, 2026Carolina Fortuna, Vid Hanžel, Tim Strnad +1LLM Inference EfficiencyEnergy-Efficient ML

  27. Dissecting GPU Utilization for LLM Inference on Nvidia Hopper

    Sep 14, 2026Mohammad Siavashi, Gerald Q. Maguire Jr., Dejan Kostic +1LLM Inference EfficiencyGPU Acceleration

  28. GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three Systems

    Sep 14, 2026Xinyu Qiu, Chuhong Xu, Bo Su +3LLM Inference Efficiency

  29. TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps

    Sep 13, 2026Rohit Patel, Susil Kumar Mohanty, Jeenal ChaudharyLLM Inference EfficiencyRetrieval-Augmented Generation

  30. REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

    Sep 11, 2026Tuan Nguyen, Qiran Hu, Banruo Liu +3LLM Inference EfficiencyRetrieval-Augmented Generation

  31. From Fixed Keys to Readable Schemas: Small Language Models for Vehicle Agent Function Calls

    Sep 10, 2026Hamed Jafarzadeh Asl, Yuanhao Yu, Vahid Partovi NiaLLM Inference EfficiencyOn-Device Language Model Inference

  32. Composable CXL Memory as a Kubernetes-Native Shared Memory for LLM Serving

    Sep 9, 2026Hongjian Fan, Kevin Zhang, David Habinsky +1LLM Inference EfficiencyKV Caching

  33. A Measurement Study of LLM Inference Trade-offs Across Edge Continuum Hardware

    Sep 8, 2026Maysam Khatib, Moysis Symeonides, Demetris Trihinas +2LLM Inference EfficiencyLLM Evaluation

  34. MetaKV: Adaptive KV Cache Compression for Constrained LLM Inference

    Sep 7, 2026Michael Wang, Keith Li, Roozbeh BostandoostLLM Inference EfficiencyLLM Inference

  35. Signed Rescue Routing: Harm-Aware Cascades for Efficient LLM Inference

    Sep 7, 2026Zheyuan Wang, Siyu Li, Peiqiao Song +3LLM Inference EfficiencyLLM Routing

  36. LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference

    Sep 2, 2026Renyuan Liu, Yuyang Leng, Kaiyan Liu +7LLM Inference EfficiencyOn-Device Language Model Inference

  37. Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems

    Sep 2, 2026Jinxi Yu, Yubei Li, Eric Hanchen Jiang +6Multi-Agent LLM SystemsLLM Inference Efficiency

  38. How Do Prompt Variations Affect Energy Consumption in On-Device LLMs?

    Sep 1, 2026Wei Hu, Xiaolong Tu, Dawei Chen +3LLM Inference EfficiencyOn-Device Language Model Inference

  39. ContextPipe: Database-Inspired Context Assembly for Long-Horizon Agents

    Sep 1, 2026Peng Xu, Zuyu Zhang, Yuze Sun +3LLM Inference EfficiencyLong-Horizon LLM Agents