LLM Inference Efficiency

LLM: Large Language Model

Latest papers 234

All topics
CardsList
  1. PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving

    Sep 17, 2026Omkar Shewale, Deepak Kumar, Divakar Kumar YadavLLM Inference EfficiencyLLM Inference

  2. Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling

    Sep 16, 2026Mobina Kashaniyan, Ali JannesariLLM Inference EfficiencyLLM Inference

  3. Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It

    Sep 16, 2026Yipeng Liu, Yingqiang Zhang, Feifei Li +1LLM Inference EfficiencyKV Caching

  4. Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning

    Sep 16, 2026Zhuo Chen, Zhen Zhang, Xinyu Wang +1LLM Inference EfficiencyLLM Agent Workflow Optimization

  5. Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum

    Sep 16, 2026Carolina Fortuna, Vid Hanžel, Tim Strnad +1LLM Inference EfficiencyEnergy-Efficient ML

  6. Dissecting GPU Utilization for LLM Inference on Nvidia Hopper

    Sep 14, 2026Mohammad Siavashi, Gerald Q. Maguire Jr., Dejan Kostic +1LLM Inference EfficiencyGPU Acceleration

  7. GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three Systems

    Sep 14, 2026Xinyu Qiu, Chuhong Xu, Bo Su +3LLM Inference Efficiency

  8. TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps

    Sep 13, 2026Rohit Patel, Susil Kumar Mohanty, Jeenal ChaudharyLLM Inference EfficiencyRetrieval-Augmented Generation

  9. REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

    Sep 11, 2026Tuan Nguyen, Qiran Hu, Banruo Liu +3LLM Inference EfficiencyRetrieval-Augmented Generation

  10. From Fixed Keys to Readable Schemas: Small Language Models for Vehicle Agent Function Calls

    Sep 10, 2026Hamed Jafarzadeh Asl, Yuanhao Yu, Vahid Partovi NiaLLM Inference EfficiencyOn-Device Language Model Inference

  11. Composable CXL Memory as a Kubernetes-Native Shared Memory for LLM Serving

    Sep 9, 2026Hongjian Fan, Kevin Zhang, David Habinsky +1LLM Inference EfficiencyKV Caching

  12. A Measurement Study of LLM Inference Trade-offs Across Edge Continuum Hardware

    Sep 8, 2026Maysam Khatib, Moysis Symeonides, Demetris Trihinas +2LLM Inference EfficiencyLLM Evaluation

  13. MetaKV: Adaptive KV Cache Compression for Constrained LLM Inference

    Sep 7, 2026Michael Wang, Keith Li, Roozbeh BostandoostLLM Inference EfficiencyLLM Inference

  14. Signed Rescue Routing: Harm-Aware Cascades for Efficient LLM Inference

    Sep 7, 2026Zheyuan Wang, Siyu Li, Peiqiao Song +3LLM Inference EfficiencyLLM Routing

  15. LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference

    Sep 2, 2026Renyuan Liu, Yuyang Leng, Kaiyan Liu +7LLM Inference EfficiencyOn-Device Language Model Inference

  16. Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems

    Sep 2, 2026Jinxi Yu, Yubei Li, Eric Hanchen Jiang +6Multi-Agent LLM SystemsLLM Inference Efficiency

  17. How Do Prompt Variations Affect Energy Consumption in On-Device LLMs?

    Sep 1, 2026Wei Hu, Xiaolong Tu, Dawei Chen +3LLM Inference EfficiencyOn-Device Language Model Inference

  18. ContextPipe: Database-Inspired Context Assembly for Long-Horizon Agents

    Sep 1, 2026Peng Xu, Zuyu Zhang, Yuze Sun +3LLM Inference EfficiencyLong-Horizon LLM Agents