Long-Context LLM Inference

LLM: Large Language Model

Latest papers 8

All topics
CardsList
  1. Strong Drafts Need Compact Memories: Long-Context Speculative Decoding with Compressed KV Cache

    Aug 31, 2026Tong Yuan, Chengxi Liao, Zeyi WenMemory-Augmented Language ModelsKV Caching

  2. KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems

    Jul 27, 2026Shuo Wang, Fang Xi, Wenyuan Huang +2LLM ServingKV-Cache Management

  3. CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference

    Jun 23, 2026Xiaolin Lin, Jingcun Wang, Olga Kondrateva +3LLM Inference EfficiencyKV Caching

  4. Inference Time Context Sparsity: Illusion or Opportunity?

    May 22, 2026Sahil Joshi, Prithvi Dixit, Agniva Chowdhury +5LLM Inference EfficiencyStructured Sparsity

  5. KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference

    May 18, 2026Jian Lin, Jiazhi Mi, Zicong Hong +5KV-Cache OffloadingKV Caching

  6. Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs

    Jan 5, 2026Amirali Ebrahimzadeh, Seyyed M. SaliliLong-Context RetrievalLLM Hallucination Mitigation