Long-Context Language Model Inference

Latest papers 166

All topics
CardsList
  1. KVEraser: Learning to Steer KV Cache for Efficient Localized Context Erasing

    Jun 15, 2026Mufei Li, Shikun Liu, Dongqi Fu +5KV CachingLong-Context Language Model Inference

  2. Doc-to-Atom: Learning to Compile and Compose Memory Atoms

    Jun 10, 2026Xingjian Diao, Wenbo Li, Yashas Malur Saidutta +3Memory-Augmented Language ModelsLong-Context Language Model Inference

  3. Express Language Modeling

    Jun 9, 2026Albert Gong, Annabelle Michael Carrell, Raaz Dwivedi +1Long-Context Language Model InferenceAttention Mechanisms

  4. Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models

    Jun 9, 2026Jing Xiong, Qi Han, Shansan Gong +5Long-Context Language Model InferenceDiffusion Language Model Inference

  5. End-to-End Context Compression at Scale

    Jun 8, 2026Ang Li, Sean McLeish, Haozhe Chen +12LLM CompressionLong-Context Language Model Inference

  6. From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs

    Jun 8, 2026Zhanchao Xu, Haoyang Li, Qingfa Xiao +4Long-Context Language Model InferenceLLM Inference Acceleration

  7. FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

    Jun 8, 2026Yan Wang, Qifan Zhang, Jiachen Yu +12KV CachingLong-Context Language Model Inference

  8. Still: Amortized KV Cache Compaction in a Single Forward Pass

    Jun 5, 2026Charles O'Neill, Alex Sandomirsky, Harry Partridge +2KV CachingLong-Context Language Model Inference

  9. You Only Index Once: Cross-Layer Sparse Attention with Shared Routing

    Jun 4, 2026Yutao Sun, Yanqi Zhang, Li Dong +2Long-Context Language Model InferenceKV-Cache Management

  10. Dense Contexts Are Hard Contexts: Lexical Density Limits Effective Context in LLMs

    Jun 4, 2026Giovanni Dettori, Matteo Boffa, Danilo Giordano +2Long-Context RetrievalLLM Evaluation

  11. SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference

    Jun 3, 2026Yaosheng Fu, Guangxuan Xiao, Xin Dong +2KV-Cache OffloadingLong-Context Language Model Inference

  12. LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional Encoding

    Jun 3, 2026Haocheng Xia, Mihir Pamnani, Hanxi Fang +2Retrieval-Augmented GenerationKV Caching

  13. MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference

    Jun 1, 2026Yu Li, Binxu Li, Tian LanKV CachingLong-Context Language Model Inference

  14. Soft-NBCE: Entropy-Weighted Chunk Fusion for Long-Context

    May 31, 2026Shihao Ji, Mingyu Li, Zihui SongLong-Context Language Model InferenceLLM Inference Acceleration

  15. WaveFilter: Enhancing the Long-Context Capability of Diffusion LLMs via Wavelet-Guided KV Cache Filtering

    May 30, 2026Jinnan Yang, Yan Wang, Zhen Bi +5Long-Context Language Model InferenceDiffusion Language Model Inference

  16. BudgetDraft: Acceptance-Aware Multi-View Training for Sparse-KV Speculative Decoding

    May 29, 2026Liang He, Jingbo Wen, Qishi Zhan +4KV CachingLong-Context Language Model Inference

  17. Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference

    May 25, 2026Sangyun Lee, Sean McLeish, Tom Goldstein +1Memory-Augmented Language ModelsLong-Context Language Model Inference

  18. IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference

    May 25, 2026Xintong Yang, Hao Gu, Binxing Xu +6Memory-Augmented Language ModelsSelf-Attention

  19. H2^{2}MT: Semantic Hierarchy-Aware Hierarchical Memory Transformer

    May 24, 2026Maryam Haghifam, Zifan He, Jason Cong +1Memory-Augmented Language ModelsLong-Context Language Model Inference

  20. Adaptive Mass-Segmented KV Compression for Long-Context Reasoning

    May 22, 2026Junzhe Yang, Xiaoyu ShenKV CachingLong-Context Language Model Inference

  21. EntmaxKV: Support-Aware Decoding for Entmax Attention

    May 20, 2026Gonçalo Duarte, Miguel Couceiro, Marcos V. TrevisoSelf-AttentionKV Caching

  22. Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs

    May 19, 2026Haiquan Lu, Zigeng Chen, Gongfan Fang +2LLM QuantizationLong-Context Language Model Inference

  23. Context Memorization for Efficient Long Context Generation

    May 18, 2026Yasuyuki Okoshi, Hao Mark Chen, Guanxi Lu +3Memory-Augmented Language ModelsLong-Context Language Model Inference

  24. VeriCache: Turning Lossy KV Cache into Lossless LLM Inference

    May 17, 2026Jiayi Yao, Samuel Shen, Kuntai Du +7KV-Cache OffloadingKV Caching