Memory-Efficient Inference

Latest papers 101

All topics
CardsList
  1. Right In-Place (RiP) Convolution: A Simple, General, and Near-Optimal Strategy for Memory-Efficient CNN Inference

    Sep 30, 2026Opegbemi Matthias Busoye, Tolulope Matthew Busoye, Eghonghon-aye EigbeMemory-Efficient InferenceConvolutional Neural Networks

  2. UltraMatch: Transport Path Routing for Ultra-Fast and Memory-Efficient Image Matching

    Sep 29, 2026Jiajun Le, Yifan Lu, Zizhuo Li +3Image MatchingMemory-Efficient Inference

  3. SANTA++: Sampling Attention through Representative Keys

    Sep 28, 2026Kyle Lee, Christian Z. Pratt, Ruoyu Fang +4Memory-Efficient InferenceLLM Inference Acceleration

  4. DuplexCadence: Exact State and Execution from a Speech Model's Declared Timelines

    Sep 28, 2026Haixiao Gao, Yimin Zheng, Linyou Xiao +1Memory-Efficient InferenceFull-Duplex Spoken Dialogue Systems

  5. DPS: Dual-Mode Precision LLM Serving with Semi-Unified Memory

    Sep 28, 2026Xuan Truong Nguyen, Tien Son Pham, Tuan Duc Chu +2LLM ServingLLM Inference

  6. MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference

    Sep 28, 2026Junfeng Wu, Zehao Fan, Hadjer Benmeziane +3Expert OffloadingMixture-of-Experts Language Models

  7. PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction

    Sep 28, 2026Hyesung Jeon, Hyeongju Ha, Jae-Joon KimMulti-Agent LLM SystemsMemory-Efficient Inference

  8. MILO: Efficient Many-shot In-Context Learning with Block-wise Low-rank Compression

    Sep 24, 2026Youpeng Zhao, Tian Tan, Liqian Peng +2Memory-Efficient InferenceLLM Inference Acceleration

  9. Risk-Controlled KV-Cache Eviction: From Memory Budgets to Risk Targets

    Sep 23, 2026Beomgu Kang, SoJin Yun, Hojoon Kim +1KV-Cache EvictionMemory-Efficient Inference

  10. Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference

    Sep 22, 2026Md Mostafizer Rahman, Md Faizul Ibne Amin, Md Shahajada Mia +2LLM Inference EfficiencyLong-Context Language Model Inference

  11. ARM: Attention with Routed-Memory for Learnable Sparse Control

    Sep 21, 2026Qiuhao Zeng, Jerry Huang, Peng Lu +7Memory-Augmented Language ModelsLong-Context Language Model Inference

  12. TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching

    Sep 18, 2026Zhihao Shu, Md Musfiqur Rahman Sanim, Jie Hu +4On-Device Language Model InferenceLong-Context Language Model Inference

  13. Jacap: Robust KV Cache Eviction via Jacobian-Based Nonlinear Information Capacity Preservation

    Sep 8, 2026Jiaming Yang, Chenwei Tang, Liangli Zhen +2KV CachingKV-Cache Eviction

  14. LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference

    Sep 2, 2026Renyuan Liu, Yuyang Leng, Kaiyan Liu +7LLM Inference EfficiencyOn-Device Language Model Inference

  15. GaLe: memory-efficient Global Approximate and Local Exact features

    Sep 2, 2026Alberto Ancilotto, Elisabetta FarellaMemory-Efficient InferenceEdge AI

  16. mzCache: On-Device LLM Memory Management under Multitasking

    Sep 1, 2026Hongseung Yu, Minsung Kim, Jongseok Park +1On-Device Language Model InferenceMemory-Efficient Inference

  17. DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving

    Aug 31, 2026Yanqi Yu, Pingwei Sun, Jianchao Tan +4KV CachingMemory-Efficient Inference

  18. DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization

    Aug 27, 2026Tao Zhang, Jianchao Tan, Pingwei Sun +6Mixed-Precision QuantizationLLM Quantization

  19. APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference

    Aug 12, 2026Alish Kanani, Layan Badawi, Umit Y. OgrasEdge InferenceMixture-of-Experts Inference

  20. MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices

    Aug 11, 2026Eunjeong Kim, Yeong Jun Jeon, Myeonggyun HanOn-Device Language Model InferenceLLM Inference Scheduling

  21. CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents

    Aug 8, 2026Weizhong Huang, Jinchao Zhang, Xiawu ZhengKV CachingLLM Agent Memory

  22. AnchorKV: Anchor-Residual KV Cache Compression

    Aug 3, 2026Malik Khalaf, Yara Shamshoum, Nitzan Hodos +2KV CachingLong-Context Language Model Inference