cs.CLOct 15, 2025

NOSA: Native and Offloadable Sparse Attention

Authors: Yuxiang Huang, Pengjie Wang, Jicheng Han, Weilin Zhao, Zhou Su, Ao Sun, Hongya Lyu, Hengyu Zhao, +4 more

Organizations: NLP Group, DCST, IAI, BNRIST, Tsinghua University, Beijing, China. · OpenBMB, China. · BUPT, Beijing, China. · School of Computer Science and Technology, Beijing Institute of Technology.

Abstract

Decoding throughput improvements from larger inference batches are limited by GPU memory, which is largely consumed by the key-value (KV) cache. Prior training-free KV cache offloading alleviates this by keeping redundant context on the CPU and fetching only a sparse subset for attention, but it often degrades long-generation quality due to training-inference mismatch on sparse patterns. Meanwhile, trainable sparse attention is incompatible with efficient offloading, as unconstrained KV accesses may force large CPU-to-GPU transfers and erase throughput gains. To this end, we propose NOSA, a trainable sparse attention mechanism natively designed for KV cache offloading. NOSA explicitly constrains the volume of CPU-GPU KV transfers, thereby achieving low communication overhead and high decoding throughput. We further build NOSI, a KV cache offloading inference system that fully unlocks NOSA's efficiency. Empirical results on 1,3,8B LLMs demonstrate that NOSA outperforms KV cache offloading baselines on general, long-input, and long-generation tasks, while boosting decoding throughput by up to 5.04x, 1.92x, and 1.83x over FullAttn, InfLLMv2, and ShadowKV, respectively. We release our code at https://github.com/thunlp/NOSA.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PulseInfer: I/O-Centric Sparse KV Cache Offloading for Efficient Long-Context LLM Decoding

    Sep 28, 2026Qiuyang Zhang, Kai Zhou, Kai Lu +6Kv-Cache ManagementLarge Language Model Decoding

  2. SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference

    Jun 3, 2026Yaosheng Fu, Guangxuan Xiao, Xin Dong +2Dynamic Sparse AttentionLLM Inference Optimization

  3. Stochastic Sparse Attention for Memory-Bound Inference

    May 3, 2026Kyle Lee, Corentin Delacour, Kevin Callahan-Coray +5Dynamic Sparse AttentionAutoregressive Decoding