NOSA: Native and Offloadable Sparse Attention
Organizations: NLP Group, DCST, IAI, BNRIST, Tsinghua University, Beijing, China. · OpenBMB, China. · BUPT, Beijing, China. · School of Computer Science and Technology, Beijing Institute of Technology.
Abstract
Decoding throughput improvements from larger inference batches are limited by GPU memory, which is largely consumed by the key-value (KV) cache. Prior training-free KV cache offloading alleviates this by keeping redundant context on the CPU and fetching only a sparse subset for attention, but it often degrades long-generation quality due to training-inference mismatch on sparse patterns. Meanwhile, trainable sparse attention is incompatible with efficient offloading, as unconstrained KV accesses may force large CPU-to-GPU transfers and erase throughput gains. To this end, we propose NOSA, a trainable sparse attention mechanism natively designed for KV cache offloading. NOSA explicitly constrains the volume of CPU-GPU KV transfers, thereby achieving low communication overhead and high decoding throughput. We further build NOSI, a KV cache offloading inference system that fully unlocks NOSA's efficiency. Empirical results on 1,3,8B LLMs demonstrate that NOSA outperforms KV cache offloading baselines on general, long-input, and long-generation tasks, while boosting decoding throughput by up to 5.04x, 1.92x, and 1.83x over FullAttn, InfLLMv2, and ShadowKV, respectively. We release our code at https://github.com/thunlp/NOSA.
Figures & tables
| Method | LongBench | Helmet | ||||||||||||||||||||
| GO | TQ | NQ | QS | MU | 2W | MQ | RB | HQ | TR | PR | PC | SA | Avg. | Recall | RAG | ICL | Cite | Rerank | LongQA | Summ. | Avg. | |
| 1B Model Full Prefill | ||||||||||||||||||||||
| FullAttn | 28.0 | 70.8 | 16.0 | 23.7 | 11.5 | 19.5 | 45.2 | 55.0 | 21.0 | 66.5 | 10.5 | 1.0 | 23.55 | 30.2 | 61.3 ±0.4 | 39.0 ±0.9 | 60.0 ±0.7 | 5.5 ±0.3 | 14.3 ±0.0 | 17.2 ±0.3 | 0.4 ±0.2 | 28.3 ±0.3 |
| ShadKV 64 | 25.2 | 70.1 | 14.9 | 23.0 | 10.7 | 18.8 | 43.4 | 54.6 | 19.2 | 67.0 | 12.0 | 0.0 | 22.6 | 29.4 | 25.9 ±1.0 | 37.5 ±1.3 | 55.4 ±0.5 | 2.7 ±0.4 | 5.8 ±0.0 | 16.4 ±0.3 | 0.1 ±0.1 | 20.6 ±0.2 |
| ShadKV 8 | 25.0 | 70.3 | 15.4 | 23.1 | 11.1 | 19.7 | 42.3 | 55.2 | 19.6 | 66.5 | 11.0 | 1.0 | 22.6 | 29.4 | 48.9 ±0.8 | 38.4 ±1.3 | 59.0 ±0.4 | 4.1 ±0.3 | 11.0 ±0.0 | 16.7 ±0.3 | 0.1 ±0.1 | 25.5 ±0.3 |
| ArkVale | 8.2 | 58.7 | 8.4 | 11.4 | 9.5 | 13.8 | 25.2 | 38.9 | 10.5 | 60.5 | 9.8 | 1.0 | 11.6 | 20.6 | 26.7 ±1.6 | 34.3 ±1.0 | 57.5 ±0.3 | 1.1 ±1.0 | 1.8 ±0.4 | 16.9 ±0.7 | 0.2 ±0.1 | 19.8 ±0.3 |
| Method | Recall | RAG | ICL | Cite | Rerank | LongQA | Summ. | Avg. |
| FullAttn | 21.7 ±0.8 | 38.1 ±3.9 | 58.0 ±0.8 | 2.9 ±0.3 | 3.3 ±1.0 | 13.9 ±2.1 | 4.9 ±1.7 | 20.4 ±0.6 |
| InfLLMv2 | 12.0 ±0.6 | 30.9 ±3.8 | 49.0 ±4.6 | 5.4 ±0.9 | 0.5 ±0.3 | 5.0 ±1.4 | 1.5 ±0.2 | 14.9 ±0.9 |
| InfLLM | 11.2 ±0.8 | 36.5 ±2.9 | 35.3 ±3.4 | 6.1 ±1.5 | 0.6 ±0.4 | 6.1 ±0.8 | 12.3 ±4.0 | 15.4 ±0.7 |
| ShadKV 64 | 10.0 ±1.4 | 41.1 ±2.5 | 52.7 ±1.9 | 1.6 ±0.4 | 0.0 ±0.0 | 3.9 ±5.5 | 2.1 ±0.7 | 15.9 ±0.8 |
| ShadKV 8 | 13.4 ±1.2 | 38.7 ±3.7 | 56.7 ±2.1 | 3.6 ±0.2 | 0.8 ±0.2 | 4.4 ±2.1 | 0.5 ±0.7 | 16.9 ±0.6 |
| ShadKV | 10.4 ±1.1 | 32.9 ±5.8 | 58.3 ±1.3 | 1.9 ±0.1 | 0.0 ±0.0 | 17.8 ±11.0 | 3.9 ±1.4 | 17.9 ±1.5 |
| Method | Impl. | Off. | Input Length: 16K | Method | Impl. | Off. | Input Length: 32K | ||||||
| EB: 16 | EB: 32 | EB: 64 | EB: 128 | EB: 16 | EB: 32 | EB: 64 | EB: 128 | ||||||
| FullAttn | HF | ✗ | 160.85 (4) | 239.94 (8) | OOM (16) | OOM (32) | FullAttn | HF | ✗ | 80.56 (2) | 120.99 (4) | OOM (8) | OOM (16) |
| FullAttn | vLLM | ✗ | 233.21 (4) | 417.99 (8) | 678.41 (16) | 927.99 (32) | FullAttn | vLLM | ✗ | 113.42 (2) | 151.58 (4) | 460.65 (8) | 818.53 (16) |
| FullAttn | SGLang | ✗ | 180.28 (4) | 358.56 (8) | 774.74 (16) | 1267.20 (32) | FullAttn | SGLang | ✗ | 116.54 (2) | 174.50 (4) | 290.29 (8) | 748.20 (16) |
| InfLLMv2 | HF | ✗ | 47.10 (4) | OOM (8) | OOM (16) | OOM (32) | InfLLMv2 | HS | ✗ | 23.17 (2) | OOM (4) | OOM (8) | OOM (16) |
| InfLLMv2 | NOSI | ✗ | 226.51 (4) | 440.76 (8) | 824.09 (16) | 1475.43 (32) | InfLLMv2 | NOSI | ✗ | 104.88 (2) | 205.74 (4) | 403.26 (8) | 710.93 (16) |
| Tasks | Full Prefill | Sparse Prefill | |||||||||||
| FullAttn | ShadKV 64 | ShadKV 8 | ArkVale | DMA F | NOSA F | InfLLMv2 | ShadKV | ShadKV | InfLLM 128 | InfLLM 64 | DMA S | NOSA S | |
| Budget Sparsity: 0.125 | |||||||||||||
| Recall | 88.0 ±0.8 | 40.3 ±4.3 | 72.6 ±1.1 | 49.3 ±3.8 | 17.5 ±1.1 | 73.9 ±0.8 | 54.2 ±0.9 | 30.5 ±3.0 | 56.0 ±2.8 | 20.3 ±2.3 | 20.4 ±2.2 | 10.7 ±2.1 | 40.4 ±2.5 |
| RAG | 54.5 ±1.9 | 53.3 ±2.5 | 53.8 ±2.2 | 54.0 ±2.6 | 52.2 ±3.2 | 56.4 ±2.0 | 51.8 ±2.3 | 44.5 ±2.9 | 43.6 ±1.4 | 46.3 ±3.6 | 46.7 ±2.2 | 28.7 ±1.8 | 46.8 ±1.4 |
| ICL | 59.7 ±2.1 | 55.7 ±0.5 | 57.3 ±2.1 | 60.7 ±2.6 | 63.7 ±2.5 | 59.0 ±2.4 | 66.3 ±3.3 | 52.3 ±2.6 | 49.3 ±5.3 | 41.3 ±3.4 | 42.7 ±2.1 | 42.3 ±3.9 | 60.7 ±5.4 |
| Cite | 13.9 ±3.7 | 9.7 ±0.6 | 8.6 ±0.5 | 10.9 ±1.1 | 9.2 ±2.8 | 11.9 ±2.3 | 12.7 ±0.9 | 4.9 ±0.9 | 6.9 ±1.5 | 8.3 ±0.8 | 9.6 ±1.2 | 7.8 ±2.9 | 12.7 ±1.0 |
| Tasks | Full Prefill | Sparse Prefill | |||||||||||
| FullAttn | ShadKV 64 | ShadKV 8 | ArkVale | DMA F | NOSA F | InfLLMv2 | ShadKV | ShadKV | InfLLM 128 | InfLLM 64 | DMA S | NOSA S | |
| Input Length: 32K | |||||||||||||
| Recall | 56.3 ±1.8 | 31.0 ±0.6 | 48.8 ±3.4 | 41.6 ±2.5 | 10.0 ±1.8 | 56.6 ±4.2 | 45.3 ±2.0 | 30.4 ±2.7 | 42.6 ±3.2 | 15.0 ±3.2 | 15.6 ±1.9 | 8.4 ±2.7 | 35.0 ±2.5 |
| RAG | 50.7 ±0.5 | 49.1 ±1.3 | 49.3 ±2.0 | 50.2 ±1.5 | 47.9 ±2.5 | 51.4 ±1.4 | 49.7 ±2.2 | 44.5 ±0.9 | 46.4 ±1.6 | 44.9 ±2.9 | 45.1 ±3.1 | 32.2 ±3.9 | 44.7 ±3.2 |
| ICL | 67.3 ±2.6 | 66.7 ±5.7 | 66.7 ±2.5 | 65.7 ±4.9 | 59.3 ±4.0 | 64.7 ±3.8 | 69.7 ±2.5 | 60.0 ±1.7 | 64.0 ±1.4 | 38.3 ±2.6 | 38.3 ±1.2 | 53.3 ±3.3 | 68.7 ±5.4 |
| Cite | 13.9 ±3.7 | 9.9 ±2.3 | 10.7 ±0.6 | 9.7 ±1.5 | 10.3 ±0.2 | 12.7 ±1.1 | 12.6 ±3.7 | 9.8 ±1.2 | 9.2 ±1.9 | 13.0 ±1.7 | 10.9 ±1.4 | 9.5 ±1.3 | 10.0 ±1.9 |
| Task | T1.1 | T1.2 | T2.1 | T2.2 | T3.1 | T3.2 |
| Context Requirement | Full | Partial | Full | Partial | Full | Partial |
| InfLLMv2 +NOSI, Offload | 561.0 | 576.7 | 590.0 | 563.5 | 568.4 | 550.1 |
| NOSA+NOSI, No Offload | 220.8 | 220.3 | 215.5 | 215.4 | 218.6 | 211.4 |
| NOSA+NOSI, Offload | 700.4 | 712.5 | 714.4 | 702.6 | 709.9 | 695.9 |
| #Core | NUMA Index | Thru. | Comment | |
| Mem. | Core | |||
| 26 | 0 | 0 | 312.39 | Main experiment. |
| 26 | 0 | 1 | 308.49 | Cross NUMA (complete). |
| 26 | 0 | Half 0, Half 1 | 309.22 | Cross NUMA (half). |
| 13 | 0 | 0 | 295.86 | 104 cores, 8 GPUs. |
| 24 | 0 | 0 | 308.13 | 192 cores, 8 GPUs. |
| Input Length | InfLLMv2 + NOSI Offload | NOSA+NOSI No Offload | NOSA+NOSI Offload |
| 16K | 1762.16 | 219.38 | 2551.66 |
| 32K | 1535.70 | 436.45 | 2024.04 |
| 64K | 1510.80 | OOM | 2034.64 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Model Size | 1B | 3B | 8B |
| Architecture | Llama-3 | Llama-3 | MiniCPM4 |
| Use P ( Yang et al., 2021 ) | ✗ | ✗ | ✓ |
| Vocabulary Size | 73448 | 73448 | 73448 |
| Number of Layers | 28 | 32 | 32 |
| Hidden Size | 2048 | 2560 | 4096 |
| Number of Attention Heads | 16 | 32 | 32 |
| Abbreviation | Full Name | Bias Calculation | Attention Calculation |
| Locret | Locret | ||
| DMA | Dynamic Mask Attention | ||
| ED-DMA | Exp-Delayed DMA | ||
| S-DMA | Simple-DMA |
| Implementation | SG1 | SG2 | SG3 | MK1 | MK2 | MK3 | MV | MQ | VT | CWE | FWE | QA1 | QA2 | Avg. |
| InfLLM-V2 | 100.0 | 100.0 | 100.0 | 84.0 | 50.0 | 18.0 | 92.5 | 84.5 | 26.0 | 0.8 | 62.7 | 36.0 | 30.0 | 60.3 |
| Locret | 100.0 | 98.0 | 86.0 | 64.0 | 40.0 | 4.0 | 77.0 | 79.5 | 41.6 | 3.6 | 64.7 | 36.0 | 32.0 | 55.9 |
| DMA | 100.0 | 100.0 | 98.0 | 70.0 | 44.0 | 8.0 | 75.5 | 70.0 | 20.8 | 2.0 | 74.7 | 36.0 | 32.0 | 56.2 |
| S-DMA | 100.0 | 96.0 | 98.0 | 60.0 | 42.0 | 8.0 | 87.0 | 77.0 | 94.8 | 3.6 | 60.7 | 30.0 | 34.0 | 60.9 |
| ED-DMA | 100.0 | 98.0 | 98.0 | 76.0 | 42.0 | 14.0 | 90.5 | 78.5 | 62.4 | 3.8 | 68.0 | 40.0 | 34.0 | 61.9 |