Sliding-window beats linear attention
Organizations: Microsoft Applied Sciences Group (ASG) · Independent
Abstract
Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy: every new token costs more than the previous one, and its keys and values must be stored in memory indefinitely, which is unsustainable. Two main lines of work address this: compressing the KV cache, e.g., by evicting or quantizing keys and values, and retrofitting LLMs to use Linear Attention, which replaces the KV cache with a fixed-size state. Retrofitting has attracted a lot of attention, given its promise to solve the quadratic scaling problem with state-of-the-art performance at low cost. However, it has not been properly compared to the simplest form of KV-cache eviction: Sliding Window Attention (SWA) with attention sinks. In this work, we show that SWA with sinks performs as well or better than most retrofitted Linear Attention models across multiple LLMs and downstream tasks, with the largest gains on long-context and generative tasks. On long-context reasoning tasks (Needle-in-a-Haystack and BABILong), SWA achieves massively higher performance (2 to 10 times higher than linear attention). SWA requires no additional training, is extremely fast, and requires little memory, making it an extremely cheap and reliable solution. When the training budget is limited, switching to SWA is a much more effective way to reduce inference memory cost than retrofitting linear attention. Linear attention models have shown promise, but they require training from scratch or extensive retrofitting to reap their architectural benefits and come close to SWA.
Figures & tables
| GSM8K | HumanEval | Avg | ||
| Model | Type | (5-shot, EM) | (0-shot, pass@1) | |
| Qwen2.5-7B-Instruct | Teacher | 75.6 | 64.0 | 69.8 |
| SWA(64,4) | 0.0 | 12.8 | 6.4 | |
| SWA(128,4) | 1.7 | 28.7 | 15.2 | |
| SWA(256,4) | 25.4 | 40.2 | 32.8 | |
| SWA(512,4) | 41.7 | 39.6 | 40.7 |
| (Base: Llama-3.1-8B) | Window | S-NIAH-1 | S-NIAH-2 | S-NIAH-3 | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | size | .5K | 1K | 2K | 4K | .5K | 1K | 2K | 4K | .5K | 1K | 2K | 4K |
| SWA(128,4) | 128 | 35.0 | 20.2 | 15.0 | 12.6 | 100 | 33.0 | 22.8 | 9.2 | 99.8 | 56.2 | 41.4 | 17.2 |
| LoLCATs(+SWA) | 128 | 29.4 | 9.6 | 3.0 | 0 | 100 | 17.4 | 7.2 | 4.2 | 98.2 | 14.6 | 3.2 | 1.6 |
| Liger-GLA(+SWA) | 128 | 28.4 | 0.2 | 0.2 | 0.2 | 100 | 1.6 | 0.6 | 1.0 | 97.6 | 2.0 | 1.2 | 0.8 |
| SWA(256,4) | 256 | 91.8 | 34.8 | 21.0 | 14.6 | 100 | 51.4 | 33.4 | 13.8 | 100 | 68.0 | 50.2 | 19.6 |
| LoLCATs(+SWA) | 256 | 84.8 | 26.2 | 10.2 | 2.2 | 100 | 37.0 | 12.2 | 8.2 | 100 | 30.6 | 6.0 | 2.2 |
| (Base: Llama-3.1-8B) | Window | BABILong (Average) | |||
|---|---|---|---|---|---|
| Model | size | 0K | 1K | 2K | 4K |
| SWA(256, 4) | 256 | 55 | 20 | 19 | 15 |
| LoLCATs(+SWA) | 256 | 56 | 22 | 10 | 3 |
| Full Attention | 74 | 70 | 67 | 60 | |
| LongBench | LongBench-v2 | RULER | BABILong | Avg | |||
| Model | Type | (subtask avg) | (subtask avg) | (NIAH/VT/CWE/FWE) | (8K) | (16K) | |
| Qwen2.5-7B-Instruct | Teacher | 36.5 | 31.8 | 98.6 | 62.2 | 61.0 | 57.1 |
| SWA(64,4) | 12.8 | 22.0 | 4.0 | 31.5 | 32.7 | 17.7 | |
| SWA(128,4) | 15.3 | 28.2 | 12.4 | 41.1 | 43.1 | 24.5 | |
| SWA(256,4) | 14.7 | 28.3 | 17.7 | 41.9 | 43.3 | 25.8 | |
| SWA(512,4) | 17.2 | 29.0 | 22.7 | 43.6 | 44.2 | 28.2 | |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Type | Retrofitting Tokens (B) | MMLU (5-shot) | ARC-C (acc-norm) | ARC-E (acc) | HellaSwag (acc-norm) | PIQA (acc) | WinoGrande (acc) | Avg. |
|---|---|---|---|---|---|---|---|---|---|
| Qwen3-8B | Teacher | 0 | 74.9 | 56.7 | 83.5 | 75.0 | 76.6 | 68.4 | 72.5 |
| SWA | 0 | 70.8 | 56.9 | 83.6 | 74.3 | 76.5 | 67.6 | 71.6 | |
| Gated DeltaNet | 0.1 | 27.5 | 46.8 | 76.6 | 59.1 | 73.9 | 52.9 | 56.1 | |
| GLA | 0.1 | 25.8 | 38.0 | 70.5 | 50.9 | 72.0 | 50.6 | 51.3 | |
| QRWKV6 | 0.1 | 24.6 | 41.5 | 74.7 | 53.7 | 72.6 | 52.0 | 53.2 | |
| Phi-4-mini-reasoning | Teacher | 0 | 57.4 | 48.0 | 71.4 | 64.8 | 69.3 | 59.1 | 61.7 |
| (Base: Llama-3.1-8B) | QA1 | QA2 | QA3 | QA4 | QA5 | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | 0K | 1K | 2K | 4K | 0K | 1K | 2K | 4K | 0K | 1K | 2K | 4K | 0K | 1K | 2K | 4K | 0K | 1K | 2K | 4K |
| SWA(256, 4) | 84 | 21 | 16 | 12 | 36 | 14 | 11 | 6 | 26 | 16 | 21 | 11 | 82 | 16 | 17 | 18 | 49 | 35 | 29 | 30 |
| LoLCATs(256) | 100 | 22 | 5 | 3 | 25 | 10 | 4 | 1 | 30 | 17 | 13 | 4 | 65 | 21 | 9 | 1 | 62 | 42 | 17 | 6 |
| Full attention | 94 | 88 | 82 | 74 | 59 | 49 | 48 | 44 | 47 | 53 | 48 | 44 | 82 | 76 | 74 | 67 | 90 | 86 | 83 | 73 |