Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy: every new token costs more than the previous one, and its keys and values must be stored in memory indefinitely, which is unsustainable. Two main lines of work address this: compressing the KV cache, e.g., by evicting or quantizing keys and values, and retrofitting LLMs to use Linear Attention, which replaces the KV cache with a fixed-size state. Retrofitting has attracted a lot of attention, given its promise to solve the quadratic scaling problem with state-of-the-art performance at low cost. However, it has not been properly compared to the simplest form of KV-cache eviction: Sliding Window Attention (SWA) with attention sinks. In this work, we show that SWA with sinks performs as well or better than most retrofitted Linear Attention models across multiple LLMs and downstream tasks, with the largest gains on long-context and generative tasks. On long-context reasoning tasks (Needle-in-a-Haystack and BABILong), SWA achieves massively higher performance (2 to 10 times higher than linear attention). SWA requires no additional training, is extremely fast, and requires little memory, making it an extremely cheap and reliable solution. When the training budget is limited, switching to SWA is a much more effective way to reduce inference memory cost than retrofitting linear attention. Linear attention models have shown promise, but they require training from scratch or extensive retrofitting to reap their architectural benefits and come close to SWA.
Figures & tables
Figure 1: Different types of attention masks: Full attention (FA), LoLCATs/Liger-GLA (mixture of linear attention and SWA), and Sliding Window Attention (SWA) with sinks (adapted from Figure 1 of Meng et al. (2026) ).
GSM8K ↑
HumanEval ↑
Avg ↑
Model
Type
(5-shot, EM)
(0-shot, pass@1)
Qwen2.5-7B-Instruct
Teacher
75.6
64.0
69.8
SWA(64,4)
0.0
12.8
6.4
SWA(128,4)
1.7
28.7
15.2
SWA(256,4)
25.4
40.2
32.8
SWA(512,4)
41.7
39.6
40.7
Table 3: Generative reasoning benchmarks (%). Avg is the mean of GSM8K and HumanEval. The best non-teacher result per base model is highlighted .
(Base: Llama-3.1-8B)
Window
S-NIAH-1
S-NIAH-2
S-NIAH-3
Model
size
.5K
1K
2K
4K
.5K
1K
2K
4K
.5K
1K
2K
4K
SWA(128,4)
128
35.0
20.2
15.0
12.6
100
33.0
22.8
9.2
99.8
56.2
41.4
17.2
LoLCATs(+SWA)
128
29.4
9.6
3.0
0
100
17.4
7.2
4.2
98.2
14.6
3.2
1.6
Liger-GLA(+SWA)
128
28.4
0.2
0.2
0.2
100
1.6
0.6
1.0
97.6
2.0
1.2
0.8
SWA(256,4)
256
91.8
34.8
21.0
14.6
100
51.4
33.4
13.8
100
68.0
50.2
19.6
LoLCATs(+SWA)
256
84.8
26.2
10.2
2.2
100
37.0
12.2
8.2
100
30.6
6.0
2.2
Table 4: Accuracy on the Single Needle-in-a-Haystack (S-NIAH) across context lengths (0.5K, 1K, 2K, 4K) and window-size (128, 256, 512). The best model at each window-size is highlighted .
(Base: Llama-3.1-8B)
Window
BABILong (Average)
Model
size
0K
1K
2K
4K
SWA(256, 4)
256
55
20
19
15
LoLCATs(+SWA)
256
56
22
10
3
Full Attention
∞
74
70
67
60
Table 5: Average accuracy (%) on the first five BABILong tasks (QA1–QA5) across context lengths, from 0K (no distractor text) to 4K tokens. Base model: Llama-3.1-8B.
LongBench ↑
LongBench-v2 ↑
RULER ↑
BABILong ↑
Avg ↑
Model
Type
(subtask avg)
(subtask avg)
(NIAH/VT/CWE/FWE)
(8K)
(16K)
Qwen2.5-7B-Instruct
Teacher
36.5
31.8
98.6
62.2
61.0
57.1
SWA(64,4)
12.8
22.0
4.0
31.5
32.7
17.7
SWA(128,4)
15.3
28.2
12.4
41.1
43.1
24.5
SWA(256,4)
14.7
28.3
17.7
41.9
43.3
25.8
SWA(512,4)
17.2
29.0
22.7
43.6
44.2
28.2
Table 6: Long-context benchmarks (accuracy %). BABILong is averaged over all 20 tasks (QA1–QA20); 8K and 16K denote the total input length in tokens. Avg is the mean of LongBench, LongBench-v2, RULER, and the BABILong average over context lengths. The best non-teacher result per base model is highlighted . † Evaluated at 8K context (LoLCATs exceeds H100 memory at 16K).
Figure 2: Speed (decoding throughput; tokens/s) and Memory cost (KV Cache or Recurrent state memory; MiB) of different attention types (Full Attention (FA), Sliding-Window Attention (SWA), Linear Attention, Linear + SWA (LoLCATs)) across context lengths (128 to 256K). Both axes are in log scale.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Type
Retrofitting Tokens (B)
MMLU (5-shot)
ARC-C (acc-norm)
ARC-E (acc)
HellaSwag (acc-norm)
PIQA (acc)
WinoGrande (acc)
Avg.
Qwen3-8B
Teacher
0
74.9
56.7
83.5
75.0
76.6
68.4
72.5
SWA (64,4)
0
70.8
56.9
83.6
74.3
76.5
67.6
71.6
Gated DeltaNet
0.1
27.5
46.8
76.6
59.1
73.9
52.9
56.1
GLA
0.1
25.8
38.0
70.5
50.9
72.0
50.6
51.3
QRWKV6
0.1
24.6
41.5
74.7
53.7
72.6
52.0
53.2
Phi-4-mini-reasoning
Teacher
0
57.4
48.0
71.4
64.8
69.3
59.1
61.7
Appendix
Table 7: Comparison of sliding-window and linear-attention methods. Teacher and SWA (64,4) are train-free baselines; linear variants are LoLCATs-style two-stage distilled on ∼ 0.1B tokens of cleaned-Alpaca ( Taori et al., 2023 ) . We use different attention variants (GLA ( Yang et al., 2023 ) , Gated DeltaNet ( Yang et al., 2025b ) , QRWKV6 ( Peng et al., 2024 ) ) on modern architectures (Qwen3 ( Yang et al., 2025a ) , Phi-4-mini-reasoning ( Abdin et al., 2024 ; Xu et al., 2025 ) , and Phi-4-reasoning-plus ( Abdin et al., 2025 ) ).
(Base: Llama-3.1-8B)
QA1
QA2
QA3
QA4
QA5
Model
0K
1K
2K
4K
0K
1K
2K
4K
0K
1K
2K
4K
0K
1K
2K
4K
0K
1K
2K
4K
SWA(256, 4)
84
21
16
12
36
14
11
6
26
16
21
11
82
16
17
18
49
35
29
30
LoLCATs(256)
100
22
5
3
25
10
4
1
30
17
13
4
65
21
9
1
62
42
17
6
Full attention
94
88
82
74
59
49
48
44
47
53
48
44
82
76
74
67
90
86
83
73
Appendix
Table 8: BABILong Benchmark Full Results. Accuracy (%) on each of the first five BABILong tasks (QA1–QA5) across context lengths, from 0K (no distractor text) to 4K tokens. Base model: Llama-3.1-8B.
Figure 3: Speed (decoding throughput; tokens/s) and Memory cost (KV Cache or Recurrent state memory; MiB) of different attention types (Full Attention (FA), Sliding-Window Attention (SWA), Linear Attention, Linear + SWA (LoLCATs)) across context lengths (128 to 256K). Both axes are in log scale.
Figure 4: Attention floating point operations (FLOPs) per decoded token of different attention types (Full Attention (FA), Sliding-Window Attention (SWA), Linear Attention, Linear + SWA (LoLCATs)) across context lengths (128 to 256K). Both axes are in log scale.