Linearizing pretrained large language models (LLMs) primarily relies on intra-layer hybrid attention mechanisms to alleviate the quadratic complexity of standard softmax attention. Existing methods perform token routing based on sliding-window partitions, resulting in position-based selection and fails to capture token-specific global importance. Meanwhile, linear attention further suffers from distribution shift caused by learnable feature maps that distort pretrained feature magnitudes. Motivated by these limitations, we propose STILL, an intra-layer hybrid linearization framework for efficiently linearizing LLMs. STILL introduces a Self-Saliency Score with strong local-global consistency, enabling accurate token selection using sliding-window computation, and retains salient tokens for sparse softmax attention while summarizing the remaining context via linear attention. To preserve pretrained representations, we design a Norm-Preserved Feature Map (NP-Map) that decouples feature direction from magnitude and reinjects pretrained norms. We further adopt a unified training-inference architecture with chunk-wise parallelization and delayed selection to improve hardware efficiency. Experiments show that STILL matches or surpasses the original pretrained model on commonsense and general reasoning tasks, and achieves up to a 86.2% relative improvement over prior linearized attention methods on long-context benchmarks. Code is available at this URL.
Figures & tables
Figure 1: Illustration of Intra-Layer Hybrid Attention. While both act as intra-layer hybrids using local Softmax Attention (SA), they differ in handling the remaining sequence. LoLCATs/Liger applies uniform Linear Attention (LA), which limits long-context capacity, while compressed sparse attention (CSA) utilizes sparse compressed SA but completely discards unselected tokens (NULL). STILL selects tokens for high-fidelity SA and preserves the context via LA without information loss.
Figure 2: Architecture of STILL for Intra-Layer Hybrid Linear Attention. The diagram shows attention computation, chunk-level routing, and hybrid aggregation to outputs. STILL first computes the Self-Saliency Score within sliding-window attention for each token. Top-scoring tokens are routed to the softmax attention, while remaining tokens are processed via linear attention. Chunk-wise selecting replaces per-token selecting, enabling parallel training and inference.
Figure 3
Model
Training
PIQA
ARC-e
ARC-c
Hella.
Wino.
MMLU
Avg.
Avg.
Tokens
acc ↑
acc ↑
acc n ↑
acc n ↑
acc ↑
(5-shot) ↑
(w. MMLU)
(wo. MMLU)
Transformer
Llama 3 8B ( Team, 2024 )
15000 +
79.5
80.0
53.3
79.1
73.1
65.3
71.8
73.0
Llama 3.1 8B ( Team, 2024 )
15000 +
81.1
81.7
55.1
79.3
73.9
68.0
74.2
73.2
Subquadratic
Mamba2 8B ( Dao and Gu, 2024 )
3500
79.8
75.9
48.1
77.1
71.6
48.7
67.0
70.6
Table 1: Comparison of the Commonsense and General Reasoning Tasks . Comparison of STILL with baseline methods under Llama 3 8B and Llama 3.1 8B models. STILL consistently demonstrates strong performance across benchmarks.
Table 5
Model
QA1
QA4
QA5
0K
1K
2K
4K
0K
1K
2K
4K
0K
1K
2K
4K
SWA
93
2
1
1
74
0
0
0
68
18
11
6
LoLCATs ( Zhang et al., 2025a )
100
22
5
3
65
21
9
1
62
42
17
6
STILL (Ours)
100 +0
45 +23
22 +17
10 +7
82 +8
59 +28
30 +21
9 +8
83 +15
77 +35
45 +28
24 +18
Table 4: BABILong Benchmark Results. Performance comparison on BABILong across increasing context lengths (0K ∼ 4K), evaluating long-context reasoning performance.
Model
S-NIAH-1
S-NIAH-2
S-NIAH-3
RULER
4K
8K
16K
4K
8K
16K
4K
8K
16K
4K
8K
16K
MambaInLlama ( Wang et al., 2024 )
-
-
-
-
-
-
-
-
-
38.75
21.55
3.88
Zebra-Llama ( Yang et al., 2025b )
-
-
-
-
-
-
-
-
-
35.75
24.80
13.37
SWA
9.4
8.2
9.0
16.4
10.2
10.6
11.4
4.4
7.2
6.00
3.96
4.66
LoLCATs ( Zhang et al., 2025a )
84.8
81.6
36.2
86.8
33.0
8.6
54.6
16.8
3.4
33.99
17.93
7.59
STILL (Ours)
100
100
100
100
99.0
74.0
70.4
67.2
25.8
53.24
35.66
27.20
Table 5: Intra-Inter Layer Hybrid Results on the RULER Benchmark. All models are linearized from Llama 3.2 and evaluated over context lengths ranging from 4K to 16K.
Table 8Table 9
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Self-saliency score comparison between local and global attention across different layers, heads, and window sizes.
Model
Training
PIQA
ARC-e
ARC-c
Hella.
Wino.
MMLU
Avg.
Avg.
Tokens (B)
acc ↑
acc ↑
acc n ↑
acc n ↑
acc ↑
(5-shot) ↑
(w. MMLU)
(wo. MMLU)
Linearized from Llama 3.2 1B
Llama 3.2 1B ( Team, 2024 )
9000
74.4
65.5
35.8
63.7
60.5
31.9
55.3
60.0
T2R ( Kasai et al., 2021 )
0.04
69.2
58.2
29.9
42.6
54.1
23.3
46.2
50.8
Hedgehog ( Zhang et al., 2024 )
0.04
70.1
55.8
29.8
47.7
50.7
23.0
46.2
50.8
LoLCATs ( Zhang et al., 2025a )
0.04
74.1
63.7
36.4
51.2
58.2
23.1
51.1
56.7
Appendix
Table 8: Different Backbone Architectures and Scales Results. STILL achieves the best performance across all models and scales.
Self-Saliency Score
NP-Map
Gate
Acc.
✓
✓
✓
27.5
✓
✓
21.7 {}_{\text{{\color[rgb]{1,0,0}-5.8}}}
✓
✓
22.4 {}_{\text{{\color[rgb]{1,0,0}-4.9}}}
✓
19.2 {}_{\text{{\color[rgb]{1,0,0}-8.3}}}
12.5 {}_{\text{{\color[rgb]{1,0,0}-15.0}}}
Appendix
Table 9: Module Ablation Results. Ablation study of key components in STILL.
Routing Metric
Avg. ↑
Δ
Self-Saliency
72.5
–
Diagonal attention
72.0
-0.5
Appendix
Table 10: Routing-metric ablation on short-context commonsense reasoning. We report the average performance over PIQA, ARC-E, ARC-C, HellaSwag, WinoGrande, and MMLU. Δ denotes the difference from Self-Saliency.
Routing Metric
0.5K ↑
Δ
1K ↑
Δ
Avg. ↑
Self-Saliency
63.2
–
22.0
–
42.6
Diagonal attention
35.6
-27.6
10.6
-11.4
23.1
qt⊤kt
35.4
-27.8
10.2
-11.8
22.8
Key norm
18.6
-44.6
7.1
-14.9
12.9
Uniform per chunk
17.8
-45.4
6.8
-15.2
12.3
Random selection
18.6
-44.6
7.0
-15.0
12.8
Appendix
Table 11: Routing-metric ablation on long-context S-NIAH retrieval. All methods use the same cache budget. Avg. denotes the mean accuracy over the 0.5K and 1K context lengths, and Δ denotes the difference from Self-Saliency.
Layer
Head
Selected
Overlap
Ratio
0
0
224
141
62.9%
0
1
224
145
64.7%
0
2
224
158
70.5%
0
3
224
190
84.8%
23
0
224
185
82.6%
23
1
224
188
83.9%
Appendix
Table 12: Overlap between the Top-224 tokens selected by locally computed Self-Saliency and the Top-224 remote keys prioritized by future full-attention queries (Llama 3.1 8B).
C
PIQA
ARC-e
ARC-c
Hella.
Wino.
MMLU
Avg. w/ MMLU
Avg. w/o MMLU
16
80.5
81.3
53.4
77.2
71.0
54.2
69.6
72.7
64
81.4
82.8
56.4
79.0
73.1
60.8
72.3
74.5
128
81.3
82.7
56.2
79.0
71.1
62.1
72.1
74.1
Appendix
Table 13: Commonsense reasoning accuracy under different chunk sizes C with a fixed selection budget.
C
4
8
16
32
64
128
Time (s/step)
OOM
20.3
14.7
11.3
9.7
9.5
Memory (GB)
OOM
46.4
37.4
32.9
30.6
29.3
Appendix
Table 14: Cost per step under different chunk sizes C .
Model
Training
PIQA
ARC-e
ARC-c
Hella.
Wino.
MMLU
Avg.
Avg.
Tokens (B)
acc ↑
acc ↑
acc n ↑
acc n ↑
acc ↑
(5-shot) ↑
(w. MMLU)
(wo. MMLU)
Transformer
Llama 3 8B ( Team, 2024 )
15000 +
79.5
80.0
53.3
79.1
73.1
65.3
71.8
73.0
Llama 3.1 8B ( Team, 2024 )
15000 +
81.1
81.7
55.1
79.3
73.9
68.0
74.2
73.2
Subquadratic
Mamba 8B ( Gu and Dao, 2023 )
1100
78.9
75.4
42.2
75.6
68.3
28.0
61.4
68.1
Appendix
Table 15: Comparison of the Commonsense and General Reasoning Tasks . Comparison of STILL with baseline methods under Llama 3 8B and Llama 3.1 8B models. STILL consistently demonstrates strong performance across benchmarks.
Model
Cache
S-NIAH-1
S-NIAH-2
Tokens
0.5K
1K
2K
4K
0.5K
1K
2K
4K
SWA
128
0
0
0
0
100
0
0
0
LoLCATs ( Zhang et al., 2025a )
128
29.4
9.6
3.0
0
100
17.4
7.2
4.2
Liger-GLA ( Lan et al., 2025 )
128
28.4
0.2
0.2
0.2
100
1.6
0.6
1.0
STILL (Ours)
≤ 128
63.2
22.0
5.8
1.4
100
33.6
3.4
2.8
SWA
256
66.4
4.2
0.2
0
100
0
0.2
0
Appendix
Table 16: Full Results of the Comparison of the Single Needle-in-a-Haystack (S-NIAH) Tasks . Accuracy on S-NIAH tasks from the RULER benchmark across context lengths from 0.5K to 4K. Methods are compared under different cache budgets and our approach evaluated under matched or smaller cache constraints.
Model
QA1
QA2
QA3
QA4
QA5
0K
1K
2K
4K
0K
1K
2K
4K
0K
1K
2K
4K
0K
1K
2K
4K
0K
1K
2K
4K
SWA
93
2
1
1
34
3
2
0
8
4
2
1
74
0
0
0
68
18
11
6
LoLCATs
100
22
5
3
25
10
4
1
30
17
13
4
65
21
9
1
62
42
17
6
STILL
100
45
22
10
57
19
9
5
37
24
15
11
82
59
30
9
83
77
45
24
Appendix
Table 17: BABILong Benchmark Full Results. Performance comparison on BABILong across increasing context lengths (0K ∼ 4K), evaluating long-context reasoning performance.
S-NIAH-1
0.5K
1K
2K
4K
8K
SWA
100
25.2
14.6
5.8
3.0
LoLCATs
100
65.6
24.6
8.8
0.0
Liger-GLA
73.4
1.0
0.0
0.0
0.0
STILL (Ours)
100
99.6
89.0
86.2
36.2
Appendix
Table 18: Extended S-NIAH-1 evaluation up to 8K on Llama 3.1 8B under the 512-token cache budget, STILL retains meaningful retrieval accuracy at 8K, while all baselines collapse to near zero.
Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy: every new token costs more than the previous one, and its keys and values must be stored in memory indefinitely, which is unsustainable. Two main lines of work address this: compressing the KV cache, e.g., by evicting or quantizing keys and values, and retrofitting LLMs to use Linear Attention, which replaces the KV cache with a fixed-size state. Retrofitting has attracted a lot of attention, given its promise to solve the quadratic scaling problem with state-of-the-art performance at low cost. However, it has not been properly compared to the simplest form of KV-cache eviction: Sliding Window Attention (SWA) with attention sinks. In this work, we show that SWA with sinks performs as well or better than most retrofitted Linear Attention models across multiple LLMs and downstream tasks, with the largest gains on long-context and generative tasks. On long-context reasoning tasks (Needle-in-a-Haystack and BABILong), SWA achieves massively higher performance (2 to 10 times higher than linear attention). SWA requires no additional training, is extremely fast, and requires little memory, making it an extremely cheap and reliable solution. When the training budget is limited, switching to SWA is a much more effective way to reduce inference memory cost than retrofitting linear attention. Linear attention models have shown promise, but they require training from scratch or extensive retrofitting to reap their architectural benefits and come close to SWA.
Large Language Models (LLMs) have become the central paradigm in artificial intelligence, yet the core computational primitive of attention has remained structurally unchanged. Local Linear Attention (LLA) is an attention mechanism derived from nonparametric statistics in the test-time regression framework. In contrast to prior research on efficient attention variants, LLA upgrades the local constant estimate in softmax attention to a local linear estimate, yielding provably superior bias-variance tradeoffs for associative memory. However, LLA has not been scaled in LLM pretraining due to computational and numerical stability concerns. We introduce Parallax, a parameterized Local Linear Attention that is scalable for LLMs. Parallax eliminates the numerical solver in LLA and learns an extra query-like projector that probes the KV covariance. We place Parallax within a family of attention mechanisms connected by the bandwidth, the probe construction and the affine structure. We propose a hardware-aware algorithm that increases the arithmetic intensity over FlashAttention, shifting attention into a more compute bound regime. Our prototype decode kernel matches or outperforms FlashAttention 2/3 across diverse batch sizes and context lengths. We pretrain Parallax at 0.6B and 1.7B scales and find consistent perplexity improvements throughout pretraining with gains that transfer to downstream benchmarks. The advantage persists under both parameter-matched and compute-matched controls, demonstrating a Pareto improvement. We perform careful pretraining ablations and identify a novel phenomenon whereby Muon unlocks the capacity of Parallax. To our knowledge, this is the first empirical demonstration of strong architecture-optimizer codesign for attention mechanisms in the architecture research literature.
Yifei Zuo, Dhruv Pai, Zhichen Zeng +3
Northwestern University · Tilde Research · University of Washington
The quadratic cost of causal self-attention severely bottlenecks long-context transformer inference. While numerous post hoc linearization pipelines exist, it is difficult to identify which components preserve model quality. This work isolates the effect of state update design in a strict frozen-backbone regime. We show that softmax relies on key-dependent, rank-1 orthogonal projections, elucidating why delta-style networks outperform purely gated accumulation. We identify a potential source of approximation errors and introduce structural interventions, specifically sink tokens, short convolutions, and fixed-budget cache routing, which reduces the remaining gap. We scale this linearization approach across LLaMA and Qwen models up to 32B parameters, outperforming prior post hoc baselines on MMLU and matching the long-context retrieval of complex adaptive-caching frameworks.
Anna Kuzina, Paul N. Whatmough, Babak Ehteshami Bejnordi