STILL: Selecting Tokens for Intra-Layer Hybrid Attention to Linearize LLMs
Organizations: SMULL Group, Harbin Institute of Technology, Shenzhen · Pengcheng Laboratory · UQMM Lab, University of Queensland · Huawei Technologies Co., Ltd.
Abstract
Linearizing pretrained large language models (LLMs) primarily relies on intra-layer hybrid attention mechanisms to alleviate the quadratic complexity of standard softmax attention. Existing methods perform token routing based on sliding-window partitions, resulting in position-based selection and fails to capture token-specific global importance. Meanwhile, linear attention further suffers from distribution shift caused by learnable feature maps that distort pretrained feature magnitudes. Motivated by these limitations, we propose STILL, an intra-layer hybrid linearization framework for efficiently linearizing LLMs. STILL introduces a Self-Saliency Score with strong local-global consistency, enabling accurate token selection using sliding-window computation, and retains salient tokens for sparse softmax attention while summarizing the remaining context via linear attention. To preserve pretrained representations, we design a Norm-Preserved Feature Map (NP-Map) that decouples feature direction from magnitude and reinjects pretrained norms. We further adopt a unified training-inference architecture with chunk-wise parallelization and delayed selection to improve hardware efficiency. Experiments show that STILL matches or surpasses the original pretrained model on commonsense and general reasoning tasks, and achieves up to a 86.2% relative improvement over prior linearized attention methods on long-context benchmarks. Code is available at this URL.
Figures & tables
| Model | Training | PIQA | ARC-e | ARC-c | Hella. | Wino. | MMLU | Avg. | Avg. |
| Tokens | acc | acc | acc n | acc n | acc | (5-shot) | (w. MMLU) | (wo. MMLU) | |
| Transformer | |||||||||
| Llama 3 8B ( Team, 2024 ) | 15000 | 79.5 | 80.0 | 53.3 | 79.1 | 73.1 | 65.3 | 71.8 | 73.0 |
| Llama 3.1 8B ( Team, 2024 ) | 15000 | 81.1 | 81.7 | 55.1 | 79.3 | 73.9 | 68.0 | 74.2 | 73.2 |
| Subquadratic | |||||||||
| Mamba2 8B ( Dao and Gu, 2024 ) | 3500 | 79.8 | 75.9 | 48.1 | 77.1 | 71.6 | 48.7 | 67.0 | 70.6 |
| Model | QA1 | QA4 | QA5 | |||||||||
| 0K | 1K | 2K | 4K | 0K | 1K | 2K | 4K | 0K | 1K | 2K | 4K | |
| SWA | 93 | 2 | 1 | 1 | 74 | 0 | 0 | 0 | 68 | 18 | 11 | 6 |
| LoLCATs ( Zhang et al., 2025a ) | 100 | 22 | 5 | 3 | 65 | 21 | 9 | 1 | 62 | 42 | 17 | 6 |
| STILL (Ours) | 100 +0 | 45 +23 | 22 +17 | 10 +7 | 82 +8 | 59 +28 | 30 +21 | 9 +8 | 83 +15 | 77 +35 | 45 +28 | 24 +18 |
| Model | S-NIAH-1 | S-NIAH-2 | S-NIAH-3 | RULER | ||||||||
| 4K | 8K | 16K | 4K | 8K | 16K | 4K | 8K | 16K | 4K | 8K | 16K | |
| MambaInLlama ( Wang et al., 2024 ) | - | - | - | - | - | - | - | - | - | 38.75 | 21.55 | 3.88 |
| Zebra-Llama ( Yang et al., 2025b ) | - | - | - | - | - | - | - | - | - | 35.75 | 24.80 | 13.37 |
| SWA | 9.4 | 8.2 | 9.0 | 16.4 | 10.2 | 10.6 | 11.4 | 4.4 | 7.2 | 6.00 | 3.96 | 4.66 |
| LoLCATs ( Zhang et al., 2025a ) | 84.8 | 81.6 | 36.2 | 86.8 | 33.0 | 8.6 | 54.6 | 16.8 | 3.4 | 33.99 | 17.93 | 7.59 |
| STILL (Ours) | 100 | 100 | 100 | 100 | 99.0 | 74.0 | 70.4 | 67.2 | 25.8 | 53.24 | 35.66 | 27.20 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Training | PIQA | ARC-e | ARC-c | Hella. | Wino. | MMLU | Avg. | Avg. |
| Tokens (B) | acc | acc | acc n | acc n | acc | (5-shot) | (w. MMLU) | (wo. MMLU) | |
| Linearized from Llama 3.2 1B | |||||||||
| Llama 3.2 1B ( Team, 2024 ) | 9000 | 74.4 | 65.5 | 35.8 | 63.7 | 60.5 | 31.9 | 55.3 | 60.0 |
| T2R ( Kasai et al., 2021 ) | 0.04 | 69.2 | 58.2 | 29.9 | 42.6 | 54.1 | 23.3 | 46.2 | 50.8 |
| Hedgehog ( Zhang et al., 2024 ) | 0.04 | 70.1 | 55.8 | 29.8 | 47.7 | 50.7 | 23.0 | 46.2 | 50.8 |
| LoLCATs ( Zhang et al., 2025a ) | 0.04 | 74.1 | 63.7 | 36.4 | 51.2 | 58.2 | 23.1 | 51.1 | 56.7 |
| Self-Saliency Score | NP-Map | Gate | Acc. |
| 27.5 | |||
| 21.7 {}_{\text{{\color[rgb]{1,0,0}-5.8}}} | |||
| 22.4 {}_{\text{{\color[rgb]{1,0,0}-4.9}}} | |||
| 19.2 {}_{\text{{\color[rgb]{1,0,0}-8.3}}} | |||
| 12.5 {}_{\text{{\color[rgb]{1,0,0}-15.0}}} |
| Routing Metric | Avg. | |
| Self-Saliency | 72.5 | – |
| Diagonal attention | 72.0 | -0.5 |
| Routing Metric | 0.5K | 1K | Avg. | ||
| Self-Saliency | 63.2 | – | 22.0 | – | 42.6 |
| Diagonal attention | 35.6 | -27.6 | 10.6 | -11.4 | 23.1 |
| 35.4 | -27.8 | 10.2 | -11.8 | 22.8 | |
| Key norm | 18.6 | -44.6 | 7.1 | -14.9 | 12.9 |
| Uniform per chunk | 17.8 | -45.4 | 6.8 | -15.2 | 12.3 |
| Random selection | 18.6 | -44.6 | 7.0 | -15.0 | 12.8 |
| Layer | Head | Selected | Overlap | Ratio |
| 0 | 0 | 224 | 141 | 62.9% |
| 0 | 1 | 224 | 145 | 64.7% |
| 0 | 2 | 224 | 158 | 70.5% |
| 0 | 3 | 224 | 190 | 84.8% |
| 23 | 0 | 224 | 185 | 82.6% |
| 23 | 1 | 224 | 188 | 83.9% |
| PIQA | ARC-e | ARC-c | Hella. | Wino. | MMLU | Avg. w/ MMLU | Avg. w/o MMLU | |
| 16 | 80.5 | 81.3 | 53.4 | 77.2 | 71.0 | 54.2 | 69.6 | 72.7 |
| 64 | 81.4 | 82.8 | 56.4 | 79.0 | 73.1 | 60.8 | 72.3 | 74.5 |
| 128 | 81.3 | 82.7 | 56.2 | 79.0 | 71.1 | 62.1 | 72.1 | 74.1 |
| 4 | 8 | 16 | 32 | 64 | 128 | |
| Time (s/step) | OOM | 20.3 | 14.7 | 11.3 | 9.7 | 9.5 |
| Memory (GB) | OOM | 46.4 | 37.4 | 32.9 | 30.6 | 29.3 |
| Model | Training | PIQA | ARC-e | ARC-c | Hella. | Wino. | MMLU | Avg. | Avg. |
| Tokens (B) | acc | acc | acc n | acc n | acc | (5-shot) | (w. MMLU) | (wo. MMLU) | |
| Transformer | |||||||||
| Llama 3 8B ( Team, 2024 ) | 15000 | 79.5 | 80.0 | 53.3 | 79.1 | 73.1 | 65.3 | 71.8 | 73.0 |
| Llama 3.1 8B ( Team, 2024 ) | 15000 | 81.1 | 81.7 | 55.1 | 79.3 | 73.9 | 68.0 | 74.2 | 73.2 |
| Subquadratic | |||||||||
| Mamba 8B ( Gu and Dao, 2023 ) | 1100 | 78.9 | 75.4 | 42.2 | 75.6 | 68.3 | 28.0 | 61.4 | 68.1 |
| Model | Cache | S-NIAH-1 | S-NIAH-2 | ||||||
| Tokens | 0.5K | 1K | 2K | 4K | 0.5K | 1K | 2K | 4K | |
| SWA | 128 | 0 | 0 | 0 | 0 | 100 | 0 | 0 | 0 |
| LoLCATs ( Zhang et al., 2025a ) | 128 | 29.4 | 9.6 | 3.0 | 0 | 100 | 17.4 | 7.2 | 4.2 |
| Liger-GLA ( Lan et al., 2025 ) | 128 | 28.4 | 0.2 | 0.2 | 0.2 | 100 | 1.6 | 0.6 | 1.0 |
| STILL (Ours) | 128 | 63.2 | 22.0 | 5.8 | 1.4 | 100 | 33.6 | 3.4 | 2.8 |
| SWA | 256 | 66.4 | 4.2 | 0.2 | 0 | 100 | 0 | 0.2 | 0 |
| Model | QA1 | QA2 | QA3 | QA4 | QA5 | |||||||||||||||
| 0K | 1K | 2K | 4K | 0K | 1K | 2K | 4K | 0K | 1K | 2K | 4K | 0K | 1K | 2K | 4K | 0K | 1K | 2K | 4K | |
| SWA | 93 | 2 | 1 | 1 | 34 | 3 | 2 | 0 | 8 | 4 | 2 | 1 | 74 | 0 | 0 | 0 | 68 | 18 | 11 | 6 |
| LoLCATs | 100 | 22 | 5 | 3 | 25 | 10 | 4 | 1 | 30 | 17 | 13 | 4 | 65 | 21 | 9 | 1 | 62 | 42 | 17 | 6 |
| STILL | 100 | 45 | 22 | 10 | 57 | 19 | 9 | 5 | 37 | 24 | 15 | 11 | 82 | 59 | 30 | 9 | 83 | 77 | 45 | 24 |
| S-NIAH-1 | 0.5K | 1K | 2K | 4K | 8K |
| SWA | 100 | 25.2 | 14.6 | 5.8 | 3.0 |
| LoLCATs | 100 | 65.6 | 24.6 | 8.8 | 0.0 |
| Liger-GLA | 73.4 | 1.0 | 0.0 | 0.0 | 0.0 |
| STILL (Ours) | 100 | 99.6 | 89.0 | 86.2 | 36.2 |