cs.CLAug 28, 2026

Sliding-window beats linear attention

Authors: Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron, Emy Gervais

Organizations: Microsoft Applied Sciences Group (ASG) · Independent

Abstract

Due to the nature of quadratic attention, Large Language Models (LLMs) consume a lot of memory and energy: every new token costs more than the previous one, and its keys and values must be stored in memory indefinitely, which is unsustainable. Two main lines of work address this: compressing the KV cache, e.g., by evicting or quantizing keys and values, and retrofitting LLMs to use Linear Attention, which replaces the KV cache with a fixed-size state. Retrofitting has attracted a lot of attention, given its promise to solve the quadratic scaling problem with state-of-the-art performance at low cost. However, it has not been properly compared to the simplest form of KV-cache eviction: Sliding Window Attention (SWA) with attention sinks. In this work, we show that SWA with sinks performs as well or better than most retrofitted Linear Attention models across multiple LLMs and downstream tasks, with the largest gains on long-context and generative tasks. On long-context reasoning tasks (Needle-in-a-Haystack and BABILong), SWA achieves massively higher performance (2 to 10 times higher than linear attention). SWA requires no additional training, is extremely fast, and requires little memory, making it an extremely cheap and reliable solution. When the training budget is limited, switching to SWA is a much more effective way to reduce inference memory cost than retrofitting linear attention. Linear attention models have shown promise, but they require training from scratch or extensive retrofitting to reap their architectural benefits and come close to SWA.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Blurry Window Attention

    May 31, 2026Axel Laborieux, Christos Sourmpis, Juan Gabriel Kostelec +1Transformer ArchitecturesKernel Method

  2. Variational Linear Attention: Stable Associative Memory for Long-Context Transformers

    May 11, 2026Vishal Pandey, Gopal SinghKimi Delta AttentionLinear Attention

  3. STILL: Selecting Tokens for Intra-Layer Hybrid Attention to Linearize LLMs

    Feb 2, 2026Weikang Meng, Liangyu Huo, Yadan Luo +4Kimi Delta Attention