cs.LGFeb 2, 2026

STILL: Selecting Tokens for Intra-Layer Hybrid Attention to Linearize LLMs

Authors: Weikang Meng, Liangyu Huo, Yadan Luo, Jiawen Guan, Jingyi Zhang, Yingjian Li, Zheng Zhang

Organizations: SMULL Group, Harbin Institute of Technology, Shenzhen · Pengcheng Laboratory · UQMM Lab, University of Queensland · Huawei Technologies Co., Ltd.

Abstract

Linearizing pretrained large language models (LLMs) primarily relies on intra-layer hybrid attention mechanisms to alleviate the quadratic complexity of standard softmax attention. Existing methods perform token routing based on sliding-window partitions, resulting in position-based selection and fails to capture token-specific global importance. Meanwhile, linear attention further suffers from distribution shift caused by learnable feature maps that distort pretrained feature magnitudes. Motivated by these limitations, we propose STILL, an intra-layer hybrid linearization framework for efficiently linearizing LLMs. STILL introduces a Self-Saliency Score with strong local-global consistency, enabling accurate token selection using sliding-window computation, and retains salient tokens for sparse softmax attention while summarizing the remaining context via linear attention. To preserve pretrained representations, we design a Norm-Preserved Feature Map (NP-Map) that decouples feature direction from magnitude and reinjects pretrained norms. We further adopt a unified training-inference architecture with chunk-wise parallelization and delayed selection to improve hardware efficiency. Experiments show that STILL matches or surpasses the original pretrained model on commonsense and general reasoning tasks, and achieves up to a 86.2% relative improvement over prior linearized attention methods on long-context benchmarks. Code is available at this URL.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Sliding-window beats linear attention

    Aug 28, 2026Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron +1Kimi Delta AttentionLarge Language Model Memory

  2. Parallax: Parameterized Local Linear Attention for Language Modeling

    May 27, 2026Yifei Zuo, Dhruv Pai, Zhichen Zeng +3Kimi Delta AttentionLanguage Modeling

  3. The Key to Going Linear: Analysis-Driven Transformer Linearization

    Jul 8, 2026Anna Kuzina, Paul N. Whatmough, Babak Ehteshami BejnordiKimi Delta AttentionTransformer Architectures