cs.LGSep 30, 2026

Switching Linear Attention

Authors: Hyun Dong Lee, Xavier Gonzalez, Nicolas Zucchet, E. Kelly Buchanan, Emily B. Fox, Scott W. Linderman

Organizations: Stanford University · Unconventional AI. Work done at Stanford University.

Abstract

Designing expressive sequence layers with efficient inference remains a central challenge in modern machine learning. Standard softmax attention achieves excellent sequence modeling performance through rich nonlinear token interactions, but it requires a key-value cache that grows linearly with sequence length, limiting its scalability. Linear attention enables efficient recurrent computation with a constant memory footprint, yet its reduced expressivity often yields inferior modeling performance. We introduce Switching Linear Attention (SwiLA), a novel sequence layer that bridges this gap by enhancing representational capacity while retaining the fixed-size recurrent state of linear attention. We derive the SwiLA recurrence from the test-time regression framework, casting the state update rule as online expectation-maximization in a mixture of linear regressions model. At test time, each output dimension dynamically selects among multiple linear attention components based on the input. Across associative recall, in-context language learning, and language modeling benchmarks, SwiLA shows strong performance and narrows the gap to softmax attention, even surpassing it in several settings.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing

    Jul 8, 2026Tommaso Cerruti, Tim Rieder, George Rowlands +2Kimi Delta AttentionLinear Attention

  2. Sliding-window beats linear attention

    Aug 28, 2026Alexia Jolicoeur-Martineau, Rhea Sanjay Sukthanker, Pashmina Cameron +1Kimi Delta AttentionLarge Language Model Memory

  3. Dynamic Linear Attention

    Jun 9, 2026Xin Wang, Hui Shen, Boyuan Zheng +7Kimi Delta AttentionEfficient Long-Context Inference