Switching Linear Attention
Organizations: Stanford University · Unconventional AI. Work done at Stanford University.
Abstract
Designing expressive sequence layers with efficient inference remains a central challenge in modern machine learning. Standard softmax attention achieves excellent sequence modeling performance through rich nonlinear token interactions, but it requires a key-value cache that grows linearly with sequence length, limiting its scalability. Linear attention enables efficient recurrent computation with a constant memory footprint, yet its reduced expressivity often yields inferior modeling performance. We introduce Switching Linear Attention (SwiLA), a novel sequence layer that bridges this gap by enhancing representational capacity while retaining the fixed-size recurrent state of linear attention. We derive the SwiLA recurrence from the test-time regression framework, casting the state update rule as online expectation-maximization in a mixture of linear regressions model. At test time, each output dimension dynamically selects among multiple linear attention components based on the input. Across associative recall, in-context language learning, and language modeling benchmarks, SwiLA shows strong performance and narrows the gap to softmax attention, even surpassing it in several settings.
Figures & tables
| Perplexity ( ) | Accuracy ( , %) | ||||||||||
| Model (374M) | Wiki. | LMB. | LMB. | PIQA | Hella. | Wino. | ARC-E | ARC-C | SIQA | BoolQ | Avg. |
| Transformer++ | 28.2 | 37.4 | 33.5 | 66.0 | 32.4 | 52.6 | 56.0 | 24.1 | 38.4 | 59.1 | 45.3 |
| FoX | 28.7 | 50.3 | 30.8 | 65.2 | 32.2 | 52.5 | 57.2 | 23.2 | 37.9 | 58.0 | 44.6 |
| DeltaNet | 29.8 | 46.3 | 27.9 | 63.9 | 31.9 | 51.6 | 55.6 | 21.8 | 38.3 | 58.3 | 43.7 |
| Gated DeltaNet | 27.7 | 35.7 | 32.0 | 66.0 | 33.3 | 51.8 | 57.5 | 24.6 | 39.0 | 55.8 | 45.0 |
| KDA | 23.4 | 26.8 | 36.4 | 67.0 | 34.9 | 51.5 | 60.7 | 27.3 | 38.9 | 56.2 | 46.6 |
| Model | SWDE | FDA | SQuAD | TriviaQA | NQ | DROP | Avg. |
|---|---|---|---|---|---|---|---|
| Transformer++ | 28.4 | 29.5 | 31.5 | 46.0 | 15.6 | 18.9 | 28.3 |
| FoX | 36.4 | 53.0 | 30.6 | 46.4 | 15.6 | 20.2 | 33.7 |
| DeltaNet | 7.9 | 5.4 | 23.4 | 41.7 | 10.5 | 14.9 | 17.3 |
| Gated DeltaNet | 9.1 | 5.5 | 25.2 | 44.4 | 12.9 | 18.0 | 19.2 |
| KDA | 13.9 | 5.5 | 25.5 | 46.8 | 12.5 | 17.6 | 20.3 |
| DN-MoM | 5.4 | 2.1 | 21.2 | 37.0 | 9.1 | 14.8 | 14.9 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.