Linear Attention

Momentum

25 papers in the last four weeks, up 92% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 183

All topics
CardsList
  1. CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection

    May 16, 2026Jiwon Song, Dongwon Jo, Beomseok Kang +1PrefillLinear Attention

  2. Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion

    May 15, 2026Kunyang Li, Mubarak Shah, Yuzhang ShangAutoregressive Video Diffusion ModelsSpatio-Temporal Attention

  3. Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation

    May 15, 2026Shuaiyi Li, Zhisong Zhang, Yan Wang +5Dynamic Sparse AttentionLinear Attention

  4. WorldParticle: Unified World Simulation of Lagrangian Particle Dynamics via Transformer

    May 14, 2026Caoliwen Wang, Minghao Guo, Siyuan Chen +13Physics SimulationParticle Dynamics

  5. Precise Verification of Transformers through ReLU-Catalyzed Abstraction Refinement

    May 14, 2026Hengjie Liu, Zhenya Zhang, Jianjun ZhaoTransformer ArchitecturesRectified Linear Unit

  6. OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention

    May 13, 2026Chenyu Zhou, Hongpei Li, Yuerou Liu +3Kimi Delta AttentionLinear Attention

  7. Composition of Memory Experts for Diffusion World Models

    May 12, 2026Sebastian Stapf, Pablo Acuaviva Huertos, Aram Davtyan +1World ModelsDiffusion Models

  8. Breaking Global Self-Attention Bottlenecks in Transformer-based Spiking Neural Networks with Local Structure-Aware Self-Attention

    May 12, 2026Lingdong Li, Hangming Zhang, Qiang YuSpiking Neural NetworksTransformer Attention

  9. Variational Linear Attention: Stable Associative Memory for Long-Context Transformers

    May 11, 2026Vishal Pandey, Gopal SinghKimi Delta AttentionLinear Attention

  10. Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime

    May 11, 2026Albert Alcalde, Leon Bungert, Konstantin Riedl +1Transformer EncoderTransformer Architectures

  11. Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions

    May 11, 2026Diancheng Kang, Zheyuan Liu, Ningshan Ma +3Linear Activation SteeringModel Activations

  12. Explanation-Aware Learning for Enhanced Interpretability in Biomedical Imaging

    May 11, 2026Zubair Faruqui, Rahul DubeyInterpretabilityGradient-Weighted Class Activation Mapping

  13. Rethinking Random Transformers as Adaptive Sequence Smoothers for Sleep Staging

    May 11, 2026Guisong Liu, Xin Gao, Martin Dresler +2SleepSpatio-Temporal Transformers

  14. Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning

    May 8, 2026Zixuan Xie, Xinyu Liu, Claire Chen +3In-Context LearningLinear Attention

  15. Federated Nested Learning: Collaborative Training of Self-Referential Memories for Test-Time Adaptation

    May 8, 2026Hong Chen, Pengcheng Wu, Yuanguo Lin +4Federated LearningParameter-Efficient Adaptation

  16. Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent

    May 7, 2026Chenyang Zhang, Yuan CaoIn-Context LearningTransformer Architectures

  17. Metonymy in vision models undermines attention-based interpretability

    May 7, 2026Ananthu Aniraj, Cassio F. Dantas, Dino Ienco +2InterpretabilityVision Foundation Models

  18. MDN: Parallelizing Stepwise Momentum for Delta Linear Attention

    May 7, 2026Yulong Huang, Xiang Liu, Hongxiang Huang +5Kimi Delta AttentionLinear Attention

  19. On the Role of Language Representations in Auto-Bidding: Findings and Implications

    May 7, 2026Guanyu Zhu, Jining Luan, Hanwen Du +11BiddingLinear Attention

  20. Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics

    May 6, 2026Dominik Dahlem, Diego Maniloff, Mac MisiuraLinear AttentionSpectral

  21. Jordan-RoPE: Non-Semisimple Relative Positional Encoding via Complex Jordan Blocks

    May 5, 2026Yaobo ZhangPositional EncodingRelative Position

  22. Cascade Token Selection for Transformer Attention Acceleration

    May 4, 2026Stephen J. ThomasTransformer AttentionLinear Attention

  23. Linearizing Vision Transformer with Test-Time Training

    May 4, 2026Yining Li, Dongchen Han, Zeyu Liu +3Self-Supervised Vision TransformersVision Transformer

  24. From Architecture to Output: Structural Origins of Hallucination in Large Language Models and the Amplifying Role of Data

    Apr 29, 2026Md. Rejaul Korim Sadi, Toufiqur Rahman Tasin, Golam Mostofa NaeemLarge Language Model HallucinationLarge Language Models Fail

  25. AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving

    Apr 28, 2026Zhongkai Yu, Haotian Ye, Chenyang Zhou +9Graphics Processing Unit MemoryLinear Attention

  26. Emergent Self-Attention from Astrocyte-Gated Associative Memory Dynamics

    Apr 28, 2026Arnau Vivet, Alex ArenasHopfield NetworksNeurons

  27. HBGSA: Hydrogen Bond Graph with Self-Attention for Drug-Target Binding Affinity Prediction

    Apr 25, 2026Junxiao Kong, Chupei Tang, Di Wang +4Interleaved Graph AttentionLinear Attention

  28. LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel

    Apr 22, 2026Zhe Feng, Sen Lian, Changwei Wang +5Linear AttentionTransformer Architectures

  29. Nexusformer: Nonlinear Attention Expansion for Stable and Inheritable Transformer Scaling

    Apr 21, 2026Weijie Zhao, Mingquan Liu, Bolun Wang +4Linear AttentionLanguage Modeling

  30. Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization

    Apr 21, 2026Andrei Andrusenko, Vladimir Bataev, Lilit Grigoryan +3Automatic Speech RecognitionStreaming

  31. Advancing Vision Transformer with Enhanced Spatial Priors

    Apr 20, 2026Qihang Fan, Huaibo Huang, Mingrui Chen +2Self-Supervised Vision TransformersVision Transformer

  32. OneDrive: Unified Multi-Paradigm Driving with Vision-Language-Action Models

    Apr 20, 2026Yiwei Zhang, Xuesong Chen, Jin Gao +5Autonomous Hard Drive DisassemblyAutoregressive Language Models

  33. Robust Diabetic Retinopathy Grading Using Dual-Resolution Attention-Based Deep Learning with Ordinal Regression

    Apr 19, 2026Afshan HashmiRetinal ImagingFundus Images

  34. Capacity-Controlled Global Attention for Graph Transformers

    Apr 19, 2026Yang Liu, Dongxin Guo, Tom Zheng +3Transformer AttentionLinear Attention

  35. CAM3DNet: Comprehensively mining the multi-scale features for 3D Object Detection with Multi-View Cameras

    Apr 18, 2026Mingxi Pang, Dingheng Wang, Zekun Li +43D Object DetectionMulti-View

  36. Improving Reasoning Capabilities in Small Models through Mixture-of-Layers Distillation with Stepwise Attention on Key Information

    Apr 17, 2026Yao Chen, Jiawei Sheng, Wenyuan Zhang +1Dataset DistillationReasoning Skills

  37. LACE: Lattice Attention for Cross-thread Exploration

    Apr 16, 2026Yang Li, Zirui Zhang, Yang Liu +1LLM Reasoning StrategiesCross-Task Interference

  38. Invertible Query-Key Coupling Composes with Attention Mechanisms

    Apr 2, 2026Barak Gahtan, Alex M. BronsteinLinear AttentionKeys And Value

  39. Beyond Acoustic Prefixes: Persistent Access to Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition

    Mar 28, 2026Hao Shi, Yuan Gao, Xugang Lu +1AcousticLinear Attention

  40. Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention

    Mar 6, 2026Haiqing Hao, Zhipeng Sui, Rong Zou +4Object DetectionDynamic Sparse Attention

  41. Quantum Attention by Overlap Interference: Predicting Classical and Many-Body Quantum Sequences

    Feb 6, 2026Alessio Pecilli, Matteo RosatiVariational Quantum AlgorithmsTransformer Attention

  42. Poly-attention: a general scheme for higher-order self-attention

    Feb 2, 2026Sayak Chakrabarti, Toniann Pitassi, Josh AlmanLinear AttentionTransformer Architectures

  43. Data-Free Pruning of Self-Attention Layers in LLMs

    Dec 3, 2025Dhananjay Saikumar, Blesson VargheseLarge Language Model CompressionLinear Attention

  44. Controllably Efficient Language Models

    Nov 7, 2025Jatin Prakash, Aahlad Puli, Rajesh RanganathTransformer ArchitecturesLinear Attention

  45. How Can Mamba Learn In Context with Outliers and Generalize Provably?

    Oct 1, 2025Hongkang Li, Songtao Lu, Xiaodong Cui +2Mamba-Based ModelsIn-Context Learning

  46. Policy Gradient with Self-Attention for Model-Free Distributed Nonlinear Multi-Agent Games

    Sep 22, 2025Eduardo Sebastián, Maitrayee Keskar, Eeman Iqbal +3Multi-Agent Reinforcement LearningPolicy Gradient

  47. SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging

    Jun 10, 2025Nhat Thanh Tran, Fanghui Xue, Shuai Zhang +4Linear AttentionMamba-Based Models