Self-Attention

Momentum

9 papers in the last four weeks, up 50% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 195

All topics
CardsList
  1. Jordan-RoPE: Non-Semisimple Relative Positional Encoding via Complex Jordan Blocks

    May 5, 2026Yaobo ZhangRotary Positional EmbeddingsSelf-Attention

  2. Cascade Token Selection for Transformer Attention Acceleration

    May 4, 2026Stephen J. ThomasSelf-AttentionTransformer Attention

  3. Linearizing Vision Transformer with Test-Time Training

    May 4, 2026Yining Li, Dongchen Han, Zeyu Liu +3Efficient Transformer InferenceSelf-Attention

  4. When Attention Collapses: Residual Evidence Modeling for Compositional Inference

    May 4, 2026Niklas HoubaGravitational-Wave AstronomySelf-Attention

  5. SAGA: A Robust Self-Attention and Goal-Aware Anchor-based Planner for Safe UAV Autonomous Navigation

    May 4, 2026Junhao Wei, Yanxiao Li, Dexing Yao +9Self-AttentionRoadmap-Based Motion Planning

  6. Projection-Free Transformers via Gaussian Kernel Attention

    May 4, 2026Debarshi Kundu, Archisman Ghosh, Swaroop Ghosh +1TransformerSelf-Attention

  7. Stochastic Sparse Attention for Memory-Bound Inference

    May 3, 2026Kyle Lee, Corentin Delacour, Kevin Callahan-Coray +5GPU AccelerationSelf-Attention

  8. Focus and Dilution: The Multi-stage Learning Process of Attention

    May 2, 2026Zheng-An Chen, Pengxiao Lin, Zhi-Qin John Xu +1Self-AttentionTransformer Attention

  9. DARE: Diffusion Language Model Activation Reuse for Efficient Inference

    May 1, 2026Natalia Frumkin, Bokun Wang, Hung-Yueh Chiang +3Self-AttentionDiffusion Language Model Inference

  10. Attention Is Where You Attack

    Apr 30, 2026Aviral Srivastava, Sourav PandaSelf-AttentionAdversarial Attacks

  11. Better Models, Faster Training: Sigmoid Attention for single-cell Foundation Models

    Apr 29, 2026Vijay Sadashivaiah, Georgios Dasoulas, Judith Mueller +1Softmax AttentionSelf-Attention

  12. Emergent Self-Attention from Astrocyte-Gated Associative Memory Dynamics

    Apr 28, 2026Arnau Vivet, Alex ArenasSoftmax AttentionDynamical Systems

  13. Learning to Rotate: Temporal and Semantic Rotary Encoding for Sequential Modeling

    Apr 27, 2026Hailing Cheng, Daqi Sun, Xinyu LuRotary Positional EmbeddingsSelf-Attention

  14. Kwai Summary Attention Technical Report

    Apr 27, 2026Chenglong Chu, Guorui Zhou, Guowang Zhang +35Self-AttentionKV Caching

  15. Pruning via Causal Attribution Preserves Reasoning Performance in Large Language Models

    Apr 27, 2026Amogh Sheth, Biruk Assefa, Yi Wen Huang +2LLM PruningSelf-Attention

  16. ELSA: Exact Linear-Scan Attention for Fast and Memory-Light Vision Transformers

    Apr 26, 2026Chih-Chung Hsu, Xin-Di Ma, Wo-Ting Liao +1Efficient Transformer InferenceSoftmax Attention

  17. HeadRouter: Dynamic Head-Weight Routing for Task-Adaptive Audio Token Pruning in Large Audio Language Models

    Apr 26, 2026Peize He, Yaodi Luo, Xiaoqian Liu +7Audio Token CompressionSelf-Attention

  18. Rank, Head-Channel Non-Identifiability, and Symmetry Breaking: A Precise Analysis of Representational Collapse in Transformers

    Apr 26, 2026Giansalvo CirrincioneTransformer InterpretabilityRepresentation Collapse

  19. Dissociating Decodability and Causal Use in Bracket-Sequence Transformers

    Apr 24, 2026Aryan Sharma, Cutter Dawes, Shivam RavalTransformer InterpretabilitySelf-Attention

  20. Sparse Forcing: Native Trainable Sparse Attention for Real-time Autoregressive Diffusion Video Generation

    Apr 23, 2026Boxun Xu, Yuming Du, Zichang Liu +7Autoregressive DiffusionSelf-Attention

  21. The Recurrent Transformer: Greater Effective Depth and Efficient Decoding

    Apr 23, 2026Costin-Andrei Oncescu, Depen Morwani, Samy Jelassi +3Efficient Transformer InferenceSelf-Attention

  22. Stream-CQSA: Exact Out-of-Memory Recovery for Attention

    Apr 22, 2026Yiming Bian, Joshua M. AkeySelf-AttentionMemory-Efficient Inference

  23. Simplified Sparse Attention via Gist Tokens

    Apr 22, 2026Yuzhen Mao, Michael Y. Li, Emily B. FoxSelf-AttentionLong-Context Language Model Inference