Self-Attention

Momentum

9 papers in the last four weeks, up 50% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 195

All topics
CardsList
  1. Tucker Bottleneck Attention for Multi-Dimensional Sequence Modeling

    Oct 6, 2026Ryan Solgi, Parsa Madinei, Zheng ZhangAttention MechanismsLow-Rank Attention

  2. Universal interpolation for deep residual self-attention networks

    Oct 1, 2026Sibylle Marcotte, Joan BrunaSoftmax AttentionTransformer

  3. Contrastive Attention Mitigates Spectral Bias in Spiking Transformers

    Oct 1, 2026Xiaoli Liu, Malu Zhang, Yang YangSpiking TransformersSelf-Attention

  4. Retrieval Capacity of Self-Attention Under Competition

    Sep 29, 2026Timur Mudarisov, Mikhail Burtsev, Radu StateSelf-AttentionLong-Context Language Modeling

  5. PQ-HSA: Reusing Product-Quantized Scores for Hybrid Sparse-Approximate Attention

    Sep 27, 2026Kunming Shao, Jierun Chen, Yanli Wang +4Self-AttentionLLM Inference Acceleration

  6. Nonequilibrium Phases of Repulsive Self-Attention: Chaos, Attention Condensation, and Emergent Locality

    Sep 23, 2026Qucheng Gao, Zuyi Yang, Xiao ChenSelf-AttentionChaotic Dynamical Systems

  7. Prescriptive SVD-Inspired Attention via Spectral Energy Retention

    Sep 21, 2026Vasileios Arampatzakis, Vasileios Sevetlidis, George PavlidisTransformer InterpretabilityLow-Rank Attention

  8. It's Not RoPE that Creates Sinks: The Role of Self-Concentration and Value-Non-Mixing in Attention

    Sep 8, 2026Raito Kiya, Satoki Ohashi, Kosuke Sato +6Self-AttentionAttention Sinks

  9. The Head Complexity of Boolean Functions in Single-Layer Attention

    Sep 3, 2026Rajmohan Rajaraman, Ravi Sundaram, Amanuel TesfayeSelf-AttentionAttention Mechanisms

  10. TANGO: Treating Tokens as Operators

    Aug 22, 2026Joshua NunleyTransformer FFNsTransformer

  11. Ask Self, Ask Others: Relation Is All You Need

    Aug 20, 2026Yuting Ge, Pengju Yang, Mingkai NieSelf-AttentionDecoder-Only Language Models

  12. Why Post-Norm Transformers Collapse: Attention Amplification and Gradient Repair Failure

    Aug 10, 2026Xingjian Wang, Qingyu Han, Xiaodong Luo +1TransformerSelf-Attention

  13. Clustered Attractor Manifolds and Dynamical Condensation in Self-Attention

    Aug 9, 2026Qucheng Gao, Zuyi Yang, Xiao ChenDynamical SystemsSelf-Attention

  14. Gated Spatial Redundancy Projection for Pathology Transformer Attentions

    Aug 8, 2026Zhiyuan Yang, Jiahao Cheng, Vincent Quoc-Huy Trinh +1Cancer Survival PredictionSelf-Attention

  15. Faster Query-Key Learning Sharpens Attention in Self-Attention Models

    Aug 7, 2026Rahul Vashisht, Harish G. RamaswamySelf-Attention

  16. Inhibited Self-Attention: Sharpening Focus in Vision Transformers

    Jul 14, 2026Peter R. D. van der Wal, Nicola Strisciuglio, George AzzopardiSelf-AttentionVision Transformer

  17. LoSA-Net: A Localized and Scale-Adaptive Network for Boundary-Sensitive Prediction of Perineural Invasion in 3D MRI

    Jul 13, 2026Youngung Han, Hyunsu Go, Kyeonghun Kim +9Self-AttentionMulti-Scale Feature Fusion

  18. From Self-Attention to Connection Laplacian: A Unified Operator View of Transformers

    Jul 12, 2026Binbin Lin, Wei Chen, Yalun Li +3Self-AttentionTransformer Attention

  19. Towards Robust EEG Decoding Based on Riemannian Self-Attention

    Jun 24, 2026Shaocheng Jin, Tao Zhou, Rui Wang +4Brain-Computer InterfacesSelf-Attention

  20. Communicability-Inspired Positional Encoding (CIPE)

    Jun 24, 2026Yipeng Zhang, Zhongtian Sun, Pietro Liò +1Graph Representation LearningSelf-Attention

  21. Attention mechanism for scalable mesh-based neural surrogates of free-surface fluids

    Jun 22, 2026Federico Lanteri, Massimiliano CremonesiNeural Surrogate ModelingSelf-Attention

  22. Kuramoto Attention: Synchronizing Self-Attention on the Torus

    Jun 10, 2026Joshua NunleyDynamical SystemsSelf-Attention

  23. RoVE: Rotary Value Embeddings Attention for Relative Position-dependent Value Pathways

    Jun 9, 2026Alejandro García-Castellanos, Maurice Weiler, Erik J BekkersRotary Positional EmbeddingsSelf-Attention

  24. Contribution Weights: A Geometrical Analysis of Self-Attention Transformers

    May 29, 2026Harry Jake Cunningham, Nicola Muca CironeTransformer InterpretabilityFeature Attribution

  25. VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion

    May 28, 2026Hidir Yesiltepe, Jiazhen Hu, Tuna Han Salih Meral +4Video Diffusion ModelsSelf-Attention

  26. Anti Mode-Collapse in Mean-Field Transformer via Auxiliary Variables

    May 28, 2026Masaaki Imaizumi, Masanori Koyama, Noboru Isobe +1TransformerSelf-Attention

  27. Attention as In-Context Empirical Bayes: A Two-Stage View via Particle Dynamics

    May 28, 2026Matthew Smart, Soumya Ganguly, Nilava Metya +2Self-AttentionEmpirical Bayes

  28. Parallax: Parameterized Local Linear Attention for Language Modeling

    May 27, 2026Yifei Zuo, Dhruv Pai, Zhichen Zeng +3Language Model PretrainingSelf-Attention

  29. Balancing Fidelity and Diversity in Diffusion Models via Symmetric Attention Decomposition: Hopfield Perspective

    May 26, 2026Hyunmin Cho, Woo Kyoung Han, Kyong Hwan JinSelf-AttentionTransformer Attention

  30. AdaMerge: Salience-Aware Adaptive Token Merging for Training-Free Acceleration of Vision Transformers

    May 26, 2026Semi Lee, Hyejin Go, Hyesong ChoiToken MergingSelf-Attention

  31. Quantized Keys Steal Attention: Bias Correction for KV-Cache Compression in Video Diffusion

    May 25, 2026Tuna Tuncer, Felix Becker, Thomas PfeilVideo Diffusion ModelsSelf-Attention

  32. IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference

    May 25, 2026Xintong Yang, Hao Gu, Binxing Xu +6Memory-Augmented Language ModelsSelf-Attention

  33. Quaternion Self-Attention with Shared Scores

    May 24, 2026Shogo Yamauchi, Tohru Nitta, Hideaki TamoriSelf-AttentionAttention Mechanisms

  34. Interdomain Attention: Beyond Token-Level Key-Value Memory

    May 23, 2026Naoki Kiyohara, Harrison Bo Hua Zhu, Riccardo El Hassanin +4Self-AttentionLong-Context Language Modeling

  35. Approaching I/O-optimality for Approximate Attention

    May 22, 2026Pál András Papp, Aleksandros Sobczyk, Anastasios ZouziasSelf-AttentionEfficient Attention

  36. Improved Belief-Attention in Vision Task

    May 22, 2026Guoqiang ZhangVisual AttentionSelf-Attention

  37. Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity

    May 21, 2026Hangyue Zhao, Paul Caillon, Erwan Fagnou +1Self-AttentionStructured Sparsity

  38. ASAP: Attention Sink Anchored Pruning

    May 21, 2026Jaehyuk Lee, Hanyoung Kim, Yanggee Kim +1Self-AttentionEfficient ViTs

  39. Tensor Cache: Eviction-conditioned Associative Memory for Transformers

    May 21, 2026Kabir Swain, Sijie Han, Daniel Karl I. Weidele +2Memory-Augmented Neural NetworksSelf-Attention

  40. Manifold-Guided Attention Steering

    May 20, 2026Ian Li, Kapilesh Guruprasad, Raunak Sengupta +3Language Model SteeringSelf-Attention

  41. EntmaxKV: Support-Aware Decoding for Entmax Attention

    May 20, 2026Gonçalo Duarte, Miguel Couceiro, Marcos V. TrevisoSelf-AttentionKV Caching

  42. Efficient Long-Context Modeling in Diffusion Language Models via Block Approximate Sparse Attention

    May 19, 2026Wenhu Zhang, Yiming Wu, Huanyu Wang +6Self-AttentionLong-Context Language Modeling