Self-Attention

Momentum

9 papers in the last four weeks, up 50% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 195

All topics
CardsList
  1. Contribution Weights: A Geometrical Analysis of Self-Attention Transformers

    May 29, 2026Harry Jake Cunningham, Nicola Muca CironeTransformer InterpretabilityFeature Attribution

  2. VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion

    May 28, 2026Hidir Yesiltepe, Jiazhen Hu, Tuna Han Salih Meral +4Video Diffusion ModelsSelf-Attention

  3. Anti Mode-Collapse in Mean-Field Transformer via Auxiliary Variables

    May 28, 2026Masaaki Imaizumi, Masanori Koyama, Noboru Isobe +1TransformerSelf-Attention

  4. Attention as In-Context Empirical Bayes: A Two-Stage View via Particle Dynamics

    May 28, 2026Matthew Smart, Soumya Ganguly, Nilava Metya +2Self-AttentionEmpirical Bayes

  5. Parallax: Parameterized Local Linear Attention for Language Modeling

    May 27, 2026Yifei Zuo, Dhruv Pai, Zhichen Zeng +3Language Model PretrainingSelf-Attention

  6. Balancing Fidelity and Diversity in Diffusion Models via Symmetric Attention Decomposition: Hopfield Perspective

    May 26, 2026Hyunmin Cho, Woo Kyoung Han, Kyong Hwan JinSelf-AttentionTransformer Attention

  7. AdaMerge: Salience-Aware Adaptive Token Merging for Training-Free Acceleration of Vision Transformers

    May 26, 2026Semi Lee, Hyejin Go, Hyesong ChoiToken MergingSelf-Attention

  8. Quantized Keys Steal Attention: Bias Correction for KV-Cache Compression in Video Diffusion

    May 25, 2026Tuna Tuncer, Felix Becker, Thomas PfeilVideo Diffusion ModelsSelf-Attention

  9. IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference

    May 25, 2026Xintong Yang, Hao Gu, Binxing Xu +6Memory-Augmented Language ModelsSelf-Attention

  10. Quaternion Self-Attention with Shared Scores

    May 24, 2026Shogo Yamauchi, Tohru Nitta, Hideaki TamoriSelf-AttentionAttention Mechanisms

  11. Interdomain Attention: Beyond Token-Level Key-Value Memory

    May 23, 2026Naoki Kiyohara, Harrison Bo Hua Zhu, Riccardo El Hassanin +4Self-AttentionLong-Context Language Modeling

  12. Approaching I/O-optimality for Approximate Attention

    May 22, 2026Pál András Papp, Aleksandros Sobczyk, Anastasios ZouziasSelf-AttentionEfficient Attention

  13. Improved Belief-Attention in Vision Task

    May 22, 2026Guoqiang ZhangVisual AttentionSelf-Attention

  14. Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity

    May 21, 2026Hangyue Zhao, Paul Caillon, Erwan Fagnou +1Self-AttentionStructured Sparsity

  15. ASAP: Attention Sink Anchored Pruning

    May 21, 2026Jaehyuk Lee, Hanyoung Kim, Yanggee Kim +1Self-AttentionEfficient ViTs

  16. Tensor Cache: Eviction-conditioned Associative Memory for Transformers

    May 21, 2026Kabir Swain, Sijie Han, Daniel Karl I. Weidele +2Memory-Augmented Neural NetworksSelf-Attention

  17. Manifold-Guided Attention Steering

    May 20, 2026Ian Li, Kapilesh Guruprasad, Raunak Sengupta +3Language Model SteeringSelf-Attention

  18. EntmaxKV: Support-Aware Decoding for Entmax Attention

    May 20, 2026Gonçalo Duarte, Miguel Couceiro, Marcos V. TrevisoSelf-AttentionKV Caching

  19. Efficient Long-Context Modeling in Diffusion Language Models via Block Approximate Sparse Attention

    May 19, 2026Wenhu Zhang, Yiming Wu, Huanyu Wang +6Self-AttentionLong-Context Language Modeling