Self-Attention

Momentum

9 papers in the last four weeks, up 50% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 195

All topics
CardsList
  1. Multi-Head Attention as Ensemble Nadaraya-Watson Estimation: Variance Reduction, Decorrelation, and Optimal Head Diversity

    May 18, 2026Ernest FokouéSelf-AttentionEnsemble Learning

  2. DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention

    May 18, 2026Yuxiang Huang, Nuno M. T. Gonçalves, Federico Alvetreti +5Self-AttentionLong-Context Modeling

  3. RAVE: Re-Allocating Visual Attention in Large Multimodal Models

    May 18, 2026Xi Leng, Xinhong Ma, Ziqiang Dong +4Visual AttentionSelf-Attention

  4. InfoFlow: A Framework for Multi-Layer Transformer Analysis

    May 18, 2026Penghao Yu, Haotian Jiang, Zeyu Bao +1Transformer ExpressivityTransformer

  5. CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection

    May 16, 2026Jiwon Song, Dongwon Jo, Beomseok Kang +1Self-AttentionKV-Cache Management

  6. Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion

    May 15, 2026Kunyang Li, Mubarak Shah, Yuzhang ShangMemory-Augmented Neural NetworksVideo Diffusion Models

  7. Grokking as Structural Inference: Transformers Need Bayesian Lottery Tickets

    May 15, 2026Kai Hidajat, Solden Stoll, Joseph AnNeural Network GeneralizationTransformer

  8. From Sparsity to Simplicity: Enabling Simpler Sequential Replacements via Sparse Attention Distillation

    May 15, 2026Yuxin Ren, Maxwell D Collins, Miao Hu +1Self-AttentionVision Transformer

  9. RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably

    May 15, 2026Yufeng Du, Phillip Harris, Minyang Tian +5Rotary Positional EmbeddingsSelf-Attention

  10. STS: Efficient Sparse Attention with Speculative Token Sparsity

    May 15, 2026Ceyu Xu, Jiangnan Yu, Yongji Wu +1Self-AttentionLong-Context Language Model Inference

  11. DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts

    May 14, 2026Jiading Gai, Shuai Zhang, Xiang Song +2GPU AccelerationSelf-Attention

  12. MambaRain: Multi-Scale Mamba-Attention Framework for 0-3 Hour Precipitation Nowcasting

    May 14, 2026Chunlei Shi, Cui Wu, Xiang Xu +10MambaPrecipitation Forecasting

  13. Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity

    May 14, 2026Jiahao Tian, Yiwei Wang, Gang Yu +1Diffusion TransformerAutoregressive Diffusion

  14. Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility

    May 13, 2026Gergely Szilvasy, Manuel Faysse, Maria Lomeli +5Self-AttentionKV Caching

  15. Rethinking Graph Convolution for 2D-to-3D Hand Pose Lifting

    May 13, 2026Chanyoung Kim, Donghyun Kim, Dong-Hyun Sim +2Graph Attention NetworksGraph Representation Learning

  16. Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation

    May 13, 2026Jiayu Chen, Junbei Tang, Wenbiao Zhao +6Self-AttentionLong-Horizon Video Generation

  17. When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction

    May 13, 2026Vardhan Dongre, Joseph Hsieh, Viet Dac Lai +3Self-AttentionLLM Interpretability

  18. ASAP: Amortized Doubly-Stochastic Attention via Sliced Dual Projection

    May 13, 2026Huy Tran, Max Milkert, David HydeTransformer InferenceSelf-Attention

  19. The Routing and Filtering Structure of Attention

    May 12, 2026Shafayeth Jamil, Rehan KapadiaAttention Head AnalysisSelf-Attention

  20. Training-Inference Consistent Segmented Execution for Long-Context LLMs

    May 12, 2026Xianpeng Shang, Jiang Li, Zehua Duo +2Self-AttentionLong-Context Language Modeling

  21. USEMA: a Scalable Efficient Mamba Like Attention for Medical Image Segmentation

    May 11, 2026Elisha Dayag, Nhat Thanh Tran, Jack XinImage SegmentationHybrid CNN-Transformer Architectures

  22. Uniform Scaling Limits in AdamW-Trained Transformers

    May 11, 2026William Gibson, Christoph ReisingerTransformerSelf-Attention

  23. RelFlexformer: Efficient Attention 3D-Transformers for Integrable Relative Positional Encodings

    May 11, 2026Byeongchan Kim, Arijit Sehanobish, Avinava Dubey +23D ViTsSelf-Attention

  24. Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions

    May 11, 2026Diancheng Kang, Zheyuan Liu, Ningshan Ma +3Language Model SteeringSelf-Attention

  25. Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining

    May 11, 2026Jinchang Zhu, Jindong Li, Yuwen Hao +3Language Model PretrainingSelf-Attention