Attention Mechanisms

Momentum

42 papers in the last four weeks, up 100% on the four weeks before. 0.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 378

All topics
CardsList
  1. Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining

    Sep 12, 2025Rupert Mitchell, Kristian KerstingLanguage Model PretrainingSoftmax Attention

  2. FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel

    Aug 25, 2025Ran Yan, Youhe Jiang, Zhuoming Chen +3GPU Kernel OptimizationGPU Acceleration

  3. MGDFIS: Multi-scale Global-detail Feature Integration Strategy for Small Object Detection

    Jun 15, 2025Yuxiang Wang, Xuecheng Bai, Chuanzhi Xu +2Remote Sensing Small Object DetectionSmall Object Detection

  4. SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging

    Jun 10, 2025Nhat Thanh Tran, Fanghui Xue, Shuai Zhang +4Self-AttentionVision Transformer

  5. ISAC: Training-Free Instance-to-Semantic Attention Control for Multi-Instance Generation

    May 27, 2025Sanghyun Jo, Wooyeol Lee, Ziseok Lee +3Diffusion Model GuidanceT2I Generation

  6. The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs

    Apr 24, 2025Piotr Nawrot, Robert Li, Renjie Huang +3Transformer InferenceLLM Evaluation

  7. Ordinary Least Squares as an Attention Mechanism

    Apr 13, 2025Philippe Goulet CoulombeTransformer AttentionLinear Regression

  8. You Do Not Fully Utilize Transformer's Representation Capacity

    Feb 13, 2025Gleb Gerasimov, Yaroslav Aksenov, Nikita Balagansky +2Representation CollapseMemory-Augmented Language Models

  9. Attention is All You Need Until You Need Retention

    Jan 15, 2025M. Murat YasliogluMemory-Augmented Language ModelsAgent Memory Poisoning

  10. A Mechanistic Study of Transformers Training Dynamics

    Oct 31, 2024Ambroise Odonnat, Wassim Bouaziz, Vivien CabannesTransformer InterpretabilityAttention Mechanisms

  11. Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models

    Oct 3, 2024Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva +5Self-AttentionMultiple-Choice Question Answering

  12. GFPack++: Attention-Driven Gradient Fields for Optimizing 2D Irregular Packing

    Jun 9, 2024Tianyang Xue, Lin Lu, Yang Liu +5Combinatorial OptimizationConstrained Optimization

  13. A retrieval conditioned rebinding circuit for dynamic entity tracking in large language models

    Date pendingSoyoung Oh, Vera DembergLLM InterpretabilityMechanistic Interpretability

  14. Attention by Synchronization in Coupled Oscillator Networks

    Date pendingFabio Pasqualetti, Taosha GuoSoftmax AttentionDynamical Systems

  15. Relevance Is Not Permission: Localizing and Controlling Metric-Facing Attention Contributions

    Date pendingMinwoo Yu, Young-guk HaFeature AttributionEvidence Selection

  16. AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

    Date pendingShaowen Wang, Yuke Zheng, Tansheng Zhu +4Rotary Positional EmbeddingsSelf-Attention

  17. Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling

    Date pendingArman Adibi, Alireza Jafari, Mohammad Ghavamzadeh +1TransformerDiffusion Sampling