Attention Mechanisms

Momentum

42 papers in the last four weeks, up 100% on the four weeks before. 0.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 378

All topics
CardsList
  1. Sink vs. diagonal patterns as mechanisms for attention switch and oversmoothing prevention

    May 8, 2026Peter Súkeník, Cristina López Amado, Christoph H. Lampert +1Self-AttentionTransformer Attention

  2. GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning

    May 8, 2026Brown Ebouky, Gabriele Carrino, Niccolo Avogaro +3Active PerceptionVLM Reasoning

  3. Revisiting Transformer Layer Parameterization Through Causal Energy Minimization

    May 8, 2026Jin Xu, Camille Couturier, Victor Rühle +2TransformerEnergy-Based Models

  4. Attention Transfer Is Not Universally Effective for Vision Transformers

    May 8, 2026Huaiyuan Qin, Muli Yang, Gabriel James Goenawan +4Self-AttentionVision Transformer

  5. The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity

    May 7, 2026Siquan Li, Kaiqi Jiang, Jiacheng Sun +1Self-AttentionAttention Sinks

  6. Long Context Pre-Training with Lighthouse Attention

    May 7, 2026Bowen Peng, Subho Ghosh, Jeffrey QuesnelleLanguage Model PretrainingSelf-Attention

  7. Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search

    May 7, 2026Faisal Aljehrai, Mohammed A. Alkhrashi, Alreem Almuhrij +6Visual Representation LearningInformation Retrieval

  8. Neuromorphic visual attention for Sign-language recognition on SpiNNaker

    May 7, 2026Sarka Liskova, Olha Vedmedenko, Mazdak Fatahi +3Event-Based VisionVisual Attention

  9. Retrieval from Within: An Intrinsic Capability of Attention-Based Models

    May 7, 2026Elad Hoffer, Yochai Blau, Edan Kinderman +3Retrieval-Augmented GenerationQuestion Answering

  10. Large Vision-Language Models Get Lost in Attention

    May 7, 2026Gongli Xi, Ye Tian, Mengyu Yang +5Transformer FFNsVision-Language Models

  11. Average Attention Transformers and Arithmetic Circuits

    May 6, 2026Lena Ehrmuth, Laura StriekerTransformer ExpressivityTransformer

  12. Angle-I2P: Angle-Consistent-Aware Hierarchical Attention for Cross-Modality Outlier Rejection

    May 6, 2026Muyao Peng, Shun Zou, Pei An +2Point Cloud RegistrationCross-Modal Feature Matching

  13. FLUID: Continuous-Time Hyperconnected Sparse Transformer for Sink-Free Learning

    May 6, 2026Waleed Razzaq, Yun-Bo ZhaoTransformerTransformer Attention

  14. How Language Models Process Negation

    May 4, 2026Zhejian Zhou, Tianyi Zhou, Robin Jia +1LLM InterpretabilityLanguage Modeling

  15. When Attention Collapses: Residual Evidence Modeling for Compositional Inference

    May 4, 2026Niklas HoubaGravitational-Wave AstronomySelf-Attention

  16. Projection-Free Transformers via Gaussian Kernel Attention

    May 4, 2026Debarshi Kundu, Archisman Ghosh, Swaroop Ghosh +1TransformerSelf-Attention

  17. Attention Is Where You Attack

    Apr 30, 2026Aviral Srivastava, Sourav PandaSelf-AttentionAdversarial Attacks

  18. Better Models, Faster Training: Sigmoid Attention for single-cell Foundation Models

    Apr 29, 2026Vijay Sadashivaiah, Georgios Dasoulas, Judith Mueller +1Softmax AttentionSelf-Attention

  19. From Architecture to Output: Structural Origins of Hallucination in Large Language Models and the Amplifying Role of Data

    Apr 29, 2026Md. Rejaul Korim Sadi, Toufiqur Rahman Tasin, Golam Mostofa NaeemHallucination in Language ModelsLLM Reliability

  20. GateMOT: Q-Gated Attention for Dense Object Tracking

    Apr 29, 2026Mingjin Lv, Zelin Liu, Feifei Shao +4Visual Object TrackingGated Attention

  21. AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving

    Apr 28, 2026Zhongkai Yu, Haotian Ye, Chenyang Zhou +9Compute-in-MemoryLLM Inference Acceleration

  22. Emergent Self-Attention from Astrocyte-Gated Associative Memory Dynamics

    Apr 28, 2026Arnau Vivet, Alex ArenasSoftmax AttentionDynamical Systems