Attention Sinks

Momentum

6 papers in the last four weeks, up 50% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 42

All topics
CardsList
  1. A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models

    May 8, 2026Zeru Shi, Zhenting Wang, Fan Yang +2LLM InterpretabilityAttention Sinks

  2. Sink vs. diagonal patterns as mechanisms for attention switch and oversmoothing prevention

    May 8, 2026Peter Súkeník, Cristina López Amado, Christoph H. Lampert +1Self-AttentionTransformer Attention

  3. The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity

    May 7, 2026Siquan Li, Kaiqi Jiang, Jiacheng Sun +1Self-AttentionAttention Sinks

  4. Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs

    Apr 22, 2026Kibum Kim, Jiwan Kim, Kyle Min +4Efficient VLM InferenceAttention Sinks

  5. SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models

    Apr 18, 2026Junnan Liu, Xinyan Liu, Peifeng Gao +4GPU AccelerationSelf-Attention

  6. When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models

    Apr 1, 2026Jiho Choi, Jaemin Kim, Sanghwan Kim +2Large Vision-Language ModelsVLM Interpretability

  7. Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling

    Jan 17, 2026Xingyue Huang, Xueying Ding, Mingxuan Ju +3Self-AttentionLong-Context Language Modeling

  8. Garbage Attention in Large Language Models: BOS Sink Heads and Sink-aware Pruning

    Jan 11, 2026Jaewon Sok, Jewon Yeom, Seonghyeon Park +2LLM PruningAttention Head Analysis

  9. RefAM: Attention Magnets for Zero-Shot Referral Segmentation

    Sep 26, 2025Anna Kukleva, Enis Simsar, Alessio Tonioni +4Zero-Shot LearningImage Segmentation