Self-Attention

Momentum

9 papers in the last four weeks, up 50% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 195

All topics
CardsList
  1. Stability Implies Redundancy: Delta Attention Selective Halting for Efficient Long-Context Prefilling

    Apr 20, 2026Yujie Chen, Tailai Chen, Yifeng Gao +4Self-AttentionLong-Context Language Model Inference

  2. HeadRank: Decoding-Free Passage Reranking via Preference-Aligned Attention Heads

    Apr 19, 2026Juyuan Wang, Chenxing Wang, Yuchen Fang +8Attention Head AnalysisSelf-Attention

  3. SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models

    Apr 18, 2026Junnan Liu, Xinyan Liu, Peifeng Gao +4GPU AccelerationSelf-Attention

  4. AdaSplash-2: Faster Differentiable Sparse Attention

    Apr 16, 2026Nuno Gonçalves, Hugo Pitorro, Vlad Niculae +4GPU Kernel OptimizationSelf-Attention

  5. Dispatch-Aware Ragged Attention for Pruned Vision Transformers

    Apr 16, 2026Seifeldin Abdellatif, Ahmad AlmasriEfficient Transformer InferenceVisual Attention

  6. Expressivity of Transformers: A Tropical Geometry Perspective

    Apr 16, 2026Ye Su, Yong LiuTransformer ExpressivityTransformer

  7. Gating Enables Curvature: A Geometric Expressivity Gap in Attention

    Apr 16, 2026Satwik Bathula, Anand A. JoshiInformation GeometryNeural Representation Geometry

  8. Sinkhorn doubly stochastic attention rank decay analysis

    Apr 9, 2026Michela Lapenna, Rita Fioresi, Bahman GharesifardSoftmax AttentionSelf-Attention

  9. Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference

    Apr 8, 2026Quantong Qiu, Zhiyi Hong, Yi Yang +5Self-AttentionLong-Context Language Model Inference

  10. Invertible Query-Key Coupling Composes with Attention Mechanisms

    Apr 2, 2026Barak Gahtan, Alex M. BronsteinInvertible Neural NetworksSelf-Attention

  11. Screening Is Enough

    Apr 1, 2026Ken M. NakanishiSelf-AttentionLong-Context Language Modeling

  12. Learning When to Attend: Conditional Memory Access for Long-Context LLMs

    Mar 18, 2026Sakshi Choudhary, Aditya Chattopadhyay, Luca Zancato +4Self-AttentionLong-Context Language Modeling

  13. WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models

    Mar 17, 2026Hanna Lee, Tan Dat Nguyen, Jaehoon Kang +1TTS SynthesisSelf-Attention

  14. Beyond Pairwise Attention: Higher-Order Modular Attention for Efficient Sequence Learning

    Mar 11, 2026Shirin Amiraslani, Xin GaoSelf-AttentionTransformer Attention

  15. StructSAM: Structure- and Spectrum-Preserving Token Merging for Segment Anything Models

    Mar 7, 2026Duy M. H. Nguyen, Tuan A. Tran, Duong Nguyen +17Token MergingImage Segmentation

  16. Incremental Learning of Sparse Attention Patterns in Transformers

    Feb 22, 2026Oğuz Kaan Yüksel, Rodrigo Alvarez Lucendo, Nicolas FlammarionSelf-AttentionTransformer Attention

  17. Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference

    Feb 9, 2026Yifei Gao, Lei Wang, Rong-Cheng Tu +3Self-AttentionKV Caching

  18. ResidualKV: Residual-Based KV Cache Compression for Efficient Long-Context Inference

    Feb 8, 2026Jitai Hao, Qiang Huang, Yaowei Wang +2Self-AttentionKV Caching

  19. Quantum Attention by Overlap Interference: Predicting Classical and Many-Body Quantum Sequences

    Feb 6, 2026Alessio Pecilli, Matteo RosatiSelf-AttentionTransformer Attention

  20. FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion

    Feb 5, 2026Zhuokun Chen, Jianfei Cai, Bohan ZhuangVideo Diffusion ModelsSelf-Attention

  21. Poly-attention: a general scheme for higher-order self-attention

    Feb 2, 2026Sayak Chakrabarti, Toniann Pitassi, Josh AlmanSelf-AttentionCompositional Reasoning

  22. Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling

    Jan 17, 2026Xingyue Huang, Xueying Ding, Mingxuan Ju +3Self-AttentionLong-Context Language Modeling

  23. Garbage Attention in Large Language Models: BOS Sink Heads and Sink-aware Pruning

    Jan 11, 2026Jaewon Sok, Jewon Yeom, Seonghyeon Park +2LLM PruningAttention Head Analysis

  24. Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference

    Dec 18, 2025Dhruv Deshmukh, Saurabh Goyal, Nipun Kwatra +1Self-AttentionLong-Context Language Model Inference

  25. Block Sparse Flash Attention

    Dec 7, 2025Daniel Ohayon, Itay Lamprecht, Itay Hubara +3GPU AccelerationSelf-Attention