Efficient Attention

Momentum

15 papers in the last four weeks, up 114% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 96

All topics
CardsList
  1. Patch Hierarchical Attention Transformer for Efficient Particle Jet Tagging

    May 20, 2026Aaron Wang, Zihan Zhao, Alan Xia +5Jet TaggingEfficient Attention

  2. DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention

    May 18, 2026Yuxiang Huang, Nuno M. T. Gonçalves, Federico Alvetreti +5Self-AttentionLong-Context Modeling

  3. CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection

    May 16, 2026Jiwon Song, Dongwon Jo, Beomseok Kang +1Self-AttentionKV-Cache Management

  4. GQA-μP: The maximal parameterization update for grouped query attention

    May 14, 2026Kyle R. Chickering, Huijuan Wang, Mengxi Wu +7Grouped-Query AttentionRepresentation Learning

  5. SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference

    May 13, 2026Anay Chauhan, Gurucharan Marthi Krishna Kumar, Arion Das +4KV CachingLong-Context Language Model Inference

  6. ASAP: Amortized Doubly-Stochastic Attention via Sliced Dual Projection

    May 13, 2026Huy Tran, Max Milkert, David HydeTransformer InferenceSelf-Attention

  7. The Routing and Filtering Structure of Attention

    May 12, 2026Shafayeth Jamil, Rehan KapadiaAttention Head AnalysisSelf-Attention

  8. Variational Linear Attention: Stable Associative Memory for Long-Context Transformers

    May 11, 2026Vishal Pandey, Gopal SinghEfficient AttentionLinear Attention

  9. USEMA: a Scalable Efficient Mamba Like Attention for Medical Image Segmentation

    May 11, 2026Elisha Dayag, Nhat Thanh Tran, Jack XinImage SegmentationHybrid CNN-Transformer Architectures

  10. Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval

    May 10, 2026Zichen Zou, Xiaosong Jia, Zuxuan Wu +13D ViTsStreaming 3D Reconstruction

  11. Positional LSH: Binary Block Matrix Approximation for Attention with Linear Biases

    May 10, 2026Daniel Wolfson, Tal WagnerSelf-AttentionLong-Context Language Modeling

  12. Long Context Pre-Training with Lighthouse Attention

    May 7, 2026Bowen Peng, Subho Ghosh, Jeffrey QuesnelleLanguage Model PretrainingSelf-Attention

  13. Nearly Optimal Attention Coresets

    May 7, 2026Edo Liberty, Alexandr Andoni, Eldar KleinerSoftmax AttentionEfficient Attention

  14. Cascade Token Selection for Transformer Attention Acceleration

    May 4, 2026Stephen J. ThomasSelf-AttentionTransformer Attention

  15. Better Models, Faster Training: Sigmoid Attention for single-cell Foundation Models

    Apr 29, 2026Vijay Sadashivaiah, Georgios Dasoulas, Judith Mueller +1Softmax AttentionSelf-Attention

  16. GateMOT: Q-Gated Attention for Dense Object Tracking

    Apr 29, 2026Mingjin Lv, Zelin Liu, Feifei Shao +4Visual Object TrackingGated Attention

  17. MixerCA: An Efficient and Accurate Model for High-Performance Hyperspectral Image Classification

    Apr 28, 2026Mohammed Q. Alkhatib, Ali JamaliHyperspectral ImagingEfficient Attention

  18. LayerBoost: Layer-Aware Attention Reduction for Efficient LLMs

    Apr 23, 2026Mohamed Ali Souibgui, Jan Fostier, Rodrigo Abadía-Heredia +3Transformer AttentionLLM Inference Acceleration

  19. The Recurrent Transformer: Greater Effective Depth and Efficient Decoding

    Apr 23, 2026Costin-Andrei Oncescu, Depen Morwani, Samy Jelassi +3Efficient Transformer InferenceSelf-Attention

  20. LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel

    Apr 22, 2026Zhe Feng, Sen Lian, Changwei Wang +5Efficient ViTsEfficient Attention

  21. DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing

    Apr 21, 2026Jinyu Guo, Zhihan Zhang, Jiehui Xie +7Long-Context Language Model InferenceLLM Inference Acceleration

  22. Efficient Video Diffusion Models: Advancements and Challenges

    Apr 17, 2026Shitong Shao, Lichen Bai, Pengfei Wan +2Model CompressionVideo Diffusion Models

  23. AdaSplash-2: Faster Differentiable Sparse Attention

    Apr 16, 2026Nuno Gonçalves, Hugo Pitorro, Vlad Niculae +4GPU Kernel OptimizationSelf-Attention

  24. Beyond Pairwise Attention: Higher-Order Modular Attention for Efficient Sequence Learning

    Mar 11, 2026Shirin Amiraslani, Xin GaoSelf-AttentionTransformer Attention

  25. FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion

    Feb 5, 2026Zhuokun Chen, Jianfei Cai, Bohan ZhuangVideo Diffusion ModelsSelf-Attention

  26. Transform Trained Transformer for Accelerating Native 4K Video Generation

    Dec 15, 2025Jiangning Zhang, Junwei Zhu, Teng Hu +9Efficient Transformer InferenceEfficient Attention

  27. Accelerating Attention with Basis Decomposition

    Oct 2, 2025Jialin ZhaoSelf-AttentionTransformer Attention

  28. SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging

    Jun 10, 2025Nhat Thanh Tran, Fanghui Xue, Shuai Zhang +4Self-AttentionVision Transformer