Self-Attention

Momentum

9 papers in the last four weeks, up 50% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 195

All topics
CardsList
  1. Data-Free Pruning of Self-Attention Layers in LLMs

    Dec 3, 2025Dhananjay Saikumar, Blesson VargheseLLM PruningLLM Compression

  2. NOSA: Native and Offloadable Sparse Attention

    Oct 15, 2025Yuxiang Huang, Pengjie Wang, Jicheng Han +9KV-Cache OffloadingGPU Acceleration

  3. Accelerating Attention with Basis Decomposition

    Oct 2, 2025Jialin ZhaoSelf-AttentionTransformer Attention

  4. Policy Gradient with Self-Attention for Model-Free Distributed Nonlinear Multi-Agent Games

    Sep 22, 2025Eduardo Sebastián, Maitrayee Keskar, Eeman Iqbal +3Policy GradientSelf-Attention

  5. Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining

    Sep 12, 2025Rupert Mitchell, Kristian KerstingLanguage Model PretrainingSoftmax Attention

  6. FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel

    Aug 25, 2025Ran Yan, Youhe Jiang, Zhuoming Chen +3GPU Kernel OptimizationGPU Acceleration

  7. MILAAP: Mobile Link Allocation via Attention-based Prediction

    Jun 24, 2025Yung-Fu Chen, Anish AroraSelf-AttentionWireless Communications

  8. SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging

    Jun 10, 2025Nhat Thanh Tran, Fanghui Xue, Shuai Zhang +4Self-AttentionVision Transformer

  9. The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs

    Apr 24, 2025Piotr Nawrot, Robert Li, Renjie Huang +3Transformer InferenceLLM Evaluation

  10. You Do Not Fully Utilize Transformer's Representation Capacity

    Feb 13, 2025Gleb Gerasimov, Yaroslav Aksenov, Nikita Balagansky +2Representation CollapseMemory-Augmented Language Models

  11. Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models

    Oct 3, 2024Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva +5Self-AttentionMultiple-Choice Question Answering

  12. AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

    Date pendingShaowen Wang, Yuke Zheng, Tansheng Zhu +4Rotary Positional EmbeddingsSelf-Attention

  13. SparSTAR: Sparse Attention for SpaceTime AutoRegressive Video Synthesis

    Date pendingJongbeom Lee, Hyunwoo Yu, Jincheol Yang +2Image-to-Video GenerationSelf-Attention