Self-Attention

Momentum

9 papers in the last four weeks, up 50% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 195

All topics
CardsList
  1. Remember to Forget: Gated Adaptive Positional Encoding

    May 11, 2026Riccardo Ali, Alessio Borgi, Mario Severino +2Rotary Positional EmbeddingsSelf-Attention

  2. FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning

    May 11, 2026Zehua Pei, Hui-Ling Zhen, Xianzhi Yu +3Supervised Fine-TuningSelf-Attention

  3. Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction

    May 10, 2026Ngoc Bui, Hieu Trung Nguyen, Arman Cohan +1Self-AttentionKV Caching

  4. A Game Theoretic Free Energy Analysis of Higher Order Synergy in Attention Heads of Large Language Models

    May 10, 2026Djamel BouchaffraLLM PruningAttention Head Analysis

  5. Positional LSH: Binary Block Matrix Approximation for Attention with Linear Biases

    May 10, 2026Daniel Wolfson, Tal WagnerSelf-AttentionLong-Context Language Modeling

  6. How LLMs Are Persuaded: A Few Attention Heads, Rerouted

    May 10, 2026Xiangkun Sun, Lingkai Kong, Aoqi Zhang +2Self-AttentionLLM Interpretability

  7. Attention Sinks in Diffusion Transformers: A Causal Analysis

    May 10, 2026Fangzheng Wu, Brian SummaTransformer InterpretabilityDiffusion Transformer

  8. Scaling Limits of Long-Context Transformers

    May 8, 2026Giuseppe Bruno, Shi Chen, Zhengjiang Lin +2Softmax AttentionSelf-Attention

  9. Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression

    May 8, 2026Mingsong Yan, Dongyang Li, Charles Kulick +1Softmax AttentionTransformer Interpretability

  10. Sink vs. diagonal patterns as mechanisms for attention switch and oversmoothing prevention

    May 8, 2026Peter Súkeník, Cristina López Amado, Christoph H. Lampert +1Self-AttentionTransformer Attention

  11. An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference

    May 8, 2026Feiyu Yao, Zhixiong Niu, Xiaqing Li +3GPU AccelerationSelf-Attention

  12. A Marine Debris Detection Framework for Ocean Robots via Self-Attention Enhancement and Feature Interaction Optimization

    May 8, 2026Yuyang Li, Jiashu Han, Yinyi Lai +2Self-AttentionSmall Object Detection

  13. Attention Transfer Is Not Universally Effective for Vision Transformers

    May 8, 2026Huaiyuan Qin, Muli Yang, Gabriel James Goenawan +4Self-AttentionVision Transformer

  14. Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache

    May 7, 2026Mohsen Dehghankar, Abolfazl AsudehSelf-AttentionKV Caching

  15. The Structural Origin of Attention Sink: Variance Discrepancy, Super Neurons, and Dimension Disparity

    May 7, 2026Siquan Li, Kaiqi Jiang, Jiacheng Sun +1Self-AttentionAttention Sinks

  16. Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent

    May 7, 2026Chenyang Zhang, Yuan CaoLogistic RegressionSelf-Attention

  17. Long Context Pre-Training with Lighthouse Attention

    May 7, 2026Bowen Peng, Subho Ghosh, Jeffrey QuesnelleLanguage Model PretrainingSelf-Attention

  18. Don't Lose Focus: Activation Steering via Key-Orthogonal Projections

    May 7, 2026Haoyan Luo, Mateo Espinosa Zarlenga, Mateja JamnikLanguage Model SteeringSelf-Attention

  19. Shallow Prefill, Deep Decoding: Efficient Long-Context Inference via Layer-Asymmetric KV Visibility

    May 7, 2026Jungsuk Oh, Hyeseo Jeon, Hyunjune Ji +2Self-AttentionLong-Context Language Model Inference

  20. Large Vision-Language Models Get Lost in Attention

    May 7, 2026Gongli Xi, Ye Tian, Mengyu Yang +5Transformer FFNsVision-Language Models

  21. CuBridge: An LLM-Based Framework for Understanding and Reconstructing High-Performance Attention Kernels

    May 6, 2026Xing Ma, Yangjie Zhou, Wu Sun +6GPU Kernel OptimizationGPU Acceleration