Attention Mechanisms

Momentum

42 papers in the last four weeks, up 100% on the four weeks before. 0.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 378

All topics
CardsList
  1. A Systematic Benchmark of Explainable Methods for Temporal Attribution in Sequential Recommendation Systems

    Sep 23, 2026Akash Pandey, Kanisha Shah, Addrish Roy +5Gradient-Based AttributionSequential Recommendation

  2. Rethinking Pairwise Token Interaction in Spiking Transformers

    Sep 22, 2026Sicheng Shen, Dongcheng Zhao, Zhiyuan Li +5Spiking TransformersTransformer Attention

  3. TACIT: Tactile Contact Supervision for Spatial Attention in Dexterous Manipulation

    Sep 21, 2026Yanhou Lai, Fucai Zhu, Ruiqiang Wang +1Contact-Rich Robotic ManipulationRobot Imitation Learning

  4. On-Demand Attention: Language Models Know When to Recall

    Sep 17, 2026Haibo Feng, Ruiqi Liang, Hanyang Peng +1Language Model DecodingLong-Context Language Model Inference

  5. SETTer: Sparse-Encoder Transformer for Long-term Multivariate Time Series Forecasting

    Sep 17, 2026Abraham Ezema, Chijioke Eze, Ferdinanda Ponci +1Multivariate Time Series ForecastingLong-Term Time Series Forecasting

  6. Beyond Flattened Tokens: Structure-Preserving EEG Decoding with Reusable TriDim Blocks

    Sep 17, 2026Shiyue Su, Song Wang, Zekai Zhan +7Cross-Subject EEG DecodingEEG Decoding

  7. Event-based Selective Attention for Multi-resolution Fast Region of Interest (ROI) Detection

    Sep 15, 2026Luca Peres, Giulia D'Angelo, Chiara Bartolozzi +1Event-Based VisionVisual Attention

  8. Training Trajectories Determine Circuit Removability in Annealable Soft-Prior Transformers

    Sep 9, 2026Zonglin Yang, Ziming Zhao, Wei Tang +5Transformer InterpretabilityAttention Mechanisms

  9. It's Not RoPE that Creates Sinks: The Role of Self-Concentration and Value-Non-Mixing in Attention

    Sep 8, 2026Raito Kiya, Satoki Ohashi, Kosuke Sato +6Self-AttentionAttention Sinks

  10. Adaptive Anisotropic Attention for Axis-Structured Signals

    Sep 8, 2026Mahir Jain, Parshva Runwal, Aditya Ray Mishra +4EEG DecodingAttention Mechanisms

  11. Why shared attention vectors fail: a case for outcome-indexed tuning

    Sep 8, 2026Lenard DomeAttention Mechanisms

  12. Do New Attention Mechanisms Actually Fix Attention Sinks at Million-Token Context?

    Sep 8, 2026Sara Rizwan, Samaanah Abdus SalamLong-Context Language Model EvaluationLong-Context Language Modeling

  13. DeepTable: Structural Attention Biases and Tree Path Encoding for Hierarchical Table Understanding

    Sep 7, 2026Jyun-Ying Yen, Cheng-Kuan Lin, Yu-Chee TsengTable QATable Structure Recognition

  14. RedKnot-MLA: Multi-Head Offline-Online Reuse for DeepSeek-V4 Long-Context Serving

    Sep 7, 2026Yang Liu, Zhaokai Luo, Huayi Jin +10KV CachingLLM Inference Acceleration

  15. RoLA: Rotary-Positioned Low-Rank Linear Attention for Efficient Diffusion Transformers

    Sep 6, 2026Zekun Zhang, Yixiang Cai, Yuxi Liu +9Rotary Positional EmbeddingsLow-Rank Attention

  16. The Head Complexity of Boolean Functions in Single-Layer Attention

    Sep 3, 2026Rajmohan Rajaraman, Ravi Sundaram, Amanuel TesfayeSelf-AttentionAttention Mechanisms

  17. High-Dimensional Learning Dynamics of Attention-Indexed Models

    Sep 3, 2026Yizhou Xu, Margarita Sagitova, Lenka Zdeborová +1Implicit BiasAttention Mechanisms

  18. Scaled Idempotence in Transformer Attention: Paired OV Geometry and Shared-Value Algebras

    Sep 1, 2026Jiming Feng, Junliang LiAttention Head AnalysisTransformer Attention

  19. Toppling the Hierarchy in Byte-level Language Modeling

    Aug 31, 2026Lukas Edman, Alexander FraserEfficient Language Model TrainingByte-Level Language Model

  20. SELECT: SELEctive Context Transfer for Class-Incremental Semantic Segmentation

    Aug 31, 2026Avi Gupta, Saurabh Yadav, Koteswar Rao Jerripothula +1Class-Incremental LearningAttention Mechanisms

  21. TPR-Attention for Combinatorial Generalization

    Aug 31, 2026Melisa Civelekoğlu, Isabeau Prémont-SchwarzAttention MechanismsCompositional Representation Learning

  22. A Lightweight Phenology-Aware YOLOv5 Framework for Tomato Growth Stage Detection in Resource-Constrained Bhutanese Greenhouse Environments

    Aug 30, 2026Sherab Gocha, Sou NobukawaPlant PhenotypingObject Detection

  23. Intrinsic Interaction Geometry Controls the Low-Rank Complexity of Softmax Attention

    Aug 28, 2026Yuhe Sui, Jianing Zhang, Yingzhi TangSoftmax AttentionLow-Rank Attention