Attention Mechanisms

Momentum

42 papers in the last four weeks, up 100% on the four weeks before. 0.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 378

All topics
CardsList
  1. On the Necessity of Attention-FFN Split in Vision Transformers

    Oct 7, 2026Junhyeok Kim, Jinyeong Kim, Jae Wan Park +1Vision TransformerAttention Mechanisms

  2. Hardware-aware Calibrated Clustered Attention for Efficient Visual Geometric Transformers

    Oct 7, 2026Weitian Wang, Shubham Rai, Cecilia De La Parra +1Efficient ViTsAttention Mechanisms

  3. Tucker Bottleneck Attention for Multi-Dimensional Sequence Modeling

    Oct 6, 2026Ryan Solgi, Parsa Madinei, Zheng ZhangAttention MechanismsLow-Rank Attention

  4. Random Feature Gaussian Process Attention: Linear-Time Probabilistic Attention with Calibrated Uncertainty

    Oct 6, 2026Amir Mohammad Mahfoozi, Zi Yang, Ying Li +1Transformer AttentionAttention Mechanisms

  5. Dynamic Positional Attention Modulation for Parameter-Efficient Fine-Tuning of Large Language Models

    Oct 6, 2026Dayan Pan, Jingyuan Wang, Xie YuAttention MechanismsParameter-Efficient Fine-Tuning

  6. HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

    Oct 5, 2026Zhuokun Chen, Xi Lin, Xiyu Wu +3Long-Context Language ModelingLinear Attention

  7. Lend Me Your Eyes: Instruction-Aware Text Embeddings via Attention Relay

    Oct 4, 2026Yiyuan Luo, Vaggos ChatziafratisText EmbeddingsAttention Mechanisms

  8. More Value per Key: Asymmetric Sparse Attention for Faster LLM Decoding

    Oct 3, 2026Noam Elata, Itay Lamprecht, Mikey Shechter +3LLM Inference AccelerationSparse Attention

  9. Attention Kernels for Learning Maps Between Heavy-Tailed Measures

    Sep 30, 2026Kailen Hargenrader, Edoardo Calvello, Bohan ChenSoftmax AttentionAttention Mechanisms

  10. Cluster Attention Neural Operators for Solving Parametric Partial Differential Equations

    Sep 30, 2026Ming Zhong, Antonio Colanera, Gianluigi Rozza +1PDE SolvingAttention Mechanisms

  11. LampAttention: Look-Ahead Mixed-Precision FlashAttention for Dedicated Accelerators

    Sep 30, 2026Stanislav Budzinskiy, Marian Gloser, Tolunay Yilmaz +5AI Accelerator InferenceLLM Inference Acceleration

  12. Attention Function as an Intrinsic Inductive Bias: How Models' Behavior Diverges in Novel Contexts

    Sep 30, 2026Dong Gyun Kang, Megha Thukral, Kwangsoo KimAttention Head AnalysisOOD Generalization

  13. When Context Changes: Understanding Update Failures in LLMs

    Sep 30, 2026Junyu Guo, Yuchen Fang, Shangding Gu +3Belief Updating in LLMsAttention Mechanisms

  14. Concept-Grounded Attention: A Controlled Evaluation of Graph-Injected Attention, Temporal Versioning, and Epistemic Status

    Sep 30, 2026Sachin Dev Duggal, Pradyumna Swarnalatha Ramanna, Alexandros VassiliadesLLM GroundingTemporal Reasoning in Language Models

  15. No Scale Left Behind: Multi-Scale Autoencoder with Bi-directional Attention for Time Series Anomaly Detection

    Sep 29, 2026Jiaheng Guo, Haochen Zhang, Yu-Chao Huang +3Reconstruction-Based Anomaly DetectionAttention Mechanisms

  16. Retrieval Capacity of Self-Attention Under Competition

    Sep 29, 2026Timur Mudarisov, Mikhail Burtsev, Radu StateSelf-AttentionLong-Context Language Modeling

  17. Pixel-Level Transformers in Remote Sensing: A Canopy Height Case Study

    Sep 29, 2026Sven Ligensa, Jan Pauls, Karsten Schrödter +2Remote Sensing Image UnderstandingRemote Sensing

  18. SMat-Attention: Structured Long-Context Sequence Modeling

    Sep 28, 2026Emile Anand, Abdullah Ateyeh, Archer Wang +1Long-Context ModelingLinear Attention

  19. Quasi Linear Kernel Attention with Infinite Capacity

    Sep 28, 2026Nicolaj Rux, Johannes Hertrich, Sebastian NeumayerTransformer ExpressivityAttention Mechanisms

  20. Dual-Stream Simultaneous Translation via 2D Grid Attention

    Sep 28, 2026Yu Pu, Wei-Qiang ZhangTransformer AttentionAttention Mechanisms

  21. TSGate: Timestep-Aware Gated Attention for Diffusion Transformers

    Sep 28, 2026Boyu Zhang, Yifan Liu, Shuxia Lin +3Diffusion TransformerGated Attention

  22. Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training

    Sep 28, 2026Junlin Chen, Daize Dong, Huanwei Di +9Softmax AttentionEfficient Attention

  23. Eyes on the Road: A Naturalistic Comparison of MTW Rider Gaze in Urban Indian Traffic

    Sep 27, 2026Prerak Srivastava, Bhaiya Vaibhaw Kumar, Kavita VemuriVisual AttentionEye Tracking

  24. SchemaMem: Schema-Indexed Recurrent Memory for Delayed State Retrieval

    Sep 27, 2026Sungwoo Goo, Hwi-yeol Yun, Sangkeun JungMemory-Augmented Neural NetworksPersistent Memory for Language Models

  25. FluidRain: Incompressible Rain Flow as an Attention Bias for Loop-in-Loop Video Deraining

    Sep 24, 2026Pu Wang, Yongcong Wang, Wenhao Li +6Attention Mechanisms

  26. Nonequilibrium Phases of Repulsive Self-Attention: Chaos, Attention Condensation, and Emergent Locality

    Sep 23, 2026Qucheng Gao, Zuyi Yang, Xiao ChenSelf-AttentionChaotic Dynamical Systems