Attention Mechanisms

Momentum

42 papers in the last four weeks, up 100% on the four weeks before. 0.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 378

All topics
CardsList
  1. Cross-Modal Attention Analysis and Optimization in Vision-Language Models: A Study on Visual Reliability

    Apr 19, 2026Lijie ZhouCross-Modal LearningVLM Robustness

  2. SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models

    Apr 18, 2026Junnan Liu, Xinyan Liu, Peifeng Gao +4GPU AccelerationSelf-Attention

  3. AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency

    Apr 17, 2026Max Henning Höth, Kristian Kersting, Björn Deiseroth +1CoT FaithfulnessFeature Attribution

  4. Improving Reasoning Capabilities in Small Models through Mixture-of-Layers Distillation with Stepwise Attention on Key Information

    Apr 17, 2026Yao Chen, Jiawei Sheng, Wenyuan Zhang +1CoT DistillationSmall Language Models

  5. LACE: Lattice Attention for Cross-thread Exploration

    Apr 16, 2026Yang Li, Zirui Zhang, Yang Liu +1LLM InferenceAttention Mechanisms

  6. AdaSplash-2: Faster Differentiable Sparse Attention

    Apr 16, 2026Nuno Gonçalves, Hugo Pitorro, Vlad Niculae +4GPU Kernel OptimizationSelf-Attention

  7. Dispatch-Aware Ragged Attention for Pruned Vision Transformers

    Apr 16, 2026Seifeldin Abdellatif, Ahmad AlmasriEfficient Transformer InferenceVisual Attention

  8. Expressivity of Transformers: A Tropical Geometry Perspective

    Apr 16, 2026Ye Su, Yong LiuTransformer ExpressivityTransformer

  9. Gating Enables Curvature: A Geometric Expressivity Gap in Attention

    Apr 16, 2026Satwik Bathula, Anand A. JoshiInformation GeometryNeural Representation Geometry

  10. What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

    Apr 9, 2026Stephen Cheng, Sarah Wiegreffe, Dinesh ManochaLanguage Model SteeringLLM Interpretability

  11. Sinkhorn doubly stochastic attention rank decay analysis

    Apr 9, 2026Michela Lapenna, Rita Fioresi, Bahman GharesifardSoftmax AttentionSelf-Attention

  12. Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference

    Apr 8, 2026Quantong Qiu, Zhiyi Hong, Yi Yang +5Self-AttentionLong-Context Language Model Inference

  13. Invertible Query-Key Coupling Composes with Attention Mechanisms

    Apr 2, 2026Barak Gahtan, Alex M. BronsteinInvertible Neural NetworksSelf-Attention

  14. Screening Is Enough

    Apr 1, 2026Ken M. NakanishiSelf-AttentionLong-Context Language Modeling

  15. When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models

    Apr 1, 2026Jiho Choi, Jaemin Kim, Sanghwan Kim +2Large Vision-Language ModelsVLM Interpretability

  16. From Dispersion to Attraction: Spectral Dynamics of Hallucination Across Whisper Model Scales

    Mar 31, 2026Ivan Viakhirev, Kirill Borodin, Grach MkrtchianTransformer InterpretabilityAutomatic Speech Recognition

  17. Pure and physics-guided deep learning approaches for spatio-temporal groundwater level prediction

    Mar 26, 2026Matteo Salis, Gabriele Sartor, Rosa Meo +2Physics-Informed MLAttention Mechanisms

  18. Learning When to Attend: Conditional Memory Access for Long-Context LLMs

    Mar 18, 2026Sakshi Choudhary, Aditya Chattopadhyay, Luca Zancato +4Self-AttentionLong-Context Language Modeling

  19. WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models

    Mar 17, 2026Hanna Lee, Tan Dat Nguyen, Jaehoon Kang +1TTS SynthesisSelf-Attention

  20. Beyond Pairwise Attention: Higher-Order Modular Attention for Efficient Sequence Learning

    Mar 11, 2026Shirin Amiraslani, Xin GaoSelf-AttentionTransformer Attention

  21. StructSAM: Structure- and Spectrum-Preserving Token Merging for Segment Anything Models

    Mar 7, 2026Duy M. H. Nguyen, Tuan A. Tran, Duong Nguyen +17Token MergingImage Segmentation

  22. SHINE: Sequential Hierarchical Integration Network for EEG and MEG

    Feb 27, 2026Xiran Xu, Yujie Yan, Songyi Li +4ElectroencephalographyNeural Processes

  23. Incremental Learning of Sparse Attention Patterns in Transformers

    Feb 22, 2026Oğuz Kaan Yüksel, Rodrigo Alvarez Lucendo, Nicolas FlammarionSelf-AttentionTransformer Attention