Attention Mechanisms

Momentum

42 papers in the last four weeks, up 100% on the four weeks before. 0.4% of all new papers.

Jul 13Week of Sep 28

Latest papers 378

All topics
CardsList
  1. VGGT-Occ: Geometry-Grounded and Density-Aware Gated Fusion for 3D Occupancy Prediction

    May 16, 2026Xun Chen, Tianchen Deng, Rui Wang +53D Semantic Segmentation3D Occupancy Prediction

  2. How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context Learning

    May 15, 2026Entang Wang, Yiwei Wang, Aleksandra Bakalova +1Transformer InterpretabilityTask Vectors

  3. Attention Dispersion in Dynamic Graph Transformers: Diagnosis and a Transferable Fix

    May 15, 2026Jinhao Zhang, Kangfei Zhao, Qiuhao Zeng +1Distribution Shift RobustnessTemporal GNNs

  4. Grokking as Structural Inference: Transformers Need Bayesian Lottery Tickets

    May 15, 2026Kai Hidajat, Solden Stoll, Joseph AnNeural Network GeneralizationTransformer

  5. Spectral Priors vs. Attention: Investigating the Utility of Attention Mechanisms in EEG-Based Diagnosis

    May 14, 2026Tawsik Jawad, Gowtham Atluri, Vikram RavindraElectroencephalographyTime Series Classification

  6. DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts

    May 14, 2026Jiading Gai, Shuai Zhang, Xiang Song +2GPU AccelerationSelf-Attention

  7. Multi-Block Attention for Efficient Channel Estimation in IRS-Assisted mmWave MIMO

    May 14, 2026Mehrdad Momen-Tayefeh, Mehrshad Momen-Tayefeh, Maryam SabbaghianWireless CommunicationsChannel Estimation

  8. AttnGen: Attention-Guided Saliency Learning for Interpretable Genomic Sequence Classification

    May 13, 2026Rayhaneh Shabani Nia, Ali KarkehabadiNeural Network InterpretabilityAttention Mechanisms

  9. Delta Attention Residuals

    May 13, 2026Cheng Luo, Zefan Cai, Junjie HuAttention Residual ConnectionsResidual Learning

  10. When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction

    May 13, 2026Vardhan Dongre, Joseph Hsieh, Viet Dac Lai +3Self-AttentionLLM Interpretability

  11. Precision Tracked Transformer via Kalman Filtering, Kriging and Process Noise

    May 12, 2026Bo Long, Deepak Agarwal, Jelena Markovic-Voronov +2Bayesian FilteringTransformer

  12. The Routing and Filtering Structure of Attention

    May 12, 2026Shafayeth Jamil, Rehan KapadiaAttention Head AnalysisSelf-Attention

  13. Phylogenetic Tree Inference with Tropical Axial Attention

    May 12, 2026Chris Teska, Kurt Pasque, Ruriko Yoshida +1Phylogenetic InferenceAttention Mechanisms

  14. Dual-Temporal LSTM with Hybrid Attention for Airline Passenger Load Factor Forecasting: Integrating Intra-Flight and Inter-Flight Booking Dynamics

    May 12, 2026ASM Nazrul Islam, Md. Hasanul Kabir, Md. Liakot Ali +1Time Series ForecastingLong Short-Term Memory Networks

  15. Breaking Global Self-Attention Bottlenecks in Transformer-based Spiking Neural Networks with Local Structure-Aware Self-Attention

    May 12, 2026Lingdong Li, Hangming Zhang, Qiang YuSpiking TransformersEfficient ViTs

  16. Block-Based Double Decoders

    May 11, 2026Asher Labovich, Benjamin Bradley, Vanessa Alexander +1Efficient Transformer InferenceLLM Inference Acceleration

  17. SLASH the Sink: Sharpening Structural Attention Inside LLMs

    May 11, 2026Yiming Liu, Bin Lu, Xinbing Wang +2LLM Reasoning with GraphsLLM Interpretability

  18. Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration

    May 11, 2026Haowen Li, Tianxiang Li, Yi Yang +2Audio EditingControllable Music Generation

  19. Positional LSH: Binary Block Matrix Approximation for Attention with Linear Biases

    May 10, 2026Daniel Wolfson, Tal WagnerSelf-AttentionLong-Context Language Modeling

  20. The Transformer as a Polar State Estimator

    May 10, 2026Peter RacioppoTransformerLayer Normalization

  21. CATO: Charted Attention for Neural PDE Operators

    May 9, 2026Chun-Wun Cheng, Sifan Wang, Carola-Bibiane Schönlieb +1PDE Operator LearningAttention Mechanisms

  22. A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models

    May 8, 2026Zeru Shi, Zhenting Wang, Fan Yang +2LLM InterpretabilityAttention Sinks