Softmax Attention

Momentum

7 papers in the last four weeks, up 133% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 47

All topics
CardsList
  1. Scaling Limits of Long-Context Transformers

    May 8, 2026Giuseppe Bruno, Shi Chen, Zhengjiang Lin +2Softmax AttentionSelf-Attention

  2. Transformers Can Implement Preconditioned Richardson Iteration for In-Context Gaussian Kernel Regression

    May 8, 2026Mingsong Yan, Dongyang Li, Charles Kulick +1Softmax AttentionTransformer Interpretability

  3. Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization

    May 8, 2026Navid Rezazadeh, Arash Gholami DavoodiSoftmax AttentionNeural Network Verification

  4. Nearly Optimal Attention Coresets

    May 7, 2026Edo Liberty, Alexandr Andoni, Eldar KleinerSoftmax AttentionEfficient Attention

  5. Better Models, Faster Training: Sigmoid Attention for single-cell Foundation Models

    Apr 29, 2026Vijay Sadashivaiah, Georgios Dasoulas, Judith Mueller +1Softmax AttentionSelf-Attention

  6. Emergent Self-Attention from Astrocyte-Gated Associative Memory Dynamics

    Apr 28, 2026Arnau Vivet, Alex ArenasSoftmax AttentionDynamical Systems

  7. QFlash: Bridging Quantization and Memory Efficiency in Vision Transformer Attention

    Apr 28, 2026Sehyeon Oh, Yongin Kwon, Jemin LeeSoftmax AttentionGPU Acceleration

  8. Transformer Approximations from ReLUs

    Apr 27, 2026Jerry Yao-Chieh Hu, Mingcheng Lu, Yi-Chen Lee +1Softmax AttentionNeural Network Approximation Theory

  9. ELSA: Exact Linear-Scan Attention for Fast and Memory-Light Vision Transformers

    Apr 26, 2026Chih-Chung Hsu, Xin-Di Ma, Wo-Ting Liao +1Efficient Transformer InferenceSoftmax Attention

  10. Sinkhorn doubly stochastic attention rank decay analysis

    Apr 9, 2026Michela Lapenna, Rita Fioresi, Bahman GharesifardSoftmax AttentionSelf-Attention

  11. Entropy-Generated Attention Beyond Softmax and Entmax: Kaniadakis and Reciprocal-Symmetric Abe Operators

    Feb 9, 2026Gunn KimSoftmax Attention

  12. Selective Rotary Position Embedding

    Nov 21, 2025Sajad Movahedi, Timur Carstensen, Arshia Afzal +3Softmax AttentionRotary Positional Embeddings

  13. Performance-Efficiency Tradeoffs in Transformers: An Approximation Theory Perspective

    Oct 4, 2025Ruoxi Yu, Haotian Jiang, Jingpu Cheng +3Softmax AttentionAttention Head Analysis

  14. Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining

    Sep 12, 2025Rupert Mitchell, Kristian KerstingLanguage Model PretrainingSoftmax Attention

  15. Attention by Synchronization in Coupled Oscillator Networks

    Date pendingFabio Pasqualetti, Taosha GuoSoftmax AttentionDynamical Systems