Softmax Attention

Momentum

7 papers in the last four weeks, up 133% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 47

All topics
CardsList
  1. Softmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It Must

    Oct 8, 2026Simon Gabet, Etienne Boursier, Claire BoyerSoftmax AttentionIn-Context Learning

  2. Learning Decision-Stump Thresholds in Context: Dynamics of Softmax Attention

    Oct 5, 2026Hong Ha Le, Jackie Lok, Atsushi Nitanda +1Language Model PretrainingSoftmax Attention

  3. Universal interpolation for deep residual self-attention networks

    Oct 1, 2026Sibylle Marcotte, Joan BrunaSoftmax AttentionTransformer

  4. Attention Kernels for Learning Maps Between Heavy-Tailed Measures

    Sep 30, 2026Kailen Hargenrader, Edoardo Calvello, Bohan ChenSoftmax AttentionAttention Mechanisms

  5. Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training

    Sep 28, 2026Junlin Chen, Daize Dong, Huanwei Di +9Softmax AttentionEfficient Attention

  6. Approximating Softmax in Pretrained LLMs: Model Sensitivity and Kernel Acceleration

    Sep 27, 2026Shangzhen Zhu, Muyan Hu, Tomasz KozlowskiEfficient Transformer InferenceSoftmax Attention

  7. Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

    Sep 18, 2026Richard Zhe WangSoftmax AttentionGated Attention

  8. Intrinsic Interaction Geometry Controls the Low-Rank Complexity of Softmax Attention

    Aug 28, 2026Yuhe Sui, Jianing Zhang, Yingzhi TangSoftmax AttentionLow-Rank Attention

  9. A Quantum Roadmap for Softmax Attention: Exact Born-Rule Analogs for Softmax Attention on the Probability Simplex

    Aug 11, 2026Eric A. F. Reinhardt, Adam J. HauserSoftmax AttentionQuantum Machine Learning

  10. Provably Learning Multi-Head Attention with Queries

    Aug 4, 2026Sunyeop Kim, Insung Kim, Jian GuoSoftmax AttentionTransformer

  11. One QK Channel, Many Sources: Tracing Low-Precision Attention Collapse

    Aug 3, 2026Shuxiao Xie, Shuyang Xie, Yuan Cao +3Softmax AttentionTransformer Attention

  12. ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression

    Jul 31, 2026Yuhang Zhan, Lisi Chen, Shuo ShangSoftmax AttentionKV Caching

  13. Self-Attention Dynamics with Rotary Position Embeddings: Twisted States and Explicit Consensus Rates on the Sphere

    Jul 27, 2026Hao YeSoftmax AttentionRotary Positional Embeddings

  14. Relative Positions Generalize, Absolute Positions Memorize: An Implicit-Bias Account of Length Generalization in Attention

    Jul 21, 2026Subham Singh, Ashutosh Mishra, Subha RautSoftmax AttentionRotary Positional Embeddings

  15. BMFA: Boundary-Minority Free-Energy Adaptive Screening

    Jul 20, 2026Wenyan Xu, Alizer WongSoftmax AttentionVision Transformer

  16. The Key to Going Linear: Analysis-Driven Transformer Linearization

    Jul 8, 2026Anna Kuzina, Paul N. Whatmough, Babak Ehteshami BejnordiSoftmax AttentionLLM Inference Acceleration

  17. Asymptotic Signal Subspace Recovery in Softmax Attention Models

    Jun 21, 2026Lan V. TruongSoftmax AttentionGradient Descent Dynamics

  18. Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence

    Jun 10, 2026Itay Lavie, Kirsten Fischer, Andrey Lekov +3Softmax AttentionPhase Transitions

  19. Forget Attention: Importance-Aware Attention Is All You Need

    Jun 1, 2026Suhyeong Shin, Yeongwook YangSoftmax AttentionLong-Context Language Modeling