Sparse Attention

Momentum

24 papers in the last four weeks, up 200% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 160

All topics
CardsList
  1. Stochastic Sparse Attention for Memory-Bound Inference

    May 3, 2026Kyle Lee, Corentin Delacour, Kevin Callahan-Coray +5GPU AccelerationSelf-Attention

  2. Spectral Dynamic Attention Network for Hyperspectral Image Super-Resolution

    Apr 30, 2026Tengya Zhang, Feng Gao, Lin Qi +2Hyperspectral Image Super-ResolutionHyperspectral Imaging

  3. Unifying Sparse Attention with Hierarchical Memory for Scalable Long-Context LLM Serving

    Apr 29, 2026Zihan Zhao, Baotong Lu, Shengjie Lin +8LLM ServingKV-Cache Offloading

  4. SparseContrast: Dynamic Sparse Attention for Efficient and Accurate Contrastive Learning in Medical Imaging

    Apr 27, 2026Paarth Prasad, Ruchika MalhotraContrastive LearningChest X-Ray Classification

  5. BSViT: A Burst Spiking Vision Transformer for Expressive and Efficient Visual Representation Learning

    Apr 25, 2026Hongxiang Peng, Dewei Bai, Hong QuSpiking TransformersEfficient ViTs

  6. SpikingBrain2.0: Brain-Inspired Foundation Models for Efficient Long-Context and Cross-Platform Inference

    Apr 24, 2026Yuqi Pan, Jinghao Zhuang, Yupeng Feng +16Long-Context Language Model InferenceSpiking Neural Networks

  7. HubRouter: A Pluggable Sub-Quadratic Routing Primitive for Hybrid Sequence Models

    Apr 24, 2026Abhinaba BasuTransformer AttentionLanguage Modeling

  8. Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion Transformers

    Apr 23, 2026Minghao Yin, Wenbo Hu, Jiale Xu +2Diffusion Transformer3D Shape Generation

  9. Sparse Forcing: Native Trainable Sparse Attention for Real-time Autoregressive Diffusion Video Generation

    Apr 23, 2026Boxun Xu, Yuming Du, Zichang Liu +7Autoregressive DiffusionSelf-Attention

  10. DynamicRad: Content-Adaptive Sparse Attention for Long Video Diffusion

    Apr 22, 2026Yongji Long, Shijun Liang, Jintao Li +1Video Diffusion ModelsLong-Horizon Video Generation

  11. Simplified Sparse Attention via Gist Tokens

    Apr 22, 2026Yuzhen Mao, Michael Y. Li, Emily B. FoxSelf-AttentionLong-Context Language Model Inference

  12. AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generation

    Apr 20, 2026Haoyue Tan, Shengnan Wang, Yulin Qiao +5Video Diffusion ModelsDiffusion Transformer

  13. AdaSplash-2: Faster Differentiable Sparse Attention

    Apr 16, 2026Nuno Gonçalves, Hugo Pitorro, Vlad Niculae +4GPU Kernel OptimizationSelf-Attention

  14. Improving Sparse Autoencoder with Dynamic Attention

    Apr 16, 2026Dongsheng Wang, Jinsen Zhang, Dawei Su +1Sparse AutoencodersNeural Network Interpretability

  15. Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference

    Apr 8, 2026Quantong Qiu, Zhiyi Hong, Yi Yang +5Self-AttentionLong-Context Language Model Inference

  16. LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows

    Apr 6, 2026Zhengqin Li, Cheng Zhang, Jakob Engel +1Inverse Rendering3D ViTs

  17. Learning When to Attend: Conditional Memory Access for Long-Context LLMs

    Mar 18, 2026Sakshi Choudhary, Aditya Chattopadhyay, Luca Zancato +4Self-AttentionLong-Context Language Modeling

  18. WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models

    Mar 17, 2026Hanna Lee, Tan Dat Nguyen, Jaehoon Kang +1TTS SynthesisSelf-Attention

  19. When Does Sparsity Mitigate the Curse of Depth in LLMs

    Mar 16, 2026Dilxat Muhtar, Xinyuan Song, Sebastian Pokutta +4Grouped-Query AttentionLanguage Model Scaling

  20. Incremental Learning of Sparse Attention Patterns in Transformers

    Feb 22, 2026Oğuz Kaan Yüksel, Rodrigo Alvarez Lucendo, Nicolas FlammarionSelf-AttentionTransformer Attention

  21. Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference

    Feb 9, 2026Yifei Gao, Lei Wang, Rong-Cheng Tu +3Self-AttentionKV Caching

  22. Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling

    Jan 17, 2026Xingyue Huang, Xueying Ding, Mingxuan Ju +3Self-AttentionLong-Context Language Modeling

  23. Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers

    Jan 14, 2026Yuxi Liu, Yipeng Hu, Zekun Zhang +2Video Diffusion ModelsDiffusion Transformer