Sparse Attention

Momentum

25 papers in the last four weeks, up 213% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 164

All topics
CardsList
  1. Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference

    Feb 9, 2026Yifei Gao, Lei Wang, Rong-Cheng Tu +3Self-AttentionKV Caching

  2. Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling

    Jan 17, 2026Xingyue Huang, Xueying Ding, Mingxuan Ju +3Self-AttentionLong-Context Language Modeling

  3. Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers

    Jan 14, 2026Yuxi Liu, Yipeng Hu, Zekun Zhang +2Video Diffusion ModelsDiffusion Transformer

  4. From Local Windows to Adaptive Candidates via Individualized Exploratory: Rethinking Attention for Image Super-Resolution

    Jan 13, 2026Chunyu Meng, Wei Long, Shuhang GuEfficient ViTsSparse Attention

  5. Kascade: A Practical Sparse Attention Method for Long-Context LLM Inference

    Dec 18, 2025Dhruv Deshmukh, Saurabh Goyal, Nipun Kwatra +1Self-AttentionLong-Context Language Model Inference

  6. Block Sparse Flash Attention

    Dec 7, 2025Daniel Ohayon, Itay Lamprecht, Itay Hubara +3GPU AccelerationSelf-Attention

  7. NOSA: Native and Offloadable Sparse Attention

    Oct 15, 2025Yuxiang Huang, Pengjie Wang, Jicheng Han +9KV-Cache OffloadingGPU Acceleration

  8. Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining

    Sep 12, 2025Rupert Mitchell, Kristian KerstingLanguage Model PretrainingSoftmax Attention

  9. FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel

    Aug 25, 2025Ran Yan, Youhe Jiang, Zhuoming Chen +3GPU Kernel OptimizationGPU Acceleration

  10. NABLA: Neighborhood Adaptive Block-Level Attention

    Jul 17, 2025Dmitrii Mikhailov, Aleksey Letunovskiy, Maria Kovaleva +6Video Diffusion ModelsDiffusion Transformer

  11. The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs

    Apr 24, 2025Piotr Nawrot, Robert Li, Renjie Huang +3Transformer InferenceLLM Evaluation

  12. Fewer yet critical: Reducing Redundant Token Dependencies for Transformer-based Time Series Forecasting

    Mar 10, 2025Jianqi Zhang, Yuchan Liu, Zeen Song +2Time Series ForecastingTransformer Attention

  13. SparSTAR: Sparse Attention for SpaceTime AutoRegressive Video Synthesis

    Date pendingJongbeom Lee, Hyunwoo Yu, Jincheol Yang +2Image-to-Video GenerationSelf-Attention