Hybrid Attention

Momentum

12 papers in the last four weeks, up 71% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 46

All topics
CardsList
  1. Morphing into Hybrid Attention Models

    Jun 29, 2026Disen Lan, Jianbin Zheng, Yuxi Ren +5Efficient AttentionNeural Architecture Search

  2. ConSA: Controllable Sparsity in Hybrid Attention via Learnable Allocation

    Jun 16, 2026Yao Chen, Yinqi Yang, Junyuan Shang +6LLM Inference AccelerationSparse Attention

  3. Taylor-Calibrate: Principled Initialization for Hybrid Linear Attention Distillation

    Jun 15, 2026Zhongzhu Zhou, Qingyang Wu, Junxiong Wang +4Neural Network InitializationLinear Attention

  4. Rethinking the Role of Efficient Attention in Hybrid Architectures

    Jun 13, 2026Ziqing Qiao, Yinuo Xu, Chaojun Xiao +6Attention Head AnalysisLong-Context Modeling

  5. Forget Attention: Importance-Aware Attention Is All You Need

    Jun 1, 2026Suhyeong Shin, Yeongwook YangSoftmax AttentionLong-Context Language Modeling

  6. DASH: Fast Differentiable Architecture Search for Hybrid Attention in Minutes on a Single GPU

    May 20, 2026Weizhe Chen, Miao Zhang, Junpeng Jiang +3LLM Inference EfficiencyAttention Mechanisms

  7. Provably Shorter Scratchpads in Hybrid DeltaNet-Attention Decoders

    May 15, 2026Tomasz SteiferGated AttentionRecurrent Transformers

  8. Mixture of Layers with Hybrid Attention

    May 10, 2026Ivan Ternovtsii, Yurii BilakMixture-of-Experts ModelsHybrid Attention

  9. Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference

    Apr 8, 2026Quantong Qiu, Zhiyi Hong, Yi Yang +5Self-AttentionLong-Context Language Model Inference

  10. STILL: Selecting Tokens for Intra-Layer Hybrid Attention to Linearize LLMs

    Feb 2, 2026Weikang Meng, Liangyu Huo, Yadan Luo +4LLM Inference AccelerationLinear Attention

  11. A Systematic Analysis of Hybrid Linear Attention

    Jul 8, 2025Dustin Wang, Rui-Jie Zhu, Steven Abreu +9Linear AttentionHybrid Attention