Grouped-Query Attention

Also known as GQA

Momentum

1 paper in the last four weeks, with none the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 10

All topics
CardsList
  1. CHASE: Channel-Aligned Structure Exploitation for Geometry-Aware Model Engineering

    Oct 7, 2026Wei Wang, Wei Jiang, Ziran LiuNeural Network CompressionKV-Cache Compression

  2. MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel

    Jul 21, 2026Lenore Mulin, Gaetan HainsEfficient Transformer InferenceGrouped-Query Attention

  3. GQA-μP: The maximal parameterization update for grouped query attention

    May 14, 2026Kyle R. Chickering, Huijuan Wang, Mengxi Wu +7Grouped-Query AttentionRepresentation Learning

  4. When Does Sparsity Mitigate the Curse of Depth in LLMs

    Mar 16, 2026Dilxat Muhtar, Xinyuan Song, Sebastian Pokutta +4Grouped-Query AttentionLanguage Model Scaling