Multi-Head Latent Attention

Momentum

2 papers in the last four weeks, against 1 the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 10

All topics
CardsList
  1. QuantMLA: Function-Aligned Dual-Path Quantization for Low-Bit MLA KV Caching

    Sep 29, 2026Zunhai Su, Yuxuan Sun, Jianchao Tan +5LLM Inference AccelerationAttention Quantization

  2. RedKnot-MLA: Multi-Head Offline-Online Reuse for DeepSeek-V4 Long-Context Serving

    Sep 7, 2026Yang Liu, Zhaokai Luo, Huayi Jin +10KV CachingLLM Inference Acceleration

  3. Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding

    Jul 29, 2026Weiye Shi, Fanxu Meng, Muhan ZhangSpeculative DecodingLLM Inference Acceleration

  4. Through the Bottleneck: How Multi-head Latent Attention Separates Content from Position in Language Models

    Jul 25, 2026Dhruvil S, Fenil Sojitra, Ravirajsinh ChauhanTransformer InterpretabilityLow-Rank Attention

  5. QK-Normed MLA: QK normalization without full key caching

    Jun 15, 2026Yizhou Han, Yao Zhao, Jun Zhou +2KV CachingTransformer Attention

  6. VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion

    May 28, 2026Hidir Yesiltepe, Jiazhen Hu, Tuna Han Salih Meral +4Video Diffusion ModelsSelf-Attention

  7. Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving

    May 7, 2026Bole Ma, Jan Eitzinger, Harald KöstlerKV CachingLLM Inference Serving