cs.CLOct 7, 2026

Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values

Authors: Fang Wan, Xufeng Liu, Fan Li, Yi Liu

Organizations: Stony Brook University · Duke University

Abstract

Large Language Models (LLMs) achieve strong performance across many domains, but their efficiency is limited by the quadratic cost of attention with respect to prompt length. Sparse attention reduces this cost by retaining only a small fraction of query-key interactions to approximate the full attention matrix. However, existing methods are trapped in a mathematically wrong view: they simply keep large scalar entries or high-mass regions of the attention matrix. This treats the attention matrix as a bag of values, ignoring that it is used as a structured matrix whose entries jointly determine the attention output through multiplication with value vectors. We argue that this is the core conceptual issue: sparse attention should be formulated as matrix approximation, not as blindly choosing the largest values from a bag of entries. Based on this view, we propose Matrix Approximation Sparse Attention (MASA). MASA replaces raw attention-mass ranking with a closed-form score that measures how much each sparse unit reduces matrix-product approximation error. As a theory-grounded plug-in correction, MASA can be added to existing sparse attention frameworks without changing their sparse kernels or budgets. Extensive experiments across multiple sparse attention methods, benchmarks, and LLM backbones show consistent accuracy gains, supporting both MASA and the matrix-approximation view of sparse attention.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. STS: Efficient Sparse Attention with Speculative Token Sparsity

    May 15, 2026Ceyu Xu, Jiangnan Yu, Yongji Wu +1Self-AttentionLong-Context Language Model Inference

  2. FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel

    Aug 25, 2025Ran Yan, Youhe Jiang, Zhuoming Chen +3GPU Kernel OptimizationGPU Acceleration

  3. The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs

    Apr 24, 2025Piotr Nawrot, Robert Li, Renjie Huang +3Transformer InferenceLLM Evaluation