Large Language Models (LLMs) achieve strong performance across many domains, but their efficiency is limited by the quadratic cost of attention with respect to prompt length. Sparse attention reduces this cost by retaining only a small fraction of query-key interactions to approximate the full attention matrix. However, existing methods are trapped in a mathematically wrong view: they simply keep large scalar entries or high-mass regions of the attention matrix. This treats the attention matrix as a bag of values, ignoring that it is used as a structured matrix whose entries jointly determine the attention output through multiplication with value vectors. We argue that this is the core conceptual issue: sparse attention should be formulated as matrix approximation, not as blindly choosing the largest values from a bag of entries. Based on this view, we propose Matrix Approximation Sparse Attention (MASA). MASA replaces raw attention-mass ranking with a closed-form score that measures how much each sparse unit reduces matrix-product approximation error. As a theory-grounded plug-in correction, MASA can be added to existing sparse attention frameworks without changing their sparse kernels or budgets. Extensive experiments across multiple sparse attention methods, benchmarks, and LLM backbones show consistent accuracy gains, supporting both MASA and the matrix-approximation view of sparse attention.
Figures & tables
Figure 1: A minimal two-column counterexample. Under the same one-column budget, conventional mass-based ranking selects u1 , whose mass is 66× larger but incurs 98% relative matrix-product error. MASA selects u2 and reduces the error to 3% . Here, X is an illustrative probe matrix.
Variant
4K
8K
16K
32K
64K
128K
Avg.
ΔFull
MInfer
Full Attention
97.2
91.8
87.3
80.8
77.4
72.2
84.4
0
Original
97.7
91.2
88.5
85.0
82.3
77.6
87.0
+2.6
MASA
97.93
91.29
88.50
85.45
81.57
78.88
87.27
+2.87
SeerAttn
Full Attention
95.53
92.37
92.01
87.63
84.39
76.26
88.01
0
Original
95.53
92.71
92.02
88.49
83.48
73.37
87.60
-0.41
MASA
95.53
92.74
92.84
89.56
84.01
73.85
88.09
+0.08
Table 1: RULER ( Hsieh et al., 2024 ) results under matched sparse-attention settings. Within each framework, Original and MASA use the same sparse implementation and differ only in the ranking rule; Full Attention is the reference. ΔFull is the signed Avg. difference from Full Attention. Higher is better. Bold marks the better Avg. between Original and MASA.
Variant
En.Sum
En.QA
En.MC
En.Dia
Zh.QA
Code Debug
Math Find
Retr. PassKey
Retr. Number
Retr. KV
Avg.
ΔFull
MInfer
Full Attention
20.2
12.4
67.3
6.0
12.9
22.1
26.6
100.0
100.0
14.4
38.2
0
Original
20.5
12.9
65.9
7.5
12.5
22.3
33.1
100.0
100.0
12.8
38.8
+0.6
MASA
19.73
12.15
66.81
8.50
12.28
22.84
33.43
100.00
100.00
16.20
39.20
+1.00
FlexPrefill
Llama
Full Attention
31.91
25.92
69.43
21.50
31.95
16.75
24.29
99.15
99.66
60.00
48.06
0
Original
31.82
24.82
69.43
19.50
35.46
16.75
31.14
98.64
99.83
44.00
47.14
-0.92
MASA
32.10
26.25
66.81
19.50
33.66
20.05
29.14
98.31
99.66
54.40
47.99
-0.07
Table 2: InfiniteBench ( Zhang et al., 2024 ) results under matched sparse-attention settings. Within each framework, Original and MASA use the same sparse implementation and differ only in the ranking rule; Full Attention is the reference. ΔFull is the signed Avg. difference from Full Attention. Higher is better. Bold marks the better Avg. between Original and MASA.
Variant
0–4K
4–8K
8K+
Avg.
ΔFull
Full Attention
55.32
53.98
52.90
54.07
0
Original
55.43
54.49
52.69
54.20
+0.13
MASA
55.45
54.48
53.16
54.37
+0.30
Table 3: LongBench ( Bai et al., 2024b ) results under matched sparse-attention settings. Within SeerAttention, Original and MASA differ only in the sparse-unit ranking rule; Full Attention is the reference. ΔFull is the signed Avg. difference from Full Attention. Higher is better. Bold marks the better Avg. and its ΔFull .
Setting
Clip p
Norm α
Avg.
RULER
MInference
–
–
87.13
MInference
90
0.6
87.27
SeerAttention
–
–
87.88
SeerAttention
90
–
88.09
SeerAttention
95
–
87.79
FlexPrefill-Llama
90
–
89.26
Table 4: Ablation of MASA probe choices. Clip p denotes norm clipping at percentile p , and Norm α denotes soft norm balancing. The case p=– and α=– is raw MASA without norm control. Bold marks the best MASA variant for each setting. See details in Sec. 3.3 .
Figure 2: Long-context kernel-level latency normalized to dense FlashAttention-2 at sequence lengths 32K, 64K, and 128K. The x-axis gives the retained sparse budget ratio. All methods use the same latency protocol; method-specific sparse kernels and selection or gating modules follow their official implementations.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Variant
4K
8K
16K
32K
64K
128K
Avg.
MInfer
Full Attention
97.2
91.8
87.3
80.8
77.4
72.2
84.4
StreamingLLM †
97.2
38.1
37.5
17.2
14.2
9.4
35.0
StreamingLLM w/ dilated †
23.4
0.7
1.4
18.8
16.5
15.6
12.7
StreamingLLM w/ strided †
2.0
0.7
0.6
0.6
0.7
1.3
1.0
InfLLM †
89.4
79.8
70.1
55.6
43.0
39.5
62.9
Original
97.7
91.2
88.5
85.0
82.3
77.6
87.0
Appendix
Table 5: RULER ( Hsieh et al., 2024 ) results with additional methods reported in the original baseline papers. Rows marked with † are included for context and may use different settings. In each framework block, Full Attention and reported methods are followed by the matched Original/MASA pair. Bold marks the better Avg. within each pair.
Variant
En.Sum
En.QA
En.MC
En.Dia
Zh.QA
Code Debug
Math Find
Retr. PassKey
Retr. Number
Retr. KV
Avg.
MInfer
Full Attention
20.2
12.4
67.3
6.0
12.9
22.1
26.6
100.0
100.0
14.4
38.2
StreamingLLM †
21.0
8.2
40.2
10.0
10.4
25.9
30.0
86.8
5.1
0.8
23.8
StreamingLLM w/ dilated †
20.1
9.4
44.5
15.5
11.2
20.5
27.5
5.0
87.5
0.5
24.2
StreamingLLM w/ strided †
17.3
8.2
27.5
14.5
11.2
19.5
27.5
4.0
2.1
1.0
13.3
InfLLM †
24.1
7.8
45.0
6.0
11.4
19.5
32.9
100.0
100.0
1.2
34.8
MInference w/ static †
19.9
8.6
43.2
3.5
8.9
20.6
25.1
92.4
96.3
0.2
31.9
Appendix
Table 6: InfiniteBench ( Zhang et al., 2024 ) results with additional methods reported in the original baseline papers. Rows marked with † are included for context and may use different settings. In each framework block, Full Attention and reported methods are followed by the matched Original/MASA pair. Bold marks the better Avg. within each pair.
Variant
0–4K
4–8K
8K+
Avg.
Full Attention
55.32
53.98
52.90
54.07
MInference †
55.23
53.78
52.18
53.73
MoA †
50.74
49.84
51.89
50.82
DuoAttention †
53.77
52.17
51.27
52.40
Original
55.43
54.49
52.69
54.20
MASA
55.45
54.48
53.16
54.37
Appendix
Table 7: LongBench ( Bai et al., 2024b ) results with additional methods reported in the original SeerAttention paper. Rows marked with † are included for context and may use different settings. Full Attention and reported methods are followed by the matched Original/MASA pair. Bold marks the better Avg. within the pair.
Framework
Backbone
Variant
E↓
Reduction (%) ↑
Paired 98.75% interval (%)
MInference
Llama
Original
1.095985
–
–
MASA
1.091532
+0.41
[+0.14,+0.70]
SeerAttention
Llama
Original
1.843×10−3
–
–
MASA
1.805×10−3
+2.10
[+0.69,+4.44]
FlexPrefill
Llama
Original
4.9059×10−2
–
–
MASA
4.7909×10−2
+2.34
[+0.50,+4.18]
Appendix
Table 8: Attention-output error decreases across all four evaluated framework–backbone pairs. Original and MASA denote the baseline and diagnostic configuration. E is the normalized squared output error in Eq. ( 27 ). Relative reductions and paired 98.75% confidence intervals are reported in percent. Bold marks the lower error and the positive reduction.
Setting
Clip p
Norm α
4K
8K
16K
32K
64K
128K
Avg.
MInference
–
–
97.66
91.46
88.57
85.21
81.76
78.11
87.13
MInference
90
0.6
97.93
91.29
88.50
85.45
81.57
78.88
87.27
SeerAttention
–
–
95.53
92.50
92.98
89.58
84.20
72.50
87.88
SeerAttention
85
–
95.63
92.40
91.98
89.55
83.64
73.34
87.76
SeerAttention
90
–
95.53
92.74
92.84
89.56
84.01
73.85
88.09
SeerAttention
95
–
95.32
92.84
92.70
89.59
83.52
72.72
87.79
Appendix
Table 9: Full RULER probe ablation. Clip p denotes norm clipping at percentile p , and Norm α denotes soft norm balancing. The case p=– and α=– is raw MASA without norm control. Bold Avg. marks the best MASA variant within each setting.
Setting
Clip p
Norm α
En.Sum
En.QA
En.MC
En.Dia
Zh.QA
Code Debug
Math Find
Retr. PassKey
Retr. Number
Retr. KV
Avg.
MInference
–
–
20.51
13.14
66.38
9.00
12.34
22.34
32.57
100.00
100.00
14.20
39.05
MInference
90
0.6
21.13
13.02
66.38
8.00
12.31
23.10
32.00
100.00
100.00
14.80
39.07
MInference
80
0.7
19.73
12.15
66.81
8.50
12.28
22.84
33.43
100.00
100.00
16.20
39.20
FlexPrefill-Llama
–
0.5
32.10
26.25
66.81
19.50
33.66
20.05
29.14
98.31
99.66
54.40
47.99
FlexPrefill-Qwen
–
–
15.55
6.38
32.31
16.50
13.09
21.07
30.29
94.41
74.58
0.00
30.42
FlexPrefill-Qwen
–
0.5
15.65
5.98
34.50
14.50
11.77
19.29
33.43
94.07
73.90
0.00
30.31
Appendix
Table 10: Full InfiniteBench probe ablation. Clip p denotes norm clipping at percentile p , and Norm α denotes soft norm balancing. The case p=– and α=– is raw MASA without norm control. Bold Avg. marks the best MASA variant within each setting.
Clip p
Norm α
0–4K
4–8K
8K+
Avg.
–
–
55.31
54.57
52.84
54.24
85
–
55.27
54.79
52.70
54.25
87.5
–
55.45
54.48
53.16
54.37
90
–
55.27
54.65
52.92
54.28
Appendix
Table 11: Full LongBench probe ablation for SeerAttention. Clip p denotes norm clipping at percentile p . The case p=– and α=– is raw MASA without norm control. Bold Avg. marks the best MASA variant.
The quadratic complexity of attention imposes severe memory and computational bottlenecks on Large Language Model (LLM) inference. This challenge is particularly acute for emerging agentic applications that require processing multi-million token sequences. We propose STS, a sparse attention mechanism that requires no model retraining. STS leverages the key insight that tokens identified as important by a smaller draft model are highly predictive of important tokens for a larger target model. By integrating into speculative decoding frameworks, STS repurposes the draft model's attention scores to dynamically construct a token-and-head-wise sparsity mask. This mask effectively prunes the expensive attention computation in the target LLM. Our evaluation shows that STS achieves a 2.67x speedup operating at approximately 90% sparsity on representative benchmark NarrativeQA, maintaining negligible accuracy degradation compared to dense attention. STS establishes a new state-of-the-art on the sparsity-accuracy trade-off, outperforming prior techniques by enabling higher sparsity levels for a given accuracy budget.
Ceyu Xu, Jiangnan Yu, Yongji Wu +1
The Hong Kong University of Science and Technology · UC Berkeley
Recent advances in sparse attention mechanisms have demonstrated strong potential for reducing the computational cost of long-context training and inference in large language models (LLMs). Native Sparse Attention (NSA), one state-of-the-art approach, introduces natively trainable, hardware-aligned sparse attention that delivers substantial system-level performance boosts while maintaining accuracy comparable to full attention. However, the kernel implementation of NSA forces a loop order that is only efficient with a relatively large number of query heads in each Grouped Query Attention (GQA) group, whereas existing LLMs widely adopt a much smaller number of query heads in each GQA group -- such an inconsistency significantly limits the applicability of this sparse algorithmic advance. In this work, we propose Flash Sparse Attention (FSA), an alternative kernel implementation that enables efficient NSA computation across a wide range of popular LLMs with a varied, smaller number of heads in each GQA group on modern GPUs. Compared to vanilla NSA kernel implementation, our empirical evaluation demonstrates that FSA achieves (i) up to 3.5x and on average 1.6x kernel-level latency reduction, (ii) up to 1.25x and 1.09x on average end-to-end training speedup on state-of-the-art LLMs, and (iii) up to 1.36x and 1.11x on average for prefill-phase speedup in LLM generative inference. The source code is open-sourced and publicly available at https://github.com/Relaxed-System-Lab/Flash-Sparse-Attention.
Ran Yan, Youhe Jiang, Zhuoming Chen +3
1. The Hong Kong University of Science and Technology · 2. Carnegie Mellon University
Sparse attention offers a promising strategy to extend long-context capabilities in Transformer LLMs, yet its efficiency-accuracy trade-offs remain unclear due to the lack of comprehensive evaluation. We address this gap with the largest-scale empirical analysis to date of training-free sparse attention, evaluating six methods across multiple model families and sizes, sequences up to 128K tokens, and sparsity levels up to 0.95 (i.e., 1/20 attention budget) on nine diverse tasks. We first organise the rapidly evolving landscape of sparse attention methods into a taxonomy along four design axes. Our analysis then yields actionable insights: 1) sparse attention is effective: larger sparse models outperform smaller dense ones at equivalent cost, improving the Pareto frontier; 2) for the training-free methods we study, fine-grained per-query importance estimation during prefilling remains impractical-due to both the cost of estimation and the lack of sparse kernels that translate fine-grained sparsity into wall-clock gains-forcing a task-dependent choice between global-to-token and block-to-block selection. Instead, during decoding, token-to-page selection becomes feasible, enabling better generalisation and higher sparsity tolerance; 3) longer sequences tolerate higher sparsity, suggesting that fixed-budget methods in production are suboptimal. Together, these findings provide practical guidance for deploying sparse attention and methodological recommendations for future evaluations. Our code is available at https://github.com/PiotrNawrot/sparse-frontier.