Large Language Models (LLMs) achieve strong performance across many domains, but their efficiency is limited by the quadratic cost of attention with respect to prompt length. Sparse attention reduces this cost by retaining only a small fraction of query-key interactions to approximate the full attention matrix. However, existing methods are trapped in a mathematically wrong view: they simply keep large scalar entries or high-mass regions of the attention matrix. This treats the attention matrix as a bag of values, ignoring that it is used as a structured matrix whose entries jointly determine the attention output through multiplication with value vectors. We argue that this is the core conceptual issue: sparse attention should be formulated as matrix approximation, not as blindly choosing the largest values from a bag of entries. Based on this view, we propose Matrix Approximation Sparse Attention (MASA). MASA replaces raw attention-mass ranking with a closed-form score that measures how much each sparse unit reduces matrix-product approximation error. As a theory-grounded plug-in correction, MASA can be added to existing sparse attention frameworks without changing their sparse kernels or budgets. Extensive experiments across multiple sparse attention methods, benchmarks, and LLM backbones show consistent accuracy gains, supporting both MASA and the matrix-approximation view of sparse attention.
Figures & tables
Figure 1: A minimal two-column counterexample. Under the same one-column budget, conventional mass-based ranking selects u1 , whose mass is 66× larger but incurs 98% relative matrix-product error. MASA selects u2 and reduces the error to 3% . Here, X is an illustrative probe matrix.
Variant
4K
8K
16K
32K
64K
128K
Avg.
ΔFull
MInfer
Full Attention
97.2
91.8
87.3
80.8
77.4
72.2
84.4
0
Original
97.7
91.2
88.5
85.0
82.3
77.6
87.0
+2.6
MASA
97.93
91.29
88.50
85.45
81.57
78.88
87.27
+2.87
SeerAttn
Full Attention
95.53
92.37
92.01
87.63
84.39
76.26
88.01
0
Original
95.53
92.71
92.02
88.49
83.48
73.37
87.60
-0.41
MASA
95.53
92.74
92.84
89.56
84.01
73.85
88.09
+0.08
Table 1: RULER ( Hsieh et al., 2024 ) results under matched sparse-attention settings. Within each framework, Original and MASA use the same sparse implementation and differ only in the ranking rule; Full Attention is the reference. ΔFull is the signed Avg. difference from Full Attention. Higher is better. Bold marks the better Avg. between Original and MASA.
Variant
En.Sum
En.QA
En.MC
En.Dia
Zh.QA
Code Debug
Math Find
Retr. PassKey
Retr. Number
Retr. KV
Avg.
ΔFull
MInfer
Full Attention
20.2
12.4
67.3
6.0
12.9
22.1
26.6
100.0
100.0
14.4
38.2
0
Original
20.5
12.9
65.9
7.5
12.5
22.3
33.1
100.0
100.0
12.8
38.8
+0.6
MASA
19.73
12.15
66.81
8.50
12.28
22.84
33.43
100.00
100.00
16.20
39.20
+1.00
FlexPrefill
Llama
Full Attention
31.91
25.92
69.43
21.50
31.95
16.75
24.29
99.15
99.66
60.00
48.06
0
Original
31.82
24.82
69.43
19.50
35.46
16.75
31.14
98.64
99.83
44.00
47.14
-0.92
MASA
32.10
26.25
66.81
19.50
33.66
20.05
29.14
98.31
99.66
54.40
47.99
-0.07
Table 2: InfiniteBench ( Zhang et al., 2024 ) results under matched sparse-attention settings. Within each framework, Original and MASA use the same sparse implementation and differ only in the ranking rule; Full Attention is the reference. ΔFull is the signed Avg. difference from Full Attention. Higher is better. Bold marks the better Avg. between Original and MASA.
Variant
0–4K
4–8K
8K+
Avg.
ΔFull
Full Attention
55.32
53.98
52.90
54.07
0
Original
55.43
54.49
52.69
54.20
+0.13
MASA
55.45
54.48
53.16
54.37
+0.30
Table 3: LongBench ( Bai et al., 2024b ) results under matched sparse-attention settings. Within SeerAttention, Original and MASA differ only in the sparse-unit ranking rule; Full Attention is the reference. ΔFull is the signed Avg. difference from Full Attention. Higher is better. Bold marks the better Avg. and its ΔFull .
Setting
Clip p
Norm α
Avg.
RULER
MInference
–
–
87.13
MInference
90
0.6
87.27
SeerAttention
–
–
87.88
SeerAttention
90
–
88.09
SeerAttention
95
–
87.79
FlexPrefill-Llama
90
–
89.26
Table 4: Ablation of MASA probe choices. Clip p denotes norm clipping at percentile p , and Norm α denotes soft norm balancing. The case p=– and α=– is raw MASA without norm control. Bold marks the best MASA variant for each setting. See details in Sec. 3.3 .
Figure 2: Long-context kernel-level latency normalized to dense FlashAttention-2 at sequence lengths 32K, 64K, and 128K. The x-axis gives the retained sparse budget ratio. All methods use the same latency protocol; method-specific sparse kernels and selection or gating modules follow their official implementations.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Variant
4K
8K
16K
32K
64K
128K
Avg.
MInfer
Full Attention
97.2
91.8
87.3
80.8
77.4
72.2
84.4
StreamingLLM †
97.2
38.1
37.5
17.2
14.2
9.4
35.0
StreamingLLM w/ dilated †
23.4
0.7
1.4
18.8
16.5
15.6
12.7
StreamingLLM w/ strided †
2.0
0.7
0.6
0.6
0.7
1.3
1.0
InfLLM †
89.4
79.8
70.1
55.6
43.0
39.5
62.9
Original
97.7
91.2
88.5
85.0
82.3
77.6
87.0
Appendix
Table 5: RULER ( Hsieh et al., 2024 ) results with additional methods reported in the original baseline papers. Rows marked with † are included for context and may use different settings. In each framework block, Full Attention and reported methods are followed by the matched Original/MASA pair. Bold marks the better Avg. within each pair.
Variant
En.Sum
En.QA
En.MC
En.Dia
Zh.QA
Code Debug
Math Find
Retr. PassKey
Retr. Number
Retr. KV
Avg.
MInfer
Full Attention
20.2
12.4
67.3
6.0
12.9
22.1
26.6
100.0
100.0
14.4
38.2
StreamingLLM †
21.0
8.2
40.2
10.0
10.4
25.9
30.0
86.8
5.1
0.8
23.8
StreamingLLM w/ dilated †
20.1
9.4
44.5
15.5
11.2
20.5
27.5
5.0
87.5
0.5
24.2
StreamingLLM w/ strided †
17.3
8.2
27.5
14.5
11.2
19.5
27.5
4.0
2.1
1.0
13.3
InfLLM †
24.1
7.8
45.0
6.0
11.4
19.5
32.9
100.0
100.0
1.2
34.8
MInference w/ static †
19.9
8.6
43.2
3.5
8.9
20.6
25.1
92.4
96.3
0.2
31.9
Appendix
Table 6: InfiniteBench ( Zhang et al., 2024 ) results with additional methods reported in the original baseline papers. Rows marked with † are included for context and may use different settings. In each framework block, Full Attention and reported methods are followed by the matched Original/MASA pair. Bold marks the better Avg. within each pair.
Variant
0–4K
4–8K
8K+
Avg.
Full Attention
55.32
53.98
52.90
54.07
MInference †
55.23
53.78
52.18
53.73
MoA †
50.74
49.84
51.89
50.82
DuoAttention †
53.77
52.17
51.27
52.40
Original
55.43
54.49
52.69
54.20
MASA
55.45
54.48
53.16
54.37
Appendix
Table 7: LongBench ( Bai et al., 2024b ) results with additional methods reported in the original SeerAttention paper. Rows marked with † are included for context and may use different settings. Full Attention and reported methods are followed by the matched Original/MASA pair. Bold marks the better Avg. within the pair.
Framework
Backbone
Variant
E↓
Reduction (%) ↑
Paired 98.75% interval (%)
MInference
Llama
Original
1.095985
–
–
MASA
1.091532
+0.41
[+0.14,+0.70]
SeerAttention
Llama
Original
1.843×10−3
–
–
MASA
1.805×10−3
+2.10
[+0.69,+4.44]
FlexPrefill
Llama
Original
4.9059×10−2
–
–
MASA
4.7909×10−2
+2.34
[+0.50,+4.18]
Appendix
Table 8: Attention-output error decreases across all four evaluated framework–backbone pairs. Original and MASA denote the baseline and diagnostic configuration. E is the normalized squared output error in Eq. ( 27 ). Relative reductions and paired 98.75% confidence intervals are reported in percent. Bold marks the lower error and the positive reduction.
Setting
Clip p
Norm α
4K
8K
16K
32K
64K
128K
Avg.
MInference
–
–
97.66
91.46
88.57
85.21
81.76
78.11
87.13
MInference
90
0.6
97.93
91.29
88.50
85.45
81.57
78.88
87.27
SeerAttention
–
–
95.53
92.50
92.98
89.58
84.20
72.50
87.88
SeerAttention
85
–
95.63
92.40
91.98
89.55
83.64
73.34
87.76
SeerAttention
90
–
95.53
92.74
92.84
89.56
84.01
73.85
88.09
SeerAttention
95
–
95.32
92.84
92.70
89.59
83.52
72.72
87.79
Appendix
Table 9: Full RULER probe ablation. Clip p denotes norm clipping at percentile p , and Norm α denotes soft norm balancing. The case p=– and α=– is raw MASA without norm control. Bold Avg. marks the best MASA variant within each setting.
Setting
Clip p
Norm α
En.Sum
En.QA
En.MC
En.Dia
Zh.QA
Code Debug
Math Find
Retr. PassKey
Retr. Number
Retr. KV
Avg.
MInference
–
–
20.51
13.14
66.38
9.00
12.34
22.34
32.57
100.00
100.00
14.20
39.05
MInference
90
0.6
21.13
13.02
66.38
8.00
12.31
23.10
32.00
100.00
100.00
14.80
39.07
MInference
80
0.7
19.73
12.15
66.81
8.50
12.28
22.84
33.43
100.00
100.00
16.20
39.20
FlexPrefill-Llama
–
0.5
32.10
26.25
66.81
19.50
33.66
20.05
29.14
98.31
99.66
54.40
47.99
FlexPrefill-Qwen
–
–
15.55
6.38
32.31
16.50
13.09
21.07
30.29
94.41
74.58
0.00
30.42
FlexPrefill-Qwen
–
0.5
15.65
5.98
34.50
14.50
11.77
19.29
33.43
94.07
73.90
0.00
30.31
Appendix
Table 10: Full InfiniteBench probe ablation. Clip p denotes norm clipping at percentile p , and Norm α denotes soft norm balancing. The case p=– and α=– is raw MASA without norm control. Bold Avg. marks the best MASA variant within each setting.
Clip p
Norm α
0–4K
4–8K
8K+
Avg.
–
–
55.31
54.57
52.84
54.24
85
–
55.27
54.79
52.70
54.25
87.5
–
55.45
54.48
53.16
54.37
90
–
55.27
54.65
52.92
54.28
Appendix
Table 11: Full LongBench probe ablation for SeerAttention. Clip p denotes norm clipping at percentile p . The case p=– and α=– is raw MASA without norm control. Bold Avg. marks the best MASA variant.