Sparse attention is widely used to accelerate long-context inference in modern large language models (LLMs), but its input-dependent execution behavior introduces previously unexplored privacy risks. We identify a new GPU micro-architectural side channel, termed Sparsity-Induced Memory Access (SIMA), which arises from secret-dependent key-value cache access patterns induced by sparse attention. Based on this observation, we present SparLeak, a phase-aware side-channel attack that extracts SIMA traces during LLM inference and enables two practical privacy extractions: query attribute inference from prefill-phase traces and autoregressive response reconstruction from decoding-phase traces. By reconstructing approximate token-level sparsity profiles from page-level observations and applying profiling-based learning, SparLeak accurately recovers sensitive information, including user-query attributes and private LLM response content. Extensive evaluation across three LLM architectures, three sparse attention mechanisms, and three privacy-sensitive datasets shows that SparLeak achieves average attack success rates of 90.9% for attribute inference and 87.3% for response reconstruction under real-world LLM serving settings, highlighting the significance to account for SIMA leakage when deploying sparse-attention-based LLM systems. We provide anonymized SIMA traces, trained attack models, evaluation scripts, and documentation as artifacts at https://anonymous.4open.science/r/Janus_artifacts/.
Figures & tables
Side-Channel Attacks
Threat Surface
Extracted Info.
Prefill-Phase Attack
Decoding-Phase Attack
Mitigation
Prevention
Detection
Time Will Tell ( Zhang et al., 2024a )
Timing (response time)
Response length
✔
✘
✘
✘
Early Bird ( Song et al., 2025 )
Serving order
KV Cache hit/miss
✔
✘
( Chu et al., 2025 ; Luo et al., 2025b )
( Chu et al., 2025 )
PromptPeek ( Wu et al., 2025 )
Timing (time-to-first-token)
KV Cache hit/miss
✔
✘
( Chu et al., 2025 ; Luo et al., 2025b )
( Chu et al., 2025 )
InputSnatch ( Zheng et al., 2024b )
Timing (time-to-first-token)
KV Cache hit/miss&hit ratio
✔
✘
( Chu et al., 2025 ; Luo et al., 2025b )
( Chu et al., 2025 )
What Was Your Prompt ( Weiss et al., 2024 )
Network traffic
Token-length sequence
✔
✔
( Mehta et al., 2022 ; Zhang et al., 2025 )
✘
Table 1. Comparison of representative side-channel attacks that expose privacy leakage in LLM serving. Prefill-Phase Attack and Decoding-Phase Attack indicate whether an attack can infer sensitive query information or recover generated outputs, respectively. Mitigation summarizes existing defense efforts: entries with citations indicate that prior prevention or detection mechanisms have been proposed and can be applied to mitigate the corresponding attack, while ✘ denotes that no effective mitigation has been reported to date.
Figure 1. The illustration of sparse attention mechanism in LLM inference and its associated GPU memory access.
Figure 2. Sparsity pattern in prefill phase. Activation frequency denotes how often a given token is selected by sparse attention during the prefill phase.
Figure 3. Sparsity pattern at a selected decoding step. Activation indicator is a binary value indicating whether a token is selected by sparse attention at the given decoding step.
Figure 4. Sparsity patterns of user queries with different attributes. In this example, we change the illness categories in different queries.
Figure 5. Sparsity patterns of different decoding tokens, including “his”, “health”, and “healthy”.
Figure 6. The illustration of the threat model.
Figure 7. The overview of the attack workflow. SparLeak proceeds in three stages: (1) phase-aware extraction of SIMA traces for prefill and decoding phases using different side-channel primitives, (2) reconstruction of token-level sparsity patterns from page-level SIMA traces, and (3) sparsity-guided inference for query attribute inference and autoregressive token recovery.
Figure 8. Spy process for estimating page-level access frequency during the prefill phase by repeatedly measuring cache replacement effects induced by sparse-attention KV accesses.
Figure 9. Spy process for recovering the sequential page-level sparsity pattern during the decoding phase. At each decoding step, the spy detects whether KV data from a prompt page are accessed by checking the persistence of its address translation state.
Figure 10. Illustration of page-level and token-level sparsity patterns in the prefill phase.
Figure 11. Illustration of page-level and token-level sparsity patterns in the decoding phase.
Attribute
Signal
LongChat-7B ( Li et al., 2023 )
LLaMA3-8B ( Grattafiori et al., 2024 )
Qwen3-8B ( Yang et al., 2025a )
Act.
Block
Learn.
Act.
Block
Learn.
Act.
Block
Learn.
Illness (60)
O(T)
89.4 ± 1.1
83.3 ± 0.7
99.8 ± 0.1
86.9 ± 0.8
92.1 ± 1.1
99.5 ± 0.4
86.3 ± 0.8
86.8 ± 0.5
98.4 ± 0.3
O(P+R)
87.3 ± 0.7
81.7 ± 1.1
99.4 ± 0.3
84.5 ± 0.5
90.9 ± 0.6
99.2 ± 0.1
84.4 ± 1.0
85.2 ± 0.2
98.1 ± 0.3
R(P)
70.4 ± 2.6
66.4 ± 1.2
78.7 ± 0.7
67.1 ± 1.9
66.51 ± 1.4
74.4 ± 1.2
60.3 ± 3.6
62.3 ± 1.8
75.3 ± 1.0
R(P+R)
85.3 ± 0.4
80.5 ± 0.3
98.8 ± 0.6
84.1 ± 0.6
88.4 ± 0.7
98.7 ± 0.5
81.8 ± 0.4
83.0 ± 0.5
94.7 ± 0.2
Age (8)
O(T)
86.2 ± 0.5
79.7 ± 0.3
99.9 ± 0.1
68.4 ± 1.0
81.3 ± 0.8
99.4 ± 0.3
99.5 ± 0.4
99.1 ± 0.2
99.8 ± 0.1
Table 2. Query attribute inference PASR (%) on Healthcare ( Prasad, 2022 ) (Illness, Age, Gender, Blood), Financial QA ( Saeedian, 2023 ) (Entity, Period, Topic), and Legal QA ( Face, 2024 ) (Domain, Role, Location). Each entry report the mean ± 95% confidence interval over 20 runs.
Figure 12. Performance of autoregressive token recovery attacks, measured by DASR.
Figure 13. The performance of sparsity reconstruction.
Model
Attribute Inference Attack
Token Recovery Attack
Per-query prefill time
#Samples
Cost
Per-token decode time
Avg. length
#Samples
Cost
LongChat-7B ( Li et al., 2023 )
0.62s
5,000
0.86h
0.05s
300
3,000
12.7h
LLaMA3-8B ( Grattafiori et al., 2024 )
0.58s
5,000
0.82h
0.06s
300
3,000
15.2h
Qwen3-8B ( Yang et al., 2025a )
0.79s
5,000
1.10h
0.07s
500
2,000
19.5h
Table 3. Offline profiling overhead for query attribute inference and autoregressive token recovery attacks.
Intensity
Baseline
10%
20%
30%
ASR (%)
Rouge
ASR (%)
Rouge
ASR (%)
Rouge
ASR (%)
Rouge
LongChat-7B ( Li et al., 2023 )
89.45
0.43
85.41
0.38
82.09
0.32
76.17
0.21
LLaMA3-8B ( Grattafiori et al., 2024 )
90.12
0.48
84.92
0.35
79.85
0.24
72.68
0.16
Qwen3-8B ( Yang et al., 2025a )
91.23
0.42
87.28
0.37
81.13
0.35
75.98
0.26
Table 4. Effectiveness of fixed sparsity pattern injection as a mitigation strategy. Mitigation intensity denotes the fraction of sparse tokens that are replaced by a random token set.