Sparse attention is widely used to accelerate long-context inference in modern large language models (LLMs), but its input-dependent execution behavior introduces previously unexplored privacy risks. We identify a new GPU micro-architectural side channel, termed Sparsity-Induced Memory Access (SIMA), which arises from secret-dependent key-value cache access patterns induced by sparse attention. Based on this observation, we present SparLeak, a phase-aware side-channel attack that extracts SIMA traces during LLM inference and enables two practical privacy extractions: query attribute inference from prefill-phase traces and autoregressive response reconstruction from decoding-phase traces. By reconstructing approximate token-level sparsity profiles from page-level observations and applying profiling-based learning, SparLeak accurately recovers sensitive information, including user-query attributes and private LLM response content. Extensive evaluation across three LLM architectures, three sparse attention mechanisms, and three privacy-sensitive datasets shows that SparLeak achieves average attack success rates of 90.9% for attribute inference and 87.3% for response reconstruction under real-world LLM serving settings, highlighting the significance to account for SIMA leakage when deploying sparse-attention-based LLM systems. We provide anonymized SIMA traces, trained attack models, evaluation scripts, and documentation as artifacts at https://anonymous.4open.science/r/Janus_artifacts/.
Figures & tables
Side-Channel Attacks
Threat Surface
Extracted Info.
Prefill-Phase Attack
Decoding-Phase Attack
Mitigation
Prevention
Detection
Time Will Tell ( Zhang et al., 2024a )
Timing (response time)
Response length
✔
✘
✘
✘
Early Bird ( Song et al., 2025 )
Serving order
KV Cache hit/miss
✔
✘
( Chu et al., 2025 ; Luo et al., 2025b )
( Chu et al., 2025 )
PromptPeek ( Wu et al., 2025 )
Timing (time-to-first-token)
KV Cache hit/miss
✔
✘
( Chu et al., 2025 ; Luo et al., 2025b )
( Chu et al., 2025 )
InputSnatch ( Zheng et al., 2024b )
Timing (time-to-first-token)
KV Cache hit/miss&hit ratio
✔
✘
( Chu et al., 2025 ; Luo et al., 2025b )
( Chu et al., 2025 )
What Was Your Prompt ( Weiss et al., 2024 )
Network traffic
Token-length sequence
✔
✔
( Mehta et al., 2022 ; Zhang et al., 2025 )
✘
Table 1. Comparison of representative side-channel attacks that expose privacy leakage in LLM serving. Prefill-Phase Attack and Decoding-Phase Attack indicate whether an attack can infer sensitive query information or recover generated outputs, respectively. Mitigation summarizes existing defense efforts: entries with citations indicate that prior prevention or detection mechanisms have been proposed and can be applied to mitigate the corresponding attack, while ✘ denotes that no effective mitigation has been reported to date.
Figure 1. The illustration of sparse attention mechanism in LLM inference and its associated GPU memory access.
Figure 2. Sparsity pattern in prefill phase. Activation frequency denotes how often a given token is selected by sparse attention during the prefill phase.
Figure 3. Sparsity pattern at a selected decoding step. Activation indicator is a binary value indicating whether a token is selected by sparse attention at the given decoding step.
Figure 4. Sparsity patterns of user queries with different attributes. In this example, we change the illness categories in different queries.
Figure 5. Sparsity patterns of different decoding tokens, including “his”, “health”, and “healthy”.
Figure 6. The illustration of the threat model.
Figure 7. The overview of the attack workflow. SparLeak proceeds in three stages: (1) phase-aware extraction of SIMA traces for prefill and decoding phases using different side-channel primitives, (2) reconstruction of token-level sparsity patterns from page-level SIMA traces, and (3) sparsity-guided inference for query attribute inference and autoregressive token recovery.
Figure 8. Spy process for estimating page-level access frequency during the prefill phase by repeatedly measuring cache replacement effects induced by sparse-attention KV accesses.
Figure 9. Spy process for recovering the sequential page-level sparsity pattern during the decoding phase. At each decoding step, the spy detects whether KV data from a prompt page are accessed by checking the persistence of its address translation state.
Figure 10. Illustration of page-level and token-level sparsity patterns in the prefill phase.
Figure 11. Illustration of page-level and token-level sparsity patterns in the decoding phase.
Attribute
Signal
LongChat-7B ( Li et al., 2023 )
LLaMA3-8B ( Grattafiori et al., 2024 )
Qwen3-8B ( Yang et al., 2025a )
Act.
Block
Learn.
Act.
Block
Learn.
Act.
Block
Learn.
Illness (60)
O(T)
89.4 ± 1.1
83.3 ± 0.7
99.8 ± 0.1
86.9 ± 0.8
92.1 ± 1.1
99.5 ± 0.4
86.3 ± 0.8
86.8 ± 0.5
98.4 ± 0.3
O(P+R)
87.3 ± 0.7
81.7 ± 1.1
99.4 ± 0.3
84.5 ± 0.5
90.9 ± 0.6
99.2 ± 0.1
84.4 ± 1.0
85.2 ± 0.2
98.1 ± 0.3
R(P)
70.4 ± 2.6
66.4 ± 1.2
78.7 ± 0.7
67.1 ± 1.9
66.51 ± 1.4
74.4 ± 1.2
60.3 ± 3.6
62.3 ± 1.8
75.3 ± 1.0
R(P+R)
85.3 ± 0.4
80.5 ± 0.3
98.8 ± 0.6
84.1 ± 0.6
88.4 ± 0.7
98.7 ± 0.5
81.8 ± 0.4
83.0 ± 0.5
94.7 ± 0.2
Age (8)
O(T)
86.2 ± 0.5
79.7 ± 0.3
99.9 ± 0.1
68.4 ± 1.0
81.3 ± 0.8
99.4 ± 0.3
99.5 ± 0.4
99.1 ± 0.2
99.8 ± 0.1
Table 2. Query attribute inference PASR (%) on Healthcare ( Prasad, 2022 ) (Illness, Age, Gender, Blood), Financial QA ( Saeedian, 2023 ) (Entity, Period, Topic), and Legal QA ( Face, 2024 ) (Domain, Role, Location). Each entry report the mean ± 95% confidence interval over 20 runs.
Figure 12. Performance of autoregressive token recovery attacks, measured by DASR.
Figure 13. The performance of sparsity reconstruction.
Model
Attribute Inference Attack
Token Recovery Attack
Per-query prefill time
#Samples
Cost
Per-token decode time
Avg. length
#Samples
Cost
LongChat-7B ( Li et al., 2023 )
0.62s
5,000
0.86h
0.05s
300
3,000
12.7h
LLaMA3-8B ( Grattafiori et al., 2024 )
0.58s
5,000
0.82h
0.06s
300
3,000
15.2h
Qwen3-8B ( Yang et al., 2025a )
0.79s
5,000
1.10h
0.07s
500
2,000
19.5h
Table 3. Offline profiling overhead for query attribute inference and autoregressive token recovery attacks.
Intensity
Baseline
10%
20%
30%
ASR (%)
Rouge
ASR (%)
Rouge
ASR (%)
Rouge
ASR (%)
Rouge
LongChat-7B ( Li et al., 2023 )
89.45
0.43
85.41
0.38
82.09
0.32
76.17
0.21
LLaMA3-8B ( Grattafiori et al., 2024 )
90.12
0.48
84.92
0.35
79.85
0.24
72.68
0.16
Qwen3-8B ( Yang et al., 2025a )
91.23
0.42
87.28
0.37
81.13
0.35
75.98
0.26
Table 4. Effectiveness of fixed sparsity pattern injection as a mitigation strategy. Mitigation intensity denotes the fraction of sparse tokens that are replaced by a random token set.
Modern large language models (LLMs) exhibit activation sparsity, wherein only a subset of their neurons is activated for given input tokens. Researchers have leveraged this property to optimize LLM serving systems by omitting weight accesses and computations pertaining to inactive neurons. Unfortunately, however, such optimizations create input-dependent weight accesses, which can be leaked over side channels. We present SparSEEty, a new token extraction attack that exploits input-dependent neuron weight accesses introduced by sparsity-exploiting LLM serving systems. SparSEEty first constructs a neuron-activation oracle using neuron weight access side channels during LLM inference, and then inverts the activation traces to reconstruct the input tokens, forming an end-to-end token extraction attack. We instantiate SparSEEty against an LLM serving system protected inside an Intel TDX confidential virtual machine (CVM), addressing three key challenges: (i) constructing a neuron-activation oracle using a combination of side channels exposed by CVMs, (ii) reducing inference-time overheads of neuron activation monitoring for covertness, and (iii) accurately inverting partial binary activation traces back to tokens. Our evaluation shows that SparSEEty can reconstruct both prompt and response tokens with consistently high BLEU scores (>0.95) across various models and datasets, while incurring monitoring overheads of 3.7% to 7.2%.
User prompts provided to large language models (LLMs) may contain sensitive or private information that can be misused by remotely deployed models, such as through inadvertent memorization during retraining. One way to protect user prompts is to execute the LLM inside a trusted execution environment (TEE), with the guarantee that the service provider has no access to computations performed within or information exchanged with the TEE. However, current TEEs are primarily CPU-based and significantly slower than GPUs optimized for LLM inference. To circumvent this, Tramer and Boneh (2019) proposed Slalom, which splits neural network inference between a TEE and an untrusted GPU and encrypts intermediate inputs sent to the GPU. We extend this split-inference architecture to LLM inference and instead protect intermediate inputs using differential privacy. We show that masking intermediate representations is necessary by showing that a prompt-reconstruction attack can recover prompts from these representations with nearly 80% accuracy. Our main contribution is a global sensitivity analysis of key LLM functions, which bounds the required scale of differentially private noise. Unlike encryption, differential privacy avoids quantization, allowing the LLM to remain in the floating-point domain. We also derive an upper bound on floating-point error from masking and noise cancellation in the TEE as a function of the privacy parameter epsilon. We implement our architecture using Intel TDX and evaluate it with two LLMs: Llama-3.2-3B and Qwen3-4B. Our split execution is nearly twice as fast as fully CPU-based inference inside TDX and 5-15 seconds faster than encryption-based Slalom while achieving higher accuracy. Finally, we demonstrate that prompt reconstruction, even with knowledge of the differential privacy mechanism, cannot recover more information than is contained in an unrelated prompt.
Shashie Dilhara Batan Arachchige, Robin Carpentier, Hassan Jameel Asghar +1
Large language models (LLMs) are increasingly deployed in privacy-sensitive domains, where users must balance the risk of data exposure through external APIs against the high computational cost of local deployment. Split learning has therefore emerged as a promising paradigm for LLM fine-tuning and inference under limited local resources. However, it introduces new privacy risks. Prior work primarily studies leakage of private input prompts, typically via inversion attacks on intermediate representations, while the potential for sensitive information leakage through generative response outputs remains largely unexplored. In this work, we unveil novel vulnerabilities of Split-LLM by presenting Patched Model Inversion with Dual-Sided Initialization (PIDI), a two-stage attack that simultaneously targets both private input prompts and output responses in Split-LLM settings. It combines dual-sided initialization with a patched inversion strategy to tackle long sequences, substantially outperforming prior inversion methods. To counter threats from both sides, we further propose the Adapter-based DualGuard with Mutual Information Defense (ADMI), which integrates an adapter-based local warmup strategy and mutual information regularization to provide a strong empirical privacy protection with minimal impact on task performance. Extensive experiments across diverse tasks and models demonstrate that ADMI effectively defends against PIDI and other state-of-the-art inversion attacks. Our code is publicly available at https://github.com/FLAIR-THU/VFLAIR-LLM.
Zixuan Gu, Xiaojun Ye, Yang Liu
School of Software, Tsinghua University, Beijing, China · the Hong Kong Polytechnic University, Hongkong, China