With the increasing capabilities of Large-Language-Models (LLMs) and LLM-based agents, users are increasingly using them to solve everyday problems, such as answering e-mails or providing programming support. Existing work has extensively investigated security and privacy risks, such as prompt injections and the disclosure of sensitive data to chatbot providers. While various solutions were developed to address these risks, including input structuring to prevent prompt injections or deploying local LLMs to avoid sharing confidential data with chatbot operators, LLMs also pose the risk of leaking confidential data to third parties. In this paper, we demonstrate with LLMLeak a novel attack vector where malicious software that runs locally but cannot communicate directly with the internet abuses LLMs to establish a covert channel. While inputs that instruct the LLM to send data directly via generated code are easy to detect and network libraries are typically restricted, LLMLeak relies only on the LLM's tool to fetch websites for further information. A malicious software component on the client side embeds a secret into a URL. It presents the referenced website as providing information required for a benign task, such as migrating a software library. When the LLM accesses the URL, the attacker receives the encoded secret through an attacker-controlled DNS or web server. We perform an extensive evaluation on eleven open-parameter models, observe an attack success rate of 79.7%, and also conduct a case study on real-world chatbots, demonstrating the relevance of LLMLeak.
Figures & tables
Figure 1 : Overview of LLMLeak ’s steps.
Figure 2 : Comparison of benign and attack stack trace containing a URL to an external page.
Model
Arch.
Params (T/A)
Ctx.
Meta-Llama-3.3-70B-Instruct
Dense
70B / 70B
128K
Meta-Llama-4-Scout-17B-16E-Instruct
MoE
109B / 17B
10M
Meta-Llama-3.1-8B-Instruct
Dense
8B / 8B
128K
Mistral-Small-24B-Instruct-2501
Dense
24B / 24B
32K
Mistral-7B-Instruct-v0.3
Dense
7.3B / 7.3B
32K
IBM-Granite-3.3-8b-instruct
Dense
8B / 8B
128K
Table 1 : Overview of evaluated language models, their architecture (Dense or Mixture-of-Experts), number of parameters (total and active number), and length of context window.
Model
S reach
S conf
Llama-3.3-70B
100.0
99.9
Llama-4-Scout-17B
95.9
95.8
Llama-3.1-8B
95.8
95.2
Mistral-Small-24B
95.2
94.0
Mistral-7B
92.8
91.0
Granite-3.3-8B
94.5
90.6
Table 2 : Effectiveness of LLMLeak in terms of attack server contacted ( S reach ) and correct payload transmitted ( S conf ) for 11 different open parameter models in %.
Model
S call
S reach
S data
S conf
r
Llama-3.3-70B
100.0
100.0
100.0
99.9
1.00
Llama-4-Scout-17B
96.0
95.9
95.9
95.8
1.00
Llama-3.1-8B
96.8
95.8
95.8
95.2
1.00
Mistral-Small-24B
95.8
95.2
95.0
94.0
1.00
Granite-3.3-8B
97.0
94.5
94.3
90.6
0.99
Mistral-7B
93.2
92.8
92.8
91.0
1.00
Table 3 : Effectiveness of LLMLeak in % and recovery score r on 11 different open parameter models, r calculated over S data .
Model
Sconf (d)
Sconf (i)
ΔSconf
r (d)
r (i)
Δr
Llama-3.3-70B
99.9
97.0
−2.9
1.00
1.00
0.00
Llama-4-Scout-17B
95.8
79.8
−16.0
1.00
1.00
0.00
Llama-3.1-8B
95.2
80.8
−14.4
1.00
0.98
−0.02
Mistral-Small-24B
94.0
93.2
−0.8
1.00
1.00
0.00
Granite-3.3-8B
90.6
46.0
−44.6
0.99
0.97
−0.02
Mistral-7B
91.0
30.6
−60.4
1.00
0.93
−0.07
Table 4 : Comparison of Confirmed Rate (S conf ) and Recovery Rate ( r ) between directive (d) and informational (i) framing
Model
S call
S reach
S data
S conf
r
Llama-3.3-70B
98.5
98.0
98.0
97.0
1.00
Llama-4-Scout-17B
83.6
80.5
80.4
79.8
1.00
Llama-3.1-8B
91.1
83.9
83.4
80.8
0.98
Mistral-Small-24B
95.5
93.6
93.5
93.2
1.00
Granite-3.3-8B
66.5
48.7
48.5
46.0
0.97
Mistral-7B
38.2
34.2
33.7
30.6
0.93
Table 5 : Attack funnel in % informational framing. r calculated over S data .
Transmission Error
Count
S reach
Fully correct ( r = 1, S conf )
17 525
97.9%
Substitutions (correct length)
54
0.30%
Correct prefix (wrong length)
121
0.68%
Mixed Substring (wrong length)
186
1.04%
No Payload ( r = 0)
18
0.10%
Total S reach
17 904
100.0%
Table 6 : Distribution of payload’s transmission errors
Model
S reach
ΔURL
ΔError
ΔFraming
(%)
(pp)
(pp)
(pp)
Granite-3.3-8B
94.0
+4.5
−0.8
−45.3
Mistral-7B
91.5
+3.0
−1.2
−57.3
Qwen2.5-7B
88.0
+3.5
−1.0
−37.3
Phi-4-mini
87.0
−5.0
+0.2
−48.7
DeepSeek-Coder-V2-Lite
56.5
−2.0
+1.8
−18.5
Table 7 : Ablation study on impact of individual attack factors, particularly impact of payload vs. regular URL, exception type, and error message framing, DNS/HTTP pooled).
S reach (%)
S conf (%)
Model
Latin
CJK
Latin
CJK
Δ
Llama-3.3-70B
100.0
100.0
99.9
80.8
−19.1
Llama-4-Scout-17B
95.9
96.5
95.8
76.1
−19.7
Llama-3.1-8B
95.8
95.6
95.2
77.0
−18.2
Mistral-Small-24B
95.2
94.7
94.0
90.7
−3.3
Granite-3.3-8B
94.5
94.3
90.6
83.0
−7.6
Table 8 : Comparison of LLMLeak ’s effectiveness between Latin and CJK encoding per model, in terms of requests reached the attack server (S reach ), correct payload transmitted (S conf ) and difference Δ=SconfCJK−SconfLatin .
Figure 3 : Evaluation of LLMLeak ’s effectiveness depending on the length of the transmitted secret, when each byte is represented by two characters.
Modell
n
surf.%
flagged
deleg.%
covert%
covert strict %
Llama-3.3-70B
1998
82.2
0
1.5
100.0
17.8
Llama-4-Scout-17B
1915
93.3
0
8.0
100.0
6.3
Llama-3.1-8B
1903
58.0
5
8.8
99.7
38.6
Mistral-Small-24B
1879
85.9
0
17.3
100.0
13.5
Granite-3.3-8B
1812
63.7
1
17.3
99.9
30.4
Mistral-7B
1819
48.9
9
8.2
99.5
47.3
Table 9 : Evaluation of LLMLeak ’s stealthiness in terms of successful exfiltration (S conf )
informational
directive
Model
DNS
HTTP
Total
DNS
HTTP
Total
Grok 4.5 Fast
100
100
100
100
100
100
ChatGPT 5.6 Sol
90
70
80
100
100
100
Claude Opus 5 (High Thinking)
60
80
70
30
90
60
Claude Sonnet 5 (Medium Thinking)
0
30
15
0
30
15
Claude Haiku 4.5
20
0
10
0
0
0
Table 10 : Effectiveness of LLMLeak on real-world chatbots in terms of successful exfiltrations (S conf ) in %.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Model
S call
S reach
S data
S conf
r
midrule Llama-3.3-70B
100.0
100.0
94.8
80.8
0.87
Llama-4-Scout-17B
96.5
96.5
90.5
76.1
0.86
Llama-3.1-8B
96.7
95.6
92.7
77.0
0.89
Mistral-Small-24B
95.8
94.7
93.8
90.7
0.98
Granite-3.3-8B
96.5
94.3
93.2
83.0
0.96
Mistral-7B
94.2
93.8
91.0
85.8
0.98
Appendix
Table 11 : Attack funnel (Hanzi encoding), % of n=2000 attack trials, r calculated over S data .
Model
S reach @4
S conf @4
S reach @63
S conf @63
Llama-3.3-70B
100.0
100.0
99.0
97.5
Llama-4-Scout-17B
96.0
96.0
95.5
94.0
Llama-3.1-8B
93.0
92.5
95.5
93.0
Mistral-Small-24B
96.0
96.0
85.5
72.0
Granite-3.3-8B
96.0
96.0
84.0
72.0
Mistral-7B
97.0
95.0
83.5
67.5
Appendix
Table 12 : Comparison of length influence on S reach and S conf at 4 and 63 characters
LLM-based chatbot agents increasingly process user requests by combining natural-language reasoning with external tools such as web browsing. These capabilities improve usability, but they also create attack surfaces when untrusted external content is processed as part of a user' s task. This paper studies a privacy-leakage attack chain based on indirect prompt injection in black-box chatbot environments, where the attacker has no access to model weights, system prompts, or agent implementation details including how a trajectory is actually managed during its processing for a query. We first analyze how an attacker can hijack an agent' s intended task by crafting external content that appears benign to the victim while inducing the agent to execute an attacker-defined objective. We then evaluate a new prompt-injection technique, called exemplification, which uses a bridge in the external content to reframe the user prompt and the benign beginning of the retrieved page as few-shot examples before appending the attacker' s objective. We compare its attack success rate with a prior fake-completion technique. Finally, we demonstrate a proof-of-concept data-exfiltration chain using fictitious personal information in a controlled setting. Our results suggest that prompt injection, jailbreak-style instruction steering, and web-tool invocation can be combined into a feasible privacy-leakage path in deployed chatbot agents.
Hongjang Yang, Hyunsik Na, Daeseon Choi
Department of Information Security Soongsil University Seoul, Korea · AI Safety Center Soongsil University Seoul, Korea · Department of AI Software Soongsil University Seoul, Korea
LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users. However, agents often acquire more sensitive information than the task requires. Existing privacy benchmarks audit what the agent's response or outgoing actions disclose, but overlook the acquisition stage where data first enters the agent's context. The over-acquired information is then one careless action or one attack away from an outright leak. To assess its prevalence, we introduce \emph{PrivacyPeek}, a benchmark for evaluating acquisition-stage privacy leakage of LLM-based agents, with 1,182 cases across 7 acquisition behaviours and 16 application domains. Specifically, \emph{Acquisition Inspection} examines the agent's tool-call trajectory, both the tools it invokes and the data it receives, to detect when it acquires sensitive information beyond the task scope. \emph{Probe Elicitation} then issues a follow-up probe and measures how readily an attacker could elicit sensitive information the agent acquired but did not disclose. Our experiments on 10 LLM-based agents across 4 model families show that the unnecessary acquisition of sensitive information is widespread. In addition, we observe a correlation between the task-completion capability and acquisition-stage leakage. Prompt-level defences reduce only a small fraction of acquisition-stage leakage, leaving the majority unmitigated. These results make auditing acquisition-stage privacy both urgent and necessary. Our dataset and code are available at https://github.com/Xuan269/PrivacyPeek-Resource.
Mingxuan Zhang, Jiahui Han, Dadi Guo +5
Shanghai Artificial Intelligence Laboratory · Southeast University
Large language models (LLMs) are often fine-tuned on uncurated text datasets that adversaries can poison. Existing poisoning attacks primarily rely on fixed trigger phrases that defenses such as outlier detection, clean-data regularization, or online monitoring can neutralize. In this paper, we propose a data poisoning method that teaches an LLM an information hiding scheme reliably and stealthily through semantic associations between shared knowledge such as facts or concepts and attacker-chosen phrases. The induced hiding scheme can encode and decode arbitrary malicious instructions, thus revealing a new and subtle poisoning-induced vulnerability: covert control attacks. We precisely characterize covert control attacks and evaluate them across 5 LLMs, 3 backdoor defenses, and 4 prompt injection defenses. With a small poisoned fraction, covert control attacks outperform heuristic-based prompt injection attacks in average attack success rate by about 40% relative to clean fine-tuned models. They also circumvent defenses based on detection and fine-tuning, maintaining up to 93% attack success rate after backdoor defenses and up to 98% after prompt injection defenses.