Multi-agent LLM systems now read documents, web pages and tool results on behalf of users, yet their resistance to prompt injection is usually reported as one number: did the attack succeed? We introduce a kill-chain canary method that plants a unique token in every injected payload and records the furthest of four stages it reaches (Exposed -> Persisted -> Relayed -> Executed), across 950 runs, five production LLMs, six attack surfaces, and five defense conditions. Exposure was 100% among runs that called the tool; the outcomes differ downstream. Claude Haiku 4.5 and Claude Sonnet 4.5 executed none of their 164 text-surface attacks, and in the text relay the canary token never appeared in a memory write (0/40); GPT-4o-mini executed 53% of its attacks. Four findings follow. (1) A Claude writer kept the canary token out of shared memory in every relay run we report; one cross-model pairing (Claude writer, GPT-4o-mini reader, n = 3) is consistent with this protecting the reader, and other pairings were not tested. (2) As readers, the Claude models executed 0/40 raw pre-seeded injections, but Claude Haiku 4.5 executed 2/3 injections relayed by GPT-4o-mini; whether relayed injections are harder to refuse than raw ones is an open question. (3) DeepSeek Chat went from 0/24 on pre-seeded memory to 8/8 on tool results, scenarios that also differ in task and payload format; white-text PDF payloads, invisible on the rendered page, succeeded at least as often as visible ones. (4) pi_detector and write_filter failed on channels they do not inspect, spotlighting failed on content it wraps, and write_filter blocked the PDF relay but not the text relay, a difference we cannot explain. Code and run logs are publicly released: https://github.com/KevinChunye/prompt_injection
Figures & tables
Figure 2 : Benchmark setup. Top: the user task and an attacker payload carrying its canary token. Middle: the attack surfaces (named as in Section 3.2 ), which reach Agent A as tool results, except pre-seeded memory, which enters shared memory directly; the two-agent pipeline (Agent A reads and writes shared memory; Agent B reads memory and calls send_report ); the four checkpoints (1 Exposed, 2 Persisted, 3 Relayed, 4 Executed); and where each defense looks. Dashed: in the single-agent tool_poison scenario the agent that reads the tool result sends the report itself. memory_poison is also single-agent (its agent is drawn as Agent B), and permission_esc passes the injection from Agent A to Agent B in a delegation message instead of shared memory. Bottom: models, attack surfaces, defense conditions and run counts: 8–36 runs per cell on text surfaces (Table A1 ), 3 per cell in the PDF relay (Table 4 and Figure 5 ), and 4 per model for the audio pilot (Table 5 ).
Model
n
Task
ASR (95% CI)
GPT-4o-mini
60
90%
53% (41–65%)
DeepSeek Chat
68
100%
25% (16–37%)
GPT-5-mini
136
94%
3% (1–7%)
Claude Haiku 4.5
80
100%
0% (0–5%)
Claude Sonnet 4.5
84
100%
0% (0–4%)
Table 1 : Attack and task success per model, text surfaces. n : no-defense attacked runs, including runs that made no tool call. ASR: share of those runs that reached Executed, with Wilson 95% CI. Task: task success on the model’s clean-control runs (all defense conditions). Exposure was 100% among runs that called the tool (Section 4.1 ).
Figure 3 : Last stage the canary token reached, as a share of attacked runs per model, no defense. Left: text relay ( propagation ), n=8 – 20 per model (Table 2 ). Middle and right: PDF relay, same-model pairs, visible text ( pdf_append ) and white text ( pdf_whitefont ), n=3 per model (Table 4 ). Green: Exposed only (the token was absent from the memory write); amber: Relayed to Agent B but not Executed; red: Executed. Hatched: GPT-5-mini never called parse_pdf , so it never read the payload; every text relay run called its tool.
Model
n
Persisted
Relayed
Executed
GPT-4o-mini
8
100%
100%
100%
DeepSeek Chat
8
100%
100%
100%
GPT-5-mini
20
15%
15%
15%
Claude Haiku 4.5
20
0%
0%
0%
Claude Sonnet 4.5
20
0%
0%
0%
Table 2 : Kill-chain stage reached, text relay ( propagation ), no-defense attacked runs. Exposure was 100% among runs that called the tool, and every text relay run called it, so Exposed is 100% in every row.
Figure 4 : Attack success by model and injection channel, no defense: runs that reached Executed over attacked runs ( k/n ), with Wilson 95% CIs in brackets. Column headers give the attack surface and the scenario. Text-surface cells ( n=8 – 36 , including runs that made no tool call) are also listed in Table A1 ; PDF relay cells are same-model pairs ( n=3 ), the Executed columns of Table 4 . Hatched: GPT-5-mini never called parse_pdf . Outlined: DeepSeek Chat’s 0/24 on memory_poison and 8/8 on tool_poison (Section 4.3 ).
Step
Tool
Clean
Attacked
Δ
1
read_memory
0.563
0.571
+0.008
2
send_report (legit)
0.714
0.747
+0.033
3
send_report (harm)
—
0.868
+0.154
Table 3 : Per-step objective drift (TF-IDF cosine distance from the task description) for one clean and one attacked GPT-4o-mini memory_poison run.
pdf_append (visible text)
pdf_whitefont (white text)
Model
Exposed
Persisted
Relayed
Executed
Exposed
Persisted
Relayed
Executed
GPT-4o-mini
100%
100%
100%
0%
100%
100%
100%
33%
DeepSeek Chat
100%
100%
100%
100%
100%
100%
100%
100%
Claude Haiku 4.5
100%
0%
0%
0%
100%
0%
0%
0%
Claude Sonnet 4.5
100%
0%
0%
0%
100%
0%
0%
0%
GPT-5-mini
0%
0%
0%
0%
0%
0%
0%
0%
Table 4 : PDF relay, same-model pairs (Agent A = Agent B), no defense, attacked runs, n=3 per cell. Exposed: canary in the extracted text; Persisted: in the memory write; Relayed: in Agent B’s read_memory result; Executed: in Agent B’s send_report arguments (= ASR).
Figure 5 : PDF relay by writer (Agent A, rows) and reader (Agent B, columns): runs that reached Executed and runs whose memory write contained the token (Persisted), out of n=3 per cell; pdf_append , no defense. Diagonal cells are same-model pairs (Table 4 ). Off-diagonal cells are cross-model pairs, in which Relayed equals Persisted in every pair; Executed with Wilson 95% CIs: GPT-4o-mini writer and Claude Haiku 4.5 reader 2/3 (21–94%); GPT-4o-mini writer and DeepSeek Chat reader 3/3 (44–100%); Claude Haiku 4.5 writer and GPT-4o-mini reader 0/3 (0–56%); DeepSeek Chat writer and GPT-4o-mini reader 3/3 (44–100%). Hatched: pairing not tested.
Surface
Model
n
ASR (95% CI)
audio_inject
GPT-4o-mini
4
0% (0–49%)
audio_inject
Claude Sonnet 4.5
4
0% (0–49%)
pdf_metadata
GPT-4o-mini
4
0% (0–49%)
pdf_metadata
GPT-5-mini
4
0% (0–49%)
pdf_metadata
Claude Sonnet 4.5
4
0% (0–49%)
Table 5 : Pilot surface results (no-defense attacked runs, n=4 per cell). These runs are not part of Table 1 or Figure 4 .
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
pre-seeded memory
tool result
web page
web page
Model
memory_poison
tool_poison
propagation
permission_esc
GPT-4o-mini
12/12
8/8
8/8
0/28
DeepSeek Chat
0/24
8/8
8/8
1/28
GPT-5-mini
0/36
0/36
3/20
1/36
Claude Haiku 4.5
0/20
0/20
0/20
0/20
Claude Sonnet 4.5
0/20
0/20
0/20
0/20
Appendix
Table A1 : Attack success k/n per model and text scenario (no-defense attacked runs, including runs that made no tool call; Figure 4 adds Wilson 95% CIs). The permission_esc total of 2/132 comprises DeepSeek Chat 1/28 and GPT-5-mini 1/36.
Figure A1 : Objective drift (Section 4.5 ). (a) Per-step drift for one clean and one attacked GPT-4o-mini memory_poison run; values are Table 3 ; shaded: the harmful send_report call. (b) Feature importances of the gradient-boosted classifier (21 features, 5-fold CV, AUC 0.853; the 14 largest shown).
LLMs see the world as a single stream of text, partitioned into roles like <user> or <tool>. We trace prompt injection to role confusion: models perceive the source of text from how it sounds, not its labeled role. A command hidden in a webpage hijacks an agent simply because it sounds like <user> text, despite its <tool> label. We design role probes to measure how LLMs internally perceive "who is speaking," and find that injected text occupies the same representational space as the trusted role it imitates. We demonstrate this with CoT Forgery, a zero-shot attack that injects fabricated reasoning into user prompts and tool outputs. Models mistake the forgery for their own thoughts, yielding 60% attack success against frontier models with near-zero baselines. Strikingly, the degree of role confusion predicts attack success before a single token is generated. This mechanism generalizes beyond CoT Forgery to standard agent prompt injections, revealing prompt injection as a measurable consequence of role perception. To the model, sounding like a role is indistinguishable from being one. Project page and writeup: https://role-confusion.github.io
Charles Ye, Jasmine Cui, Dylan Hadfield-Menell
*Equal contribution 1Independent · 2Massachusetts Institute of Technology, Cambridge, MA, United States
Prompt-injection benchmarks for LLM agents typically test attacks through a single injection surface and report the resulting attack success rate as a property of the model. We ask whether those robustness conclusions remain stable when the same adversarial content enters through a different part of the agent interface. Using AgentDojo, we evaluate 13 LLMs across four task suites and place a byte-identical payload either in a tool output or in the tool description. This small change produces large differences in comparative robustness: 44.9% of all model pairs change their relative ordering across the two surfaces, with substantial ranking instability in every suite. The effect is especially pronounced for a small number of models, showing that a benchmark can substantially underestimate vulnerability when it tests only one surface. We further find that this behavior has predictive structure. Using three suites to identify the riskier surface for each model predicts the more vulnerable surface on an unseen suite with 76.9% accuracy. Defense results show the same dependence, as mitigations effective against tool-output attacks can leave substantial exposure through tool descriptions. Our results show that prompt-injection robustness is not surface-invariant and that agent evaluations should test the surfaces on which their security conclusions depend.
Syed Nazmus Sakib, Nafiul Haque, Shahrear Bin Amin +1
Department of Robotics and Mechatronics Engineering, University of Dhaka · Department of Computer Science and Engineering, University of Dhaka
A growing class of agentic systems maintain persistent state across sessions through memory files, behavioral preferences, and knowledge bases. While this makes agents more useful and self-improving, it also creates a new attack surface for prompt injections in which malicious instructions can be embedded within persistent files and influence future behavior. In this work, we study prompt injection attacks in memory-based agentic systems using a sandboxed synthetic workspace. We evaluate two agentic systems, Anthropic Claude Code and OpenAI Codex, across four models: Claude Haiku 4.5, Claude Opus 4.7, GPT-5.2, and GPT-5.5. Our results show that although it is difficult to make an agent overwrite its own memory files using untrusted external content, payloads already planted in those files can successfully attack current and future sessions. Attack success and payload persistence vary substantially across systems, models, adversarial goals, and multi-session attack sequences. These findings show that persistent memory changes the threat model for prompt injection and motivate defenses that protect memory updates without removing useful agent adaptation.