Indirect prompt injection causes LLM agents to follow commands embedded in external data. A probe may distinguish instructions from data without identifying a state edit that changes the next action. We study this gap through counterfactual role probes, component-wise activation patching, and separate interventions on AgentDojo trajectories. Role decoding survives changes in content and format. In controlled Qwen tests, it precedes strong tool-choice effects from patches along an independently estimated role direction. On AgentDojo, directions estimated from hijacked and resisted training trajectories reduce attack success at pre-action and injected-span positions, but have little effect at random positions. In longer Qwen trajectories, single-position edits become less effective at later layers; span-wide and repeated edits reduce attack success on the same evaluation set. Removing the learned channel subspace preserves role decoding, yet effective intervention directions transfer poorly across the tested channels. These findings distinguish a readable role signal from an effective behavioral intervention: depth matters in controlled tool choice, while position and context also matter in attack trajectories.
Figures & tables
Figure 1: Role readout precedes strong patching effects in controlled Qwen tests. (A) Schematic role-component patches at the pre-action token; L16 and L20 are separate runs. (B) Measured role readout and tool-choice flips, evaluated on separate datasets ( 100 tool-choice pairs). Lines connect measured layers.
Hijacked vs. resisted trajectories in AgentDojo suite c
Attack outcomes in Slack or Workspace; within-/cross-suite tests.
dtool
Source-role pairs in synthetic tool outputs
Tool-output attacks; cross-channel transfer.
dmemory
Trusted system vs. untrusted retrieved memory
Memory poisoning; cross-channel transfer.
Table 2: Direction definitions. All contrasts use training data. Slack and Workspace subscripts denote suite-specific estimates.
Evaluation split
Qwen hidden
Llama hidden
Structural only
Content-grouped 5-fold
0.962
0.955
0.534
Leave-one-family-out
0.951
0.943
0.521
Counterfactual role–format
0.924
0.918
0.508
Shuffled-label null
0.501
0.498
0.499
Table 3: Source-role AUROC under held-out and counterfactual splits. The structural classifier uses only length, position, punctuation, wrapper, and format features. Hidden-state results are reported separately for each model; the structural baseline is pooled across the balanced model sets.
Figure 2: Role readout and intervention effects. (A,D) Role readout versus structural and format controls. (B,C) Position and channel comparisons on separate evaluation sets. (E,F) Component-wise and cross-model tool-choice patching. Left branches denote separate analyses; lines connect measured estimates.
Figure 3: Channel-matched interventions on held-out AgentDojo trajectories. (A,B) Slack controls and Workspace/Llama replications with separately estimated directions ( n=240 /test set; 95% bootstrap CIs). (C) Leave-attack-family-out point estimates. (D) Training-sample efficiency; bars show SD over five subsets, except the fixed full-pool endpoint.
Figure 4: Channel removal and intervention transfer. (A,B) Qwen role AUROC ( n=120 /condition) and ASR ( n=96 ) before/after channel removal; point estimates, with Instruct conditions only in B. (C) ASR reductions (percentage points) from each column’s baseline; outlines mark matched channels. The first three columns re-express Figure 2 C; the fourth tests obfuscation.
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Split
Structural (95% CI)
Qwen struct.
Llama struct.
Qwen hidden
Llama hidden
Content 5-fold
0.534 [0.478, 0.589]
0.536
0.532
0.962
0.955
Leave-family-out
0.521 [0.465, 0.576]
0.523
0.519
0.951
0.943
Counterfactual
0.508 [0.452, 0.563]
0.509
0.507
0.924
0.918
Shuffled null
0.499 [0.441, 0.558]
0.500
0.498
0.501
0.498
Appendix
Table 4: Structural-only and hidden-state source-role AUROC. Structural results are pooled across the balanced Qwen and Llama sets; model-level values show the same near-chance pattern in each model.
Condition
Layer
Channel acc.
Raw role
Residual
Instruct Native
L20
0.985
0.972
0.958
Instruct Wrapper
L20
0.981
0.965
0.951
Instruct Plain
L20
0.972
0.948
0.932
Base Plain
L24
0.952
0.912
0.887
Appendix
Table 5: Role readout after channel removal across Qwen-2.5-7B tuning and serialization conditions. “Residual” is role AUROC after removing the learned channel-identity subspace.
Position
Probe AUROC
ASR
Δ
95% CI
b/c
Exact p
No intervention
—
84.38%
—
—
—
—
Pre-action
0.972
31.25%
−53.13 pp
[ −63.5,−42.7 ]
51/0
8.88×10−16
Injected-span mean
0.842
28.12%
−56.25 pp
[ −65.6,−45.8 ]
54/0
1.11×10−16
Random context
0.504
83.33%
−1.04 pp
[ −3.1,0.0 ]
1/0
1.000
Appendix
Table 6: Qwen-2.5-7B position comparisons at L20 ( n=96 ). CI is the paired-bootstrap interval for the ASR difference, computed from the discordant-pair counts. Here b counts baseline successes prevented by intervention; c counts new successes caused by intervention.
Position
Probe AUROC
ASR
Δ
95% CI
b/c
Exact p
No intervention
—
83.33%
—
—
—
—
Pre-action
0.955
33.33%
−50.00 pp
[ −60.4,−39.6 ]
48/0
7.11×10−15
Injected-span mean
0.835
30.21%
−53.12 pp
[ −63.5,−42.7 ]
51/0
8.88×10−16
Tool-output final
0.858
44.79%
−38.54 pp
[ −47.9,−29.2 ]
37/0
1.46×10−11
Random context
0.502
82.29%
−1.04 pp
[ −3.1,0.0 ]
1/0
1.000
Appendix
Table 7: Llama-3.1-8B position sweep at L22 ( n=96 ). CI is the paired-bootstrap interval for the ASR difference, computed from the discordant-pair counts. Definitions of b/c match Table 6 .
Condition
n
L20 ASR
L24 ASR
Step 1
40
22.5%
25.0%
Step 2
12
33.3%
41.7%
Step 3+
44
38.6%
84.1%
Single-token patch
96
31.3%
54.2%
Tool-output-span patch
96
28.1%
32.3%
Repeated L20+L22
96
27.1%
Appendix
Table 8: AgentDojo rebound breakdown on the matched position set. The first three rows stratify single-token patching by trajectory length. The final rows compare patch spans on all 96 trajectories.
Shift (LD)
Flip rate
Layer
Full h
h∥
h⊥
Full h
h∥
h⊥
8
0.14
0.08
0.06
0%
0%
0%
12
0.22
0.11
0.12
0%
0%
0%
16
1.21
0.84
0.41
0%
0%
0%
18
34.42
28.51
5.23
100%
82%
4%
20
41.05
33.26
6.81
100%
96%
6%
Appendix
Table 9: Component-wise patching on Qwen-2.5-7B ( n=100 ). At L18–L22, parallel-component patches yield much higher flip rates than perpendicular-component patches.
Figure 5: Per-attack results on Qwen-2.5-7B (L24, λ=32 ). (A) Baseline versus intervention ASR; area encodes sample size, color attack family. (B) Ten largest ASR reductions; blue marks reductions of at least 10 percentage points.
λ
ASR (Qwen-7B, L24)
Benign acc.
ASR (Qwen-1.5B, L22)
Benign acc.
0
46% [36, 56]
100%
34% [25, 44]
92% [86, 97]
4
38% [29, 47]
100%
32% [23, 41]
92% [86, 97]
8
38% [29, 48]
100%
28% [19, 37]
92% [86, 97]
16
36% [27, 46]
100%
30% [21, 39]
92% [86, 97]
32
26% [18, 35]
100%
28% [20, 37]
92% [86, 97]
Appendix
Table 10: One-sided directional intervention (Eq.( 2 )). ASR is attack-success rate; benign acc. is benign-task accuracy. Brackets give 95% bootstrap CIs ( B=10,000 , n=100 ). At λ=32 , Qwen-7B ASR falls by 20 percentage points; the Qwen-1.5B estimate falls by 6 points and remains within sampling uncertainty.
Suite
Condition
n
ASR (95% CI)
Paired p
Slack
No intervention
240
34.6% [28.5, 40.8]
—
Slack
Controlled-role direction
240
33.8% [27.8, 40.0]
—
Slack
Random direction
240
34.2% [28.1, 40.4]
—
Slack
Out-of-band L8
240
33.3% [27.3, 39.5]
—
Slack
Suite-matched
240
21.2% [16.2, 26.6]
<0.001
Slack
Suite-matched, LAF-out
240
24.6% [19.3, 30.2]
0.003
Appendix
Table 11: Larger-scale AgentDojo intervention (Qwen-2.5-7B, L20, λ=32 ). Confidence intervals resample trajectories. Reported p -values are paired-bootstrap comparisons with the baseline.
Layer
Qwen-1.5B
Qwen-7B
Shift (LD)
Flip
Shift (LD)
Flip
14
0.3
0%
<0.3
0%
16
0.6
0%
1.2
0%
18
25.5
75%
34.4
100%
22
28.9
100%
47.1
100%
Appendix
Table 12: Full-state activation patching on matched tool-choice prompts. Mean logit-difference shift (LD) and sign-flip rate for the two Qwen models.
Layer
Two-tool
Six-tool
Flip rate (%)
Mean shift (LD)
Top-1 target (%)
Flip rate (%)
Mean shift (LD)
Top-1 target (%)
L10
0.0%
−0.08
0.0%
0.8%
+0.42
1.5%
L14
0.0%
+0.29
0.0%
4.2%
+1.15
8.3%
L16
0.0%
+1.20
0.0%
25.6%
+3.80
34.0%
L18
100.0%
+34.44
100.0%
68.2%
+8.45
71.5%
L20
100.0%
+41.01
100.0%
95.4%
+12.10
96.8%
Appendix
Table 13: Multi-tool activation patching across layers (Qwen-2.5-7B; two-tool n=100 pairs, six-tool n=100 per tool). Mean shift = change in target-vs.-current logit difference (LD). Top-1 target = fraction where the patched-in target tool becomes top-1.
Layer
Mean shift (LD)
SD across pairs (LD)
0
+0.001
0.088
2
+0.013
0.099
4
+0.099
0.116
6
+0.217
0.162
8
+0.255
0.193
10
+0.073
0.182
Appendix
Table 14: Llama-3.1-8B per-layer activation-patching shift ( n=100 pairs, weather → calc direction). The columns report the mean and SD across pairs. The mean patched calc logit-difference crosses zero at L16 (base =−7.95 LD; patched =+6.36 LD at L16).
Figure 6: Qwen-2.5-7B intervention at L24 under greedy and stochastic decoding ( T=0.3,0.7 ). (A) ASR and (B) benign utility before and after intervention.
Decoding
Temperature
Top- p
Baseline ASR
L24 ASR
Baseline utility
L24 utility
Greedy
0.0
–
46.0 [40.3, 51.7]
26.0 [21.0, 31.0]
88.3 [84.6, 92.0]
87.0 [83.1, 90.9]
Mild sampling
0.3
0.90
48.4 [42.7, 54.1]
27.5 [22.4, 32.6]
87.8 [84.1, 91.5]
86.4 [82.5, 90.3]
Standard sampling
0.7
0.95
52.1 [46.4, 57.8]
29.8 [24.6, 35.0]
86.5 [82.6, 90.4]
85.1 [81.1, 89.1]
Appendix
Table 15: Intervention robustness across decoding strategies (Qwen-2.5-7B, n=300 ). Brackets are 95% bootstrap CIs.
Decoding strategy
Inflection layer band
Peak intervention layer
Greedy ( T=0.0 )
L18–L22
L24
Mild sampling ( T=0.3 )
L18–L22
L24
Standard sampling ( T=0.7 )
L18–L22
L24
Appendix
Table 16: Intervention-layer sweep on controlled Qwen-2.5-7B prompts across sampling strategies.
Model
Params
Layers
Benign acc.
Baseline ASR
First perfect layer
Qwen-2.5-0.5B
0.49B
24
75%
32.1%
0.083 (L2)
Qwen-2.5-1.5B
1.5B
28
91.7%
37.5%
0.286 (L8)
Qwen-2.5-3B
3.1B
36
100%
37.1%
0.111 (L4)
Qwen-2.5-7B
7.6B
28
100%
47.0%
0.143 (L4)
Llama-3.1-8B
8.0B
32
100%
27.1%
0.125 (L4)
Qwen-2.5-14B (4-bit)
14.8B
48
100%
42.5%
0.167 (L8)
Appendix
Table 17: Six-model scan under the controlled cross-template protocol. “Benign acc.” is benign-task accuracy, and “Baseline ASR” is attack success after adversarial text replaces benign data. “First perfect layer” is the first layer with cross-template probe accuracy 1.00 , divided by model depth. Qwen-2.5-14B uses 4-bit weight quantization.
Figure 7: Scale scan under the controlled protocol. (A) Baseline ASR versus parameter count. (B) Relative depth of the first layer with perfect cross-template probe accuracy. Measurements are from Table 17 .
Attack
d
No int.
dtool
dobf.
Direct
2.854
85.1%
8.5%
8.5%
Paraphrase
2.412
82.4%
14.2%
10.5%
Leetspeak
1.251
62.1%
48.5%
22.4%
Base64
0.985
45.2%
39.1%
25.1%
Multi-hop
1.054
55.4%
49.2%
28.5%
Semantic concealment
0.822
74.2%
62.1%
34.2%
Appendix
Table 18: Obfuscation ladder. Cohen’s d measures standardized separation of attacked and benign state projections. The tool-output direction weakens as d decreases (Qwen-2.5-7B, L24, λ=32 ).
Layer
Memory
Tool output
CV
Held-out template
DOM
CV
Held-out template
DOM
Qwen-2.5-1.5B
0
0.619
0.508
0.553
0.999
0.883
0.886
4
0.753
0.625
0.686
0.997
0.883
0.988
8
0.875
0.783
0.828
0.999
1.000
0.994
12
0.942
0.858
0.908
1.000
1.000
0.999
Appendix
Table 19: Memory and tool-output source-role classification ( n=360 pairs per channel). CV: five content-grouped folds; held-out template: train on five templates, test on the sixth; DOM: difference-of-means threshold. Entries are accuracies. Bold marks peak memory-probe accuracy; dashes denote unreported measurements.
Figure 8: Within-channel strength sweeps using channel-matched directions. Markers show tested strengths; the dotted line is the tool-output random-direction control. Settings and evaluation sets differ from Figure 4 .
Layer
clean proj
poisoned proj
Cohen’s d (H − R)
0
+0.74
+0.79
+0.09
4
−1.24
−1.14
+0.54
14
+5.14
+5.02
+0.45
22
+7.63
+7.35
+0.56
26
+17.14
+18.10
+0.62
Appendix
Table 20: Per-layer pre-action residual projected onto the source-role direction d , on 240 memory-poisoning prompts (Qwen-2.5-1.5B). Cohen’s d summarizes separation between hijacked and resisted projections; it is not the patch-induced logit shift measured in § 5 .
Result
GPU-hours
Probe extraction and grouped/counterfactual fits (3 models)
≈13
Two-/six-tool patching plus layer, direction, and position controls
AgentDojo position, direction, and sampling controls
≈13
Memory-channel probe and intervention
≈4
Model-scale scan (5 Qwen sizes plus Llama-3.1-8B)
≈27
Appendix
Table 21: Approximate GPU-hours per result type (single-node, mixed 24 GB and 80 GB GPUs).
ntrain
Cosine to full d
ASR change (pp)
Utility drop (pp)
8
0.452±0.185
−2.1±1.8
+5.2±3.4
16
0.684±0.092
−5.4±2.2
+2.8±1.5
32
0.852±0.041
−9.8±1.5
+1.8±0.8
64
0.925±0.022
−11.5±0.8
+1.5±0.4
128
0.968±0.012
−12.8±0.5
+1.6±0.3
240
1.000
−13.4
+1.6
Appendix
Table 22: Sample-efficiency sweep for dslack (Qwen-2.5-7B, L20, λ=32 ). Mean and standard deviation across five random subset seeds for ntrain<240 . The full-pool endpoint uses one fixed direction, so no subset SD is reported.
Figure 9: Bidirectional state patching on the matched Qwen-2.5-7B AgentDojo-Slack trajectories. Pre-action swaps have their largest effect at L18–L22; ASR moves back toward the unpatched outcome at L24.
Direction
Training source
Evaluation source
Content
Family
Trajectory
dtool
Synthetic role pairs
Synthetic attacks
Yes
Partial
Yes
dslack
Slack train (LAF-out)
Slack held-out
Yes
Yes
Yes
dworkspace
Workspace train (LAF-out)
Workspace eval
Yes
Yes
Yes
dmemory
Memory role pairs
Memory poisoning
Yes
Yes
Yes
Appendix
Table 23: Train–test separation for direction estimation. The last three columns indicate whether the corresponding unit is held out.
Intervention
Slack ASR
Workspace ASR
No intervention
34.6%
31.2%
Rank-1 (in-channel)
21.8%
20.2%
Rank-2 (in-channel)
20.4%
19.1%
Rank-4 (in-channel)
19.8%
18.9%
Rank-8 (in-channel)
19.5%
18.8%
Rank-1 Slack → Workspace
—
30.5%
Appendix
Table 24: Rank- k SVD intervention on AgentDojo (Qwen-2.5-7B, L20, λ=32 ).
LLMs see the world as a single stream of text, partitioned into roles like <user> or <tool>. We trace prompt injection to role confusion: models perceive the source of text from how it sounds, not its labeled role. A command hidden in a webpage hijacks an agent simply because it sounds like <user> text, despite its <tool> label. We design role probes to measure how LLMs internally perceive "who is speaking," and find that injected text occupies the same representational space as the trusted role it imitates. We demonstrate this with CoT Forgery, a zero-shot attack that injects fabricated reasoning into user prompts and tool outputs. Models mistake the forgery for their own thoughts, yielding 60% attack success against frontier models with near-zero baselines. Strikingly, the degree of role confusion predicts attack success before a single token is generated. This mechanism generalizes beyond CoT Forgery to standard agent prompt injections, revealing prompt injection as a measurable consequence of role perception. To the model, sounding like a role is indistinguishable from being one. Project page and writeup: https://role-confusion.github.io
Charles Ye, Jasmine Cui, Dylan Hadfield-Menell
*Equal contribution 1Independent · 2Massachusetts Institute of Technology, Cambridge, MA, United States
As LLM agents take real-world actions through tools, indirect prompt injection (IPI) has emerged as a serious threat. The standard metric, Attack Success Rate (ASR), counts whether an injection succeeds but ignores what the user notices in the agent's final response. Looking at successful injection traces, we find two distinct outcomes: the agent executes the injection while returning an otherwise normal response, or reports the injected action in its final response, giving the user a chance to notice. We call these covert and overt successes. From the user's perspective, we decompose ASR into the Covert Success Rate (CSR), counting successes leaving no trace in the final response, and the Overt Success Rate (OSR), counting successes the user can detect. To understand what drives the gap, we analyze successful trajectories and find that the agent's behavior after the injection separates covert from overt: covert traces hand control back to the user task before ending, while overt traces end at the attack itself. This split follows from the ReAct format, where the final response summarizes the most recent action. Building on this observation, we propose ICoA (Induced Covert Attack), an IPI attack designed to induce covert outcomes by steering the agent back to the user task after executing the injection. Across four target models on AgentDojo, ICoA achieves the highest CSR, with gains of 3.79-12.01 percentage points over the strongest baseline.
Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While many efforts have sought to address the threats, little is known about the internals of agentic LLMs when they are exposed to IPI attacks, a condition which we call IPI exposure. In this paper, we study this problem in depth from three aspects. (1) Probing: Across six models, including the giant 753B-parameter GLM-5.2, simple linear probes trained on pre-generation hidden states can predict LLMs' IPI exposure. These probes achieve 90%+ AUROC on unseen attacks, agent instructions, and task suites; they exhibit high robustness under adaptive attacks and in cross-lingual settings. (2) Defense: Our CoT measurement reveals a recognition--action gap: though models encode such signals, they often fail to translate them into safe actions. We then introduce AGRI, a probe-gated reasoning-based defense that prepends anti-injection reasoning on demand. On difficult AgentDojo settings, AGRI substantially reduces attack success rate, e.g., from 34.6% to 0% on Qwen3.5-27B, while largely maintaining clean-task utility. (3) Explanation: We introduce an analysis framework that identifies natural-language explanations most strongly correlated with probe-captured signals. The resulting profiles differ across models: latent signals can align with either direct IPI-exposure claims or indirect operational cues. Code is available: https://github.com/jianshuod/IPI-exposure-signal.
Jianshuo Dong, Yiming Liu, Maosen Zhang +6
Tsinghua University · MatrixOrigin · Nanyang Technological University +1