Readable Before Actionable: Causal Tracing of Indirect Prompt Injection
Organizations: Zhejiang University · Binjiang Institute of Zhejiang University · Hong Kong Baptist University
Abstract
Indirect prompt injection causes LLM agents to follow commands embedded in external data. A probe may distinguish instructions from data without identifying a state edit that changes the next action. We study this gap through counterfactual role probes, component-wise activation patching, and separate interventions on AgentDojo trajectories. Role decoding survives changes in content and format. In controlled Qwen tests, it precedes strong tool-choice effects from patches along an independently estimated role direction. On AgentDojo, directions estimated from hijacked and resisted training trajectories reduce attack success at pre-action and injected-span positions, but have little effect at random positions. In longer Qwen trajectories, single-position edits become less effective at later layers; span-wide and repeated edits reduce attack success on the same evaluation set. Removing the learned channel subspace preserves role decoding, yet effective intervention directions transfer poorly across the tested channels. These findings distinguish a readable role signal from an effective behavioral intervention: depth matters in controlled tool choice, while position and context also matter in attack trajectories.
Figures & tables
| Purpose | Data | Metric | Control |
|---|---|---|---|
| Decode source role | 1,440 prompts/model | Role AUROC | Held-out content; reversed role–format correlation. |
| Change tool choice | 100 binary pairs; 100 examples/tool for six tools | Logit shift; choice changes | Same recipient prompt and weights; different state components. |
| Read attack outcomes | 192 balanced AgentDojo trajectories | Outcome AUROC | Controlled role direction, without refitting. |
| Compare edit sites | 96 AgentDojo trajectories/model | Paired ASR difference | Same direction and trajectories; different positions. |
| Test held-out attacks | 240 AgentDojo trajectories/test set | ASR; benign utility | Held-out trajectories and attack families. |
| Separate role from channel | 120 examples/condition; 96 behavioral trajectories | Role AUROC; ASR | With/without the learned channel subspace. |
| Direction | Estimation contrast | Evaluation |
|---|---|---|
| Instruction vs. data in controlled role pairs | Controlled patches/edits; fixed-direction outcome readout. | |
| ( , ) | Hijacked vs. resisted trajectories in AgentDojo suite | Attack outcomes in Slack or Workspace; within-/cross-suite tests. |
| Source-role pairs in synthetic tool outputs | Tool-output attacks; cross-channel transfer. | |
| Trusted system vs. untrusted retrieved memory | Memory poisoning; cross-channel transfer. |
| Evaluation split | Qwen hidden | Llama hidden | Structural only |
|---|---|---|---|
| Content-grouped 5-fold | 0.962 | 0.955 | 0.534 |
| Leave-one-family-out | 0.951 | 0.943 | 0.521 |
| Counterfactual role–format | 0.924 | 0.918 | 0.508 |
| Shuffled-label null | 0.501 | 0.498 | 0.499 |
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
| Split | Structural (95% CI) | Qwen struct. | Llama struct. | Qwen hidden | Llama hidden |
|---|---|---|---|---|---|
| Content 5-fold | 0.534 [0.478, 0.589] | 0.536 | 0.532 | 0.962 | 0.955 |
| Leave-family-out | 0.521 [0.465, 0.576] | 0.523 | 0.519 | 0.951 | 0.943 |
| Counterfactual | 0.508 [0.452, 0.563] | 0.509 | 0.507 | 0.924 | 0.918 |
| Shuffled null | 0.499 [0.441, 0.558] | 0.500 | 0.498 | 0.501 | 0.498 |
| Condition | Layer | Channel acc. | Raw role | Residual |
|---|---|---|---|---|
| Instruct Native | L20 | 0.985 | 0.972 | 0.958 |
| Instruct Wrapper | L20 | 0.981 | 0.965 | 0.951 |
| Instruct Plain | L20 | 0.972 | 0.948 | 0.932 |
| Base Plain | L24 | 0.952 | 0.912 | 0.887 |
| Position | Probe AUROC | ASR | 95% CI | Exact | ||
|---|---|---|---|---|---|---|
| No intervention | — | 84.38% | — | — | — | — |
| Pre-action | 0.972 | 31.25% | pp | [ ] | 51/0 | |
| Injected-span mean | 0.842 | 28.12% | pp | [ ] | 54/0 | |
| Random context | 0.504 | 83.33% | pp | [ ] | 1/0 | 1.000 |
| Position | Probe AUROC | ASR | 95% CI | Exact | ||
|---|---|---|---|---|---|---|
| No intervention | — | 83.33% | — | — | — | — |
| Pre-action | 0.955 | 33.33% | pp | [ ] | 48/0 | |
| Injected-span mean | 0.835 | 30.21% | pp | [ ] | 51/0 | |
| Tool-output final | 0.858 | 44.79% | pp | [ ] | 37/0 | |
| Random context | 0.502 | 82.29% | pp | [ ] | 1/0 | 1.000 |
| Condition | L20 ASR | L24 ASR | |
|---|---|---|---|
| Step 1 | 40 | 22.5% | 25.0% |
| Step 2 | 12 | 33.3% | 41.7% |
| Step 3+ | 44 | 38.6% | 84.1% |
| Single-token patch | 96 | 31.3% | 54.2% |
| Tool-output-span patch | 96 | 28.1% | 32.3% |
| Repeated L20+L22 | 96 | 27.1% | |
| Shift (LD) | Flip rate | |||||
|---|---|---|---|---|---|---|
| Layer | Full | Full | ||||
| 8 | 0.14 | 0.08 | 0.06 | 0% | 0% | 0% |
| 12 | 0.22 | 0.11 | 0.12 | 0% | 0% | 0% |
| 16 | 1.21 | 0.84 | 0.41 | 0% | 0% | 0% |
| 18 | 34.42 | 28.51 | 5.23 | 100% | 82% | 4% |
| 20 | 41.05 | 33.26 | 6.81 | 100% | 96% | 6% |
| ASR (Qwen-7B, L24) | Benign acc. | ASR (Qwen-1.5B, L22) | Benign acc. | |
|---|---|---|---|---|
| 0 | 46% [36, 56] | 100% | 34% [25, 44] | 92% [86, 97] |
| 4 | 38% [29, 47] | 100% | 32% [23, 41] | 92% [86, 97] |
| 8 | 38% [29, 48] | 100% | 28% [19, 37] | 92% [86, 97] |
| 16 | 36% [27, 46] | 100% | 30% [21, 39] | 92% [86, 97] |
| 32 | 26% [18, 35] | 100% | 28% [20, 37] | 92% [86, 97] |
| Suite | Condition | ASR (95% CI) | Paired | |
| Slack | No intervention | 240 | 34.6% [28.5, 40.8] | — |
| Slack | Controlled-role direction | 240 | 33.8% [27.8, 40.0] | — |
| Slack | Random direction | 240 | 34.2% [28.1, 40.4] | — |
| Slack | Out-of-band L8 | 240 | 33.3% [27.3, 39.5] | — |
| Slack | Suite-matched | 240 | 21.2% [16.2, 26.6] | |
| Slack | Suite-matched, LAF-out | 240 | 24.6% [19.3, 30.2] | 0.003 |
| Layer | Qwen-1.5B | Qwen-7B | ||
|---|---|---|---|---|
| Shift (LD) | Flip | Shift (LD) | Flip | |
| 14 | 0% | 0% | ||
| 16 | 0.6 | 0% | 1.2 | 0% |
| 18 | 25.5 | 75% | 34.4 | 100% |
| 22 | 28.9 | 100% | 47.1 | 100% |
| Layer | Two-tool | Six-tool | ||||
|---|---|---|---|---|---|---|
| Flip rate (%) | Mean shift (LD) | Top-1 target (%) | Flip rate (%) | Mean shift (LD) | Top-1 target (%) | |
| L10 | 0.0% | 0.0% | 0.8% | 1.5% | ||
| L14 | 0.0% | 0.0% | 4.2% | 8.3% | ||
| L16 | 0.0% | 0.0% | 25.6% | 34.0% | ||
| L18 | 100.0% | 100.0% | 68.2% | 71.5% | ||
| L20 | 100.0% | 100.0% | 95.4% | 96.8% | ||
| Layer | Mean shift (LD) | SD across pairs (LD) |
|---|---|---|
| 0 | 0.088 | |
| 2 | 0.099 | |
| 4 | 0.116 | |
| 6 | 0.162 | |
| 8 | 0.193 | |
| 10 | 0.182 |
| Decoding | Temperature | Top- | Baseline ASR | L24 ASR | Baseline utility | L24 utility |
|---|---|---|---|---|---|---|
| Greedy | 0.0 | – | 46.0 [40.3, 51.7] | 26.0 [21.0, 31.0] | 88.3 [84.6, 92.0] | 87.0 [83.1, 90.9] |
| Mild sampling | 0.3 | 0.90 | 48.4 [42.7, 54.1] | 27.5 [22.4, 32.6] | 87.8 [84.1, 91.5] | 86.4 [82.5, 90.3] |
| Standard sampling | 0.7 | 0.95 | 52.1 [46.4, 57.8] | 29.8 [24.6, 35.0] | 86.5 [82.6, 90.4] | 85.1 [81.1, 89.1] |
| Decoding strategy | Inflection layer band | Peak intervention layer |
|---|---|---|
| Greedy ( ) | L18–L22 | L24 |
| Mild sampling ( ) | L18–L22 | L24 |
| Standard sampling ( ) | L18–L22 | L24 |
| Model | Params | Layers | Benign acc. | Baseline ASR | First perfect layer |
|---|---|---|---|---|---|
| Qwen-2.5-0.5B | 0.49B | 24 | 75% | 32.1% | 0.083 (L2) |
| Qwen-2.5-1.5B | 1.5B | 28 | 91.7% | 37.5% | 0.286 (L8) |
| Qwen-2.5-3B | 3.1B | 36 | 100% | 37.1% | 0.111 (L4) |
| Qwen-2.5-7B | 7.6B | 28 | 100% | 47.0% | 0.143 (L4) |
| Llama-3.1-8B | 8.0B | 32 | 100% | 27.1% | 0.125 (L4) |
| Qwen-2.5-14B (4-bit) | 14.8B | 48 | 100% | 42.5% | 0.167 (L8) |
| Attack | No int. | |||
|---|---|---|---|---|
| Direct | 2.854 | 85.1% | 8.5% | 8.5% |
| Paraphrase | 2.412 | 82.4% | 14.2% | 10.5% |
| Leetspeak | 1.251 | 62.1% | 48.5% | 22.4% |
| Base64 | 0.985 | 45.2% | 39.1% | 25.1% |
| Multi-hop | 1.054 | 55.4% | 49.2% | 28.5% |
| Semantic concealment | 0.822 | 74.2% | 62.1% | 34.2% |
| Layer | Memory | Tool output | ||||
|---|---|---|---|---|---|---|
| CV | Held-out template | DOM | CV | Held-out template | DOM | |
| Qwen-2.5-1.5B | ||||||
| 0 | 0.619 | 0.508 | 0.553 | 0.999 | 0.883 | 0.886 |
| 4 | 0.753 | 0.625 | 0.686 | 0.997 | 0.883 | 0.988 |
| 8 | 0.875 | 0.783 | 0.828 | 0.999 | 1.000 | 0.994 |
| 12 | 0.942 | 0.858 | 0.908 | 1.000 | 1.000 | 0.999 |
| Layer | clean proj | poisoned proj | Cohen’s (H R) |
|---|---|---|---|
| 0 | |||
| 4 | |||
| 14 | |||
| 22 | |||
| 26 |
| Result | GPU-hours |
|---|---|
| Probe extraction and grouped/counterfactual fits (3 models) | |
| Two-/six-tool patching plus layer, direction, and position controls | |
| Directional-intervention strength sweeps (3 model–layer settings) | |
| AgentDojo position, direction, and sampling controls | |
| Memory-channel probe and intervention | |
| Model-scale scan (5 Qwen sizes plus Llama-3.1-8B) |
| Cosine to full | ASR change (pp) | Utility drop (pp) | |
|---|---|---|---|
| Direction | Training source | Evaluation source | Content | Family | Trajectory |
|---|---|---|---|---|---|
| Synthetic role pairs | Synthetic attacks | Yes | Partial | Yes | |
| Slack train (LAF-out) | Slack held-out | Yes | Yes | Yes | |
| Workspace train (LAF-out) | Workspace eval | Yes | Yes | Yes | |
| Memory role pairs | Memory poisoning | Yes | Yes | Yes |
| Intervention | Slack ASR | Workspace ASR |
|---|---|---|
| No intervention | 34.6% | 31.2% |
| Rank-1 (in-channel) | 21.8% | 20.2% |
| Rank-2 (in-channel) | 20.4% | 19.1% |
| Rank-4 (in-channel) | 19.8% | 18.9% |
| Rank-8 (in-channel) | 19.5% | 18.8% |
| Rank-1 Slack Workspace | — | 30.5% |