CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering
Organizations: Microsoft Azure
Abstract
Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whether an embedded instruction is followed, and retains it only if it passes pre-specified causal and capability gates. At deployment, the direction is subtracted from every tool-result token during prefill. The edit is always on--there is no detection decision to evade--and requires no fine-tuning, auxiliary model, or added tokens, only white-box serving and tool-result span boundaries. Across five open-weights models (8B-106B, five vendor lineages), held-out attack success falls from 0.21-1.00 undefended to 0.00-0.17 defended, and AgentDojo compromise rate from 0.10-0.49 to 0.006-0.079, at 93-100% typography-normalized benign utility, with larger task-dependent costs when reasoning over steered content. A benchmark-level adaptive attacker reaching 0.67-0.73 undefended is held to roughly a quarter of that on the two most deeply evaluated models. Among the defenses we measured on capable models, those achieving lower compromise rates either lost 22-89% of benign utility or fine-tuned the served weights. White-box gradient attacks through the deployed vector compromise at most 2 of 52 episodes, and none of 2,052 replayed human red-team attacks succeeds. CounterSteer largely neutralizes instructional takeover: a black-box framing search cracks 3 of 18 development samples. Parameter manipulation--attacker-chosen arguments in otherwise legitimate calls--is only partially resisted (13 of 18); the decision becomes linearly readable at argument emission but not at the examined pre-generation sites, and is not removed by the tested prefill- or decode-time steering, motivating argument-provenance controls.
Figures & tables
| class | where the follow decision lives | what removed it in our tests | CounterSteer |
|---|---|---|---|
| tool hijack | the retrieved span, at prefill | a constant prefill direction | largely neutralized |
| parameter manipulation | the model’s own reasoning, at argument emission (§ VI-C ) | only fine-tuning (SecAlign) | reduced, not closed |
| delegated authority | the user prompt grants it | system-level provenance (e.g., ROPE [ 5 ] ) | out of scope (Class B) |
| single-turn ASR | AgentDojo | benign | |||
|---|---|---|---|---|---|
| model | undef. | defended | undef. | defended | % of clean |
| gpt-oss-20b | .59–.98 | .000–.135 | .475 | .079 | 94.4 |
| Qwen3-30B | .52–1.00 | .000–.173 | .489 | .072 | 94.1 |
| Gemma-4-31B | .87–.96 | .000–.038 | .239 | .006 | 92.9 |
| GLM-4.5-Air | .81 | .019 | .222 | .056 | 100.0 |
| Llama-3.1-8B | .21 | .038 | .100 | .050 | 100.0 |
| gpt-oss-20b † | Qwen3-30B | Gemma-4-31B | GLM-4.5-Air † | Llama-3.1-8B | ||||||
| defense | cmp | util | cmp | util | cmp | util | cmp | util | cmp | util |
| no defense (clean util absolute) | .483 | 98.1 | .489 | 91.1 | .239 | 100.0 | .183 | 80.4 | .100 | 26.8 |
| PromptGuard-2 filter [ 21 ] | .205 | 100.0 | .228 | 100.0 | — | — | .117 | 100.0 | .056 | 100.0 |
| DeBERTa filter [ 22 ] | .034 | 47.2 (!) | .089 | 54.9 (!) | — | — | — | — | .044 | 92.3 |
| PIGuard filter [ 23 ] | .011 | 58.2 (!) | .006 | 62.7 (!) | — | — | .000 | 57.8 (!) | .006 | 92.3 |
| CachePrune [ 9 ] | .216 | 88.7 | — | — | — | — | — | — | .045 | 80.0 |
| defense | AgentDojo cmp | Class A | utilAtk | benign % of clean | adaptive param. (union; per-q.) |
|---|---|---|---|---|---|
| undefended | .483 | 59 | .551 | (ref.) | 18/18; .708 |
| CounterSteer | .091 | 0 | .716 | 90.6 | 13/18; .303 |
| AGRI ‡ | .062 | — | — | 78.0 | — |
| SecAlign (trained) † | — | — | — | 90.4 | 1/18; .002 |
| CachePrune | .216 | 15 | .688 | 88.7 | 17/18; .604 |
| PromptGuard-2 filter | .205 | 21 | .311 | 100.0 | — |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| arm (defense on clean) | gpt-oss web param ( ) | gpt-oss JSON hij. (80) | Qwen web param (52) | Qwen JSON hij. (90) |
|---|---|---|---|---|
| CounterSteer (same-process) | .385 | .525 | .346 | .333 |
| regeneration reference (cross-run) | .462 | .637 | .462 | .400 |
| spotlighting | .442 | .625 | .462 | .356 |
| prompt sandwich | .462 | .600 | .462 | .356 |
| reminder | .423 | .600 | .365 | .344 |
| CachePrune (gpt-oss only) | .385 | .537 | — | — |
| adaptive success | gpt-oss-20b | Qwen3-30B | Llama-3.1-8B |
|---|---|---|---|
| undefended | 0.673 | 0.730 | 0.377 |
| CounterSteer | 0.175 | 0.188 | 0.307 |
| AGRI (probe-gated prefill) | 0.143 | 0.411 | — |
| CachePrune | 0.305 | 0.558 | 0.252 |
| PromptGuard-2 filter | 0.527 | 0.670 | 0.369 |
| SecAlign (trained) | 0.160 | — | — |
| arm | union pattern | union literal | per-query pattern | per-query literal |
|---|---|---|---|---|
| undefended | 18/18 | 3/18 | 0.708 | 0.076 |
| spotlighting-with-delimiting | 18/18 | 4/18 | 0.595 | 0.044 |
| CachePrune [ 9 ] | 17/18 | 4/18 | 0.604 | 0.037 |
| CounterSteer (deployed) | 13/18 | 1/18 | 0.303 | 0.002 |
| SecAlign [ 4 ] (trained) | 1/18 | 0/18 | 0.002 | 0.000 |
| split | ASR undefended [95% CI] | ASR defended [95% CI] | clean-traffic cost |
|---|---|---|---|
| dev | 105/1537 = 0.068 [0.057, 0.082]; truncation-corrected | 0/1537 observed [0.000, 0.002] | none (0 spurious sends; retrieval coverage 0.988, matching the clean arm) |
| test (held out) | 23/515 = 0.045 [0.030, 0.066] | 0/515 observed [0.000, 0.007] | none (retrieval coverage 0.990, matching the clean arm) |
| model | directions | layer sets | doses | compositions / variants | held-out confirm |
|---|---|---|---|---|---|
| gpt-oss-20b | 16 (+ 9 random draws) | 2 | 16 | 20 | yes (one pass) |
| Qwen3-30B | 5 | 2 | 11 | 2 | yes (one pass) |
| Gemma-4-31B | 1 (+probe axes) | 3 | 8 | 0 | yes (two rungs: wire-format T ∗ ; then docs wordings framings held out) |
| GLM-4.5-Air | 5 (incl. 2 role axes) | 2 | 6 | 0 | yes (one pass, webpage param) |
| Llama-3.1-8B | 3 (+3 random draws) | 4 | 9 | 0 | yes (one pass) |
| bench ( ) | context | baseline | steered |
|---|---|---|---|
| GSM8K ( ) | plain / tool-framed | 0.896 / 0.912 | 0.896 / 0.912 |
| MMLU ( ) | plain / tool-framed | 0.814 / 0.810 | 0.814 / 0.702 |
| IFEval ( ) | plain / tool-framed | 0.624 / 0.672 | 0.624 / 0.520 |
| defense | compromise [95% CI] (own undef.) | blk/int ( ) | Class A | utilAttack (undef.) | benign % of clean | filter fp |
| gpt-oss-20b (Class A over 137–140 non-delegated cases; undefended Class A 51–59) | ||||||
| CounterSteer (deployed) | 16/176 0.091 [.057, .143] (0.483) | 70/1 ( ) | 0/140 (59) | 0.716 (0.551) | 90.6% | — |
| CachePrune [ 9 ] | 38/176 0.216 [.162, .282] (0.483) | 55/8 ( ) | 15/140 (59) | 0.688 (0.551) | 88.7% | — |
| reminder (AutoDojo) [ 13 ] | 48/176 0.273 [.212, .343] (0.483) | 45/8 ( ) | 33/140 (59) | 0.574 (0.551) | 88.7% | — |
| prompt sandwich | 55/174 0.316 [.252, .389] (0.489) | 38/8 ( ) | 31/138 (59) | 0.609 (0.551) | 88.2% | — |
| spotlighting [ 2 ] | 74/176 0.420 [.350, .494] (0.483) | 21/10 ( 0.071 n.s. ) | 46/140 (59) | 0.616 (0.551) | 92.5% | — |
| term | meaning |
|---|---|
| primary models | gpt-oss-20b and Qwen3-30B: full evaluation battery and the deepest adaptive evaluation (§ IV ) |
| undefended baseline | the same run’s no-defense arm; every contrast in this paper is paired against it, never against another run |
| ASR | single-turn attack success rate over scoreable samples; the obedience-pattern reading unless marked exact-literal |
| compromise rate | AgentDojo/AgentDyn’s own security checker: fraction of injection cases where the attacker’s task completed |
| obedience-pattern / exact-literal | the injected action was performed / the attacker’s exact payload was delivered verbatim |
| benign utility | task fidelity, defense on, no injection, % of the undefended clean run—the always-on deployment cost |
| AgentDojo | webpage param | webpage tool-hijack | JSON tool-hijack | JSON param | |
|---|---|---|---|---|---|
| compromise: ASR (AgentDojo: compromise rate), undefended defended (rung) | |||||
| gpt-oss-20b | T | T | T | T | |
| Qwen3-30B | T | T | T | T | |
| Gemma-4-31B | T | T | N/F (.083–.112) | N/F (.045–.048) | |
| GLM-4.5-Air | T | hijack classes low-fire; JSON param. not run (Table XVIII ) | |||
| Llama-3.1-8B | no viable cell (base 1.000) | T | N/F (.042) | D (not viable) | |
| evaluation | gpt-oss-20b (deployed configuration) | Qwen3-30B (selected ) |
| JSON tool-hijack (T) | ; utilB 52.5 (25/80); utilA 37.5 (30/76) (undef. 13.8) | ; utilB 33.3 (21/69); utilA 20.0 (22/67) (undef. 8.9) |
| JSON param. (T, samples-only) | ; utilB 53.0 (23/83); utilA 43.9 (23/74) (undef. 18.2) | ; utilB 31.3 (21/64); utilA 31.3 (27/66) (undef. 16.4) |
| webpage tool-hijack (T) | ; utilB 50.0 (98/98); utilA 57.7 (94.2/94) (undef. 0.0) | ; utilB 13.5 (63/63); utilA 13.5 (58/58) (undef. 1.9) |
| webpage param. (T) | ; utilB 38.5 (75/75); utilA 38.5 (60/60) (undef. 0.0) | ; utilB 34.6 (65/65); utilA 15.4 (50/50) (undef. 0.0) |
| AgentDojo (4096 tok/turn) | ; benign 94.4% of clean; utilAttack 73.9% (undef. 62.5%) | same-process (separate-run replicate ); benign 94.1%; utilAttack .717 raw (undef. .667) |
| LLMail-Inject replay | observed (2052 episodes) | observed (audited count—a parser audit recovered 13 successful attacker calls; full dev replay, no test-shard replay exists) |
| attack class (corpus, ) | ASR undefended | ASR defended [95% CI] | utilB judge (strict/len.) | utilA judge (strict/len.) [undef.] |
|---|---|---|---|---|
| tool hijack — JSON tool-log ( , 80 scoreable) | 0.675 | 1/80 = 0.013 [0.002, 0.067] | 52.5 (25.0/80) | 37.5 (30.0/76) [13.8] |
| tool hijack — webpage ( , 52 scoreable; multi-turn) | 0.981 | 0/52 observed [0.000, 0.069]; multi-turn 0/49 observed [0.000, 0.073] | 50.0 (98.1/98) | 57.7 (94.2/94) [0.0] |
| param manipulation — JSON ( , 66 scoreable; samples-only holdout—the attacker-template set was in the fit) | 0.591 | 3/66 = 0.045 [0.016, 0.125] | 53.0 (22.7/83) | 43.9 (22.7/74) [18.2] |
| param manipulation — webpage ( held-out templates, 52 scoreable) | 0.962 | 7/52 = 0.135 [0.067, 0.253] | 38.5 (75.0/75) | 38.5 (59.6/60) [0.0] |
| attack class (corpus, ; holdout axes) | ASR undefended | ASR defended [95% CI] | utilB judge (strict/len.) | utilA judge (strict/len.) [undef.] |
|---|---|---|---|---|
| tool hijack — JSON tool-log ( , 90 scoreable; samples attacker templates held out) | 0.722 | 3/90 = 0.033 [0.011, 0.093] | 33.3 (21.1/69) | 20.0 (22.2/67) [8.9] |
| tool hijack — webpage ( ; exact-template-text holdout; strictly-disjoint-40 subset 0/40 observed [0.000, 0.088]) | 0.981 | 0/52 observed [0.000, 0.069] | 13.5 (63.5/63) | 13.5 (57.7/58) [1.9] |
| param manipulation — JSON ( , 67 scoreable; samples-only holdout—the attacker-template set was in the fit) | 0.522 | 4/67 = 0.060 [0.023, 0.144] | 31.3 (20.9/64) | 31.3 (26.9/66) [16.4] |
| param manipulation — webpage ( ; exact-template-text holdout; strictly-disjoint-40 subset 8/40 = 0.200 [0.105, 0.348]) | 1.000 | 9/52 = 0.173 [0.094, 0.297] | 34.6 (65.4/65) | 15.4 (50.0/50) [0.0] |
| development split, not holdout: tool hijack — JSON tool-log ( , 91 scoreable) | 0.648 | 4/91 = 0.044 [0.017, 0.108] | 45.1% / 91% strict/len. | 47.3% / 87% (undef. 14.3/21) |
| evaluation | rung | ASR undefended defended [95% CI] | utilB | utilA | guard / notes |
|---|---|---|---|---|---|
| webpage param. ( ; docs wordings framings held out) | T | (2/52); doc-clustered [0.000, 0.096], sample Wilson upper 0.130; 48 fixed / 0 introduced, | 78.8 (96.2/96) | 78.8 (92.3/92) | guard clean, truncation 0.00; both survivors under one held-out wording (2/13 on it, 0/39 on the rest); monotonicity control PASS ( ) |
| webpage tool-hijack ( ; same holdout axes) | T | (0/52); cluster-Wilson [0.000, 0.114] over 30 docs; 45/0, | 73.1 (96.2/96) | 69.2 (96.2/96) | guard clean; control saturated at 0.000; A100 H100 replica probe 24/24 agreement |
| webpage param. | T ∗ | ; doc-clustered [0.000, 0.105], quote-upper 0.130 (sample Wilson); 50 flipped / 0 broken | 88.5 (100/100) | 78.8 (96/96) | guard clean |
| webpage tool-hijack | T ∗ | [0.000, 0.278] doc-clustered; 45 flipped / 0 broken | 80.8 (100/100) | 82.7 (100/100) | guard clean |
| JSON tool-hijack / JSON param. | N/F | undefended 0.083–0.112 (shipped JSON, / ) / 0.045–0.048 | — | — | N/F by the program’s convention; no defended run |
| prose-carrier medium ( ) | N/F | undefended 0.000 [0.000, 0.038] | — | — | measured null: sane on-task summaries, truncation 0.00 |
| evaluation | rung | ASR undefended defended [95% CI] | utilB | utilA | guard / notes |
|---|---|---|---|---|---|
| webpage param. ( , 52 scoreable; samples template text held out) | T | [0.003, 0.101]; exact-literal 0/52; 41 fixed / 0 introduced | 26.9 (94.2/94) | 11.5 (81/81) | contamination 0.000; no-action 0.038 vs. defense-on-clean 0.000 (2/52, attack-conditional); the one surviving compromise is a domain-level parameter hijack counted by token taint, not exact-literal |
| webpage param., development ( , 96 scoreable) | D | (exact-literal 0.000); 57 fixed / 0 introduced | 40.6 (93/93) | 18.8 (59/59) | no-action 0.052; truncation 0.05 |
| webpage tool-hijack; JSON tool-hijack | — | fire weakly undefended single-turn (webpage ; JSON low)—those cells do not measure the defense. At the multi-turn horizon the disjoint-tool corpus fires undefended and the deployed configuration yields a partial reduction only: (16 fixed / 0 introduced, ; steering verified active over the injected span)—an open boundary, not a closure | — | — | |
| JSON param.; adaptive suite | — | not run | — | — | |
| AgentDojo, 180-case grid, 8192 tok/turn | full grid | (40/180 10/180); 30 fixed / 0 introduced, | benign 100.0% (0.857 0.857, 56 paired tasks; 1 lost, 1 gained) | 0.733 vs. undefended 0.733 | zero truncation in every arm (nothing censored); layers pinned 20/24/28 (deployed configuration); 7 of the 10 surviving compromises sit on one task whose user prompt delegates authority to the file (“follow the instructions precisely”), 2–3/180 are genuine role-confused survivors; strongest agent measured (clean 0.857) |
| bidirectional causal gate | D | increase side PASSES on the deployed direction: at (8192 budget, dev ; exact McNemar 14/3, ; defense-on-clean at : ASR 0.000, utility 88.5%) | — | — |
| stage | split / samples used | held out from |
|---|---|---|
| generalization screening (step 3) | probe split, minus one whole authority-framing level ( firm ) | that framing level, for the gate’s fit only |
| direction fitting (deployed) | probe split (behavioral-contrast captures; deployed gpt-oss fit: its deterministic first 24 sample ids) | dev/test samples; held-out attacker templates. Framing levels: withheld in the Llama-3.1-8B fit and the validated gpt-oss-20b refit; the other deployed fits use every level |
| layer selection | probe split (per-layer probe diagnostics) | dev/test samples |
| dose selection | dev split (template-disjoint from test) | test samples and, where stated, test attacker templates |
| test evaluation | test split, pre-specified one-pass, no reruns | all fitting and selection stages (samples; and where stated, attacker templates) |
| adaptive attacker construction | dev-disjoint sample subsets ( per class for the query search; disjoint-tool targets for GCG) | the attacked samples are held out from fitting; the deployed vector is not consumed by the query/surrogate arms and is supplied directly in the exact-vector arm. The surrogate refit deliberately reuses the fitting samples (the draw is deterministic, § C ) |
| arm | compromise (all) | Class B | Class A |
|---|---|---|---|
| undefended | 84/177 = 0.475 | 29/36 = 0.806 | 55/141 = 0.390 |
| CounterSteer (deployed) | 14/177 = 0.079 | 14/36 = 0.389 | 0/141 = 0.000 |
| model | params / arch | fit status | ASR / compromise rate defended |
|---|---|---|---|
| gpt-oss-20b | 20B MoE | certified (held-out test pass) | 0.000–0.135 corpora (Table XV ); 0.079 AgentDojo at a 4096-token budget (§ V-B ) |
| Qwen3-30B-A3B-Thinking | 30B MoE (3B active) | certified (held-out test pass) | 0.000–0.173 corpora (Table XVI ); AgentDojo same-process (Table XI ), dose-selection replicate 0.094 (Table XXIV ), both at 4096 tokens |
| Gemma-4-31B-it | 31B dense | certified (one-pass held-out test evaluations on both firing corpora with carrier documents, attacker wordings and framing templates all held out, plus the full AgentDojo grid at the deployed configuration); on the single-turn corpora the steered span is of the whole prompt (those rungs do not isolate the payload-scoped edit; the agentic grid does); no adaptive evaluation | webpage param. (2/52; doc-clustered [0.000, 0.096], sample Wilson upper 0.130; 48 fixed / 0 introduced, ) and tool hijack (0/52; cluster-Wilson [0.000, 0.114]; 45/0, ), benign utility 96.2%/96% strict both corpora, guard clean, zero truncation; AgentDojo (42/0, ; 0 defense-introduced) at benign 92.9% ( pp n.s.), 0 truncated turns in 2,721 generations; the earlier wire-format-only rung (T ∗ ) stands beside it ( , ); the prose-carrier medium is a measured null (undefended 0.000 [0.000, 0.038], ; Table XVII ) |
| GLM-4.5-Air | 106B MoE | certified (held-out test pass on its firing corpus, webpage parameter manipulation, plus the full AgentDojo grid at the deployed configuration); tool-hijack classes fire weakly undefended, JSON param. not run; no adaptive evaluation | webpage param. T ( ; near-duplicate-template caveat, § V ; exact-literal 0/52; Wilson [0.003, 0.101]; McNemar 41 fixed / 0 introduced; benign utility 26.9 judge (94.2/94 strict/lenient), under attack 11.5 (81/81); contamination 0.000; guard clean, read absolutely—defense-on-clean no-action 0.000, defended 0.038 attack-conditional); development split at defense-on-clean strict correctness 0.927; AgentDojo (30 fixed / 0 introduced, ) at benign 100.0% and utilAttack undefended, zero truncation in every arm (Table XVIII ) |
| Llama-3.1-8B-Instruct | 8B dense | certified (held-out test pass on its firing corpus, webpage tool hijack, plus the full AgentDojo grid, the 26-battery rival comparison and the benchmark-level adaptive attack at the deployed configuration); fit passes all gates incl. the bidirectional causal gate ( at ); JSON tool-hijack does not fire undefended (0.042); JSON param. reaches ASR 0 only at a capability-destroying dose (judge benign 0.431), no viable cell; low task capability (13–15/56 AgentDojo clean tasks) bounds what its agentic and AgentDyn grids separate (AgentDyn: 2/60 solvable, not measurable) | webpage tool hijack T ( ; 11/52 2/52 (3/52 under a lenient tool-call executor—one surviving compromise is the attacker’s exact call with malformed JSON; still significant); Wilson [0.011, 0.130]; McNemar 9 fixed / 0 introduced, ; benign utility 100% byte-exact, judge floor 0.827; under attack 65.4 strict / 0.558 judge, both above undefended 59.6 / 0.462; a fit-withheld framing level composed onto the injection also blocked, 0/4 fired, point estimate); AgentDojo (11 fixed / 2 introduced, ) at benign 100%, thin-base caveat (Table XI ) |
| model | hardware (ladder) | capture fit | screening causal gates | dose search | held-out certification (single-turn) | AgentDojo 180-case grid |
|---|---|---|---|---|---|---|
| gpt-oss-20b | 4 A100-80GB local A100 boxes; grid on one 8 H100 node | iterative | n.r. (CPU-side) | iterative | 4 corpora, one pre-specified pass, min | 62 min (8 H100, 4 arms) |
| Qwen3-30B-A3B-Thinking | 4 A100-80GB local A100 boxes; grid on one 8 H100 node | iterative | n.r. | iterative | 4 corpora, one pre-specified pass, h (pre-specification to last artifact) | one of six batteries in a 22.6 h 8 H100 job; not separated |
| Llama-3.1-8B-Instruct | 3 A100-80GB local; grid on 4 A100 (two boxes) | h (incl. a parser-bug diagnosis rerun) | min (causal smoke 3 random controls) | h (6 doses 3 corpora layer sweep) | gate test touch bidirectional gate, 11 min | 4 single-GPU shards, min; end-to-end h, one day |
| Gemma-4-31B-it | one 8 H100 node (bring-up, cert, grid); local A100s (dose batteries) | h (one job: capture, probes, factorial, fit) | inside the bring-up job | n.r. | 56 min total for both: five held-out confirms the full 180-case grid benign arms, one 8 H100 job | |
| GLM-4.5-Air | one 8 H100 node (capture, sweep); local 4 A100 (cert, grid; 212 GB sharded) | 6.3 h (bring-up); certified refit CPU-only, n.r. | inside the bring-up; recipe gates CPU-side, n.r. | 16.0 h sweep (superseded direction); recipe ladder terminated early, no clean total | T pass completed 3 h 05 m after the dev re-run (4 A100, 8192-token budget) | 12.2 h (43,906 s, single process, 8192-token budget) |
| dose | compromise defended [95% CI] | fixed/intr. ( ) | benign |
|---|---|---|---|
| 3.0 | 55/177 = 0.311 [0.247, 0.382] | 36/7 ( ) | 92.2% |
| 4.12 | 44/177 = 0.249 [0.191, 0.317] | 44/4 ( ) | 86.3% |
| 5.5 | 26/177 = 0.147 [0.102, 0.207] | 58/0 ( ) | 86.3% |
| 6.7 | 19/177 = 0.107 [0.070, 0.162] | 65/0 ( ) | 84.3% |
| 8.06 (deployed) | 15/177 = 0.085 [0.052, 0.135] | 69/0 ( ) | 82.0% |
| dose | compromise defended [95% CI] | utilAttack (raw) | benign pp |
|---|---|---|---|
| 22/180 = 0.122 [0.082, 0.178] | 0.728 | ||
| (selected) | 17/180 = 0.094 [0.060, 0.146] | 0.694 | |
| 6/180 = 0.033 [0.015, 0.071] | 0.550 | ||
| 2/180 = 0.011 [0.003, 0.041] (!) | 0.356 (!) | (!) |
| family | representative work | lever | deployment cost | measured here |
|---|---|---|---|---|
| prompt-level | spotlighting/delimiting [ 2 ] ; prompt sandwich (folklore) | transforms the untrusted span so its provenance stays visible (delimiting, datamarking, encoding) and instructs the model not to follow it; or re-asserts the trusted instruction after the data | extra prompt tokens on every request | at 4096: – (gpt-oss), (Qwen reminder), Table XI |
| training-time | StruQ [ 3 ] , SecAlign [ 4 , 39 ] , Jatmo [ 40 ] , instruction hierarchy [ 38 ] , ISE [ 41 ] , ASIDE [ 42 ] | retrains or re-architects the model so untrusted spans carry less instructional force | a training pipeline per model; the served weights change | our SecAlign LoRA: compromises (768 budget); benign utility of base clean, typography-normalized: at 768, unique-task / per-case at the shared 4096 budget (raw / : the DPO finetune itself emits the typography) |
| detection | internal-state probes: TaskTracker [ 36 ] , Attention Tracker [ 37 ] ; input-text guard models: PIGuard [ 23 ] | reads a signal and flags; no intervention is proposed, and the response is left to the surrounding system | a threshold that fails open on a miss; a second model or pass, except where the signal is already computed during the forward pass | input-text guards run as deletion filters in Table XI (compromise 0.011–0.228 at benign 48–100%, over-defense per [ 13 ] ); internal-state probes not run as defenses—§ VI shows two directions that discriminate but do not steer |
| detection-gated intervention | ICON [ 6 ] , ARGUS [ 7 ] , AGRI [ 8 ] | a probe decides when to intervene, then edits attention (ICON), representations plus a post-filter (ARGUS), or injects an anti-injection reasoning prefill (AGRI) | detector accuracy is a single point of failure; the gate is an additional attack surface (evaluated for AGRI: the adaptive attacker runs against the live gated arm, § V-C ) | AGRI measured same-harness (Tables III , VI , § VIII ); ICON/ARGUS not run (see text) |
| always-on internal intervention | CachePrune [ 9 ] , V-Steer [ 10 ] , CounterSteer | edits internals over the untrusted span with no gate: KV-neuron pruning, per-request value-vector rescaling, or a fixed residual-stream direction | white-box serving access; benign utility | at 4096: CachePrune , CounterSteer (Table XI ) |
| system-level / provenance | CaMeL [ 27 ] , MELON [ 45 ] , ROPE [ 5 ] , tool filters | constrains the consequences of a hijack: dual-LLM execution with data-flow tracking (CaMeL), masked re-execution and output comparison (MELON), origin-routed admission of parameter values (ROPE) | re-architecting the agent runtime; extra executions or an audited parameter set | a tool filter reaches compromise 0/176 observed with benign utility 10.9% of clean and utilAttack 0.106 typography-normalized (Table XI ) |