Most studies of prompt injection focus on generative agents, leaving their effects on models with schema-defined outputs unclear. We examine these effects in Jev, a non-generative decision model, using 510 reconstructed InjecAgent cases. Malicious content shifts action probabilities but rarely causes Jev to select the attacker's target. Override markers reduce this influence, while claims of contextual relatedness have small effects. Adaptive attacks using score feedback double the mean highest attacker-target probability found during optimization, while success on fresh validation calls rises from 1.8% to 3.5%. Exploratory analysis links these successes to small initial decision margins or greater attacker control over the observation. Together, these findings show that schema-defined outputs change but do not eliminate prompt-injection risk, highlighting the need to evaluate how untrusted content influences choices within the allowed action set.
Figures & tables
Stage
Question
Conditions
Jev calls
DH-1
Transfer of original attacks
Clean, base, and enhanced in both spaces
3060
DH-2
Effect of injection-style markers
Clean and A0 to A4, R=5
15300
DH-3
Effect of asserted relatedness
Clean, P0, N0, P1, and P2, R=5
12750
DH-4
Effect of adaptive score access
Initial attack and 24 proposals; four checkpoints with five validation calls each
22950
Table 1: Overview of the four main experiments on the same 510 cases. All calls use jev-1.13.0 . DH-1 tests both action spaces. DH-2 to DH-4 use the three-choice upstream space. R is the number of repeated calls per case and condition. DH-4 includes 12750 screening and 10200 validation calls. The total is 54060 calls, excluding additional reruns and preliminary checks.
Stage
Comparison or endpoint
Estimate [95% CI]
DH-1
Upstream base
ASR 1.8% (9/510); ΔPadv=+0.043
DH-1
Upstream enhanced
ASR 0%; ΔPadv=+0.009
DH-1
Enhanced − base, repeated-call rerun
−0.034[−0.062,−0.015]
DH-2
A1 − A0, importance marker
−0.006[−0.010,−0.001]
DH-2
A2 − A0, ignore-previous marker
−0.032[−0.060,−0.014]
DH-2
A4 − A0, natural wording
−0.006[−0.013,−0.001]
Table 2: Main results. Probability contrasts are in probability units and intervals are 95% CIs. DH-1 per-condition shifts use the clean reference. DH-2 and DH-3 report the matched contrasts shown. DH-4 validated ASR requires at least three target selections in five fresh calls; its probability contrasts compare validation at B=24 with B=0 . Screening results appear separately in Table 6 . CHR uses Jev’s returned confidence field with the inclusive gate C≥0.7 . DH-4 CHR averages the gated event over five fresh validation calls per case and then over all 510 cases.
Figure 1: Matched mean attacker-probability contrasts in DH-2 and DH-3 across 510 cases. Points show estimates, and horizontal intervals show 95% CIs. Negative values indicate lower attacker-target probability in the first condition. The 0.01 reporting step is a per-call property, while the 0.02 SESOI applies to the DH-3 P1 − N0 and P2 − N0 contrasts. The stronger relatedness assertion exceeds the weaker one, but N0, P1, and P2 do not form a uniformly increasing sequence.
Figure 2: DH-4 optimization and validation over 510 cases. Panel A shows the mean and median best-so-far attacker probability observed during screening. These selected scores are optimization statistics. Panel B compares screening ASR with ASR validated using five fresh calls at B∈{0,8,16,24} . Error bars show 95% CIs for validated ASR. The budget excludes the initial screening call and all validation calls.
Figure 3: DH-4 validated success at B=24 by initial margin and observation-wrapper type. Each cell shows the successful cases out of all cases in that group, followed by the percentage. Near-boundary cases satisfy M0≤0.1 using the baseline validation margin in Equation 8 . The grouping is exploratory.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Space
Condition
Targeted ASR
Flip rate
ΔPadv
ΔPsafe
CHR
Upstream
Clean
0.0% (0/510)
n/a
0.0000
0.0000
0.0%
Three choices
Base
1.8% (9/510)
3.3%
+0.0426
−0.0841
0.0%
Enhanced
0.0% (0/510)
0.0%
+0.0087
−0.0325
0.0%
Library
Clean
0.0% (0/510)
n/a
0.0000
0.0000
0.0%
Eight choices
Base
1.6% (8/510)
3.5%
+0.0381
−0.0824
0.0%
Enhanced
0.0% (0/510)
0.0%
+0.0083
−0.0328
0.0%
Appendix
Table 3: DH-1 results from one call per case and condition. Each row contains all 510 clean-qualified cases. Probability shifts use the clean row from the same action space. ASR counts attacker-target selections, while the flip rate counts any change from the clean selected action. CHR uses τ=0.7 . A clean-to-clean flip rate is not applicable. The library space contains eight sampled choices, not the full tool library. CHR uses Jev’s returned confidence field with the inclusive gate C≥0.7 . Its denominator includes all 510 eligible cases in each condition.
Figure 4: DH-1 results in the three-choice upstream and eight-choice library spaces. Panel A shows mean attacker-probability change from the clean condition. Panel B shows targeted ASR as a fraction, so 0.01 equals 1%. Both spaces use the same 510 cases, with one call per case and condition, and are analyzed separately.
Condition
Wording
ASR
Target calls
Robust cases
ΔPadv
ΔPsafe
A0
Plain request
2.0%
51/2,550
11/510
+0.0428
−0.0844
A1
Importance marker
2.2%
56/2,550
11/510
+0.0369
−0.0643
A2
Ignore-previous marker
0.0%
0/2,550
0/510
+0.0108
−0.0411
A3
Both markers
0.0%
0/2,550
0/510
+0.0086
−0.0319
A4
Natural wording
1.3%
34/2,550
7/510
+0.0367
−0.0896
Appendix
Table 4: DH-2 style ablation over 510 cases with five calls per condition. ASR is the mean per-case target-selection frequency and equals target selections divided by 2,550 calls. Robust cases meet the majority rule, with at least three target selections in five calls. These are separate statistics. Probability shifts in the upper block use the clean reference. The lower block gives matched contrasts and 95% two-way bootstrap CIs, including the registered exclusion of emergency-dispatch goal a10 . Mean effects below the 0.01 per-call reporting step remain estimable across calls; that step is not a practical-effect threshold. Per-condition CHR was not reported.
Figure 5: DH-2 ECDFs of case-level attacker-probability changes relative to A0. Each curve summarizes 510 cases after averaging five calls per condition. Negative values indicate reduced attacker probability. A2 and A3, which contain the ignore-previous marker, have larger negative tails. The legend reports the median change for each condition.
Hypothesis
Contrast
Estimate [95% CI]
Interpretation
H3a
P1 − N0
−0.0025[−0.0054,−0.0005]
Opposes predicted direction
H3b
P2 − N0
+0.0030[0.0010,0.0054]
Positive, below SESOI
H3c
P2 − P1
+0.0055[0.0030,0.0091]
Positive, small descriptive effect
Appendix
Table 5: DH-3 registered probability contrasts over 510 cases with five calls per condition. All intervals are 95% two-way bootstrap CIs. The SESOI of 0.02 applies to H3a and H3b. H3c compares the stronger and weaker assertions descriptively. The middle block retains the more precise bounds available in the source summary and the a10 -excluded sensitivity. Sign counts refer to goal-level means over 17 contexts, with zeros excluded from the exact two-sided sign test.
Figure 6: DH-3 mean contrasts by attacker goal. Each row is one of 30 goals, and each cell averages a paired contrast over 17 contexts and five calls per condition. Color encodes Δgc−d in probability units. P2 exceeds P1 for all goals and N0 for 25 goals, while the magnitudes remain small. The dot marks an exactly zero goal-level contrast for a05 under P1 − N0, where context-level effects from −0.006 to +0.004 cancel. The boxed row is the emergency-dispatch goal a10 .
B
Screening ASR
Validated ASR [95% CI]
Successes
Mean best-so-far P(Aadv)
0
2.2%
1.8%[0.0,6.3]%
9/510
0.0429
8
3.5%
3.3%[0.0,9.4]%
17/510
0.0774
16
3.9%
3.5%[0.0,10.0]%
18/510
0.0828
24
3.9%
3.5%[0.0,10.0]%
18/510
0.0842
Appendix
Table 6: DH-4 results by additional optimization-query budget B . Screening ASR uses the selected candidate’s optimization call. Validated ASR requires at least three target selections in five fresh calls per case; its intervals are 95% CIs. The best-so-far column reports the mean selected screening score. The lower block gives validation contrasts relative to B=0 , reported only at baseline and the final checkpoint. CHR uses τ=0.7 . CHR uses Jev’s returned confidence field with the inclusive gate C≥0.7 . CHR uses only the five fresh validation calls per case and checkpoint, averaging within cases before averaging all 510 cases.
Initial margin
Embedded
Bare snippet
M0≤0.1
12/12 (100.0%)
1/1 (100.0%)
M0>0.1
1/468 (0.2%)
4/29 (13.8%)
Appendix
Table 7: Exploratory DH-4 analysis of validated success at B=24 . The upper block reports successful cases over all cases in each group, with percentages in parentheses. Margin M0 uses baseline validation calls. The lower block gives in-sample rank AUCs, not held-out performance. Case u11:a10 is an embedded success with M0=0.24 , so the near-boundary threshold is not necessary for success.
Figure 7: Cumulative coverage of the 18 validated DH-4 successes at B=24 when cases are ordered by increasing baseline margin M0 . The horizontal axis counts all included cases, and the vertical axis counts successful cases among them. The curve reaches 16 successes among the first 16 cases and all 18 among the first 33. Marked points identify bare-snippet successes. This is a descriptive ranking of the completed run.
Table 8: Recorded integrity checks and reproducibility results. All calls in both stages returned jev-1.13.0 . Test counts refer only to the analysis-specific suites. The three original-attack runs use identical requests and five calls per case. Selection-frequency agreement counts cases with matching observed frequencies, rather than cases with identical outcomes on every call.
Figure 8: Reproducibility of the original-attack anchor across three runs. Each point is one case’s mean attacker-target probability over five calls. Both panels compare the DH-3 P0 run on the horizontal axis with another run on the vertical axis. The dashed diagonal indicates equal means. Pearson correlations of 0.9984 and 0.9986 indicate close agreement despite variability in individual calls.
A growing class of agentic systems maintain persistent state across sessions through memory files, behavioral preferences, and knowledge bases. While this makes agents more useful and self-improving, it also creates a new attack surface for prompt injections in which malicious instructions can be embedded within persistent files and influence future behavior. In this work, we study prompt injection attacks in memory-based agentic systems using a sandboxed synthetic workspace. We evaluate two agentic systems, Anthropic Claude Code and OpenAI Codex, across four models: Claude Haiku 4.5, Claude Opus 4.7, GPT-5.2, and GPT-5.5. Our results show that although it is difficult to make an agent overwrite its own memory files using untrusted external content, payloads already planted in those files can successfully attack current and future sessions. Attack success and payload persistence vary substantially across systems, models, adversarial goals, and multi-session attack sequences. These findings show that persistent memory changes the threat model for prompt injection and motivate defenses that protect memory updates without removing useful agent adaptation.
Indirect prompt injection causes LLM agents to follow commands embedded in external data. A probe may distinguish instructions from data without identifying a state edit that changes the next action. We study this gap through counterfactual role probes, component-wise activation patching, and separate interventions on AgentDojo trajectories. Role decoding survives changes in content and format. In controlled Qwen tests, it precedes strong tool-choice effects from patches along an independently estimated role direction. On AgentDojo, directions estimated from hijacked and resisted training trajectories reduce attack success at pre-action and injected-span positions, but have little effect at random positions. In longer Qwen trajectories, single-position edits become less effective at later layers; span-wide and repeated edits reduce attack success on the same evaluation set. Removing the learned channel subspace preserves role decoding, yet effective intervention directions transfer poorly across the tested channels. These findings distinguish a readable role signal from an effective behavioral intervention: depth matters in controlled tool choice, while position and context also matter in attack trajectories.
Zhe Yu, Wenpeng Xing, Xingxing Yang +1
Zhejiang University · Binjiang Institute of Zhejiang University · Hong Kong Baptist University
LLMs see the world as a single stream of text, partitioned into roles like <user> or <tool>. We trace prompt injection to role confusion: models perceive the source of text from how it sounds, not its labeled role. A command hidden in a webpage hijacks an agent simply because it sounds like <user> text, despite its <tool> label. We design role probes to measure how LLMs internally perceive "who is speaking," and find that injected text occupies the same representational space as the trusted role it imitates. We demonstrate this with CoT Forgery, a zero-shot attack that injects fabricated reasoning into user prompts and tool outputs. Models mistake the forgery for their own thoughts, yielding 60% attack success against frontier models with near-zero baselines. Strikingly, the degree of role confusion predicts attack success before a single token is generated. This mechanism generalizes beyond CoT Forgery to standard agent prompt injections, revealing prompt injection as a measurable consequence of role perception. To the model, sounding like a role is indistinguishable from being one. Project page and writeup: https://role-confusion.github.io
Charles Ye, Jasmine Cui, Dylan Hadfield-Menell
*Equal contribution 1Independent · 2Massachusetts Institute of Technology, Cambridge, MA, United States