Decision Hijacking: Prompt Injection Attacks on Jev's Typed Probabilistic Decisions
Organizations: Nanyang Technological University Singapore
Abstract
Most studies of prompt injection focus on generative agents, leaving their effects on models with schema-defined outputs unclear. We examine these effects in Jev, a non-generative decision model, using 510 reconstructed InjecAgent cases. Malicious content shifts action probabilities but rarely causes Jev to select the attacker's target. Override markers reduce this influence, while claims of contextual relatedness have small effects. Adaptive attacks using score feedback double the mean highest attacker-target probability found during optimization, while success on fresh validation calls rises from 1.8% to 3.5%. Exploratory analysis links these successes to small initial decision margins or greater attacker control over the observation. Together, these findings show that schema-defined outputs change but do not eliminate prompt-injection risk, highlighting the need to evaluate how untrusted content influences choices within the allowed action set.
Figures & tables
| Stage | Question | Conditions | Jev calls |
| DH-1 | Transfer of original attacks | Clean, base, and enhanced in both spaces | 3060 |
| DH-2 | Effect of injection-style markers | Clean and A0 to A4, | 15300 |
| DH-3 | Effect of asserted relatedness | Clean, P0, N0, P1, and P2, | 12750 |
| DH-4 | Effect of adaptive score access | Initial attack and 24 proposals; four checkpoints with five validation calls each | 22950 |
| Stage | Comparison or endpoint | Estimate [95% CI] |
| DH-1 | Upstream base | ASR 1.8% (9/510); |
| DH-1 | Upstream enhanced | ASR 0%; |
| DH-1 | Enhanced base, repeated-call rerun | |
| DH-2 | A1 A0, importance marker | |
| DH-2 | A2 A0, ignore-previous marker | |
| DH-2 | A4 A0, natural wording |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Space | Condition | Targeted ASR | Flip rate | CHR | ||
| Upstream | Clean | 0.0% (0/510) | n/a | 0.0% | ||
| Three choices | Base | 1.8% (9/510) | 3.3% | 0.0% | ||
| Enhanced | 0.0% (0/510) | 0.0% | 0.0% | |||
| Library | Clean | 0.0% (0/510) | n/a | 0.0% | ||
| Eight choices | Base | 1.6% (8/510) | 3.5% | 0.0% | ||
| Enhanced | 0.0% (0/510) | 0.0% | 0.0% |
| Condition | Wording | ASR | Target calls | Robust cases | ||
| A0 | Plain request | 2.0% | 51/2,550 | 11/510 | ||
| A1 | Importance marker | 2.2% | 56/2,550 | 11/510 | ||
| A2 | Ignore-previous marker | 0.0% | 0/2,550 | 0/510 | ||
| A3 | Both markers | 0.0% | 0/2,550 | 0/510 | ||
| A4 | Natural wording | 1.3% | 34/2,550 | 7/510 |
| Hypothesis | Contrast | Estimate [95% CI] | Interpretation |
| H3a | P1 N0 | Opposes predicted direction | |
| H3b | P2 N0 | Positive, below SESOI | |
| H3c | P2 P1 | Positive, small descriptive effect |
| Screening ASR | Validated ASR [95% CI] | Successes | Mean best-so-far | |
| 0 | 2.2% | 9/510 | 0.0429 | |
| 8 | 3.5% | 17/510 | 0.0774 | |
| 16 | 3.9% | 18/510 | 0.0828 | |
| 24 | 3.9% | 18/510 | 0.0842 |
| Initial margin | Embedded | Bare snippet |
| 12/12 (100.0%) | 1/1 (100.0%) | |
| 1/468 (0.2%) | 4/29 (13.8%) |
| Stage | Expected | Completed | Failures | Identity and replay checks | Tests |
| DH-3 | 12,750 | 12,750 | 0 | P0 request identity 510/510; design hash recorded | 5 |
| DH-4 | 22,950 | 22,950 | 0 | Validation identity 10,200/10,200; policy replay 510/510 | 5 |