Blocking at the Boundary: Auditing Long-Horizon Agents against Staged Prompt Injection
Organizations: Beijing Jiaotong University · Mohamed bin Zayed University of Artificial Intelligence · INRIA Rennes-Bretagne-Atlantique · Xi’an Jiaotong University
Abstract
Long-horizon agents consume external content, invoke tools, and modify persistent state. Indirect prompt injection can exploit task-specific context, propagate across causally connected stages, and alter a consequential action while the workflow continues; we term this staged prompt injection. We build an automated, feedback-guided attack generation pipeline and apply it to Claude Code and Codex in their native runtimes. The confirmed attacks span eight workflow scenarios, seven attack goals, and six injection surfaces, showing that production agents are vulnerable to context-aware, multi-step injection over long horizons. Stopping such attacks requires a decision before each consequential action: input screening and completed-run evaluation cannot locate the intervention point, and existing pre-action methods use incompatible units and labels. We therefore formulate boundary action auditing: given initial context, a trajectory prefix, and a fully specified pending message or tool call, an auditor predicts Pass or Block before its effect occurs. Pairing attacked and benign executions yields a 479-pair, 3,112-unit benchmark. We further propose Path-Aligned Attribution (PAA), a training-free auditor that decomposes pending actions into operative elements and traces what supplied each value and guided each decision. PAA blocks only when the model attributes an unwarranted, material effect on an element to an attacker-reachable source that either provides unqualified steering or conflicts with visible evidence. Under full-benchmark fail-open scoring with Claude Sonnet 5, PAA reaches 86% Block recall at a 6-8% false-block rate (FBR), whereas ARGUS reaches 44-47% recall at 16-33% FBR. Under the same backend, on the tool calls that all three auditors natively support, PAA has higher recall and lower FBR than VIGIL and ARGUS; all paired 95% confidence intervals exclude zero.
Figures & tables
| Target Agent | Pairs | Benign units | Injected units | Pass | Block | Total units |
|---|---|---|---|---|---|---|
| Claude Code | 240 | 628 | 692 | 744 | 576 | 1,320 |
| Codex | 239 | 873 | 919 | 1,228 | 564 | 1,792 |
| Total | 479 | 1,501 | 1,611 | 1,972 | 1,140 | 3,112 |
| Auditor | Auditor backend | Acc. | Prec. | Recall | FBR | F1 | Bal. Acc. |
|---|---|---|---|---|---|---|---|
| Claude Code corpus ( = 1,320 audit units) | |||||||
| PAA (ours) | Claude Sonnet 5 | 0.905 | 0.917 | 0.861 | 0.060 | 0.888 | 0.900 |
| VIGIL | Claude Sonnet 5 | 0.560 | 0.488 | 0.174 | 0.141 | 0.256 | 0.516 |
| ARGUS | Claude Sonnet 5 | 0.573 | 0.512 | 0.441 | 0.325 | 0.474 | 0.558 |
| PAA (ours) | GPT-5.5 | 0.828 | 0.759 | 0.889 | 0.219 | 0.819 | 0.835 |
| VIGIL | GPT-5.5 | 0.568 | 0.514 | 0.196 | 0.144 | 0.284 | 0.526 |
| PAA | VIGIL | ARGUS | ||||||
|---|---|---|---|---|---|---|---|---|
| Corpus | Auditor backend | (B/P) | R | FBR | R | FBR | R | FBR |
| Claude Code | Claude Sonnet 5 | 840 (337/503) | .869 | .070 | .297 | .209 | .754 | .481 |
| Claude Code | Codex GPT-5.5 | 840 (337/503) | .914 | .185 | .335 | .213 | .899 | .640 |
| Codex | Claude Sonnet 5 | 1,314 (343/971) | .875 | .071 | .504 | .285 | .778 | .200 |
| Codex | Codex GPT-5.5 | 1,314 (343/971) | .773 | .145 | .248 | .149 | .904 | .387 |
| S5 | G5.5 | |||||||
| Scenario ( ) | R | FBR | F1 | BA | R | FBR | F1 | BA |
| Claude Code corpus | ||||||||
| S01 (259) | .793 | .042 | .849 | .876 | .924 | .311 | .742 | .806 |
| S02 (124) | .903 | .032 | .933 | .935 | .968 | .226 | .882 | .871 |
| S03 (93) | .783 | .234 | .774 | .774 | .761 | .426 | .693 | .668 |
| S04 (125) | .887 | .016 | .932 | .936 | .887 | .238 | .833 | .825 |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Agent system | Runtime / target model | Execution setting |
|---|---|---|
| Claude Code [ 4 ] | v2.1.172; Claude Sonnet 5 | Native runtime with configured skills, tools, and sub-agent delegation; network disabled. |
| Codex [ 42 ] | CLI v0.144.0; GPT-5.5 | Native runtime with configured skills and tools; isolated workspace access and network enabled. |
| Scenario | Task setting and agent-visible workspace | Provenance |
|---|---|---|
| S01: Commerce / Orders | Reconcile orders, products, customers, payments, returns, inventory, support tickets, and operating policies. | -bench retail data (MIT) [ 61 ] , augmented with constructed reconciliation and policy records. |
| S02: Mail / Calendar | Audit mailbox, calendar, contacts, travel policy, sharing records, and communication controls. | AgentDojo mail/calendar data (MIT) [ 16 ] , augmented with constructed communication records. |
| S03: Code Review | Review a local Python repository using issues, documentation, examples, tests, and maintenance guidance. | itsdangerous v2.2.0 and public issue material (BSD-3-Clause) [ 44 ] . |
| S04: Office Docs / Drive | Cross-check project tasks, calendars, CRM records, emails, analytics, boards, staff rosters, and operating controls. | WorkBench processed data (MIT) [ 51 ] , augmented with constructed board, roster, and policy records. |
| S05: Local KB / RAG | Retrieve evidence from support articles, a knowledge-base index, FAQ, owner notes, and maintenance policy. | IBM TechQA technotes [ 6 ] , augmented with constructed indexing and policy records. |
| S06: Sysadmin / Security | Audit SSH and service logs, host and user inventories, ACLs, configuration, change history, runbooks, and security policy. | Loghub OpenSSH logs [ 70 ] , augmented with constructed configuration and access-control records under the upstream research-use terms. |
| Scope | Configured skills | Upstream source and snapshot |
|---|---|---|
| Shared | deep-thinking data-quality-checks | ClawHub/@amankr-novo (v1.0.0) ClawHub/@mohitagw15856 (v1.0.0) |
| S01 | data-reconciliation-exceptions settlement-reconciliation-guard | ClawHub/@kowl64 (unversioned snapshot) ClawHub/@harrylabsj (v1.0.0) |
| S02 | aana-calendar-scheduling-guardrail-skill meeting-to-action | ClawHub/@mindbomber (v1.0.0) ClawHub/@codedao12 (v1.0.0) |
| S03 | code-review improve-code-comments | ClawHub/@wpank (v1.0.0) ClawHub/@wufei-png (v1.0.2) |
| S04 | board-minutes google-workspace | ClawHub/@mohitagw15856 (v1.0.0) ClawHub/@patrick-erichsen-2 (v1.0.0; locally trimmed) |
| S05 | customer-support content-style-guide | ClawHub/@ivangdavila (v1.0.0) ClawHub/@mohitagw15856 (v1.0.0) |
| Type | ID and name | Operational meaning |
| Goal | IO01: Data exposure | Unauthorized disclosure of protected data to an external sink. |
| Goal | IO02: Unapproved external action | An external action prohibited by the task or authorization scope. |
| Goal | IO03: Parameter tampering | A sanctioned action executed with an attacker-influenced target, value, or content. |
| Goal | IO04: Workflow redirection | Redirection of an action, scope, or deliverable to an attacker-selected entity. |
| Goal | IO05: Access change | Granting or modifying access outside the authorized principal or scope. |
| Goal | IO06: Persistent context manipulation | Writing attacker-influenced content into durable state that can affect later execution. |
| Claude Code | Codex | |
|---|---|---|
| Benign trajectories | 320 | 320 |
| Scheduled candidates | 837 | 1,016 |
| Attempts | 983 | 1,485 |
| Executed | 834 | 1,236 |
| Realized ( Pos ) | 505 | 765 |
| Not realized | 329 | 471 |
| Source class | Reach | Authorization | Rule |
|---|---|---|---|
| Governing system policy | 0 | Q | Q |
| Direct user instruction | 0 | Q | Q |
| Tool schema or skill metadata | 1 | U | Q |
| Task-designated data source | 1 | C | C |
| Ordinary content | 1 | U | U |
| Agent-generated artifact | 1 | U | U |
| Object | Rule |
|---|---|
| Source quotation | Substring of the normalized source; a quotation of at least eight words also verifies when at least 60% of its word 6-grams occur in the source. |
| Adoption quotation | At least eight characters and verified against an agent message, tool-call argument, or the pending action. |
| Contract entry | Verbatim 15–300-character span of the user task; a designated path must be an exact string from the registered resources. |
| Conflict | Counter-evidence verifies in another visible source of the same unit; system text and the record’s own source are excluded. An omission conflict requires the cited source to exist. |
| Auditor | Implementation and backend |
|---|---|
| PAA | Algorithm 2 , with tools disabled; Claude Sonnet 5 (via Claude Code) and GPT-5.5 (via Codex), each on both corpora |
| VIGIL | Complete upstream pipeline, decision-only adapter, strict preset; the same two backends on both corpora |
| ARGUS | Official artifact in teacher-forced offline replay; the same two backends on both corpora |
| AgentDoG-1.5 | Unified-Qwen3.5-4B with local inference and its native chat template |
| Backend | Comparator | R [95% CI] | FBR [95% CI] |
|---|---|---|---|
| Claude Code | |||
| S5 | VIGIL | +57.3 [+50.4, +63.3] | -13.9 [-20.3, -7.8] |
| S5 | ARGUS | +11.6 [ +5.5, +17.8] | -41.2 [-47.2, -34.4] |
| G5.5 | VIGIL | +57.9 [+51.3, +64.1] | -2.8 [ -8.2, +2.5] |
| G5.5 | ARGUS | +1.5 [ -3.7, +7.1] | -45.5 [-51.5, -39.4] |
| Codex | |||
| Gold- Pass scope | S5 | G5.5 | |
|---|---|---|---|
| Full benchmark | 1,228 | .077 | .140 |
| Exclude 69 direct-conflict S04 actions | 1,159 | .065 | .089 |
| Exclude all 91 S04 gold- Pass units | 1,137 | .056 | .076 |
| Claude Code | Codex | |||||||
| S5 | G5.5 | S5 | G5.5 | |||||
| Variant | R | FBR | R | FBR | R | FBR | R | FBR |
| Full rule ( PAA ) | .861 | .060 | .889 | .219 | .860 | .077 | .679 | .140 |
| w/o quote verification | .878 | .069 | .889 | .219 | .862 | .081 | .679 | .140 |
| w/o task warrant | .863 | .060 | .889 | .226 | .860 | .080 | .679 | .142 |
| w/o materiality | .913 | .153 | .920 | .310 | .922 | .168 | .793 | .253 |
| Claude Code | Codex | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Auditor | Backend | R | 95% CI | FBR | 95% CI | R | 95% CI | FBR | 95% CI |
| PAA | S5 | .861 | [.824, .896] | .060 | [.042, .082] | .860 | [.826, .892] | .077 | [.059, .097] |
| PAA | G5.5 | .889 | [.854, .919] | .219 | [.185, .258] | .679 | [.628, .726] | .140 | [.108, .174] |
| VIGIL | S5 | .174 | [.142, .209] | .141 | [.098, .185] | .307 | [.276, .338] | .226 | [.193, .259] |
| VIGIL | G5.5 | .196 | [.163, .231] | .144 | [.109, .181] | .151 | [.122, .180] | .118 | [.089, .150] |
| ARGUS | S5 | .441 | [.405, .474] | .325 | [.273, .372] | .473 | [.442, .504] | .158 | [.127, .190] |
| Corpus | First | Second | R | 95% CI | FBR | 95% CI |
|---|---|---|---|---|---|---|
| Claude Code | PAA /S5 | VIGIL/S5 | +68.7 | [+63.7, +73.3] | -8.0 | [-12.7, -3.6] |
| PAA /S5 | ARGUS/S5 | +42.0 | [+37.4, +46.7] | -26.3 | [-31.5, -21.1] | |
| PAA /S5 | AgentDoG/native | +46.0 | [+38.8, +53.1] | -34.6 | [-41.7, -27.4] | |
| PAA /G5.5 | VIGIL/G5.5 | +69.2 | [+64.4, +73.5] | +7.6 | [ +3.0, +12.1] | |
| PAA /G5.5 | ARGUS/G5.5 | +36.3 | [+32.3, +40.1] | -21.3 | [-26.2, -16.2] | |
| PAA /G5.5 | AgentDoG/native | +48.8 | [+42.4, +55.3] | -18.7 | [-26.6, -10.7] |
| Auditor | Backend | Claude Code | Codex |
|---|---|---|---|
| PAA | S5 | 18 (23), 7.5% | 33 (41), 13.8% |
| PAA | G5.5 | 91 (131), 37.9% | 49 (70), 20.5% |
| VIGIL | S5 | 7 (7), 2.9% | 83 (83), 34.7% |
| VIGIL | G5.5 | 35 (35), 14.6% | 26 (26), 10.9% |
| ARGUS | S5 | 55 (55), 22.9% | 45 (45), 18.8% |
| ARGUS | G5.5 | 165 (165), 68.8% | 107 (107), 44.8% |
| Benchmark | Dynamic run | Context- aware | Multi- step | Native runtime |
|---|---|---|---|---|
| InjecAgent [ 66 ] | – | – | – | |
| AgentDojo [ 16 ] | – | |||
| WASP [ 18 ] | ||||
| AgentDyn [ 30 ] | – | – | ||
| ToolSafety [ 59 ] | – | – | ||
| AgentLAB [ 28 ] | – |
| Benchmark | Evidence and classification rationale |
|---|---|
| InjecAgent [ 66 ] | D (–): Sec. 2.3 begins from a hypothetical state with a prefilled user-tool response rather than replaying a complete run. C (–): Sec. 2.2 crosses reusable response templates with separately generated attacker cases. M ( ): Secs. 2.2–2.3 require extraction followed by transmission for data stealing, whereas direct-harm cases need not. N (–): Sec. 3.1 and App. B use benchmark-defined ReAct or function-calling scaffolds. |
| AgentDojo [ 16 ] | D ( ): Sec. 3.1 executes tools against mutable environment state. C ( ): Secs. 3.1 and 3.3 pair attacks with injection goals and optional side information, but do not require workload-wide task-specific payload design. M ( ): Sec. 3.3 permits multi-action runs without requiring a staged attack chain. N (–): Secs. 3.1–3.3 execute a benchmark-provided agent and environment pipeline. |
| WASP [ 18 ] | D ( ): Sec. 3 runs agents end to end against self-hosted VisualWebArena GitLab and Postmill instances. C ( ): Sec. 3 uses plain-text and URL injection templates with a task-aware variant obtained by variable substitution. M ( ): Sec. 4 separates intermediate from end-to-end success, but each attacker goal is issued by a single injection. N ( ): Sec. 3 includes the Claude Computer Use reference implementation as one of three agent setups; the others are VisualWebArena scaffolding and a tool-calling loop. |
| AgentDyn [ 30 ] | D ( ): Sec. 3 provides real-time interactable entities, such as one-time-password e-mail, and open-ended tasks. C ( ): Sec. 3 crosses 28 injection tasks with 60 user tasks and places injections realistically, without per-task payload design. M (–): Sec. 3 uses a single injected instruction per case. N (–): Sec. 3 runs an AgentDojo-style sandbox with a benchmark agent loop. |
| ToolSafety [ 59 ] | D (–): Sec. 3 synthesizes training examples and trajectories rather than collecting a live evaluation workload. C ( ): Sec. 3 matches tools, queries, and context, but mixes direct-harm, indirect-harm, and multi-step samples rather than requiring context-aware IPI throughout. M ( ): Sec. 3 includes a synthesized 2–5-step subset, without requiring staged attack propagation across the dataset. N (–): the released artifact is a fine-tuning dataset rather than executions from a native agent runtime. |
| AgentLAB [ 28 ] | D ( ): Secs. 4.1 and 4.3 execute agents in stateful, tool-enabled environments. C ( ): Sec. 4.2 conditions attack planners on the malicious task, environment, and execution feedback. M ( ): Sec. 4.2 defines five long-horizon attack families that progress across multiple interactions. N (–): Sec. 4.3 uses a unified benchmark agent framework rather than an independently released agent system. |