APEX: Active Protection at Execution Boundaries for LLM Agents
Organizations: Tsinghua University · Imperial College London · Nanjing University · The Chinese University of Hong Kong · University College London · A*STAR · University of British Columbia · Shenzhen University · The University of Hong Kong
Abstract
Indirect prompt injection (IPI) hides adversarial instructions in content that large language model (LLM) agents read at runtime. As agents compose heterogeneous capability units, including Tools, MCP servers, and Skills, the carriers of injection multiply, and defenses built to recognize attack patterns fall behind them. We instead shift defense from covering attack patterns to one stable point: whatever the carrier and however the injection propagates, harm materializes only at the \emph{execution boundary}, where the agent turns internal state into an external action or released output. Safety there turns on two conditions, both settled by the trusted task rather than by the run: whether the proposed effect is authorized, and whether the runtime information reaching it is endorsed by that task. We present APEX, an active defense that enforces both at this boundary from a single authorization contract compiled before untrusted execution: \emph{evidence-gated prevention} admits an effect only when the contract justifies it, while \emph{deception-based exposure} makes unendorsed use reveal itself before the effect commits. Protection therefore follows from what the task permits rather than from how an attack is built, and applies uniformly across capability units without attack-specific policies or taint tracking. Against 13 baselines, APEX attains 0% attack success on five of six benchmarks and 0.56% on the sixth, holds 0% under adaptive attacks on all three capability-unit types, and remains effective across defender backbones. Code is available at https://github.com/ZhengXR930/APEX_official/tree/official.
Figures & tables
| Method | Agentdojo | ASB-OPI | ||||
|---|---|---|---|---|---|---|
| BU | AU | ASR | BU | AU | ASR | |
| Undefended | 88.66 | 73.77 | 17.01 | 88.24 | 71.23 | 59.80 |
| Spotlighting Hines et al. (2024) | 94.85 | 86.96 | 0.16 | 92.16 | 62.06 | 51.13 |
| Tool Filter Debenedetti et al. (2024) | 21.65 | 22.42 | 0.16 | 47.06 | 50.88 | 15.29 |
| TaskShield Jia et al. (2025) | 43.30 | 49.76 | 0.32 | 78.43 | 29.36 | 14.51 |
| MELON Zhu et al. (2025) | 51.55 | 38.95 | 0.64 | 84.31 | 80.39 | 13.48 |
| Benchmark | Method | AU (Orig.) | AU (Adapt.) | ASR (Orig.) | ASR (Adapt.) | ASR |
|---|---|---|---|---|---|---|
| ASB-OPI ∗ | Undefended | 83.25 | 97.00 | 41.00 | 86.25 | +45.25 |
| MELON | 83.00 | 99.75 | 12.50 | 29.50 | +17.00 | |
| CaMeL | 76.00 | 98.75 | 10.75 | 26.25 | +15.50 | |
| APEX | 86.00 | 96.00 | 0.00 | 0.00 | 0.00 | |
| MCPTox ∗ | Undefended | 36.70 | 30.77 | 31.87 | 53.85 | +21.98 |
| MCPGuard | 52.75 | 67.25 | 0.22 | 7.91 | +7.69 |
| Mode | ASB-OPI | MCPTox | SkillInject | ||||||
|---|---|---|---|---|---|---|---|---|---|
| BU | AU | ASR | BU | AU | ASR | BU | AU | ASR | |
| APEX | 90.20 | 85.05 | 0.00 | 66.67 | 75.52 | 0.00 | 81.11 | 78.89 | 0.56 |
| APEX w/o PLANT | 88.24 | 81.47 | 0.00 | 62.46 | 66.62 | 6.45 | 72.78 | 48.33 | 3.33 |
| APEX w/o WRAP | 86.27 | 70.49 | 57.60 | 67.23 | 58.16 | 4.15 | 71.11 | 68.33 | 2.22 |
| APEX w/o Continuation | 92.16 | 62.20 | 0.00 | 60.78 | 48.15 | 1.11 | 74.44 | 77.78 | 3.89 |
| Undefended | 88.24 | 71.23 | 59.80 | 69.19 | 41.25 | 36.20 | 82.22 | 83.33 | 28.89 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Operators | Semantics |
|---|---|---|
| Arithmetic | add , multiply , percent_of | Compute sums, products, and percentages. |
| Structured data | field , keys , project , object_set | Extract fields or keys, project a field across records, or update an object field. |
| Collections | count , map_count , frequency , flatten , union , difference , sort_by | Count elements or group sizes, aggregate frequencies, combine or subtract collections, and sort records. |
| Selection | argmin , argmax , select_eq , aligned_lookup | Select by an extremal score, a unique field-value match, or an aligned key–value correspondence. |
| Predicate gating | gt , lt | Return the designated candidate only when its score passes the declared threshold. |
| Temporal | normalize_date , datetime_combine , add_duration , interval_free | Normalize or construct timestamps, add durations, and return an interval’s start only when no recorded event overlaps it. |
| Benchmark | Attack cases | Cases with deployment (%) | Plane checks | Abstained planes | Triggered / deployed probes (%) |
|---|---|---|---|---|---|
| AgentDojo | 629 | 606 (96.3%) | 3,728 | 2,595 | 93/1,155 (8.1%) |
| ASB-OPI | 2,040 | 1,849 (90.6%) | 4,363 | 1,181 | 91/5,872 (1.5%) |
| MCPTox | 1,348 | 1,256 (93.2%) | 16,065 | 14,755 | 200/1,772 (11.3%) |
| MSB | 622 | 256 (41.2%) | 11,337 | 10,911 | 181/702 (25.8%) |
| SkillInject | 180 | 180 (100.0%) | 1,362 | 1,205 | 25/397 (6.3%) |
| SCR-Bench | 667 | 250 (37.5%) | 1,663 | 1,396 | 189/295 (64.1%) |
| Model | Method | Agentdojo | ASB-OPI | MCPTox | MSB | SkillInject | SCR-bench |
|---|---|---|---|---|---|---|---|
| AU / ASR | AU / ASR | AU / ASR | AU / ASR | AU / ASR | AU / ASR | ||
| Undefended | 73.77/17.01 | 71.23/59.80 | 41.25/36.20 | 90.60/48.71 | 83.33/28.89 | 91.45/86.21 | |
| DeepSeek-V4-Flash | APEX | 80.29/0.00 | 85.05/0.00 | 75.52/0.00 | 96.39/0.00 | 78.89/0.56 | 99.85/0.00 |
| GLM-5.2 | APEX | 71.51/0.00 | 74.56/0.00 | 67.21/0.00 | 97.03/0.00 | 66.11/2.78 | 98.65/1.20 |
| GPT-5.6-sol | APEX | 82.88/0.00 | 85.55/0.00 | 80.92/0.00 | 97.12/0.00 | 88.60/0.56 | 99.85/0.00 |
| Benchmark | Method | AU (Orig.) | AU (Adapt.) | ASR (Orig.) | ASR (Adapt.) | ASR |
|---|---|---|---|---|---|---|
| MSB ∗ | Undefended | 82.08 | 30.19 | 37.74 | 46.70 | +8.96 |
| MCPGuard | 74.06 | 33.49 | 18.87 | 42.92 | +24.05 | |
| StackOne | 77.83 | 32.08 | 23.11 | 12.26 | -10.85 | |
| APEX | 82.55 | 61.32 | 0.00 | 0.00 | 0.00 | |
| SCR-Bench ∗ | Undefended | 34.00 | 14.00 | 64.00 | 83.33 | +19.33 |
| ClawGuard | 32.67 | 23.33 | 65.33 | 76.00 | +10.67 |
| Method | CapFlow | AuthBlur | TrustLift | Aggregate | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| BU | AU | ASR | BU | AU | ASR | BU | AU | ASR | BU | AU | ASR | |
| Undefended | 80.00 | 62.00 | 60.00 | 100.00 | 100.00 | 72.41 | 100.00 | 100.00 | 100.00 | 95.50 | 91.45 | 86.21 |
| ClawGuard | 100.00 | 32.67 | 65.33 | – | – | – | 6.48 | 6.48 | 6.48 | 31.94 | 13.61 | 22.50 |
| Progent | 96.00 | 37.33 | 61.33 | – | – | – | 100.00 | 100.00 | 100.00 | 98.91 | 82.94 | 89.47 |
| TaskShield | 78.67 | 50.00 | 42.00 | 100.00 | 99.14 | 70.69 | 81.05 | 78.55 | 78.55 | 83.81 | 75.71 | 68.97 |
| DynamicGuardian | 98.00 | 25.33 | 73.33 | 100.00 | 100.00 | 72.41 | 100.00 | 100.00 | 100.00 | 99.55 | 83.21 | 89.21 |
| Benchmark | Attack cases | Utility errors | Task-agent limitation | Contract error | WRAP false positive | PLANT false positive | Continuation failure |
| AgentDojo | 629 | 124 | 50 | 31 | 74 | 0 | 97 |
| ASB-OPI | 2,040 | 305 | 305 | 0 | 0 | 0 | 11 |
| MCPTox | 1,348 | 330 | 237 | 2 | 96 | 0 | 164 |
| MSB | 622 | 15 ∗ | 3 | 0 | 11 | 0 | 12 |
| SkillInject | 180 | 38 | 2 | 0 | 36 | 0 | 37 |
| SCR-Bench | 667 | 1 | 0 | 1 | 1 | 0 | 1 |