ActionGuard: Tool Call Authorization under Poisoned Skills
Organizations: Korea University Republic of Korea
Abstract
LLM-based agents extend their capabilities through third-party skills that provide task-specific instructions, scripts, and tool-use procedures. However, malicious instructions inserted into an otherwise benign skill can cause a benign user request to trigger dangerous Tool Calls, including data exfiltration, file deletion, or unauthorized code execution. This paper presents ActionGuard, which inspects skill-influenced Tool Calls immediately before execution. ActionGuard separates the target agent's action-generation context from the safeguard's authorization context. The target agent may use the original skill for planning, but the Reviewer does not receive the potentially poisoned raw skill text. Instead, it determines whether each action is justified by the trusted user request using a balanced skill profile, current and recent Tool Calls, and local script contents. ActionGuard intercepts each Tool Call at OpenClaw's before-tool-call stage and enforces the Reviewer's ALLOW or DENY decision under a fail-closed policy. We evaluate ActionGuard on 139 contextual and 180 obvious injections in a SKILL-INJECT-based setting against Dynamic Guardian and SkillGuard, using three open-source and two commercial Reviewer models. Each condition is repeated three times and evaluated using Attack Success Rate (ASR) and Task Success Rate (TSR). Overall, ActionGuard reduced ASR by 35.54 to 46.11 percent relative to existing safeguards and by 70.44 percent relative to No Safeguard, while maintaining high benign-task completion. These results show that execution-boundary authorization grounded in trusted user intent and runtime evidence can restrict unauthorized Tool Calls induced by skill injection.
Figures & tables
| Field | Description |
|---|---|
| Decision | An ALLOW or DENY verdict that indicates whether the current Tool Call may execute. |
| Reason | An explanation of the main evidence used in the decision and the basis for authorization or blocking. |
| Reason Code | A normalized category for consistent analysis of the decision rationale. |
| Type | Model | Model Identifier |
|---|---|---|
| Open Source | Gemma 4 | gemma4:e4b-it-q4_K_M |
| Qwen 3.5 | qwen3.5:9b | |
| Ministral 3 | ministral-3:8b-instruct-2512-q4_K_M | |
| Commercial | GPT-5.4 mini | gpt-5.4-mini-2026-03-17 |
| Claude Haiku | claude-haiku-4-5-20251001 |
| Gemma 4 | Qwen 3.5 | Ministral 3 | GPT-5.4 mini | Claude Haiku | ||||||
| Defense | ASR | TSR | ASR | TSR | ASR | TSR | ASR | TSR | ASR | TSR |
| Contextual | ||||||||||
| Dynamic Guardian ( Fujinuma et al., 2026 ) | 14.39% | 86.67% | 12.71% | 88.67% | 7.19% | 87.67% | 20.62% | 89.00% | 19.66% | 89.33% |
| SkillGuard ( Pan et al., 2026 ) | 22.06% | 90.00% | 19.18% | 89.67% | 18.71% | 88.33% | 18.94% | 89.00% | 20.14% | 88.67% |
| ActionGuard | 8.87% | 87.33% | 8.63% | 86.00% | 8.87% | 90.67% | 7.67% | 89.67% | 6.95% | 85.00% |
| Obvious | ||||||||||
| Method | ASR (%) | TSR (%) |
|---|---|---|
| Dynamic Guardian ( Fujinuma et al., 2026 ) | 16.05 | 92.08 |
| SkillGuard ( Pan et al., 2026 ) | 13.42 | 88.55 |
| ActionGuard | 8.65 | 90.38 |
| No Safeguard | 29.05 | 93.46 |
| Item | Configuration |
|---|---|
| Host operating system | Windows 11 Pro, version 10.0.26200 (build 26200) |
| Processor and memory | Intel Core i7-13700F; 16 physical cores, 24 logical processors, and 63.8 GiB host memory |
| GPU | NVIDIA GeForce RTX 4070 Ti SUPER; 16,376 MiB VRAM; driver 591.86 |
| Container engine | Docker Desktop / Docker Engine 27.3.1 |
| Task image | instruct-bench-agent , derived from python:3.11-slim |
| In-container environment | Debian GNU/Linux 13, Python 3.11.16, Node.js 24.19.0, OpenClaw 2026.5.28 (commit e932160 ), and Codex CLI 0.147.0 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Setting |
|---|---|
| Interception point | OpenClaw before_tool_call hook |
| Profile-database scope | Empty at task start; reused across subsequent Tool Calls in the same sandbox |
| Skill-package snapshot | At most 32 text files and 100,000 source characters |
| Path handling | Symbolic links skipped; resolved paths restricted to the configured Skill root |
| Package identity | SHA-256 digest over collected relative paths and source content |
| Maximum profile entries | 40 entries per profile field |