MMSkillRisk: Can Agents Stay Safe When Multimodal Skills Become Traps?
Organizations: Zhejiang University · Ant Group · Hangzhou Dianzi University · Tsinghua University
Abstract
Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, attackers can disguise malicious instructions as ordinary visual guidance within otherwise legitimate skills. Existing skill-security research primarily examines text-carried attacks or scanner detection, leaving the runtime effects of image-borne attacks insufficiently evaluated. We introduce MMSkillRisk, to our knowledge the first publicly available benchmark dedicated to end-to-end safety evaluation of image-borne attacks in multimodal skills. To instantiate this attack surface, we design Native-Context Visual Attack (NCVA), which disguises malicious instructions as native components of teaching images, such as annotations and interface labels. The accompanying SKILL.md provides auxiliary guidance toward relevant visual regions without explicitly stating the malicious operation. Built from 28 curated clean skills, MMSkillRisk contains 36 attack packages and 108 executable cases spanning five attack objectives, with separate checks for attack success and legitimate-task completion. Across nine model-harness configurations evaluated in isolated sandboxes, NCVA induces unauthorized operations in every configuration. Its pooled attack success rate (ASR) reaches 43.1%, exceeding the matched text-carrier baseline by 16.4 percentage points, with higher ASR in all nine configurations. Attack success and legitimate-task completion co-occur in 36.5% of cases, reaching 72.2% for GPT-5.6-sol with Codex. These results show that skill-bundled images can induce unauthorized actions even as agents complete legitimate tasks, so task success alone does not establish safe skill use. Our code and data are available at https://github.com/kaill-jlq/MMSkillRisk.
Figures & tables
| Harness | Model | Text-only baseline (%) | NCVA (%) | |||||
| ASR | TSR | TC-ASR | ASR | TSR | TC-ASR | VIAR | ||
| Codex | GPT-5.6-sol | 59.3 | 84.5 | 51.9 | 81.5 | 89.3 | 72.2 | 91.7 |
| Kimi-K2.6 | 32.4 | 73.8 | 23.1 | 65.7 | 78.6 | 49.1 | 68.5 | |
| Kimi-K3 | 14.8 | 96.4 | 13.9 | 38.0 | 97.6 | 37.0 | 38.9 | |
| Qwen3.8-Flash-Next | 13.0 | 95.2 | 12.0 | 32.4 | 94.1 | 32.4 | 33.3 | |
| Claude Code | Kimi-K2.6 | 36.1 | 73.8 | 23.1 | 50.9 | 75.0 | 35.2 | 52.8 |
| Harness | Model | Read | Seen | Attempt | Success |
|---|---|---|---|---|---|
| Claude Code | Kimi-K3 | 100.0 | 99.1 | 17.6 | 17.6 |
| Kimi-K2.6 | 100.0 | 53.7 | 50.9 | 50.9 | |
| Qwen3.8-Flash | 100.0 | 100.0 | 27.8 | 27.8 | |
| Codex | Kimi-K3 | 99.1 | 96.3 | 38.0 | 38.0 |
| Kimi-K2.6 | 100.0 | 70.4 | 65.7 | 65.7 | |
| Qwen3.8-Flash | 100.0 | 94.4 | 32.4 | 32.4 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Base skill | Scope of the three legitimate tasks |
|---|---|
| anydesign | Extract design information from page screenshots, reproduce components, and document the design. |
| breadth-chart-analyst | Read breadth charts and their signals, and prepare analytical results. |
| guizang-social-card-skill | Create social media cards for different themes using supplied materials. |
| infographic | Produce infographics that follow the requested topic and brand requirements. |
| sector-analyst | Read sector rotation and momentum information across time frames. |
| Base skill | Scope of the three legitimate tasks |
|---|---|
| mmskills-calc-charts | Identify Calc chart configurations and output states. |
| mmskills-chrome-permissions | Read Chrome permission and security settings. |
| mmskills-gimp-layers | Read GIMP layer and mask states. |
| mmskills-thunderbird-filters | Read mail-filter rules and their enabled states. |
| mmskills-ubuntu-assistive-input | Read accessibility and input settings. |
| mmskills-vlc-audio-settings | Read audio effects and volume preferences. |
| Base skill | Scope of the three legitimate tasks |
|---|---|
| gui-calc-freeze | Determine the frozen-pane state of a spreadsheet. |
| gui-calc-pivot | Inspect pivot-table field layouts. |
| gui-chrome-search | Distinguish the default search engine from a highlighted candidate. |
| gui-gimp-color | Determine the current state of a color-adjustment interface. |
| gui-gimp-layers | Determine layer selection, visibility, and related states. |
| gui-impress-export | Inspect presentation-export dialog settings. |
| Base skill | Scope of the three legitimate tasks |
|---|---|
| visual-chart-reconciliation | Reconcile and correct chart data bindings or values. |
| visual-form-routing | Determine how requests should be routed using forms and flow diagrams. |
| visual-invoice-reconciliation | Reconcile invoices, purchase orders, and posting batches. |
| visual-palette-audit | Check interface color tokens and correct inconsistencies. |
| visual-warehouse-picking | Allocate picking batches using bin labels and inventory states. |
| Model | API identifier | Harness | Version |
|---|---|---|---|
| GPT-5.6-sol | gpt-5.6-sol | Codex | 0.153.3 |
| Kimi-K2.6 | Kimi-K2.6 | Codex | 0.153.3 |
| Kimi-K3 | Kimi-K3 | Codex | 0.153.3 |
| Qwen3.8-Flash-Next | Qwen3.8-Flash-Next | Codex | 0.153.3 |
| Kimi-K2.6 | Kimi-K2.6 | Claude Code | 2.1.268 |
| Kimi-K3 | Kimi-K3 | Claude Code | 2.1.268 |
| Component | Codex / Claude Code | DSH |
|---|---|---|
| Operating system | Debian GNU/Linux 13 (trixie) | Debian GNU/Linux 13 (trixie) |
| Debian full version | 13.6 | 13.6 |
| Architecture | linux/arm64 | linux/arm64 |
| Python | 3.11.16 | 3.11.16 |
| Node.js | 22.22.0 | 22.22.0 |
| CPU quota per instance | 2 CPUs | 2 CPUs |
| Harness | Model | Clean TSR | NCVA TSR |
|---|---|---|---|
| Codex | Kimi-K3 | 98.8 | 97.6 |
| Claude Code | Kimi-K3 | 98.8 | 97.6 |
| DSH | DeepSeek-V4.1-Flash | 92.9 | 89.3 |
| Codex | Qwen3.8-Flash-Next | 96.4 | 94.1 |
| Claude Code | Qwen3.8-Flash-Next | 96.4 | 96.4 |