cs.CRSep 28, 2026

MMSkillRisk: Can Agents Stay Safe When Multimodal Skills Become Traps?

Authors: Lingqi Jiang, Jialuo Chen, Jianan Ma, Xinhao Deng, Xiaohu Du, Sibo Yi, Yuqi Qing, Zhenguang Liu, +3 more

Organizations: Zhejiang University · Ant Group · Hangzhou Dianzi University · Tsinghua University

Abstract

Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, attackers can disguise malicious instructions as ordinary visual guidance within otherwise legitimate skills. Existing skill-security research primarily examines text-carried attacks or scanner detection, leaving the runtime effects of image-borne attacks insufficiently evaluated. We introduce MMSkillRisk, to our knowledge the first publicly available benchmark dedicated to end-to-end safety evaluation of image-borne attacks in multimodal skills. To instantiate this attack surface, we design Native-Context Visual Attack (NCVA), which disguises malicious instructions as native components of teaching images, such as annotations and interface labels. The accompanying SKILL.md provides auxiliary guidance toward relevant visual regions without explicitly stating the malicious operation. Built from 28 curated clean skills, MMSkillRisk contains 36 attack packages and 108 executable cases spanning five attack objectives, with separate checks for attack success and legitimate-task completion. Across nine model-harness configurations evaluated in isolated sandboxes, NCVA induces unauthorized operations in every configuration. Its pooled attack success rate (ASR) reaches 43.1%, exceeding the matched text-carrier baseline by 16.4 percentage points, with higher ASR in all nine configurations. Attack success and legitimate-task completion co-occur in 36.5% of cases, reaching 72.2% for GPT-5.6-sol with Codex. These results show that skill-bundled images can induce unauthorized actions even as agents complete legitimate tasks, so task success alone does not establish safe skill use. Our code and data are available at https://github.com/kaill-jlq/MMSkillRisk.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

    May 12, 2026Chang Jin, An Wang, Zeming Wei +7Large Language Model AgentsAdversarial Robustness

  2. Seeing Is Not Screening: Multimodal Hidden Instruction Attacks on Agent Skill Scanners

    Jun 16, 2026Xiaojun Jia, Jie Liao, Simeng Qin +5White-Box Spectral-Subspace-Guided AttackMulti-Llm Agents

  3. SkillHarm: Lifecycle-Aware Skill-Based Attacks via Automated Construction

    Jun 1, 2026Yuting Ning, Zhehao Zhang, Yash Kumar Lal +8Malicious AgentsAgentic Workflows