cs.CRSep 28, 2026

MMSkillRisk: Can Agents Stay Safe When Multimodal Skills Become Traps?

Authors: Lingqi Jiang, Jialuo Chen, Jianan Ma, Xinhao Deng, Xiaohu Du, Sibo Yi, Yuqi Qing, Zhenguang Liu, +3 more

Organizations: Zhejiang University · Ant Group · Hangzhou Dianzi University · Tsinghua University

Abstract

Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, attackers can disguise malicious instructions as ordinary visual guidance within otherwise legitimate skills. Existing skill-security research primarily examines text-carried attacks or scanner detection, leaving the runtime effects of image-borne attacks insufficiently evaluated. We introduce MMSkillRisk, to our knowledge the first publicly available benchmark dedicated to end-to-end safety evaluation of image-borne attacks in multimodal skills. To instantiate this attack surface, we design Native-Context Visual Attack (NCVA), which disguises malicious instructions as native components of teaching images, such as annotations and interface labels. The accompanying SKILL.md provides auxiliary guidance toward relevant visual regions without explicitly stating the malicious operation. Built from 28 curated clean skills, MMSkillRisk contains 36 attack packages and 108 executable cases spanning five attack objectives, with separate checks for attack success and legitimate-task completion. Across nine model-harness configurations evaluated in isolated sandboxes, NCVA induces unauthorized operations in every configuration. Its pooled attack success rate (ASR) reaches 43.1%, exceeding the matched text-carrier baseline by 16.4 percentage points, with higher ASR in all nine configurations. Attack success and legitimate-task completion co-occur in 36.5% of cases, reaching 72.2% for GPT-5.6-sol with Codex. These results show that skill-bundled images can induce unauthorized actions even as agents complete legitimate tasks, so task success alone does not establish safe skill use. Our code and data are available at https://github.com/kaill-jlq/MMSkillRisk.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

May 12, 2026cs.CR

SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools, memory, and execution environments. However, this modularity introduces attack surfaces that are largely missed by existing safety evaluations: even when the user request is benign, unsafe influence may reside in skill guidance, local artifacts, or execution-environment files that steer the agent toward unsafe actions. We present SkillSafetyBench, a runnable benchmark for evaluating such skill-mediated safety failures. SkillSafetyBench includes 155 adversarial cases across 47 tasks, 6 risk domains, and 30 safety categories, each evaluated with a case-specific rule-based verifier. Experiments with multiple CLI agents and model backends show that non-user attacks can consistently induce unsafe behavior, with distinct failure patterns across domains, attack methods, and scaffold-model pairings. Our findings suggest that agent safety depends not only on model-level alignment, but also on how agents interpret skills, trust workflow context, and act through executable environments.
Jun 16, 2026cs.CR

Seeing Is Not Screening: Multimodal Hidden Instruction Attacks on Agent Skill Scanners

Agent skills are emerging as an important attack surface in LLM-based systems. Through an empirical study of existing skill scanners, we find that current defenses primarily rely on textual descriptions, manifests, and source code as the main signals for security analysis, which can leave visually conveyed malicious intent insufficiently examined. This creates a practical blind spot: harmful operational instructions hidden in images may bypass scanning while still being recoverable by multimodal agents during deployment. To systematically investigate this threat, we propose SkillCamo, a document-mediated multimodal instruction attack that conceals malicious instructions within images bundled with a skill while rewriting the surrounding documentation to naturally reference those images as part of the normal workflow. Thus, the attack does not rely on the image alone, but on the joint interpretation of textual guidance and visual payload at execution time. To defend against such attacks, we further propose ExecScan, an execution-grounded multimodal scanning module that performs intent extraction, behavior reconstruction, abuse assessment, and deliberative execution simulation over skill artifacts. ExecScan jointly analyzes documentation, code, referenced resources, and visual content to recover hidden instructions, reconstruct executable behavior chains, and identify downstream risks such as exfiltration, destruction, persistence, deception, and privilege escalation. Extensive experiments show that image-hidden malicious instructions challenge existing skill scanners, while ExecScan can improve the skill scanning performance.
Jun 1, 2026cs.CL

SkillHarm: Lifecycle-Aware Skill-Based Attacks via Automated Construction

Agent skills occupy a privileged position in the agent workflow, as agents are expected to implicitly follow and execute them, rendering third-party skills a vulnerable attack surface. Existing studies have revealed unsafe agent behaviors induced by skill-based attacks, but they primarily evaluate poisoned skills within a single task execution and enumerate harms through ad-hoc risk lists. To bridge these gaps, we introduce SkillHarm, a benchmark of skill-based attacks across the skill-use lifecycle, paired with a systematic taxonomy of skill-relevant risks. SkillHarm evaluates two attack scenarios: Fixed-Payload Poisoning (FPP), where a fixed poisoned skill package directly compromises any task session that invokes it, and Self-Mutating Poisoning (SMP), where an initially benign execution silently mutates persistent skill content, deferring harm until a subsequent reuse. It further defines 12 risk types based on the agent workflow component targeted by the harm: data pipelines, system environments, and agent autonomy. To instantiate these attacks at scale, we build AutoSkillHarm, an automated construction pipeline with coding agents driven by natural-language harnesses. The resulting benchmark contains 879 attack samples across 71 skills. Experiments show that current agents remain vulnerable with attack success rates up to 86.3% in FPP and 69.3% in SMP. Our analysis further reveals a latent risk: many apparent attack failures stem from the agent failing to engage with the poisoned file rather than genuine resistance, and current defenses still fail to reliably mitigate the threat.