cs.CRSep 30, 2026

Refusals That Bend: Measuring and Predicting Task Malleability in Embodied VLM Planners

Authors: Leo Y. Lin, Mikhail Kuznetsov, Muslum Ozgur Ozmen, Z. Berkay Celik

Organizations: Purdue University · Amazon · Arizona State University

Abstract

Embodied vision-language models (VLMs) are increasingly deployed as high-level planners for robots because they generalize across diverse environments. However, this requires their safety alignment to also hold in unseen environments. Existing red-teaming assumes an adversary who optimizes the prompt, the pixels, or text in the environment, and existing benchmarks ask whether a planner recognizes or mitigates a hazard in a fixed scene. Neither asks whether a refusal the planner has already given survives an ordinary change to the environment. We ask that question by placing a single everyday object into the environment, with no pixel, gradient, or prompt under adversarial control. On 846846 tasks that a constitution-guarded planner initially refuses, we find 20.2%20.2\% of tasks can be flipped to compliance by one or more objects, and the number of objects differs from one task to another. In addition, the object need not be chosen for the task, i.e., items drawn from a fixed list, with no knowledge of the environment or the instruction, bypass safety about as often as items proposed for the specific task. We qualitatively contrast the tasks bypassed most and least often and find that the distinction lies in how conspicuous the hazard is in the instruction and environment. Susceptibility to safety bypass is therefore a property of the task, which we call its \emph{malleability}, and we show that it can be predicted before the target is ever queried. A composite of signals read from a small open-source VLM identifies malleable tasks 2.4×2.4\times as often as picking at random. Everyday objects, whether placed by an adversary or introduced by ordinary rearrangement of the environment, are thus sufficient to overturn a refusal. Because susceptibility is determined by how a task is specified, we recommend assessing malleability per task prior to deployment.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Using large language models for embodied planning introduces systematic safety risks

    Apr 20, 2026Tao Zhang, Kaixian Qu, Zhibin Li +4Large Language Model PlanningPlanning

  2. Safe Task Planning with Long-Term Graph Memory for Embodied Agents

    Sep 8, 2026Siyuan Li, Taiyan Lang, Aoqi Yan +6Embodied AgentsGeneralizable Vision-Language-Action Policies

  3. The Yes-Man Syndrome: Benchmarking Abstention in Embodied Robotic Agents

    May 19, 2026Doguhan Yeke, Elif Su Temirel, Ananth Shreekumar +3Abstention