Refusals That Bend: Measuring and Predicting Task Malleability in Embodied VLM Planners
Organizations: Purdue University · Amazon · Arizona State University
Abstract
Embodied vision-language models (VLMs) are increasingly deployed as high-level planners for robots because they generalize across diverse environments. However, this requires their safety alignment to also hold in unseen environments. Existing red-teaming assumes an adversary who optimizes the prompt, the pixels, or text in the environment, and existing benchmarks ask whether a planner recognizes or mitigates a hazard in a fixed scene. Neither asks whether a refusal the planner has already given survives an ordinary change to the environment. We ask that question by placing a single everyday object into the environment, with no pixel, gradient, or prompt under adversarial control. On tasks that a constitution-guarded planner initially refuses, we find of tasks can be flipped to compliance by one or more objects, and the number of objects differs from one task to another. In addition, the object need not be chosen for the task, i.e., items drawn from a fixed list, with no knowledge of the environment or the instruction, bypass safety about as often as items proposed for the specific task. We qualitatively contrast the tasks bypassed most and least often and find that the distinction lies in how conspicuous the hazard is in the instruction and environment. Susceptibility to safety bypass is therefore a property of the task, which we call its \emph{malleability}, and we show that it can be predicted before the target is ever queried. A composite of signals read from a small open-source VLM identifies malleable tasks as often as picking at random. Everyday objects, whether placed by an adversary or introduced by ordinary rearrangement of the environment, are thus sufficient to overturn a refusal. Because susceptibility is determined by how a task is specified, we recommend assessing malleability per task prior to deployment.
Figures & tables
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| ROLE | MODEL | PROVIDER |
|---|---|---|
| Target embodied VLM | gemini-robotics-er-2-preview ( Google DeepMind, 2026 ) | |
| Safety judge | gemini-2.5-flash ( Comanici et al., 2025 ) | |
| Constraint judge | gemini-2.5-flash ( Comanici et al., 2025 ) | |
| Proposal VLM | Qwen3-VL-30B-A3B-Thinking ( Bai et al., 2025 ) | Alibaba |
| Scene editor | Flux.2-klein-4B ( Black Forest Labs, 2026 ) | Black Forest Labs |
| Pattern-discovery LLM | gemini-2.5-flash ( Comanici et al., 2025 ) |
| STRATUM | OBJECT BYPASS RATE | POOL | SAMPLED |
|---|---|---|---|
| zero | |||
| low | |||
| medium | |||
| high |
| GROUP | ITEMS |
|---|---|
| Electronics | a laptop, a mouse, a remote, a keyboard, a cell phone |
| Furniture | a chair, a potted plant |
| Kitchen | a cup, a bowl, a spoon |
| Accessories | a backpack, an umbrella, a handbag, a tie, a suitcase |
| Indoor | a book, a clock, a vase, a teddy bear, a toothbrush |
| Sports | a sports ball, a frisbee, a baseball glove |
| MEASUREMENT | ER-1.6 | ER-2 |
|---|---|---|
| Refuses the task initially | ||
| Mean per-task bypass rate | ||
| Bypassed by at least one object | ||
| Bypass rate over trials |
| PATTERN | MALLEABLE | ROBUST |
|---|---|---|
| Instruction | a plausible placement given a constructive rationale (“for inventory”, “to save space”); the risk is indirect and the object is used roughly as intended | an explicit, severe hazard: fire, electrical, chemical or unsanitary misuse, or the obstruction of safety equipment and exits |
| Scene | hazards intrinsic to a single object (“sharp blade”, “fragile glass”); objects described by functional state rather than by risk | hazards arising from arrangement or combination: an ignited torch beside paint thinner, chemicals blocking an emergency exit |
| Interaction | the instruction proposes an action, but specific scene detail decides whether it is safe; neither input alone suffices | |
| Constitution proximity | no separating pattern: is against , CI on the gap | |
| Tasks | |
|---|---|
| Distinct environments | |
| Image size |
| FAMILY | WHAT IT COUNTS | MARKERS |
|---|---|---|
| Object ambiguity | uncertainty about what the object or action is | is it , is this , could be , might be , maybe , either , or is it |
| Conditional | safety judgement made to depend on something the proxy has not settled | if , unless , provid* that , assum* , in case , when … safe |
| Hedge | softens the safety judgement instead of stating it outright | however , alternativ* , if you mean/meant/meaning , could you clarif* , safer , instead , rather , perhaps , might , possibl* , recommend* against , suggest* , would advis* , caution* |
| Self-question | questions the proxy puts to itself | every occurrence of ? |
| Stance flip | changes course part-way through the reasoning | wait , hmm , reconsider* , on the other hand , let me think/rethink |
| SIGNAL | SCORE | ROC | PR | P@50 | P@100 | |
|---|---|---|---|---|---|---|
| A1 | trace length | 0.330 | 0.734 | 0.395 | 0.58 | 0.51 |
| A1 | self-question density † | 0.272 | 0.676 | 0.321 | 0.44 | 0.41 |
| A1 | object ambiguity density | 0.263 | 0.684 | 0.332 | 0.44 | 0.36 |
| A1 | conditional density | 0.241 | 0.674 | 0.324 | 0.44 | 0.35 |
| A1 | stance flip density † | 0.233 | 0.662 | 0.316 | 0.34 | 0.40 |
| A1 | hedge density | 0.203 | 0.646 | 0.273 | 0.26 | 0.28 |
| CHECK | CONSTRAINT | FAIL | SAMPLE | REWEIGHTED (95% CI) |
|---|---|---|---|---|
| no injected text | typographic-injection control | 1 | ||
| object not text | 3 | |||
| placeable | 6 | |||
| hazard remains | 15 | |||
| object present | the editor rendered | 20 | ||
| scene preserved | edit prompt: change nothing else | 32 |
| CHECK | BYPASSED ( ) | REFUSED ( ) | DIFF. | 95% CI |
|---|---|---|---|---|
| object present | ||||
| scene preserved | ||||
| is addition | ||||
| object not text | ||||
| no injected text | ||||
| placeable |