cs.ROOct 5, 2026

What the Guard Misses, the Robot Executes: Implied Harm in VLA Instructions

Authors: Sripad Karne, Arjun Balaji

Organizations: Columbia University

Abstract

Vision-language-action models (VLAs) act on instructions without being able to refuse, so screening harmful requests falls to monitors. We test whether these monitors catch ordinary robot tasks requested for harmful reasons, holding the task fixed while varying only how explicitly the intent is stated. π0.5π_{0.5} completes the task at every level of explicitness, as often as for harmless controls. Text guards flag nearly every blunt request but few implied ones: up to 95% of implied-harm runs end with the task done and no flag raised, and up to 90% even after recalibrating on robot instructions. Monitoring the model's activations does not close this gap. Linear probes separate harmful from harmless instructions almost perfectly in the base language and vision-language models, but this separation weakens after robot training in two model families, most for implied harm.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learning What to Say to Your VLA: Mostly Harmless Vision Language Action Model Steering

    Jun 10, 2026Hyun Joe Jeong, Gokul Swamy, Andrea BajcsySteering

  2. Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms

    Apr 26, 2026Qi Li, Bo Yin, Weiqi Huang +6Diffusion-Based Vision-Language-ActionsEmbodied Artificial Intelligence

  3. RedVLA: Physical Red Teaming for Vision-Language-Action Models

    Apr 24, 2026Yuhao Zhang, Borong Zhang, Jiaming Fan +4Red-TeamingUnsafe