Vision-Language-Action (VLA) models have become a major paradigm for Vision-and-Language Navigation (VLN). However, in safety-critical facilities, invisible risks such as radiation or temperature spikes cannot be detected by an RGB camera, and handling each risk is expensive, requiring a new encoder, new data, and model retraining. We propose Physics-Guided Visual Prompting (PG-VP), a plug-and-play multimodal perception module that instead reuses what a frozen VLA model already does well: avoiding visible obstacles. Given a proximal radiation or thermal source, PG-VP performs a physics-guided risk assessment to determine the avoidance direction and overlays a corresponding virtual obstacle that moves across consecutive frames (Dynamic Visual Prompting). The navigation policy then naturally detours around this invisible hazard. The identical virtual obstacle is used regardless of hazard type, so the visual prompting pattern remains fixed as sensors are added. When no hazard is detected, nothing is rendered, and the policy behaves exactly as it would without PG-VP. We evaluate PG-VP on OmniNav using the val-unseen splits of R2R-CE and RxR-CE, where it guides the policy toward intended low-risk actions in 84.9% and 83.2% of cases, at a cost of 6.8 and 7.9 percentage points in navigation success rate. We further test it with distinct scenarios on a real robot in the presence of actual thermal and radiation sources, all without any retraining. The real test shows that PG-VP effectively avoids these invisible hazards, improving worst-10% average trajectory safety by 63.45% and 32.59% against thermal and radiation sources, respectively.
Figures & tables
Fig. 1: Overview of PG-VP framework.
Fig. 2: Hardware configuration of the robotic multimodal perception system.
Fig. 3: Dynamic visual prompting process in R2R-CE val-unseen sample. (Right wall, Sguide=0.6 , Ng=Nv=4 )
Fig. 4: Trade-off between PG-VP trajectory steering effectiveness and VLN success rate. The baseline model is OmniNav [ 3 ] . (a) Correct-guidance ratio from PG-VP, (b) Maximum displacement in the x-direction from PG-VP, (c) Success rate drop of VLN
Method
Obs.
R2R-CE
RxR-CE
SR ↑
SPL ↑
SR ↑
SPL ↑
NaVILA (2025) [ 2 ]
S.RGB
54.0
49.0
49.3
44.0
StreamVLN (2025) [ 18 ]
S.RGB
56.4
50.2
54.4
45.4
CorrectNav (2026) [ 26 ]
S.RGB
65.1
62.3
69.3
63.3
JanusVLN (2026) [ 21 ]
S.RGB
60.5
56.8
56.2
47.5
NavFoM (2026) [ 16 ]
Pano.
61.7
55.3
64.4
56.2
TABLE I: Navigation performance (%) on R2R-CE and RxR-CE Val-Unseen splits.
Fig. 5: Cartesian average of PG-VP trajectories on the R2R-CE and RxR-CE val-unseen splits. ( Sguide=0.6 , Ng=Nv=4 )
TABLE II: Safety metrics comparison between the baseline and PG-VP.
Fig. 7: Validation of low-risk navigation trajectories from the safety metrics comparison between the baseline and PG-VP-applied model. All are plotted as a function of distance traveled.
The real-world deployment of Vision-Language-Action (VLA) models remains limited by the risk of unpredictable and irreversible physical harm. However, we currently lack effective mechanisms to proactively detect these physical safety risks before deployment. To address this gap, we propose \textbf{RedVLA}, the first red teaming framework for physical safety in VLA models. We systematically uncover unsafe behaviors through a two-stage process: (I) \textbf{Risk Scenario Synthesis} constructs a valid and task-feasible initial risk scene. Specifically, it identifies critical interaction regions from benign trajectories and positions the risk factor within these regions, aiming to entangle it with the VLA's execution flow and elicit a target unsafe behavior. (II) \textbf{Risk Amplification} ensures stable elicitation across heterogeneous models. It iteratively refines the risk factor state through gradient-free optimization guided by trajectory features. Experiments on six representative VLA models show that RedVLA uncovers diverse unsafe behaviors and achieves the ASR up to 95.5% within 10 optimization iterations. To mitigate these risks, we further propose SimpleVLA-Guard, a lightweight safety guard built from RedVLA-generated data. Our data, assets, and code are available here.
Yuhao Zhang, Borong Zhang, Jiaming Fan +4
Institute for AI, Peking University · State Key Laboratory of General Artificial Intelligence, Peking University
Current Vision-Language-Action (VLA) models rely primarily on RGB perception, preventing them from capturing modalities such as thermal signals that are imperceptible to conventional visual sensors. Moreover, end-to-end generative policies lack explicit safety constraints, making them fragile when encountering obstacles and novel scenarios outside the training distribution. To address these limitations, we propose Safe-Night VLA, a multimodal manipulation framework that enables robots to see the unseen while enforcing rigorous safety constraints for thermal-aware manipulation in unstructured environments. Specifically, Safe-Night VLA integrates long-wave infrared thermal perception into a pre-trained vision-language backbone, enabling semantic reasoning grounded in thermodynamic properties. To ensure safe execution under out-of-distribution conditions, we incorporate a safety filter via control barrier functions, which provide deterministic workspace constraint enforcement during policy execution. We validate our framework through real-world experiments on a Franka manipulator, introducing a novel evaluation paradigm featuring temperature-conditioned manipulation, subsurface target localization, and reflection disambiguation, while maintaining constrained execution at inference time. Results demonstrate that Safe-Night VLA outperforms RGB-only baselines and provide empirical evidence that foundation models can effectively leverage non-visible physical modalities for robust manipulation.
Dian Yu, Qingchuan Zhou, Bingkun Huang +2
Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich (TUM), 80992 Munich, Germany
Vision-Language-Action (VLA) models have demonstrated impressive end-to-end performance across a variety of robotic manipulation tasks. However, these policies offer no guarantees against collisions with task-irrelevant objects in the scene. Existing safety filters sidestep this problem by querying a vision-language model (VLM) to identify obstacles and their locations. This, however, is too slow to run in the control loop and can only be invoked at episode initialization, leaving the filter unable to track moving obstacles. We discover that a small number of attention heads within a VLA model reliably localize the object the policy intends to approach. These heads can be exploited within a training-free safety framework that obtains the active target from the attention heads at every step, treats the remainder of the scene as obstacles, and feeds these into a Control Barrier Function (CBF) filter. Together with a lightweight real-time object tracker, this allows for collision avoidance for non-static obstacles. We evaluate our framework on SafeLIBERO, which we extend with moving obstacles. On the original static benchmark, our method performs comparably to an oracle that uses privileged simulator state to identify the target, emulating a VLM-based identification step run once at episode initialization. On the dynamic variant, where the oracle's init-time target assignment becomes stale, our method substantially outperforms it by 43%, on average. Our findings suggest that the perceptual signals needed for real-time safety filtering are already present within VLA policies and can be exploited without additional training or heavy auxiliary models.
Seongbin Park, Fan Zhang, Baharan Mirzasoleiman +2
University of California Los Angeles United States