A vision-language-action (VLA) policy can finish a manipulation task while knocking over objects unrelated to it, so task success alone does not show that the policy is safe to deploy in clutter. We study how to keep a pretrained VLA policy clear of such hazards at run time without retraining it, which requires guarding more of the arm than the end effector, following the hazard as it moves, and sharing onboard compute with the policy. Our training-free shield covers the gripper, wrist, and forearm with five ellipsoids and filters every commanded motion through one barrier program against a keep-out ellipsoid fitted from RGB-D perception at reset. Sparse optical flow then carries that ellipsoid's center along with the hazard, with no repeated detection or refitting. Over six simulated hazard-motion conditions, the shield lowers collision from 65.62% to 27.27% and raises safe-success, task completion without collision, from 29.35% to 50.43%. Ablations show that guarding the arm links protects beyond end-effector shielding, and that tracking recovers most of the protection lost when the hazard estimate is frozen at reset. On heterogeneous edge hardware, the five-ellipsoid barrier runs on the CPU in 2.2~ms at the 99th percentile, and trimming the vision--language prefix and taking fewer flow-matching steps shortens each π0.5 policy call on the integrated GPU from 343 to 177.3~ms. On a physical SO-101 arm across four tasks, the arm touched the hazard in 3 of 16 shielded episodes versus 11 of 16 unshielded ones. Project page: https://yathag.github.io/multilink-safety-filter/
Figures & tables
Fig. 1: Multi-link shielding around a pretrained policy. (a,b) The unshielded policy and our shield on the moving-hazard benchmark in simulation. (c) Policy optimization and a measured device mapping place the complete loop on one heterogeneous edge platform.
Fig. 2: The shield runs perception and control at distinct cadences. (a) At reset, perception names the hazard, detects it, and fits its ellipsoid. (b) Sparse Lucas–Kanade (LK) optical flow transports the ellipsoid center every five steps. (c) The policy returns 10 actions per call; the panel shows the optimized configuration. (d) Five guarded ellipsoids constrain each command at nominal 20 Hz. Perception and control exchange only the hazard ellipsoid.
Condition
π0.5 (unshielded)
AEGIS §
OSCBF §
Ours
Escape 300
45.99 / 67.63 / 42.15
29.17 / 62.66 / 46.63
31.73 / 69.87 / 53.53
22.44 / 64.58 / 56.41
Escape 150
53.04 / 65.38 / 39.58
40.22 / 53.69 / 37.50
41.83 / 64.90 / 45.67
24.52 / 58.49 / 49.04
Escape 50
72.76 / 56.41 / 23.88
42.47 / 58.17 / 39.42
57.69 / 58.65 / 33.49
25.32 / 55.29 / 44.07
Shuttle 300
59.62 / 68.59 / 37.18
41.67 / 63.14 / 43.91
44.71 / 68.75 / 47.92
36.54 / 62.82 / 51.92
Orbit 25
78.85 / 66.03 / 18.59
38.14 / 74.68 / 50.32
59.94 / 68.43 / 36.06
33.65 / 66.83 / 48.72
Stationary
83.49 / 53.53 / 14.74
22.76 / 62.98 / 51.60
57.85 / 60.42 / 34.62
21.15 / 60.26 / 52.40
TABLE I: Static and moving-hazard performance. C/S/SS: collision/task success/safe-success (%); bold: best per row in (a); underline: best per block in (b). EE: end-effector; both coverage variants hold the hazard estimate frozen at reset. Shield-inactive episodes stay in every denominator. Distances: mm; Mean: over the six conditions. § Adapted from the authors’ code ( Sec. IV ).
Fig. 3: The shield in simulation and on hardware. (a,b) MuJoCo, 300 mm escape condition: both executions complete the task, but only the shielded execution leaves the moving hazard untouched, a safe-success. (c–f) One of the four tasks on a physical SO-101 arm driven by the same edge platform.
Configuration
Coll. (%) ↓
Succ. (%) ↑
Safe-succ. (%) ↑
Unshielded π0.5
45.99
67.63
42.15
Frozen hazard estimate
30.29
58.81
45.51
Fixed-cadence refresh
30.13
58.97
45.83
Exact geometry + true motion ‡
18.43
65.87
60.10
Detector + true motion †
20.19
68.59
60.90
Detector + flow †
22.44
64.74
56.57
TABLE III: Where perception and tracking lose protection. 300 mm escape condition. Bold: best rate per column; † : detector prompted with the oracle-corrected caption.
Fig. 4: (a) Per-step optical flow is equivalent to stride 5 within the shaded ±3 pp margin (pp: percentage points). (b) Motion increases placement sensitivity: the hazard leaves an ellipsoid frozen at reset; both displacement sweeps use the static benchmark. Thin bars: 95% CIs; thick: 90% .
Vision-Language-Action (VLA) models have demonstrated impressive end-to-end performance across a variety of robotic manipulation tasks. However, these policies offer no guarantees against collisions with task-irrelevant objects in the scene. Existing safety filters sidestep this problem by querying a vision-language model (VLM) to identify obstacles and their locations. This, however, is too slow to run in the control loop and can only be invoked at episode initialization, leaving the filter unable to track moving obstacles. We discover that a small number of attention heads within a VLA model reliably localize the object the policy intends to approach. These heads can be exploited within a training-free safety framework that obtains the active target from the attention heads at every step, treats the remainder of the scene as obstacles, and feeds these into a Control Barrier Function (CBF) filter. Together with a lightweight real-time object tracker, this allows for collision avoidance for non-static obstacles. We evaluate our framework on SafeLIBERO, which we extend with moving obstacles. On the original static benchmark, our method performs comparably to an oracle that uses privileged simulator state to identify the target, emulating a VLM-based identification step run once at episode initialization. On the dynamic variant, where the oracle's init-time target assignment becomes stale, our method substantially outperforms it by 43%, on average. Our findings suggest that the perceptual signals needed for real-time safety filtering are already present within VLA policies and can be exploited without additional training or heavy auxiliary models.
Seongbin Park, Fan Zhang, Baharan Mirzasoleiman +2
University of California Los Angeles United States
Recent vision-language-action (VLA) models are promising for general-purpose manipulation, but long-horizon execution remains fragile. Small state-estimation or control errors can lead to irreversible failures (e.g., collisions and object drops). Avoiding these risks requires a proactive safety mechanism capable of anticipating hazards. In this paper, we introduce SafeLoop, a non-invasive external wrapper that adds hazard prediction and rollback-based recovery to a VLA model without changing its parameters. SafeLoop trains a risk predictor from vision and proprioception to output four values: the probability and time-to-hazard for body collisions and for object failures. A lightweight controller then chooses one of three actions based on the predicted risk: continue execution (noop), save a safety checkpoint (record), or retreat in joint space (rollback). Rollback moves the robot back to a recent safe waypoint and queries the base policy again, which may yield an alternative continuation. Across 24 LIBERO tasks (16 random seeds each) and three real-robot tasks (25 rollouts each), SafeLoop achieves a stronger overall safety-success trade-off than alternative methods, reducing hazard cases by roughly 70% while preserving task success and the base-policy control rate. Project code is available at https://github.com/Loule0-0/SafeLoop/tree/release/safeloop.
Zeyu Lou, Tianran Zhang, Xinquan Yue +2
Nanjing University, Nanjing, China · The Hong Kong University of Science and Technology (Guangzhou), Guangdong, China · Beijing University of Technology, Beijing, China
Vision-language-action (VLA) benchmarks measure whether a policy completes a requested manipulation task, but binary success can hide safety violations along the trajectory: a policy may reach the goal while applying excessive contact, disturbing bystander objects, destabilizing a held object, or entering robot self-contact. We present SafeVLA-Bench, a post-hoc safety-evaluation framework for existing simulator-based VLA benchmarks that reveals violations missed by success-only evaluation. It encodes task-aware safety requirements as Signal Temporal Logic (STL) invariants with quantitative robustness semantics. Alongside native success, it reports the safety rate and the success-but-unsafe rate (SBU) used in prior safety evaluations, and introduces the Violation Severity Index (VSI), a bounded worst-violation depth score. We instantiate SafeVLA-Bench on LIBERO and RoboCasa-365, evaluating twenty-seven policy-benchmark entries across tabletop and kitchen manipulation tasks. High task success does not imply safe execution: the fifteen tabletop policies above 90% mean success still have 18-28% unsafe-episode rates, and 38-56% of successful RoboCasa-365 rollouts violate at least one active safety clause. A post-training case study further shows that SafeVLA-Bench can be used to improve policy safety. Project page: https://safevla.org
Jialiang Fan, Weizhe Xu, Zijun Wang +3
University of Notre Dame · University of Pennsylvania