WBAG: A Whole-Body and Attached-Geometry Safety Framework for Vision-Language-Action Manipulation
Authors: Samuel Zhen, Siwon Jo, Yanze Zhang, Wenhao Luo
Organizations: Department of Computer Science and Engineering, Texas A&M University, College Station, TX 77843, USA. · GRASP Laboratory, University of Pennsylvania, Philadelphia, PA 19104, USA. · Department of Computer Science, University of Illinois Chicago, Chicago, IL 60607, USA.
Vision-language-action (VLA) policies have demonstrated impressive capabilities in generalizable robotic manipulation, but their deployment in the real world remains challenging due to potential collisions involving different parts of the robot, manipulated objects, and the surrounding environment. Existing inference-time VLA safety frameworks typically rely on simplified end-effector-centered representations that do not explicitly model the full articulated robot and attached-object geometry. In this paper, we present WBAG, a safety framework that models the robot's whole-body and grasp-dependent attached geometry. WBAG constructs a grasp-conditioned safe set that adapts the protected geometry as objects are grasped, then converts this evolving geometry into differentiable CBF constraints that minimally modify the VLA's native six-dimensional operational-space action for collision avoidance across robot, scene, and attached geometry. On the SafeLIBERO benchmark, a variant of LIBERO augmented with obstacles for safety evaluation, WBAG achieves the best overall safety and safe task success among the evaluated methods under a scene-level safety evaluator that monitors all eligible non-task objects, reaching 97.38% aggregate Scene Safety and 59.38% Safe Success.
Figures & tables
Fig. 1: Example of our WBAG execution. Green shows the constructed protected whole-body robot geometry, and orange shows the attached target geometry after grasping, with blue showing the extracted scene obstacle geometry. The yellow dotted curve is the filtered safe end-effector trajectory.
Fig. 2: Overview of WBAG. RGB-D observations reconstruct scene geometry, while robot BP-SDFs are fitted offline. Task-role exclusion defines scene obstacles, and the protected set expands from the articulated robot to include attached geometry after grasp confirmation. The resulting composite clearances are enforced by the grasp-conditioned geometric CBF-QP.
Fig. 3: Attached-geometry construction and refinement. (a) The task target is reconstructed from RGB-D observations and represented by a fitted BP-SDF. (b) After grasp confirmation, target RGB-D points are registered to the initial model with ICP, refining the hand-relative attachment pose before rigid propagation with the gripper.
Method
Perception
Obstacle geometry
Protected robot geometry
Attached geometry representation
AEGIS baseline [ 6 ]
Grounding DINO [ 15 ]
single MVEE
EEF ellipsoid
handcrafted EEF proxy
SAM3 EEF-MVEE
SAM3 [ 12 ]
single MVEE
EEF ellipsoid
handcrafted EEF proxy
SAM3 Scene EEF-MVEE
SAM3 [ 12 ]
scene MVEEs
EEF ellipsoid
handcrafted EEF proxy
WBAG EEF-only + attached geometry
SAM3 [ 12 ]
scene BP-SDFs
EEF-only
attached BP-SDF
WBAG w/o attached geometry
SAM3 [ 12 ]
scene BP-SDFs
whole-body BP-SDF
none
WBAG
SAM3 [ 12 ]
scene BP-SDFs
whole-body BP-SDF
attached BP-SDF
TABLE I: Experimental progression from EEF-MVEE references to scene-aware EEF protection and WBAG.
Fig. 4: Visual comparison of the configurations in Table I : (a) designated-obstacle EEF-MVEE, (b) scene-wide EEF-MVEE, (c) EEF-only with attached geometry, (d) whole-body without attached geometry, and (e) full WBAG.
Suite
Metric
Policy only
AEGIS baseline
SAM3 EEF-MVEE
SAM3 Scene EEF-MVEE
WBAG EEF + attached
WBAG w/o attached
WBAG
Overall
Scene Safety
21.25%
70.87%
72.94%
82.75%
68.40%
77.50%
97.38%
Safe Success
19.00%
51.06%
55.56%
46.25%
46.96%
46.31%
59.38%
Unsafe Success
39.56%
11.25%
8.13%
4.31%
17.58%
14.69%
0.68%
Goal
Scene Safety
23.75%
83.50%
86.25%
77.25%
74.00%
88.75%
98.00%
Safe Success
19.50%
72.75%
76.75%
52.25%
47.50%
53.25%
58.00%
Unsafe Success
40.25%
7.00%
8.50%
8.25%
10.00%
6.75%
0.25%
TABLE II: Performance on SafeLIBERO, aggregated over Levels I and II.
Fig. 5: Representative failure modes of simplified safety geometry: (a) unintended pickup of a non-task object excluded from the designated-obstacle constraint set, (b) upper-link collision under EEF-only protection, and (c) attached-object collision under a handcrafted EEF proxy.
Component
Mean (ms)
Contact search
20.41
Robot BP-SDF eval.
0.26
QP
1.58
Complete WBAG filter
26.76
Policy request
15.43
TABLE III: Mean per-step runtime for WBAG and policy inference.
Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in generalizing across diverse robotic manipulation tasks. However, deploying these models in unstructured environments remains challenging due to the critical need for simultaneous task compliance and safety assurance, particularly in preventing potential collisions during physical interactions. In this work, we introduce a Vision-Language-Safe Action (VLSA) architecture, named AEGIS, which contains a plug-and-play safety constraint (SC) layer formulated via control barrier functions. AEGIS integrates directly with existing VLA models to improve safety with theoretical guarantees, while maintaining their original instruction-following performance. To evaluate the efficacy of our architecture, we construct a comprehensive safety-critical benchmark SafeLIBERO, spanning distinct manipulation scenarios characterized by varying degrees of spatial complexity and obstacle intervention. Extensive experiments demonstrate the superiority of our method over state-of-the-art baselines. Notably, AEGIS achieves over 50% improvement in obstacle avoidance rate while substantially increasing the task success rate by nearly 10%. All benchmark datasets, code, and supplementary materials are publicly available at https://vlsa-aegis.github.io/.
Songqiao Hu, Zeyi Liu, Shuang Liu +5
Department of Automation, Tsinghua University, Beijing 100084, China. · Institute for Embodied Intelligence and Robotics, Tsinghua University, Beijing 100084, China. · 3TetraBOT. +1
Vision-language-action (VLA) benchmarks measure whether a policy completes a requested manipulation task, but binary success can hide safety violations along the trajectory: a policy may reach the goal while applying excessive contact, disturbing bystander objects, destabilizing a held object, or entering robot self-contact. We present SafeVLA-Bench, a post-hoc safety-evaluation framework for existing simulator-based VLA benchmarks that reveals violations missed by success-only evaluation. It encodes task-aware safety requirements as Signal Temporal Logic (STL) invariants with quantitative robustness semantics. Alongside native success, it reports the safety rate and the success-but-unsafe rate (SBU) used in prior safety evaluations, and introduces the Violation Severity Index (VSI), a bounded worst-violation depth score. We instantiate SafeVLA-Bench on LIBERO and RoboCasa-365, evaluating twenty-seven policy-benchmark entries across tabletop and kitchen manipulation tasks. High task success does not imply safe execution: the fifteen tabletop policies above 90% mean success still have 18-28% unsafe-episode rates, and 38-56% of successful RoboCasa-365 rollouts violate at least one active safety clause. A post-training case study further shows that SafeVLA-Bench can be used to improve policy safety. Project page: https://safevla.org
Jialiang Fan, Weizhe Xu, Zijun Wang +3
University of Notre Dame · University of Pennsylvania
Vision-Language-Action (VLA) models have demonstrated impressive end-to-end performance across a variety of robotic manipulation tasks. However, these policies offer no guarantees against collisions with task-irrelevant objects in the scene. Existing safety filters sidestep this problem by querying a vision-language model (VLM) to identify obstacles and their locations. This, however, is too slow to run in the control loop and can only be invoked at episode initialization, leaving the filter unable to track moving obstacles. We discover that a small number of attention heads within a VLA model reliably localize the object the policy intends to approach. These heads can be exploited within a training-free safety framework that obtains the active target from the attention heads at every step, treats the remainder of the scene as obstacles, and feeds these into a Control Barrier Function (CBF) filter. Together with a lightweight real-time object tracker, this allows for collision avoidance for non-static obstacles. We evaluate our framework on SafeLIBERO, which we extend with moving obstacles. On the original static benchmark, our method performs comparably to an oracle that uses privileged simulator state to identify the target, emulating a VLM-based identification step run once at episode initialization. On the dynamic variant, where the oracle's init-time target assignment becomes stale, our method substantially outperforms it by 43%, on average. Our findings suggest that the perceptual signals needed for real-time safety filtering are already present within VLA policies and can be exploited without additional training or heavy auxiliary models.
Seongbin Park, Fan Zhang, Baharan Mirzasoleiman +2
University of California Los Angeles United States