WBAG: A Whole-Body and Attached-Geometry Safety Framework for Vision-Language-Action Manipulation
Authors: Samuel Zhen, Siwon Jo, Yanze Zhang, Wenhao Luo
Organizations: Department of Computer Science and Engineering, Texas A&M University, College Station, TX 77843, USA. · GRASP Laboratory, University of Pennsylvania, Philadelphia, PA 19104, USA. · Department of Computer Science, University of Illinois Chicago, Chicago, IL 60607, USA.
Vision-language-action (VLA) policies have demonstrated impressive capabilities in generalizable robotic manipulation, but their deployment in the real world remains challenging due to potential collisions involving different parts of the robot, manipulated objects, and the surrounding environment. Existing inference-time VLA safety frameworks typically rely on simplified end-effector-centered representations that do not explicitly model the full articulated robot and attached-object geometry. In this paper, we present WBAG, a safety framework that models the robot's whole-body and grasp-dependent attached geometry. WBAG constructs a grasp-conditioned safe set that adapts the protected geometry as objects are grasped, then converts this evolving geometry into differentiable CBF constraints that minimally modify the VLA's native six-dimensional operational-space action for collision avoidance across robot, scene, and attached geometry. On the SafeLIBERO benchmark, a variant of LIBERO augmented with obstacles for safety evaluation, WBAG achieves the best overall safety and safe task success among the evaluated methods under a scene-level safety evaluator that monitors all eligible non-task objects, reaching 97.38% aggregate Scene Safety and 59.38% Safe Success.
Figures & tables
Fig. 1: Example of our WBAG execution. Green shows the constructed protected whole-body robot geometry, and orange shows the attached target geometry after grasping, with blue showing the extracted scene obstacle geometry. The yellow dotted curve is the filtered safe end-effector trajectory.
Fig. 2: Overview of WBAG. RGB-D observations reconstruct scene geometry, while robot BP-SDFs are fitted offline. Task-role exclusion defines scene obstacles, and the protected set expands from the articulated robot to include attached geometry after grasp confirmation. The resulting composite clearances are enforced by the grasp-conditioned geometric CBF-QP.
Fig. 3: Attached-geometry construction and refinement. (a) The task target is reconstructed from RGB-D observations and represented by a fitted BP-SDF. (b) After grasp confirmation, target RGB-D points are registered to the initial model with ICP, refining the hand-relative attachment pose before rigid propagation with the gripper.
Method
Perception
Obstacle geometry
Protected robot geometry
Attached geometry representation
AEGIS baseline [ 6 ]
Grounding DINO [ 15 ]
single MVEE
EEF ellipsoid
handcrafted EEF proxy
SAM3 EEF-MVEE
SAM3 [ 12 ]
single MVEE
EEF ellipsoid
handcrafted EEF proxy
SAM3 Scene EEF-MVEE
SAM3 [ 12 ]
scene MVEEs
EEF ellipsoid
handcrafted EEF proxy
WBAG EEF-only + attached geometry
SAM3 [ 12 ]
scene BP-SDFs
EEF-only
attached BP-SDF
WBAG w/o attached geometry
SAM3 [ 12 ]
scene BP-SDFs
whole-body BP-SDF
none
WBAG
SAM3 [ 12 ]
scene BP-SDFs
whole-body BP-SDF
attached BP-SDF
TABLE I: Experimental progression from EEF-MVEE references to scene-aware EEF protection and WBAG.
Fig. 4: Visual comparison of the configurations in Table I : (a) designated-obstacle EEF-MVEE, (b) scene-wide EEF-MVEE, (c) EEF-only with attached geometry, (d) whole-body without attached geometry, and (e) full WBAG.
Suite
Metric
Policy only
AEGIS baseline
SAM3 EEF-MVEE
SAM3 Scene EEF-MVEE
WBAG EEF + attached
WBAG w/o attached
WBAG
Overall
Scene Safety
21.25%
70.87%
72.94%
82.75%
68.40%
77.50%
97.38%
Safe Success
19.00%
51.06%
55.56%
46.25%
46.96%
46.31%
59.38%
Unsafe Success
39.56%
11.25%
8.13%
4.31%
17.58%
14.69%
0.68%
Goal
Scene Safety
23.75%
83.50%
86.25%
77.25%
74.00%
88.75%
98.00%
Safe Success
19.50%
72.75%
76.75%
52.25%
47.50%
53.25%
58.00%
Unsafe Success
40.25%
7.00%
8.50%
8.25%
10.00%
6.75%
0.25%
TABLE II: Performance on SafeLIBERO, aggregated over Levels I and II.
Fig. 5: Representative failure modes of simplified safety geometry: (a) unintended pickup of a non-task object excluded from the designated-obstacle constraint set, (b) upper-link collision under EEF-only protection, and (c) attached-object collision under a handcrafted EEF proxy.
Component
Mean (ms)
Contact search
20.41
Robot BP-SDF eval.
0.26
QP
1.58
Complete WBAG filter
26.76
Policy request
15.43
TABLE III: Mean per-step runtime for WBAG and policy inference.
Department of Automation, Tsinghua University, Beijing 100084, China. · Institute for Embodied Intelligence and Robotics, Tsinghua University, Beijing 100084, China. · 3TetraBOT. +1