Predictive Semantic Safety: From Visual Physical Reasoning to Safety-Critical Control
Authors: Taekyung Kim, Salem Fradi, Yanning Dai, Mateusz Ostaszewski, Jürgen Schmidhuber
Organizations: Department of Robotics, University of Michigan, Ann Arbor, MI 48109, USA · Center of Excellence in Generative AI, King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia. · Dalle Molle Institute for Artificial Intelligence Research (IDSIA), Switzerland. · Universit`a della Svizzera italiana (USI), Switzerland. · Scuola universitaria professionale della Svizzera italiana (SUPSI), Switzerland.
Physical interactions can create future hazards that are not apparent from the robot's current geometric surroundings. We present a framework termed Predictive Semantic Safety (PSS), which connects visual physical reasoning to backup-based safety filtering. A vision-language model (VLM) predicts physical events and their timing or directly predicts object displacements. An explicit motion model converts event hypotheses into object trajectories. Split conformal prediction calibrates position errors jointly across specified objects, observation times, and future times; geometric shape bounds convert the resulting position regions into predicted object occupancy. PSS evaluates a prescribed backup maneuver against this occupancy and derives input-affine constraints for minimally modifying the nominal input while preserving backup feasibility under the robot dynamics and input limits. MuJoCo experiments with a Unitree Go1 consider falling fixtures, impact-driven support loss, and contact propagation. PSS achieves a safe episode rate of 99.3%, compared with 43.3% for a Backup Control Barrier Function baseline that only uses current obstacle geometry.
Figures & tables
Fig. 1: Motivating example of predictive semantic safety. A partially detached ceiling fixture remains above an apparently clear route, but its anticipated fall may leave the robot with insufficient control authority to avoid collision.
Fig. 2: Overview of PSS, illustrated for the ceiling fixture. RGB-D observations and measured motion are used to predict a physical event and its timing with a VLM. The event hypothesis and an explicit motion model determine future object motion, which is combined with geometric shape bounds and calibrated uncertainty to obtain unsafe occupancy. A Backup CBF safety filter evaluates a prescribed backup maneuver against this occupancy and modifies the nominal command when necessary. The VLM panel shows an excerpt from a recorded model output. MuJoCo trials are shown in Fig. 4 .
Fig. 3: Backup feasibility with and without physical-event prediction. (a) Semantic-Off evaluates the backup using current obstacle geometry, so the resulting motion may continue toward the object’s future occupancy. (b) Semantic-On evaluates the same backup against predicted future occupancy and modifies the nominal motion while the backup remains feasible. Predicted occupancy is shown in both panels but is used only by Semantic-On . Generated conceptual illustration; geometry, trajectories, and timing are schematic.
Fig. 4: Selected MuJoCo trials with physical-event prediction disabled (Semantic-Off) and enabled (Semantic-On). (a) Ceiling fixture: Semantic-Off continues beneath the falling fixture and collides, whereas Semantic-On brakes before full detachment and later navigates around the debris. The red region shows the unsafe occupancy used by the filter at 7.0 s. (b) Impact-induced loss of support: Semantic-Off continues as the stack begins to fall and is struck by a block, whereas Semantic-On brakes and subsequently reaches the goal. Insets show synchronized robot-view images. These qualitative trials are separate from the aggregate benchmark.
Nanjing University, Nanjing, China · The Hong Kong University of Science and Technology (Guangzhou), Guangdong, China · Beijing University of Technology, Beijing, China