Cross-embodiment learning has shown that a single model, such as a vision-language-action (VLA) model, can learn state representations and manipulation skills that can be applied across heterogeneous robots to accomplish various tasks. We hypothesize that the same holds for safety enforcement. The reasoning required to satisfy a safety constraint, such as detecting an obstacle, recognizing that it should be avoided, and selecting a safe abstract action, is largely shared across robots. What differs across embodiments is how the abstract safe action is realized: morphology, kinematics, and dynamics determine which actions are safe and feasible. Consequently, the same action can be safe for one robot and unsafe for another. This is especially important for generalist manipulation policies that operate in a common end-effector action space without explicitly capturing how safety depends on the robot's morphology and kinematics. We propose embodiment-conditioned safety filtering, in which a Hamilton-Jacobi reachability-based value function and its corresponding safety-maximizing policy are shared across robots. Using a morphology-aware latent representation of the robot and its environment, we perform Hamilton-Jacobi reachability analysis directly in latent space so that the learned safety concepts can generalize across embodiments while remaining explicitly conditioned on each robot's morphology and kinematics. We evaluate our approach across five bimanual robot embodiments and five manipulation tasks with whole-body collision-avoidance constraints. Our results show that a single policy, jointly trained across five manipulation tasks and four embodiments, exhibits zero-shot generalization to a held-out embodiment, reducing the nominal policy's collision rate. They also show that training using more embodiments improves generalization.
Figures & tables
Fig. 1 : Overview of CrossSafe. A frozen HoloBrain-0 produces scene and per-link robot tokens, which we augment with safety-related per-link features not encoded by HoloBrain-0. Geometry-aware cross-attention produces a safety-aware latent state with one token per robot link. The safety critic and actor use this latent state to estimate safety and output per-joint angle displacements, respectively, generalizing across embodiments with different degrees of freedom (DoFs).
Fig. 2 : Tasks from left to right: Place Bread in Basket, Place Container on Plate, Stack Two Blocks, Place Burger & Fries, and Stack Two Bowls. Embodiments from left to right: Aloha-AgileX, ARX-X5, Franka-Panda, Piper, and UR5-WSG.
In-distribution Embodiments
Out-of-distribution Embodiment
Method
CR (%) ↓
SR (%) ↑
IR (%)
Force [N] ↓
CR (%) ↓
SR (%) ↑
IR (%)
Force [N] ↓
Nominal
64.1
32.2
—
179.6
64.1
32.2
—
179.6
X-VLA-Safe
50.4 ± 0.7
17.0 ± 4.1
5.2 ± 0.8
189.7 ± 30.3
53.8 ± 17.4
16.7 ± 12.1
4.8 ± 1.5
247.0 ± 144.6
CrossSafe w/o Geo, Aug
38.3 ± 7.7
31.3 ± 4.1
11.1 ± 7.0
81.3 ± 18.0
51.5 ± 18.6
35.0 ± 15.7
7.6 ± 4.3
126.2 ± 94.8
CrossSafe w/o Geo
45.4 ± 7.4
23.8 ± 5.9
20.2 ± 4.6
60.2 ± 17.4
53.4 ± 17.7
30.0 ± 11.5
15.5 ± 10.4
79.0 ± 53.7
CrossSafe w/o Aug
37.1 ± 5.0
26.0 ± 7.5
7.7 ± 1.8
150.1 ± 35.8
49.7 ± 13.3
23.3 ± 13.1
9.2 ± 8.8
133.5 ± 31.8
TABLE I : Comparison of methods using CR, SR, IR, and Force. Each learned method has five trained models, one per leave-one-embodiment-out fold. Each model is evaluated on its four training embodiments (InD) and held-out fifth (OOD).
Model
CR (%) ↓
SR (%) ↑
IR (%)
Force [N] ↓
InD Embodiment(s)
Specialist
37.6 ± 21.3
18.7 ± 10.2
11.2 ± 4.8
113.0 ± 87.1
Generalist
39.8 ± 3.8
23.7 ± 2.8
13.2 ± 1.2
113.6 ± 41.9
OOD Embodiment(s)
Specialist
54.4 ± 8.2
27.0 ± 5.5
5.6 ± 3.4
144.8 ± 24.5
Generalist
49.8 ± 8.3
29.4 ± 15.7
9.2 ± 6.1
102.9 ± 44.3
TABLE II : CrossSafe specialists vs. generalists.
Module
Parameters
Fusion of per-link tokens
202,752
Geometry-aware transformer ( 3 blocks)
3,177,720
Safe actor head
514
Twin safety critic ( 2× [ 2 blocks + value head])
4,373,666
CrossSafe (total)
7,754,652
CrossSafe w/o Geo
7,708,676
TABLE III : Trainable parameters (frozen HoloBrain-0 excluded).
Problem definition
Physics frequency
250 Hz
Control frequency
25 Hz
Control period Δt
0.04 s
Action-limit fraction κ
1.0
Soft-min temperature
T=0.4
Discount γ
0.9
TABLE IV : Hyperparameters.
Task
Emb.
Nominal
X-VLA
Specialist
CrossSafe w/o Geo, Aug
CrossSafe w/o Geo
CrossSafe w/o Aug
CrossSafe
Place Bread in Basket
Piper
74.0/28.0
46.0/8.5
38.0/16.0
47.5/15.0
53.5/ 17.0
48.5/16.5
37.5 /15.0
Franka-Panda
98.0/10.0
85.5/4.0
66.0/0.0
62.5/ 7.0
74.0/4.5
60.0 /1.5
80.0/0.5
ARX-X5
88.0/24.0
66.0/5.5
30.0 /6.0
40.0/ 34.5
63.0/26.0
46.5/25.0
63.5/14.5
UR5-WSG
98.0/10.0
87.0/12.0
96.0/10.0
77.0 / 46.5
82.0/34.0
83.5/27.5
93.0/16.5
Aloha-AgileX
76.0/16.0
73.0/6.0
0.0 /0.0
43.0/ 12.0
51.0/7.0
44.5/5.0
13.0/0.0
Task avg.
86.8/17.6
71.5/7.2
46.0 /6.4
54.0/ 23.0
64.7/17.7
56.6/15.1
57.4/9.3
TABLE V: In-distribution collision and success rate per task and embodiment, shown as CR/SR (both %). Nominal is the unfiltered cuRobo planner and is training-independent. X-VLA-Safe and the four CrossSafe variants are the leave-one-embodiment-out models: each cell averages the four models that had that embodiment in their training pool ( 4×50=200 episodes). Specialist is a CrossSafe model trained on that embodiment alone and evaluated on it ( 50 episodes). Among the learned methods, bold marks the lowest CR and the highest SR in each row. Task avg. averages the five embodiments; Overall average averages all 25 pairs.
Task
Emb.
Nominal
X-VLA
Specialist
CrossSafe w/o Geo, Aug
CrossSafe w/o Geo
CrossSafe w/o Aug
CrossSafe
Place Bread in Basket
Piper
74.0/28.0
24.0 /6.0
37.0/9.5
68.0/12.0
76.0/ 26.0
64.0/16.0
74.0/18.0
Franka-Panda
98.0/10.0
94.0 /18.0
95.5/11.5
98.0/ 22.0
98.0/6.0
98.0/10.0
96.0/6.0
ARX-X5
88.0/24.0
64.0/4.0
66.5/25.0
32.0 /40.0
36.0/28.0
38.0/18.0
42.0/ 44.0
UR5-WSG
98.0/10.0
82.0/8.0
82.0/9.0
92.0/4.0
66.0 / 22.0
72.0/4.0
82.0/0.0
Aloha-AgileX
76.0/16.0
82.0/8.0
24.5 /0.0
60.0/8.0
54.0/ 22.0
60.0/4.0
66.0/4.0
Task avg.
86.8/17.6
69.2/8.8
61.1 /11.0
70.0/17.2
66.0/ 20.8
66.4/10.4
72.0/14.4
TABLE VI: Zero-shot collision and success rate on held-out embodiments, shown as CR/SR (both %). Conditions are as in Table V , but each X-VLA-Safe and CrossSafe cell comes from the single model for which that embodiment was excluded from safety-filter training ( 50 episodes), and each Specialist cell averages the four single-embodiment models that did not train on it ( 4×50=200 episodes). Among the learned methods, bold marks the lowest CR and the highest SR in each row.
Ensuring safety of learning-enabled robotic manipulation across diverse embodiments and tasks still requires significant manual engineering. Existing approaches typically rely on heuristically designed fallback controllers or complex forward invariance assessments. These methods are often too conservative for task success, too computationally expensive for real-time execution, too heuristic to provide useful safety guarantees, or too engineering-heavy to transfer between setups. In this paper, we propose a universal safeguarding approach, X-Safe, which reasons directly in the robot's configuration space to provide formal probabilistic guarantees for collision avoidance. By operating in the configuration space, our method transfers across embodiments while relying solely on an object-based, quasi-static scene representation and a forward kinematics model of the robotic manipulator. Thus, X-Safe provides useful formal safety guarantees without requiring additional data, or engineering effort for different embodiments or scenes. We demonstrate X-Safe for diverse embodiments and policies, both in simulation and on hardware. We observe less degradation in task performance compared to state-of-the-art safeguarding, no collisions on hardware experiments, and empirically corroborate our formal guarantees.
Robot policies are becoming increasingly general, with vision-language-action (VLA) models enabling a single policy to execute diverse tasks specified in natural language. Safe deployment, however, requires adapting not only to new tasks but also to varying safety requirements across users, environments, and applications. Existing safety filters remain largely constraint-specific and thus must be redesigned or relearned when safety requirements change. In this paper, we investigate language-conditioned safety filtering, in which a Hamilton-Jacobi safety actor and critic are conditioned on language-specified constraints. We evaluate this formulation across pick-and-place, table-wiping, and block-stacking tasks in the vision-based setting, examining its ability to enforce language-specified constraints and transfer to unseen constraint instances within the evaluated constraint families. Our experiments provide evidence that language-conditioned safety filters reduce constraint violations and exhibit partial transfer to unseen constraint instances.
Ihab Tabbara, Yuxuan Yang, Hussein Sibai
Department of Computer Science and Engineering, Washington University in St. Louis, MO 63130, USA
Scalable robot imitation learning relies on large-scale heterogeneous data from diverse robots or body-free data, making Cartesian end-effector actions a key interface for embodiment-agnostic policy learning. However, end-effector-only abstraction leaves Cartesian policies unaware of the deployed robot body, making them brittle under robot-specific constraints such as whole-body collision avoidance. To overcome this limitation, we present EmbodiSteer, a training-free framework that steers embodiment-agnostic visuomotor policies toward zero-shot, embodiment-aware deployment. EmbodiSteer keeps policy learning in Cartesian space while efficiently lifting inference-time diffusion sampling into the target robot's joint space via forward kinematics and Jacobian-based updates. With whole-body collision-aware guidance over joint trajectories after each denoising step, the arm can be steered away from collisions while preserving learned end-effector behavior. Compared with Cartesian-only execution, EmbodiSteer reduces collision rate by 46.1% and improves task success rate by 28.5% across 9 simulated robots, and further achieves 90.0% collision rate reduction and 36.7% success rate increase on two physical robots in highly constrained scenarios. Our project page is at https://frankwang67.github.io/EmbodiSteer-Page.
Shihefeng Wang, Kangchen Lv, Mingrui Yu +1
Department of Automation, Tsinghua University · Beijing Key Laboratory of Embodied Intelligence Systems · Institute for Embodied Intelligence and Robotics, Tsinghua University