Cross-embodiment learning has shown that a single model, such as a vision-language-action (VLA) model, can learn state representations and manipulation skills that can be applied across heterogeneous robots to accomplish various tasks. We hypothesize that the same holds for safety enforcement. The reasoning required to satisfy a safety constraint, such as detecting an obstacle, recognizing that it should be avoided, and selecting a safe abstract action, is largely shared across robots. What differs across embodiments is how the abstract safe action is realized: morphology, kinematics, and dynamics determine which actions are safe and feasible. Consequently, the same action can be safe for one robot and unsafe for another. This is especially important for generalist manipulation policies that operate in a common end-effector action space without explicitly capturing how safety depends on the robot's morphology and kinematics. We propose embodiment-conditioned safety filtering, in which a Hamilton-Jacobi reachability-based value function and its corresponding safety-maximizing policy are shared across robots. Using a morphology-aware latent representation of the robot and its environment, we perform Hamilton-Jacobi reachability analysis directly in latent space so that the learned safety concepts can generalize across embodiments while remaining explicitly conditioned on each robot's morphology and kinematics. We evaluate our approach across five bimanual robot embodiments and five manipulation tasks with whole-body collision-avoidance constraints. Our results show that a single policy, jointly trained across five manipulation tasks and four embodiments, exhibits zero-shot generalization to a held-out embodiment, reducing the nominal policy's collision rate. They also show that training using more embodiments improves generalization.
Figures & tables
Fig. 1 : Overview of CrossSafe. A frozen HoloBrain-0 produces scene and per-link robot tokens, which we augment with safety-related per-link features not encoded by HoloBrain-0. Geometry-aware cross-attention produces a safety-aware latent state with one token per robot link. The safety critic and actor use this latent state to estimate safety and output per-joint angle displacements, respectively, generalizing across embodiments with different degrees of freedom (DoFs).
Fig. 2 : Tasks from left to right: Place Bread in Basket, Place Container on Plate, Stack Two Blocks, Place Burger & Fries, and Stack Two Bowls. Embodiments from left to right: Aloha-AgileX, ARX-X5, Franka-Panda, Piper, and UR5-WSG.
In-distribution Embodiments
Out-of-distribution Embodiment
Method
CR (%) ↓
SR (%) ↑
IR (%)
Force [N] ↓
CR (%) ↓
SR (%) ↑
IR (%)
Force [N] ↓
Nominal
64.1
32.2
—
179.6
64.1
32.2
—
179.6
X-VLA-Safe
50.4 ± 0.7
17.0 ± 4.1
5.2 ± 0.8
189.7 ± 30.3
53.8 ± 17.4
16.7 ± 12.1
4.8 ± 1.5
247.0 ± 144.6
CrossSafe w/o Geo, Aug
38.3 ± 7.7
31.3 ± 4.1
11.1 ± 7.0
81.3 ± 18.0
51.5 ± 18.6
35.0 ± 15.7
7.6 ± 4.3
126.2 ± 94.8
CrossSafe w/o Geo
45.4 ± 7.4
23.8 ± 5.9
20.2 ± 4.6
60.2 ± 17.4
53.4 ± 17.7
30.0 ± 11.5
15.5 ± 10.4
79.0 ± 53.7
CrossSafe w/o Aug
37.1 ± 5.0
26.0 ± 7.5
7.7 ± 1.8
150.1 ± 35.8
49.7 ± 13.3
23.3 ± 13.1
9.2 ± 8.8
133.5 ± 31.8
TABLE I : Comparison of methods using CR, SR, IR, and Force. Each learned method has five trained models, one per leave-one-embodiment-out fold. Each model is evaluated on its four training embodiments (InD) and held-out fifth (OOD).
Model
CR (%) ↓
SR (%) ↑
IR (%)
Force [N] ↓
InD Embodiment(s)
Specialist
37.6 ± 21.3
18.7 ± 10.2
11.2 ± 4.8
113.0 ± 87.1
Generalist
39.8 ± 3.8
23.7 ± 2.8
13.2 ± 1.2
113.6 ± 41.9
OOD Embodiment(s)
Specialist
54.4 ± 8.2
27.0 ± 5.5
5.6 ± 3.4
144.8 ± 24.5
Generalist
49.8 ± 8.3
29.4 ± 15.7
9.2 ± 6.1
102.9 ± 44.3
TABLE II : CrossSafe specialists vs. generalists.
Module
Parameters
Fusion of per-link tokens
202,752
Geometry-aware transformer ( 3 blocks)
3,177,720
Safe actor head
514
Twin safety critic ( 2× [ 2 blocks + value head])
4,373,666
CrossSafe (total)
7,754,652
CrossSafe w/o Geo
7,708,676
TABLE III : Trainable parameters (frozen HoloBrain-0 excluded).
Problem definition
Physics frequency
250 Hz
Control frequency
25 Hz
Control period Δt
0.04 s
Action-limit fraction κ
1.0
Soft-min temperature
T=0.4
Discount γ
0.9
TABLE IV : Hyperparameters.
Task
Emb.
Nominal
X-VLA
Specialist
CrossSafe w/o Geo, Aug
CrossSafe w/o Geo
CrossSafe w/o Aug
CrossSafe
Place Bread in Basket
Piper
74.0/28.0
46.0/8.5
38.0/16.0
47.5/15.0
53.5/ 17.0
48.5/16.5
37.5 /15.0
Franka-Panda
98.0/10.0
85.5/4.0
66.0/0.0
62.5/ 7.0
74.0/4.5
60.0 /1.5
80.0/0.5
ARX-X5
88.0/24.0
66.0/5.5
30.0 /6.0
40.0/ 34.5
63.0/26.0
46.5/25.0
63.5/14.5
UR5-WSG
98.0/10.0
87.0/12.0
96.0/10.0
77.0 / 46.5
82.0/34.0
83.5/27.5
93.0/16.5
Aloha-AgileX
76.0/16.0
73.0/6.0
0.0 /0.0
43.0/ 12.0
51.0/7.0
44.5/5.0
13.0/0.0
Task avg.
86.8/17.6
71.5/7.2
46.0 /6.4
54.0/ 23.0
64.7/17.7
56.6/15.1
57.4/9.3
TABLE V: In-distribution collision and success rate per task and embodiment, shown as CR/SR (both %). Nominal is the unfiltered cuRobo planner and is training-independent. X-VLA-Safe and the four CrossSafe variants are the leave-one-embodiment-out models: each cell averages the four models that had that embodiment in their training pool ( 4×50=200 episodes). Specialist is a CrossSafe model trained on that embodiment alone and evaluated on it ( 50 episodes). Among the learned methods, bold marks the lowest CR and the highest SR in each row. Task avg. averages the five embodiments; Overall average averages all 25 pairs.
Task
Emb.
Nominal
X-VLA
Specialist
CrossSafe w/o Geo, Aug
CrossSafe w/o Geo
CrossSafe w/o Aug
CrossSafe
Place Bread in Basket
Piper
74.0/28.0
24.0 /6.0
37.0/9.5
68.0/12.0
76.0/ 26.0
64.0/16.0
74.0/18.0
Franka-Panda
98.0/10.0
94.0 /18.0
95.5/11.5
98.0/ 22.0
98.0/6.0
98.0/10.0
96.0/6.0
ARX-X5
88.0/24.0
64.0/4.0
66.5/25.0
32.0 /40.0
36.0/28.0
38.0/18.0
42.0/ 44.0
UR5-WSG
98.0/10.0
82.0/8.0
82.0/9.0
92.0/4.0
66.0 / 22.0
72.0/4.0
82.0/0.0
Aloha-AgileX
76.0/16.0
82.0/8.0
24.5 /0.0
60.0/8.0
54.0/ 22.0
60.0/4.0
66.0/4.0
Task avg.
86.8/17.6
69.2/8.8
61.1 /11.0
70.0/17.2
66.0/ 20.8
66.4/10.4
72.0/14.4
TABLE VI: Zero-shot collision and success rate on held-out embodiments, shown as CR/SR (both %). Conditions are as in Table V , but each X-VLA-Safe and CrossSafe cell comes from the single model for which that embodiment was excluded from safety-filter training ( 50 episodes), and each Specialist cell averages the four single-embodiment models that did not train on it ( 4×50=200 episodes). Among the learned methods, bold marks the lowest CR and the highest SR in each row.
Department of Automation, Tsinghua University · Beijing Key Laboratory of Embodied Intelligence Systems · Institute for Embodied Intelligence and Robotics, Tsinghua University