Robot motion planning in everyday environments must satisfy hard geometric constraints while accounting for context-dependent semantic risk. We present a foundation-model-guided, topology-aware semantic risk field that extends manipulation safety beyond collision avoidance. For each manipulated-object/scene-object pair, a foundation model provides six directional risk weights and a pair-specific spatial decay scale. The method combines these priors with voxelized 3D scene geometry using topology-aware shielding and geodesic spatial decay. A GPU-parallel backend batches object-level distance and risk computations to construct a dense 3D field that serves as a modular cost for downstream motion planning. We evaluate the field's shielding behavior under full and partial barriers and compare its 3D workspace representation with a pixel-wise semantic-prior baseline. Across three household simulation scenarios, trajectories optimized with the proposed field have lower semantic exposure than collision-only trajectories under the same geometric constraints. We also evaluate the computational practicality and reliability of the supporting pipeline. Together, these results support the proposed field as a practical topology-aware semantic cost representation for manipulation planning beyond collision avoidance.
Figures & tables
Fig. 1: Motivating example of semantic and topology-aware manipulation safety with a UR5 carrying a cup of water. (a) Carrying water above a laptop is collision-free but semantically risky. (b) An intervening shelf shields the laptop, reducing the corresponding semantic risk.
Fig. 2: Overview of the proposed semantic risk field construction pipeline. Semantic priors for each manipulated-object/scene-object pair are combined with the observed 3D scene geometry to construct pairwise risk fields Vo(m,p∣E) , which are aggregated into the scene-level field V(m,p∣E) .
Symbol
Meaning
Symbol
Meaning
m
Manipulated object
o
Scene object
O
Scene-object set
p
Query position
Ω
3D workspace
E
Occupied geometry
Ωfree
Free workspace
Bo
Boundary seed voxels
V
Scene risk
Vo
Object risk
D
Six directions
wm,o,d
Directional weight
TABLE I: Notation used in the semantic risk field formulation.
Fig. 3: Each wm,o,d controls the direction-dependent semantic severity, while σm,o controls the Gaussian decay of risk with geodesic distance.
Fig. 4: Oracle-labeled OOPSIEVERSE case studies for the fragile-cluster, laptop, and hot-stove scenarios. For each scenario, the columns show the human reference risk field, the pixel-wise semantic-prior baseline, the proposed 3D semantic risk field projected into the camera view, and the CHOMP trajectory comparison. The human reference risk fields are used only for evaluation and are not provided to the planners. Field colors range from higher risk in red/orange to lower risk in blue/violet. The red trajectory represents a collision-only path, and the green trajectory represents a risk-aware path.
Fig. 5: Topology-aware shielding under full and partial coverage. Full-shielding configurations are (a) an overhead table above a laptop with a cup of water manipulated, (b) a side wall beside a power drill with a hot soldering iron manipulated, and (c) a multi-rack shelf containing a soccer ball, with a knife as the manipulated object. Partial-shielding configurations shorten the side wall in (d) and shift the overhead table inward in (f). (e) is provided for better understanding of (d).
Scene/condition
deuc (m)
dgeo (m)
A
Laptop: table present
0.230
0.395
0.582
Laptop: table removed
0.230
0.230
1.000
Power drill: side wall present
0.230
0.351
0.655
Power drill: side wall removed
0.230
0.230
1.000
Soccer ball: rack present
0.230
0.849
0.271
Soccer ball: rack removed
0.230
0.230
1.000
TABLE II: Counterfactual topology comparison at matched query positions. The straight-line distance is fixed at deuc=0.230m . A=deuc/dgeo is the topology-attenuation ratio. Smaller values indicate stronger shielding.
Field construction
Risk-field value
Residual ratio
Euclidean-distance decay only
3.080×105
1.000
Geodesic-distance decay only
3.114×104
0.101
Full topology-aware formulation
9.085×10−1
2.949×10−6
TABLE III: Component ablation at the matched laptop query.
Scenario
Planner
Mean path risk ↓
Max risk ↓
Path length (m)
Fragile Cluster
Collision-only
0.654
0.850
0.96
Fragile Cluster
Risk-aware
0.014
0.153
1.86
Laptop
Collision-only
0.580
0.999
0.68
Laptop
Risk-aware
0.408
0.691
0.91
Hot Stove
Collision-only
0.275
0.361
1.45
Hot Stove
Risk-aware
0.014
0.083
1.75
TABLE IV: CHOMP-planned EE trajectories evaluated on human reference risk fields.
# Objects
8 mm
1 cm
2 cm
4 cm
1
63.4±0.9
40.0±0.9
16.0±0.6
12.2±0.7
2
82.3±1.4
49.2±0.8
17.7±0.7
13.2±0.6
4
137.7±2.1
80.3±1.6
23.7±0.9
15.9±0.6
6
177.9±2.8
102.4±2.2
28.6±0.9
18.0±0.5
8
256.0±4.7
142.3±3.8
35.5±1.6
20.1±0.7
10
265.0±4.9
149.5±3.5
36.8±0.9
20.8±0.6
TABLE V: Full GPU backend runtime (in milliseconds). Each entry reports the mean ± standard deviation over 100 recorded runs after one unrecorded warm-up.
Fig. 6: Semantic risk-field overlays for scenes in which a hot soldering iron is the manipulated object.
N
SAM 2.1
MobileSAMv2
FastSAM-s
1
1276.5±25.6
62.2±1.8
7.6±1.1
2
1296.8±26.0
62.3±1.7
7.0±0.4
4
1346.4±43.8
66.6±1.3
7.5±0.3
6
1273.2±21.5
67.3±2.2
8.5±0.6
8
1267.2±28.5
69.2±2.2
8.7±0.7
10
1238.6±77.9
71.8±1.9
8.6±0.6
TABLE VI: Runtime comparison of mask-generation methods. Latency is reported in milliseconds as mean ± standard deviation over 100 runs. N is the number of objects.
Fig. 7: Mask-generation results using three different models on a tabletop scene containing 6 objects.
Conventional motion planning treats collision as a binary constraint, although contact with different objects can have drastically different consequences. A robot may safely brush against a cardboard box while even minor contact with a glass, laptop, or unstable object may be undesirable. Moreover, a direct robot--object collision can move the contacted object and trigger secondary object--object collisions, making the risk of a motion depend on the physical evolution of the scene rather than only on the robot's geometric path. We present CaSCo, a cascade-aware soft-collision motion planning framework in which a vision-language or language model assigns semantic risk to objects and a physics simulator predicts the consequences of candidate robot motions. CaSCo searches for a path that minimizes the total semantic risk of the unique objects displaced either directly by the robot or indirectly through cascaded collisions. Because collisions change the environment, we augment roadmap states with the predicted object arrangement and the set of objects whose risk has already been incurred. We develop an optimal graph-search algorithm with an admissible and consistent cascade-relaxed heuristic and caching and pruning mechanisms for efficient search. Experiments in cluttered manipulation environments evaluate semantic risk, cascade reasoning, planning efficiency, and real-robot operation.
We present MC-Risk, a planner-aligned, multi-component risk field on a bird's-eye-view grid that yields early, calibrated, and class-aware risk localization. MC-Risk linearly composes three interpretable modules: (i) a motorized-agent field that fuses a black-box multimodal trajectory predictor with an analytic Gaussian-torus construction whose lateral width grows with speed/curvature and whose height attenuates with look-ahead; (ii) a VRU risk field that replaces isotropic pedestrian blobs with a forward-biased anisotropic kernel aligned to heading and speed; and (iii) a road penalty field that exploits full HD-map topology, imposing an off-road penalty and lane-aware risk exposure for same/opposite directions. We conduct, to our knowledge, the first standardized quantitative evaluation of a risk-field formulation on RiskBench's collision subset. MC-Risk attains the best overall risk localization and the earliest hazard indication. Finally, we demonstrate a plug-and-play planning interface by using the field as an MPC cost density, enabling risk-aware trajectory generation without additional training.
Maximilian Link, Yingjie Xu, Yingbai Hu +1
Technical University of Munich · Munich, Germany · The Chinese University of Hong Kong +3
We propose an online monocular perception-to-control framework that embeds semantic risk into the distance field used by Control Barrier Function (CBF)-based safe navigation and teleoperation. Many perception-based safety filters assign the same distance-based safety margin to all mapped obstacles or use semantics only as a downstream controller adjustment, rather than encoding semantic risk in the spatial representation. Our framework instead reasons online about obstacle geometry and class-dependent risk by embedding semantic information directly into the Euclidean Signed Distance Field (ESDF). This design encodes semantic risk before control optimization, so high-risk objects exert a larger spatial influence in the safety field while retaining efficient ESDF queries at runtime. Specifically, a foundation-model-based SLAM front end reconstructs dense 3-D geometry from monocular RGB video, while per-frame semantic segmentation provides pixel-level class labels that are fused into the reconstructed geometry. The resulting geometric-semantic representation is then converted into an ESDF, where semantic labels identify safety-relevant regions and impose class-dependent inflation before field computation. The semantic-aware ESDF provides the local distance values and spatial derivatives required by the CBF controller, while class-dependent gains further regulate the controller response. Extensive simulation and hardware experiments demonstrate online operation at 10--20 Hz and semantic-aware safe behavior in both teleoperation and autonomous navigation.
Dawei Zhang, Nuo Chen, Shuo Liu +2
Division of Systems Engineering, Boston University, United States · Department of Electrical and Computer Engineering at Texas A&M University, United States · Department of Mechanical Engineering, Boston University, United States