Organizations: Field Robotics Engineering and Science Hub (FRESH), Illinois Autonomous Farm, University of Illinois at Urbana-Champaign (UIUC), IL · Mobile Robotics Group, S˜ao Carlos School of Engineering, University of S˜ao Paulo (EESC-USP), S˜ao Carlos, SP, Brazil
The advancement of robotics and autonomous navigation systems hinges on the ability to accurately predict terrain traversability. Traditional methods for generating datasets to train these prediction models often involve putting robots into potentially hazardous environments, posing risks to equipment and safety. To solve this problem, we present ZeST, a novel approach that treats repeated VLM outputs as stochastic measurements and fuses them into an uncertainty-aware posterior. Our approach not only performs zero-shot traversability and mitigates the risks associated with real-world data collection but also accelerates the development of advanced navigation systems, offering a cost-effective and scalable solution. To support our findings, we present navigation results, in both controlled indoor and unstructured outdoor environments. As shown in the experiments, ZeST provides safer navigation with 90-100% success rate with up to 4s inference delays when compared to other state-of-the-art methods, constantly reaching the final goal.
Figures & tables
Fig. 1: ZeST. A zero-shot navigation framework that treats a multimodal foundation model as a stochastic traversability sensor, fuses visual-language predictions into an uncertainty-aware map, and plans risk-aware trajectories.
Fig. 2: ZeST System Overview . (1) VLM Prediction : SLIC-based numbered image used by VLM to generate traversability samples. (2) Statistical Traversability Mapping : Samples are fused into a probabilistic map. (3) Risk Assessment : Navigation risk is quantified via expected shortfall. (4) Zero-Shot Navigation : The robot navigates unfamiliar environments using the risk-aware map.
Fig. 3: VLM Traversability Query. A structured prompt combining physical robot characteristics, numeric terrain anchors, and segmented image regions to condition the foundation model for continuous traversability outputs.
Fig. 4: Traversability Map Examples. Qualitative examples of 3D traversability maps generated by ZeST. Each row shows the RGB observation on the left and the corresponding traversability map on the right. (Top) Indoor Environment: ZeST identifies the open floor region as traversable while assigning lower traversability to cluttered areas around desks, chairs, and obstacles. (Bottom) Outdoor Environment: ZeST produces a traversability map for a grassy scene, highlighting feasible regions through dense vegetation and uneven terrain. Warmer colors represent more traversable regions.
Fig. 5: Qualitative Navigation Results. Navigation results for DWA, CoNVOI, NoMaD, and ZeST across four cluttered indoor and unstructured outdoor scenarios. (Top) Trajectories: Each method’s executed path from one representative run is overlaid, with red and yellow stars marking the start and goal locations, respectively; indoor paths ( a , b ) are shown over the traversability Octomap generated by ZeST, where warmer regions are more traversable, and outdoor paths ( c , d ) over top-down views. (Bottom) Environments. Scene observations: ( a ) an indoor room of desks, chairs, and fragile objects whose goal is reachable only through a single passage; ( b ) a tight corridor with furniture, wire mesh, boxes, other robots, and a thin rod on the floor; ( c ) outdoor terrain mixing grass, uneven soil, exposed roots, and bushes; and ( d ) a snow region that appears flat yet is non-traversable for the robot.
Method
Success Rate (%) ↑
Norm. Traj. Length (→1)
Time to Goal (s) ↓
Min. Clearance (m) ↑
1
DWA [ 34 ]
40
1.053 ± 0.126
50.6 ± 1.7
0.41 ± 0.029
CoNVOI [ 10 ]
30
1.061 ± 0.143
64.9 ± 3.9
0.33 ± 0.056
NoMaD [ 35 ]
50
1.202 ± 0.178
58.5 ± 2.1
0.48 ± 0.031
ZeST (Ours)
100
1.136 ± 0.029
61.6 ± 1.8
0.53 ± 0.059
2
DWA [ 34 ]
30
1.012 ± 0.041
58.7 ± 2.1
0.23 ± 0.051
CoNVOI [ 10 ]
20
1.081 ± 0.072
60.9 ± 5.4
0.30 ± 0.031
TABLE I: Comparative navigation results over 10 runs
Variant
Success Rate (%) ↑
Norm. Traj. Length (→1)
Time to Goal (s) ↓
Min. Clearance (m) ↑
Indoor
ZeST (Full)
80
1.546 ± 0.254
56.1 ± 2.7
0.32 ± 0.154
w/o Fusion
30
1.665 ± 0.301
54.9 ± 3.3
0.22 ± 0.129
w/o Risk
10
1.271
51.1
0.23
Raw Oracle
10
1.031
46.6
1.42
Outdoor
ZeST (Full)
90
1.509 ± 0.159
62.3 ± 2.7
0.90 ± 0.074
w/o Fusion
30
1.438 ± 0.027
64.6 ± 1.4
0.76 ± 0.321
TABLE II: Component ablation results
Fig. 6: VLM Query Variability. Traversability scores from repeated VLM queries on fixed regions and prompts. (a) Per-region scores from 30 queries on one held-out frame, grouped by terrain type; black bars mark medians. Scores stay tight and repeatable on confident terrain like pavement, trees, and buildings, and spread widely on ambiguous terrain like the open lot. (b) Per-region score standard deviation versus mean score across 83 regions from 17 frames; the outlined marker is the open-lot region from (a). The spread is lowest at both ends of the score range and peaks at intermediate scores, so the per-query variability concentrates on genuinely ambiguous terrain rather than scattering at random.
Added Delay (s)
Avg. Latency (s)
Success Rate (%) ↑
Time to Goal (s) ↓
Avg. Speed (% nominal)
None
1.82
100
55.3 ± 2.3
61
+1.0
2.81
100
57.8 ± 1.8
57
+2.0
3.63
100
61.8 ± 2.4
49
+3.0
5.15
0 80
64.7 ± 2.3
45
+4.0
6.12
0 70
72.4 ± 2.7
38
TABLE III: Injected VLM query delays in Indoor Scenario 2
Method
HDR@0.1 ↓
HDR@0.25 ↓
HDR@0.5 ↓
ZeST
0.32[0.22,0.41]
0.31[0.21,0.40]
0.37[0.27,0.47]
WayFAST
0.29[0.23,0.35]
0.48[0.42,0.55]
0.71[0.66,0.78]
TABLE IV: Human disagreement rate for zero-shot and in-domain traversability scores
Visual traversability estimation is central to autonomous navigation, yet most approaches either rely on prompt-driven Vision-Language Model (VLM) or decouple traversability from trajectory planning, requiring separate planners with heavy mapping, manual tuning, and extended deployment time. We propose EmbodiedDiffusion, a diffusion-based framework that simultaneously predicts traversability maps and generates feasible trajectories from RGB images using planner-free synthetic supervision and embodiment conditioning for cross-platform transfer. The framework distills category-level traversability semantics from a VLM teacher into a lightweight student model during training, enabling prompt-free, real-time inference at deployment. A modular FiLM-based conditioning mechanism isolates embodiment-specific reasoning into a compact trainable subset of the network, allowing rapid adaptation to new robot platforms without retraining the visual backbone or the trajectory diffusion model. Across indoor environments with quadruped and aerial robots, EmbodiedDiffusion achieves 80-100% navigation success in the full-data regime with real-time inference (90 ms) and adapts to new platforms using only 10 min of visual data collection, demonstrating scalable, unified traversability reasoning and trajectory generation for heterogeneous robots.
Iana Zhura, Sausar Karaf, Faryal Batool +7
Intelligent Space Robotics Laboratory, Center for Digital Engineering, Skolkovo Institute of Science and Technology.
Traversability prediction is a critical component of autonomous navigation in unstructured environments, where complex and uncertain robot-terrain interactions pose significant challenges such as traction loss and dynamic instability. Despite recent progress in learning-based traversability prediction, these methods often fail to adapt to novel terrains. Even when adaptation is achieved, retaining experience from previously trained environments remains a challenge, a problem known as catastrophic forgetting. To address this challenge, we propose a continual learning framework for traversability prediction that incrementally adapts to new terrains using a generative experience recall model. A key virtue of the proposed framework is two folds: i) retain prior experience without storing past data; and ii) incorporate the uncertainty of the generated samples from the recall model, enabling uncertainty-aware adaptation. Real-world experiments with a skid-steering robot validate the effectiveness of the proposed framework, demonstrating its ability to adapt across a series of diverse environments while mitigating catastrophic forgetting.
Hojin Lee, Yunho Lee, Daniel A Duecker +1
Ulsan National Institute of Science and Technology, Ulsan, 44919, Republic of Korea · Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich (TUM), Germany
Zero-shot Object Navigation (ZSON) has shown promise for open-vocabulary target search in unseen environments, yet most existing systems remain tied to planar representations and single-floor assumptions. These assumptions become inadequate in real buildings, where navigation involves floors, stairs, landings, and vertically overlapping spaces. This article presents TravExplorer, a cross-floor embodied exploration framework that couples zero-shot semantic guidance with traversability-aware 3-D planning. TravExplorer maintains a unified volumetric map that distinguishes occupied structures from robot-reachable support surfaces and extracts traversable frontiers from connected support surfaces, including floors, stairs, and landings. A FOV-aware active perception strategy further resolves incomplete observations during cross-floor traversal. To reduce semantic-reasoning latency, a lightweight guidance module aligns a probabilistic instance map from online open-vocabulary segmentation with a spatial value map from fast image-to-text matching. Based on these geometric and semantic memories, a hierarchical planner performs target-aware frontier touring over object hypotheses, traversable frontiers, and stair landmarks, and generates executable cross-floor motions through foothold-guided 3-D search and vertically constrained local trajectory optimization. Experiments over 4,195 simulated episodes on HM3D and MP3D demonstrate consistent advantages over representative ObjectNav baselines. Fifty real-world trials on a Unitree Go2 further validate open-vocabulary target search across single-floor and cross-floor indoor environments without prior maps or human intervention. The code will be released at https://github.com/wuyi2121/TravExplorer.