Organizations: Field Robotics Engineering and Science Hub (FRESH), Illinois Autonomous Farm, University of Illinois at Urbana-Champaign (UIUC), IL · Mobile Robotics Group, S˜ao Carlos School of Engineering, University of S˜ao Paulo (EESC-USP), S˜ao Carlos, SP, Brazil
The advancement of robotics and autonomous navigation systems hinges on the ability to accurately predict terrain traversability. Traditional methods for generating datasets to train these prediction models often involve putting robots into potentially hazardous environments, posing risks to equipment and safety. To solve this problem, we present ZeST, a novel approach that treats repeated VLM outputs as stochastic measurements and fuses them into an uncertainty-aware posterior. Our approach not only performs zero-shot traversability and mitigates the risks associated with real-world data collection but also accelerates the development of advanced navigation systems, offering a cost-effective and scalable solution. To support our findings, we present navigation results, in both controlled indoor and unstructured outdoor environments. As shown in the experiments, ZeST provides safer navigation with 90-100% success rate with up to 4s inference delays when compared to other state-of-the-art methods, constantly reaching the final goal.
Figures & tables
Fig. 1: ZeST. A zero-shot navigation framework that treats a multimodal foundation model as a stochastic traversability sensor, fuses visual-language predictions into an uncertainty-aware map, and plans risk-aware trajectories.
Fig. 2: ZeST System Overview . (1) VLM Prediction : SLIC-based numbered image used by VLM to generate traversability samples. (2) Statistical Traversability Mapping : Samples are fused into a probabilistic map. (3) Risk Assessment : Navigation risk is quantified via expected shortfall. (4) Zero-Shot Navigation : The robot navigates unfamiliar environments using the risk-aware map.
Fig. 3: VLM Traversability Query. A structured prompt combining physical robot characteristics, numeric terrain anchors, and segmented image regions to condition the foundation model for continuous traversability outputs.
Fig. 4: Traversability Map Examples. Qualitative examples of 3D traversability maps generated by ZeST. Each row shows the RGB observation on the left and the corresponding traversability map on the right. (Top) Indoor Environment: ZeST identifies the open floor region as traversable while assigning lower traversability to cluttered areas around desks, chairs, and obstacles. (Bottom) Outdoor Environment: ZeST produces a traversability map for a grassy scene, highlighting feasible regions through dense vegetation and uneven terrain. Warmer colors represent more traversable regions.
Fig. 5: Qualitative Navigation Results. Navigation results for DWA, CoNVOI, NoMaD, and ZeST across four cluttered indoor and unstructured outdoor scenarios. (Top) Trajectories: Each method’s executed path from one representative run is overlaid, with red and yellow stars marking the start and goal locations, respectively; indoor paths ( a , b ) are shown over the traversability Octomap generated by ZeST, where warmer regions are more traversable, and outdoor paths ( c , d ) over top-down views. (Bottom) Environments. Scene observations: ( a ) an indoor room of desks, chairs, and fragile objects whose goal is reachable only through a single passage; ( b ) a tight corridor with furniture, wire mesh, boxes, other robots, and a thin rod on the floor; ( c ) outdoor terrain mixing grass, uneven soil, exposed roots, and bushes; and ( d ) a snow region that appears flat yet is non-traversable for the robot.
Method
Success Rate (%) ↑
Norm. Traj. Length (→1)
Time to Goal (s) ↓
Min. Clearance (m) ↑
1
DWA [ 34 ]
40
1.053 ± 0.126
50.6 ± 1.7
0.41 ± 0.029
CoNVOI [ 10 ]
30
1.061 ± 0.143
64.9 ± 3.9
0.33 ± 0.056
NoMaD [ 35 ]
50
1.202 ± 0.178
58.5 ± 2.1
0.48 ± 0.031
ZeST (Ours)
100
1.136 ± 0.029
61.6 ± 1.8
0.53 ± 0.059
2
DWA [ 34 ]
30
1.012 ± 0.041
58.7 ± 2.1
0.23 ± 0.051
CoNVOI [ 10 ]
20
1.081 ± 0.072
60.9 ± 5.4
0.30 ± 0.031
TABLE I: Comparative navigation results over 10 runs
Variant
Success Rate (%) ↑
Norm. Traj. Length (→1)
Time to Goal (s) ↓
Min. Clearance (m) ↑
Indoor
ZeST (Full)
80
1.546 ± 0.254
56.1 ± 2.7
0.32 ± 0.154
w/o Fusion
30
1.665 ± 0.301
54.9 ± 3.3
0.22 ± 0.129
w/o Risk
10
1.271
51.1
0.23
Raw Oracle
10
1.031
46.6
1.42
Outdoor
ZeST (Full)
90
1.509 ± 0.159
62.3 ± 2.7
0.90 ± 0.074
w/o Fusion
30
1.438 ± 0.027
64.6 ± 1.4
0.76 ± 0.321
TABLE II: Component ablation results
Fig. 6: VLM Query Variability. Traversability scores from repeated VLM queries on fixed regions and prompts. (a) Per-region scores from 30 queries on one held-out frame, grouped by terrain type; black bars mark medians. Scores stay tight and repeatable on confident terrain like pavement, trees, and buildings, and spread widely on ambiguous terrain like the open lot. (b) Per-region score standard deviation versus mean score across 83 regions from 17 frames; the outlined marker is the open-lot region from (a). The spread is lowest at both ends of the score range and peaks at intermediate scores, so the per-query variability concentrates on genuinely ambiguous terrain rather than scattering at random.
Added Delay (s)
Avg. Latency (s)
Success Rate (%) ↑
Time to Goal (s) ↓
Avg. Speed (% nominal)
None
1.82
100
55.3 ± 2.3
61
+1.0
2.81
100
57.8 ± 1.8
57
+2.0
3.63
100
61.8 ± 2.4
49
+3.0
5.15
0 80
64.7 ± 2.3
45
+4.0
6.12
0 70
72.4 ± 2.7
38
TABLE III: Injected VLM query delays in Indoor Scenario 2
Method
HDR@0.1 ↓
HDR@0.25 ↓
HDR@0.5 ↓
ZeST
0.32[0.22,0.41]
0.31[0.21,0.40]
0.37[0.27,0.47]
WayFAST
0.29[0.23,0.35]
0.48[0.42,0.55]
0.71[0.66,0.78]
TABLE IV: Human disagreement rate for zero-shot and in-domain traversability scores
Ulsan National Institute of Science and Technology, Ulsan, 44919, Republic of Korea · Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich (TUM), Germany