Large language models (LLMs) are increasingly used as high-level planners in robot navigation, but their outputs may become unreliable when instructions are ambiguous, unsupported by the environment, or semantically inconsistent. This paper presents a Risk-Aware Semantic Grounding framework for trustworthy LLM-based robot planning. Unlike existing LLM-based planners that primarily optimize plan generation, we formulate semantic grounding reliability as a multi-dimensional risk estimation problem. The proposed architecture explicitly models grounding uncertainty through ambiguity, hallucination and semantic-conflict risks before planning occurs, enabling the system to decide whether to execute the instruction, request clarification, or reject it. To evaluate the approach, we introduce TRUST-NAV, a benchmark containing both standard navigation tasks and risk-inducing instruction scenarios. Experimental results show that while conventional LLM planners achieve strong performance on valid navigation tasks, the proposed framework substantially improves ambiguity detection and semantic conflict rejection. These findings suggest that trustworthy robot planning should be evaluated not only by task completion, but also by the ability to recognize when execution should not occur.
Figures & tables
Figure 1 : Architecture of the proposed Risk-Aware Semantic Grounding framework. Grounding risk is estimated prior to planning to support execute, clarify, or reject decisions.
Figure 2 : Smart home environment for a home assistant robot
Category
Number of Queries
Percentage (%)
Single-Step Planning
34
16.50
Multi-Step Planning
61
29.61
Ambiguous Instructions
41
19.90
Hallucination Scenarios
40
19.42
Semantic Conflict Scenarios
30
14.56
Total
206
100
Table 1 : Composition of the TRUST-NAV benchmark.
Figure 3 : Planning performance of all evaluated planners.
Planner
Single-Step
Multi-Step
Baseline 1: NSOP
35.29
0.00
Baseline 2: RBSGP
58.82
14.75
Baseline 3: SLLmP
97.06
80.33
Baseline 4: TA-LLmPA
100.00
65.57
Proposed Method: RA-SGF
91.18
42.62
Table 2 : Overall planning accuracy (%) on the TRUST-NAV benchmark.
Planner
Single Step
Multi Step
Ambi- guous
Halluci- nation
Semantic Conflict
Overall
Baseline 1: NSOP
55.88
65.57
0.00
82.50
13.33
46.60
Baseline 2: RBSGP
85.29
100.00
0.00
77.50
13.33
60.68
Baseline 3: SLLmP
97.06
93.44
78.05
67.50
86.67
84.95
Baseline 4: TA-LLmPA
100.00
78.69
58.54
92.50
86.67
82.04
Proposed Method: RA-SGF
91.18
72.13
87.80
82.50
96.67
83.98
Table 3 : Decision accuracy (%) across different TRUST-NAV benchmark categories.