Interactive robot planning requires robots to infer and execute human intentions from natural language instructions that are often ambiguous, incomplete, or underspecified. Although large language models (LLMs) provide a powerful interface for clarification, relying on the generative model to drive an multi-turn conversation can introduce systematic failures. We propose a Bayesian framework that treats clarification as an active learning problem over grounded Signal Temporal Logic (STL) task specifications. Our method uses LLMs to initialize candidate formal specifications and translate informative contrasts into natural-language clarification questions, while Bayesian optimization maintains uncertainty estimation over user intent and selects queries that maximize information gain. After convergence, the inferred STL specification is passed to a formal planner to synthesize a verifiable robot trajectory. Across four simulated and real-world task domains, our approach generally achieves higher task satisfaction and requires fewer clarification rounds than LLM baselines, while helping smaller models close the performance gap against larger reasoning models.
Figures & tables
Fig. 1 : Overview of the proposed language-grounded Bayesian active learning framework. Given an ambiguous instruction and scene representation, an LLM proposes candidate user intents as STL task specifications with prior utility estimates. A grammar VAE maps these formulas into a continuous latent space, where a latent utility model (a Gaussian process in our implementation) is initialized from the LLM prior. An acquisition function contrasts the current best specification with an informative challenger, which may be an LLM candidate or a novel formula sampled directly in the latent space, and the contrast is converted into a clarification question. After each user response, the utility model is updated and the LLM regenerates the candidate set. After convergence, the inferred specification is decoded and passed to a formal planner to synthesize a feasible robot trajectory.
Fig. 2 : Benchmark environments used in simulation and real-world deployment.
Base Model
Panda
City
BAL
LM
LM+P
BAL
LM
LM+P
Success Rate (%) ↑
GPT-5.4
91.1 ± 4.2
48.9 ± 7.5
82.2 ± 5.7
62.5 ± 9.9
29.2 ± 9.3
20.8 ± 8.3
GPT-5.4-mini
93.2 ± 3.8
64.4 ± 7.1
86.7 ± 5.1
64.7 ± 11.6
20.8 ± 8.3
12.5 ± 6.8
Claude-Opus-4-6
91.1 ± 4.2
64.4 ± 7.1
82.2 ± 5.7
50.0 ± 10.2
29.2 ± 9.3
20.8 ± 8.3
Claude-Sonnet-4-6
91.1 ± 4.2
73.3 ± 6.6
86.7 ± 5.1
41.7 ± 10.1
29.2 ± 9.3
8.3 ± 5.6
TABLE I : Success rate (%, ↑ ) and rounds of clarification ( ↓ ) on Panda and City. Mean ± standard error. Bold marks the per-domain best.
Fig. 3 : (a, b) Comparison between human-approval and auto-stop modes on the Panda domain. BAL maintains high task success under auto-stop, while LLM-based baselines terminate earlier but suffer from substantially lower success rates. (c) Pareto scatter plot showing the trade-off between time cost and performance. BAL works best with smaller models near the top-left corner.
Method
Success rate ↑
Avg. rounds ↓
BAL
0.65±0.07
3.67±0.21
LM
0.48±0.07
4.38±0.14
LM+P
0.37±0.06
4.25±0.15
TABLE II : User study results (mean ± SE over N=30 participants).
Domain
Metric
BAL
LM
LM+P
School
Success rate (%) ↑
90.0 ± 9.5
60.0 ± 15.5
80.0 ± 12.6
Avg. rounds ↓
3.50 ± 1.13
6.30 ± 1.17
3.50 ± 1.16
Factory
Success rate (%) ↑
90.0 ± 9.5
80.0 ± 12.6
90.0 ± 9.5
Avg. rounds ↓
2.90 ± 0.90
4.30 ± 1.27
4.10 ± 1.04
TABLE III : Real-world deployment results (mean ± SE over 10 runs per site).
Large Language Models (LLMs) possess extensive foundational knowledge and moderate reasoning abilities, making them suitable for general task planning in open-world scenarios. However, it is challenging to ground a LLM-generated plan to be executable for the specified robot with certain restrictions. This paper introduces CLMASP, an approach that couples LLMs with Answer Set Programming (ASP) to overcome the limitations, where ASP is a non-monotonic logic programming formalism renowned for its capacity to represent and reason about a robot's action knowledge. CLMASP initiates with a LLM generating a basic skeleton plan, which is subsequently tailored to the specific scenario using a vector database. This plan is then refined by an ASP program with a robot's action knowledge, which integrates implementation details into the skeleton, grounding the LLM's abstract outputs in practical robot contexts. Our experiments conducted on the VirtualHome platform demonstrate CLMASP's efficacy. Compared to the baseline executable rate of under 2% with LLM approaches, CLMASP significantly improves this to over 90%.
Xinrui Lin, Yangfan Wu, Huanyu Yang +3
University of Science and Technology of China, China · Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, China
Large language models (LLMs) are increasingly used as high-level planners in robot navigation, but their outputs may become unreliable when instructions are ambiguous, unsupported by the environment, or semantically inconsistent. This paper presents a Risk-Aware Semantic Grounding framework for trustworthy LLM-based robot planning. Unlike existing LLM-based planners that primarily optimize plan generation, we formulate semantic grounding reliability as a multi-dimensional risk estimation problem. The proposed architecture explicitly models grounding uncertainty through ambiguity, hallucination and semantic-conflict risks before planning occurs, enabling the system to decide whether to execute the instruction, request clarification, or reject it. To evaluate the approach, we introduce TRUST-NAV, a benchmark containing both standard navigation tasks and risk-inducing instruction scenarios. Experimental results show that while conventional LLM planners achieve strong performance on valid navigation tasks, the proposed framework substantially improves ambiguity detection and semantic conflict rejection. These findings suggest that trustworthy robot planning should be evaluated not only by task completion, but also by the ability to recognize when execution should not occur.
Łukasz Sobczak, Nur Keleşoğlu, Sławomir Piotr Nowak
Institute of Theoretical and Applied Informatics, Polish Academy of Sciences, Gliwice 44-100, Poland
Large Language Models (LLMs) have recently shown strong promise for robotic task planning, particularly through automatic planning domain generation. However, prior approaches largely treat generated planning domains as planning utilities, which are brittle under imperfect logical states and perception noise, overlooking their potential as scalable sources of reasoning supervision and structured reward signals. At the same time, reasoning LLMs depend on chain-of-thought (CoT) supervision that is expensive to collect for robotic tasks, and reinforcement learning (RL) faces challenges in reward engineering. We propose Self-CriTeach, an LLM self-teaching and self-critiquing framework in which an LLM autonomously generates symbolic planning domains that serve a dual role: (1) enabling large-scale generation of robotic planning problem-plan pairs, and (2) providing structured reward functions. First, the self-written domains enable large-scale generation of symbolic task plans, which are automatically transformed into extended CoT trajectories for supervised fine-tuning. Second, the self-written domains are reused as structured reward functions, providing dense feedback for reinforcement learning without manual reward engineering. This unified training pipeline yields a planning-enhanced LLM with higher planning success rates, stronger cross-task generalization, reduced inference cost, and resistance to imperfect logical states. GitHub Page: https://markli1hoshipu.github.io/Plan_LLM/
Jinbang Huang, Zhiyuan Li, Yuanzhao Hu +4
Huawei Noah’s Ark Lab · University of Toronto · University of British Columbia +1