Interactive robot planning requires robots to infer and execute human intentions from natural language instructions that are often ambiguous, incomplete, or underspecified. Although large language models (LLMs) provide a powerful interface for clarification, relying on the generative model to drive an multi-turn conversation can introduce systematic failures. We propose a Bayesian framework that treats clarification as an active learning problem over grounded Signal Temporal Logic (STL) task specifications. Our method uses LLMs to initialize candidate formal specifications and translate informative contrasts into natural-language clarification questions, while Bayesian optimization maintains uncertainty estimation over user intent and selects queries that maximize information gain. After convergence, the inferred STL specification is passed to a formal planner to synthesize a verifiable robot trajectory. Across four simulated and real-world task domains, our approach generally achieves higher task satisfaction and requires fewer clarification rounds than LLM baselines, while helping smaller models close the performance gap against larger reasoning models.
Figures & tables
Fig. 1 : Overview of the proposed language-grounded Bayesian active learning framework. Given an ambiguous instruction and scene representation, an LLM proposes candidate user intents as STL task specifications with prior utility estimates. A grammar VAE maps these formulas into a continuous latent space, where a latent utility model (a Gaussian process in our implementation) is initialized from the LLM prior. An acquisition function contrasts the current best specification with an informative challenger, which may be an LLM candidate or a novel formula sampled directly in the latent space, and the contrast is converted into a clarification question. After each user response, the utility model is updated and the LLM regenerates the candidate set. After convergence, the inferred specification is decoded and passed to a formal planner to synthesize a feasible robot trajectory.
Fig. 2 : Benchmark environments used in simulation and real-world deployment.
Base Model
Panda
City
BAL
LM
LM+P
BAL
LM
LM+P
Success Rate (%) ↑
GPT-5.4
91.1 ± 4.2
48.9 ± 7.5
82.2 ± 5.7
62.5 ± 9.9
29.2 ± 9.3
20.8 ± 8.3
GPT-5.4-mini
93.2 ± 3.8
64.4 ± 7.1
86.7 ± 5.1
64.7 ± 11.6
20.8 ± 8.3
12.5 ± 6.8
Claude-Opus-4-6
91.1 ± 4.2
64.4 ± 7.1
82.2 ± 5.7
50.0 ± 10.2
29.2 ± 9.3
20.8 ± 8.3
Claude-Sonnet-4-6
91.1 ± 4.2
73.3 ± 6.6
86.7 ± 5.1
41.7 ± 10.1
29.2 ± 9.3
8.3 ± 5.6
TABLE I : Success rate (%, ↑ ) and rounds of clarification ( ↓ ) on Panda and City. Mean ± standard error. Bold marks the per-domain best.
Fig. 3 : (a, b) Comparison between human-approval and auto-stop modes on the Panda domain. BAL maintains high task success under auto-stop, while LLM-based baselines terminate earlier but suffer from substantially lower success rates. (c) Pareto scatter plot showing the trade-off between time cost and performance. BAL works best with smaller models near the top-left corner.
Method
Success rate ↑
Avg. rounds ↓
BAL
0.65±0.07
3.67±0.21
LM
0.48±0.07
4.38±0.14
LM+P
0.37±0.06
4.25±0.15
TABLE II : User study results (mean ± SE over N=30 participants).
Domain
Metric
BAL
LM
LM+P
School
Success rate (%) ↑
90.0 ± 9.5
60.0 ± 15.5
80.0 ± 12.6
Avg. rounds ↓
3.50 ± 1.13
6.30 ± 1.17
3.50 ± 1.16
Factory
Success rate (%) ↑
90.0 ± 9.5
80.0 ± 12.6
90.0 ± 9.5
Avg. rounds ↓
2.90 ± 0.90
4.30 ± 1.27
4.10 ± 1.04
TABLE III : Real-world deployment results (mean ± SE over 10 runs per site).