Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation. We develop an inference-time monitor that exposes and verifies intermediate plans without disrupting the original decoding trajectory. Building on this monitor, we propose SafeInferCom, a formal verifier-guided framework that preserves valid intermediate plans and directs error correction during generation. Experiments across multiple LRLMs and planning domains reveal reasoning-response inconsistency and limited self-correction under one-shot inference. SafeInferCom improves planning success and accelerates error correction relative to one-shot inference. When combined with iterative refinement, it further improves success while reducing token usage compared with refinement alone. We additionally evaluate SafeInferCom in VirtualHome and provide a real-world robotic-arm demonstration.
Figures & tables
Fig. 1: Overall workflow of the inference-time monitor. At a selected checkpoint, primary decoding is paused and an auxiliary branch appends the induction prompt pind to the reasoning prefix r<k to produce an induced candidate plan πk . The external verifier returns its label yk , after which the auxiliary branch is discarded and primary decoding resumes until the final plan π is produced.
Fig. 2: The workflow of our method SafeInferCom.
Method
Int. Plan
Verifier
Intervention
Early Stop
One
✗
✗
✗
✗
IR
✗
✓
Post-hoc
✗
Early
✓
✓
✗
✓
S
✓
✓
Mid-gen
✓
S+IR
✓
✓
Both
✓
FD
N/A
N/A
Search
N/A
TABLE I: Comparison of inference strategies. Int. denotes intermediate plan extraction. Verifier denotes the use of external verification within the solution procedure.
Metric
Model
BW
Log
Dep
Grip
Mic
Avg.
Suc. (%) ↑
7B
0.0
10.0
0.0
20.0
22.0
10.4
14B
46.0
46.0
65.0
58.0
59.0
54.8
32B
79.0
84.0
98.0
85.0
68.0
82.8
Tok. ↓
7B
5006
2277
2302
2479
1536
2720
14B
4380
1730
2550
2797
2581
2808
32B
2485
1464
1186
2768
1872
1955
TABLE II: Reasoning evaluation across LRLMs and planning domains
7B
14B
32B
Method
Succ. ↑
Tok. ↓
Succ. ↑
Tok. ↓
Succ. ↑
Tok. ↓
One
10.4
2720
54.8
2808
82.8
1955
IR
13.6
4861
83.2
4014
92.6
2290
Early
13.0
3075
60.0
2943
88.6
1920
S
14.2
3953
78.2
3220
93.2
2337
S+IR
19.6
4463
90.8
3952
98.4
2119
TABLE III: Final success rate (%) and token consumption (average).
One
IR
S
S+IR
Mean
132.03
150.69
130.01
116.10
Median
109.69
119.58
95.59
83.61
P95
275.57
329.03
356.62
288.33
TABLE IV: End-to-end wall-clock time (s) for the 32B model.
R1-14B
Qwen
Phi
Method
Succ. ↑
Tok. ↓
Succ. ↑
Tok. ↓
Succ. ↑
Tok. ↓
One
46.0
1730
64.0
5186
6.0
6411
S
80.0
2460
67.0
4610
7.0
4943
S+IR
96.0
2778
97.0
7674
19.0
15587
TABLE V: Performance on additional LRLM families.
7B
14B
32B
Method
Succ. ↑
Tok. ↓
Succ. ↑
Tok. ↓
Succ. ↑
Tok. ↓
One
3.0
1489
20.0
1227
34.0
1282
S
8.0
7779
59.0
4411
52.0
3520
S+IR
16.0
20746
69.0
7823
61.0
7988
TABLE VI: Performance of SafeInferCom in VirtualHome.
Large language models are increasingly used as planners for robotic systems, yet how safely they plan remains an open question. To evaluate safe planning systematically, we introduce DESPITE, a benchmark of 12,279 tasks spanning physical and normative dangers with fully deterministic validation. Across 23 models, even near-perfect planning ability does not ensure safety: the best-planning model fails to produce a valid plan on only 0.4% of tasks but produces dangerous plans on 28.3%. Among 18 open-source models from 3B to 671B parameters, planning ability improves substantially with scale (0.4-99.3%) while safety awareness remains relatively flat (38-57%). We identify a multiplicative relationship between these two capacities, showing that larger models complete more tasks safely primarily through improved planning, not through better danger avoidance. Three proprietary reasoning models reach notably higher safety awareness (71-81%), while non-reasoning proprietary models and open-source reasoning models remain below 57%. As planning ability approaches saturation for frontier models, improving safety awareness becomes a central challenge for deploying language-model planners in robotic systems.
Tao Zhang, Kaixian Qu, Zhibin Li +4
ETH Zurich, Zurich, Switzerland. · University College London, London, United Kingdom. · Stanford University, Stanford, California, United States. +2
Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on execution rather than verifying the feasibility actions planning models propose. Like general-purpose LLMs, robotics planning models carry risks: biased toward user-specified goals, they may suggest actions misaligned with scientific ethics, they may be unsafe due to an inability to "remember" prior safety risks, or they may be vulnerable to adversarial attacks on the autonomy ecosystem. We propose a LLM-driven verification layer between planning and execution to evaluate action permissibility. Our LLM-as-a-Judge ensemble combines chain-of-thought reasoning across models and synthesizes those expert judge outputs, mirroring a combination of a mixture of experts and self-consistency approach. This layer serves as middleware, gating plans from the server's planning module before they reach the MCP server and therefore the robot's low-level controls: plans are approved, rejected for reformulation, or escalated for human review. With this system, we achieve near 85% precision across accept/escalate/reject categories 97% containment of adversarial attacks, with negligible errors between accepting and rejecting tasks, and errors mostly manifesting at the escalate boundary.
Large Language Models enable flexible natural-language planning but remain unreliable in determinism-critical domains due to their probabilistic nature. This limitation is especially problematic in running planning, where violating safety rules can lead to safety risks. We propose SafeRun, a framework for deterministic LLM-based planning via a decoupled architecture. SafeRun separates soft interpretation by an LLM from hard constraint enforcement by a deterministic solver, ensuring strict safety constraints while preserving natural-language flexibility. To validate SafeRun, we build a comprehensive benchmark for running planning under realistic physiological and safety constraints. Experiments across five LLMs show that SafeRun achieves 100% safety score (vs.\ 79.1% PE average and 97.6% CodeAct average) while maintaining competitive instruction-following scores. The SafeRun benchmark is publicly available at huggingface.
Meilin Chen, Zepeng Zhai, Jiaxuan Zhao +1
Xiaohongshu Inc. · Institute of Information Engineering, Chinese Academy of Sciences