Large language model (LLM)-based robot task planning is promising for open-ended instruction following, but degrades on long-horizon tasks in large environments. When spatial information is conveyed to the LLM through text, the model can fail to capture spatial context, and token cost grows with environment size. Generating action sequences directly with an LLM also makes it difficult to satisfy the current world state and action preconditions. We address this with an ontology-grounded scene representation that aligns objects, spaces, relations, and states in a shared symbolic vocabulary for spatial reasoning and task planning, and with OntoPlan, an agentic framework that interprets instructions, selectively retrieves task-relevant information, formalizes goals and constraints, and produces executable plans. Across 150 general tasks spanning five indoor environments and three scene scales, OntoPlan achieves 0.89 average task success, compared with 0.27 for the strongest baseline, while using 18.1k total tokens per task on average, about 5.6× fewer than the most efficient baseline. These advantages persist as scene scale increases, whereas prior methods degrade more sharply in success and remain far more costly in tokens. OntoPlan also responds appropriately to ambiguous or infeasible instructions by asking follow-up questions or reporting insufficient information rather than committing to invalid plans. Code available at https://github.com/namhyeongwoo/OntoPlan.
Figures & tables
Figure 1 : Overview of OntoPlan . From a user instruction, the Flow Orchestrator routes scene grounding, clarification, task formalization, and planning across specialized agents, while the Scene query and PDDL plan tools operate over a shared ontology-grounded world model.
Figure 2 : World Model. Ontology-grounded representation with ontology, static, and dynamic layers; ask supports scene inspection and planner-state extraction, and tell updates dynamic facts from symbolic observations and action effects. Red edges denote facts inferred by the OWL reasoner.
Figure 3 : System Architecture of OntoPlan . Flow Orchestrator, Scene Explorer, Task Formalizer, and Planning Manager coordinate through shared memory, while the Scene query and PDDL plan tools access the shared world model for scene grounding, subgoal planning, and revision. Each agent box lists the memory fields the agent reads (Input) and writes (Output), its role in the system prompt, and its routing options.
Environment
Method
Plan Returned
Plan Executable
Goal Reached
Task Success
Input Tokens per Call
Total Tokens per Task
Klickitat
SayPlan
0.67
0.47
0.33
0.27
30.9
221.7
DELTA
0.63
0.23
0.23
0.20
23.9
101.3
OntoPlan
1.00
1.00
0.93
0.90
2.6
17.3
Lakeville
SayPlan
0.57
0.40
0.33
0.33
33.3
205.5
DELTA
0.47
0.23
0.17
0.17
23.4
98.8
OntoPlan
1.00
1.00
0.97
0.97
2.6
18.8
Table 3: General-Task Performance by Environment. Results are averaged over the three scene scales for each environment. All four performance metrics are rates, and token metrics are reported in ×103 . Bold marks the best value in each environment.
Table 5
Figure 4 : Scalability Across Scene Scales. OntoPlan degrades more gracefully than the baselines while remaining more token-efficient as scene complexity increases. Each scale comprises 50 tasks, and dashed lines connect a method’s rates across scales.
Setting
Task Success
Total Tokens per Task
Peak Input Tokens per Call
S
M
L
Avg.
S
M
L
L/S
S
M
L
L/S
Full system
0.92
0.90
0.84
0.89
17.1
17.5
19.7
1.15×
3.6
3.7
3.9
1.06×
w/o Scene query
0.86
0.90
0.84
0.87
18.1
32.5
72.7
4.02×
7.6
22.3
62.2
8.13×
w/o PDDL plan
0.32
0.26
0.36
0.31
23.6
25.0
26.9
1.14×
8.4
8.5
8.4
1.01×
Table 6 : Ablation of Scene query and PDDL plan across scene scales. The w/o Scene query and w/o PDDL plan variants replace selective retrieval with full-scene serialization and the integrated planning path with direct LLM action-sequence generation, respectively. Avg. denotes the mean across scales; token metrics are reported in ×103 , and L/S denotes the Large-to-Small ratio.
Figure 5 : Qualitative Examples of OntoPlan . A trace of spatial question answering, commonsense task grounding, executable planning, and sequential execution with world model updates.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
(a) Explicit (2,908 triples).
(b) Reasoned (19,394 triples; 16,486 inferred).
(a) Small. Original source scene graph (121 objects).
(b) Medium. Small with manually added objects (423 objects).
(c) Large. Medium with random augmentation to 3 × objects (1269 objects).
Large Language Models (LLMs) can reason over complex instructions but often fail to satisfy the physical and spatial constraints required for robotic task planning. Recent LLM-based planners directly translate text into action sequences, yet they lack structured reasoning about feasibility, reachability, and logical order, resulting in invalid or incomplete plans. We present a heterogeneous multi-LLM framework that decomposes instructions into atomic reasoning tasks and allocates them to role-specialized expert agents under a token budget for real-world computational and communicational constraints. By combining role-oriented reasoning from heterogeneous agents followed by constraint-driven plan synthesis, HEART validates capability, reachability, and constraint conditions before planning and helps produce physically executable plans while maintaining efficiency. Experiments across different household benchmarks show that HEART consistently improves plan success compared to single-LLM and rule-based planners, demonstrating that heterogeneous LLM collaboration enables robust and scalable robotic task planning under resource constraints.
Junho Lee, Seabin Lee, Wonjong Lee +3
Dept. of Electronic Engineering, Sogang University, Seoul, Korea
We present a hierarchical language-driven framework for robotic task and motion planning to improve natural, intuitive human-robot interaction in service and assistance scenarios. The proposed system employs two large language model (LLM) modules: a high-level planning agent and a low-level spatial reasoning sub-module. The primary agent processes natural language commands and generates action sequences using a ReAct-style prompt, interacting with tools for object perception and manipulation (e.g., pick, place, release). For precise spatial placement, such as interpreting "place the mug next to the plate", a separate sub-prompting module handles 3D reasoning based on object geometry and scene layout. The system integrates YOLOX-GDRNet for object detection and pose estimation, along with a motion execution stub. We evaluated the system in 24 test scenarios, ranging from simple spatial commands to high-level instructions and infeasible requests. The system achieved an overall task success rate of 86%.
Karolina Źróbek, Tessa Pulli, Paweł Gajewski +2
IBM, Krakow, Poland · TU Wien, Vienna, Austria · Jagiellonian University, Krakow, Poland +1
Large Language Models (LLMs) possess extensive foundational knowledge and moderate reasoning abilities, making them suitable for general task planning in open-world scenarios. However, it is challenging to ground a LLM-generated plan to be executable for the specified robot with certain restrictions. This paper introduces CLMASP, an approach that couples LLMs with Answer Set Programming (ASP) to overcome the limitations, where ASP is a non-monotonic logic programming formalism renowned for its capacity to represent and reason about a robot's action knowledge. CLMASP initiates with a LLM generating a basic skeleton plan, which is subsequently tailored to the specific scenario using a vector database. This plan is then refined by an ASP program with a robot's action knowledge, which integrates implementation details into the skeleton, grounding the LLM's abstract outputs in practical robot contexts. Our experiments conducted on the VirtualHome platform demonstrate CLMASP's efficacy. Compared to the baseline executable rate of under 2% with LLM approaches, CLMASP significantly improves this to over 90%.
Xinrui Lin, Yangfan Wu, Huanyu Yang +3
University of Science and Technology of China, China · Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, China