In dynamic open-world environments, autonomous agents often encounter novelties that hinder their ability to find plans to achieve their goals. Specifically, traditional symbolic planners fail to generate plans when the robot's planning domain lacks the operators that enable it to interact appropriately with novel objects in the environment. We propose a neuro-symbolic architecture that integrates symbolic planning, reinforcement learning, and a large language model (LLM) to learn how to handle novel objects. In particular, we leverage the common sense reasoning capability of the LLM to identify missing operators, generate plans with the symbolic AI planner, and write reward functions to guide the reinforcement learning agent in learning control policies for newly identified operators. Our method outperforms the state-of-the-art methods in operator discovery as well as operator learning in continuous robotic domains.Our webpage and code can be access here: helenlu66.github.io/hybridLLMguided/
Figures & tables
Fig. 1: The plan-learn-execute loop. The Hybrid LLM Symbolic planner parses the domain and problem PDDL files, prompts the LLM to define the missing operator(s) in PDDL, and finds a plan with grounded operators. It lifts the grounded operators, and outputs the modified domain PDDL file with the added lifted operator definitions and the plan. The LLM’s newly defined operators are in blue while the existing ones are in black. The learn-execute loop starts by executing the operators in the plan. When it encounters a new operator for which an executor policy does not yet exist, it prompts the LLM to generate dense reward shaping function candidates and launches RL agents to learn the policy. The effects of the operator are treated as sub-goals for the RL agents. One agent is launched per dense reward function candidate and the worst performing agent is eliminated periodically based on sub-goal success rates. The sub-goals are trained in phases. Once training is complete, the best performing policy is saved as the operator’s executor. The pseudocode for the algorithm is shown in Algorithm 1 .
Fig. 2: Problem domains and hybrid LLM-symbolic planner outputs. Domains are ordered by the difficulty of discovering a plannable state via random exploration. Green arrows indicate injected novelties: a lid (Kitchen), a round peg (Nut Assembly), and a drawer or box (Coffee). In the planning graph, green nodes represent states reached via LLM-suggested operators, blue nodes denote existing operators, and orange nodes indicate where the search-ahead algorithm finds a valid plan. Identified missing operators (shown in blue text) include: pick-up-lid-from-pot (Kitchen), pick-up-nut-from-peg (Nut Assembly), pick-up-from-open-box (Coffee-Box), and open-drawer and pick-up-from-open-drawer (Coffee-Drawer).
Fig. 3: The Prompt-LLM-for-New-Operator Pipeline. We sample five operator candidates and select the majority [ 42 ] . The prompt is dynamically filled in with the current state, goal, and existing operators. The LLM suggests names and parameters for missing operators. Preconditions are automatically filled with grounded predicates involving these parameters, while the LLM defines and orders the effects. For example, if preconditions for open-drawer include not open drawer1 and not grasped drawer1 , the LLM generates grasped drawer1 and open drawer1 as ordered effects, which subsequently serve as sub-goals for the guided learning stage. We observe that errors in effects ordering and operator generation are greatly reduced through self-consistency [ 42 ] . Finally, the grounded operator is lifted into a general PDDL definition by mapping specific entities to their variable types (e.g., drawer1 to ?d - drawer), enabling domain-wide generalization.
Fig. 4: The LLM Guided Sub-goal Learning Pipeline. Three reward shaping function candidates are sampled from the LLM. We use a temperature of 1.0 to encourage diversity in the samples. We filter out semantically identical candidates by passing in test values to the candidates and eliminating duplicates that output the same reward. The prompt contains the template for a reward shaping function. The definition of the grounded operator and the observation space of the robot are dynamically filled into the prompt. The LLM writes a function to compute a velocity based progress using the relevant observations in the robot’s observation space. For example, the function computes the progress for open drawer1 using the distance between the drawer handle and the cabinet. During each sub-goal’s training phase, the reward shaping for the sub-goal is unlocked. The worst performing candidate is periodically eliminated.
Hybrid
OD
Hybrid
OD
Domain
Successes
Successes
Time
Time
Kitchen
10/10
10/10
75.19 sec
85.51 sec
Nut Assembly
10/10
0/10
120.76 sec
>7 h
Coffee (box)
10/10
0/10
45.26 sec
>7 h
Coffee (drawer)
7/10
0/10
681.83 sec
>7 h
TABLE I: Hybrid LLM Symbolic Planner vs Operator Discovery (OD)
TABLE II: Diversity in Generated Reward Function Candidates ( Temperature=1 ). Values denote the number of semantically unique candidates generated per sub-goal.
Success Rate p-values
Operator
LS vs LG (Ours)
RM vs LG (Ours)
pick-up-lid-from-pot
0.0010*
0.1094
pick-up-nut-from-peg
0.0010*
0.0107*
pick-up-from-open-box
0.0010*
0.0010*
open-drawer
0.0020*
0.0010*
pick-up-from-open-drawer
0.0010*
0.0010*
TABLE III: Wilcoxon signed tests p-values for Success Rate and Progress. Comparisons are made between our LLM-Guided (LG) Subgoal agent and the baselines-the LEAGUE-sparse agent (LS) and the Reward Machine (RM). Values marked with * indicate statistical significance ( p<0.05 ).
Large Language Models (LLMs) have recently shown strong promise for robotic task planning, particularly through automatic planning domain generation. However, prior approaches largely treat generated planning domains as planning utilities, which are brittle under imperfect logical states and perception noise, overlooking their potential as scalable sources of reasoning supervision and structured reward signals. At the same time, reasoning LLMs depend on chain-of-thought (CoT) supervision that is expensive to collect for robotic tasks, and reinforcement learning (RL) faces challenges in reward engineering. We propose Self-CriTeach, an LLM self-teaching and self-critiquing framework in which an LLM autonomously generates symbolic planning domains that serve a dual role: (1) enabling large-scale generation of robotic planning problem-plan pairs, and (2) providing structured reward functions. First, the self-written domains enable large-scale generation of symbolic task plans, which are automatically transformed into extended CoT trajectories for supervised fine-tuning. Second, the self-written domains are reused as structured reward functions, providing dense feedback for reinforcement learning without manual reward engineering. This unified training pipeline yields a planning-enhanced LLM with higher planning success rates, stronger cross-task generalization, reduced inference cost, and resistance to imperfect logical states. GitHub Page: https://markli1hoshipu.github.io/Plan_LLM/
Jinbang Huang, Zhiyuan Li, Yuanzhao Hu +4
Huawei Noah’s Ark Lab · University of Toronto · University of British Columbia +1
Recent works have explored integrating Vision-Language Models (VLMs) with classical planners that rely on symbolic representations of planning problems to generate long-horizon plans for complex embodied tasks. However, in open-ended environments, these symbolic representations obtained from perception are often incomplete, leading to suboptimal performance. To address this, we introduce SCOPE, a self-adaptive symbolic planning framework that supports refining action plans and evolving the symbolic world, i.e., the symbolic representations of open-ended environments. SCOPE comprises two synergistic modules: a Symbolic Execution Simulator (SESim) that conducts symbolic validation and real execution of action plans, leveraging the feedback to refine the plans and evolve the symbolic world; and a Self-Adaptive Symbolic Memory (SASMem) that further distills feedback into evolving symbolic knowledge to enhance long-horizon planning and modeling of the symbolic world. Experiments in open-ended environments show that SCOPE significantly improves the completeness of the symbolic world, the success rate of plans under environment perturbations, and cross-task grounding and adaptability across diverse embodied scenarios.
Yundaichuan Zhan, Minghe Gao, Zhongqi Yue +7
Zhejiang University, Hangzhou, China · Chalmers University of Technology, Gothenburg, Sweden · Lanzhou University, Lanzhou, China
Deploying a neuro-symbolic task planner on a new domain today requires significant manual effort: a domain expert must author relaxation and complementary rules, and hundreds of training problems must be solved to supervise a Graph Neural Network (GNN) object scorer. We propose LLM-Flax, a three-stage framework that eliminates all three sources of manual effort using a locally hosted LLM given only a PDDL domain file. Stage 1 automatically generates relaxation and complementary rules via structured prompting with format validation and self-correction. Stage 2 introduces LLM-guided failure recovery with a feasibility-gated budget policy that explicitly reserves API latency cost before each LLM call, preventing the downstream relaxation fallback from being starved. Stage 3 replaces the domain-trained GNN entirely with zero-shot LLM object importance scoring, requiring no training data. We evaluate all three stages on the MazeNamo benchmark across 10x10, 12x12, and 15x15 grids (8 benchmarks total). LLM-Flax achieves average SR 0.945 versus the manual baseline's 0.828 (+0.117), matching or outperforming manual rules on every one of the eight benchmarks. On 12x12 Expert, LLM-Flax attains SR 0.733 where the manual planner fails entirely (SR 0.000); on 15x15 Hard, it achieves SR 1.000 versus Manual's 0.900. Stage 3 demonstrates feasibility (SR 0.720 on 12x12 Hard with no training data) but faces a context-window bottleneck at scale, pointing to the primary open challenge for future work.
Seongmin Kim, Daegyu Lee
Department of Physical AI Engineering, Jeonbuk National University(JBNU) · Faculty of Department of Advanced Defense Technology and Industry, Jeonbuk National University(JBNU)