In dynamic open-world environments, autonomous agents often encounter novelties that hinder their ability to find plans to achieve their goals. Specifically, traditional symbolic planners fail to generate plans when the robot's planning domain lacks the operators that enable it to interact appropriately with novel objects in the environment. We propose a neuro-symbolic architecture that integrates symbolic planning, reinforcement learning, and a large language model (LLM) to learn how to handle novel objects. In particular, we leverage the common sense reasoning capability of the LLM to identify missing operators, generate plans with the symbolic AI planner, and write reward functions to guide the reinforcement learning agent in learning control policies for newly identified operators. Our method outperforms the state-of-the-art methods in operator discovery as well as operator learning in continuous robotic domains.Our webpage and code can be access here: helenlu66.github.io/hybridLLMguided/
Figures & tables
Fig. 1: The plan-learn-execute loop. The Hybrid LLM Symbolic planner parses the domain and problem PDDL files, prompts the LLM to define the missing operator(s) in PDDL, and finds a plan with grounded operators. It lifts the grounded operators, and outputs the modified domain PDDL file with the added lifted operator definitions and the plan. The LLM’s newly defined operators are in blue while the existing ones are in black. The learn-execute loop starts by executing the operators in the plan. When it encounters a new operator for which an executor policy does not yet exist, it prompts the LLM to generate dense reward shaping function candidates and launches RL agents to learn the policy. The effects of the operator are treated as sub-goals for the RL agents. One agent is launched per dense reward function candidate and the worst performing agent is eliminated periodically based on sub-goal success rates. The sub-goals are trained in phases. Once training is complete, the best performing policy is saved as the operator’s executor. The pseudocode for the algorithm is shown in Algorithm 1 .
Fig. 2: Problem domains and hybrid LLM-symbolic planner outputs. Domains are ordered by the difficulty of discovering a plannable state via random exploration. Green arrows indicate injected novelties: a lid (Kitchen), a round peg (Nut Assembly), and a drawer or box (Coffee). In the planning graph, green nodes represent states reached via LLM-suggested operators, blue nodes denote existing operators, and orange nodes indicate where the search-ahead algorithm finds a valid plan. Identified missing operators (shown in blue text) include: pick-up-lid-from-pot (Kitchen), pick-up-nut-from-peg (Nut Assembly), pick-up-from-open-box (Coffee-Box), and open-drawer and pick-up-from-open-drawer (Coffee-Drawer).
Fig. 3: The Prompt-LLM-for-New-Operator Pipeline. We sample five operator candidates and select the majority [ 42 ] . The prompt is dynamically filled in with the current state, goal, and existing operators. The LLM suggests names and parameters for missing operators. Preconditions are automatically filled with grounded predicates involving these parameters, while the LLM defines and orders the effects. For example, if preconditions for open-drawer include not open drawer1 and not grasped drawer1 , the LLM generates grasped drawer1 and open drawer1 as ordered effects, which subsequently serve as sub-goals for the guided learning stage. We observe that errors in effects ordering and operator generation are greatly reduced through self-consistency [ 42 ] . Finally, the grounded operator is lifted into a general PDDL definition by mapping specific entities to their variable types (e.g., drawer1 to ?d - drawer), enabling domain-wide generalization.
Fig. 4: The LLM Guided Sub-goal Learning Pipeline. Three reward shaping function candidates are sampled from the LLM. We use a temperature of 1.0 to encourage diversity in the samples. We filter out semantically identical candidates by passing in test values to the candidates and eliminating duplicates that output the same reward. The prompt contains the template for a reward shaping function. The definition of the grounded operator and the observation space of the robot are dynamically filled into the prompt. The LLM writes a function to compute a velocity based progress using the relevant observations in the robot’s observation space. For example, the function computes the progress for open drawer1 using the distance between the drawer handle and the cabinet. During each sub-goal’s training phase, the reward shaping for the sub-goal is unlocked. The worst performing candidate is periodically eliminated.
Hybrid
OD
Hybrid
OD
Domain
Successes
Successes
Time
Time
Kitchen
10/10
10/10
75.19 sec
85.51 sec
Nut Assembly
10/10
0/10
120.76 sec
>7 h
Coffee (box)
10/10
0/10
45.26 sec
>7 h
Coffee (drawer)
7/10
0/10
681.83 sec
>7 h
TABLE I: Hybrid LLM Symbolic Planner vs Operator Discovery (OD)
TABLE II: Diversity in Generated Reward Function Candidates ( Temperature=1 ). Values denote the number of semantically unique candidates generated per sub-goal.
Success Rate p-values
Operator
LS vs LG (Ours)
RM vs LG (Ours)
pick-up-lid-from-pot
0.0010*
0.1094
pick-up-nut-from-peg
0.0010*
0.0107*
pick-up-from-open-box
0.0010*
0.0010*
open-drawer
0.0020*
0.0010*
pick-up-from-open-drawer
0.0010*
0.0010*
TABLE III: Wilcoxon signed tests p-values for Success Rate and Progress. Comparisons are made between our LLM-Guided (LG) Subgoal agent and the baselines-the LEAGUE-sparse agent (LS) and the Reward Machine (RM). Values marked with * indicate statistical significance ( p<0.05 ).
Department of Physical AI Engineering, Jeonbuk National University(JBNU) · Faculty of Department of Advanced Defense Technology and Industry, Jeonbuk National University(JBNU)