cs.AIAug 22, 2026

From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Planning

Authors: Chenghao Zhang, Yikai Mao, Haoyu Gao, Saisai Hu, Yuxi Cheng, Shanqi Liu, Dan Roth

Organizations: University of Pennsylvania · Georgia Institute of Technology · Pace University · Unaffiliated

Abstract

Reliable planning requires converting natural-language instructions into executable symbolic specifications, yet large language models remain brittle without costly PDDL annotations and may exploit solver success in semantically unfaithful ways. We study how to learn faithful natural-language-to-PDDL formalization using only solver feedback, without human-written demonstrations. We propose a solvergrounded multi-role reinforcement learning framework where a single language model acts as an Actor, Judge, and Editor for generation, verification, and repair. The Actor proposes PDDL specifications, the Judge provides a solver-calibrated quality signal, and the Editor performs bounded diagnostic-conditioned refinement. On PlanBench, our method improves average success from 35.5% for LLM+P to 70.8%, achieves 66.3% faithful success, and reduces semantic drift to 6.4%. These results show that organizing solver feedback into generation, verification, and repair roles enables more scalable and faithful annotation-free symbolic planning

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Grounded Evaluation and Repair for NL-to-PDDL Problem Generation

    Sep 9, 2026Joana Rosa, Pedro Santos, Valdemar Oliveira +3Large Language Model PlanningLarge Language Model Generation

  2. Toward Secure and Reliable PDDL Formalization of Large Language Models with Planner-in-the-Loop Feedback

    Jun 29, 2026Jiamei Jiang, Jiajing Zhang, Feifei Mo +2Large Language Model PlanningFormalization

  3. Plan Right, Then Plan Tight: Symbolic RL for Efficient Embodied Reasoning

    Jun 30, 2026Xiangli Shi, Xiaomeng Zhu, Ye Tian +5