cs.AISep 29, 2026

Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution

Authors: Subba Reddy Oota, Francisco Herrera, Jordi Cabot Sagrera, Marcos López de Prado, Shadab Khan

Organizations: ADIA Lab, Abu Dhabi, United Arab Emirates · University of Granada, Granada, Spain · Luxembourg Institute of Science and Technology, Luxembourg · Cornell University, Ithaca, USA · Lawrence Berkeley National Laboratory, Berkeley, CA

Abstract

Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner--executor systems can fail at either stage, while final task success alone cannot distinguish selection from execution failures. We therefore study the Plan Declaration--Execution Gap and introduce Planning-as-Routing, where an LLM declares one of four planning modes: Predefined, Sequential, Hierarchical, or Search, and a deterministic router dispatches the task to the corresponding pattern-specific executor. Across four benchmarks and three LLMs, we find three consistent patterns. First, generic Plan+ReAct often fails to preserve declared planning structure, especially for longer plans: across three benchmarks, only (22)--(45%) of trajectories preserve it, whereas pattern-specific executors enforce the intended structure. Second, planning-mode effectiveness varies across environments and models: Search performs best on ALFWorld, Hierarchical on SWE-bench, and the strongest pattern can vary across models within the same benchmark. Third, the largest gains come from execution: pattern-specific executors improve task success from (0.48) to (0.92) on ALFWorld and from (0.36) to (0.44) on SWE-bench Verified over Plan+ReAct. Current LLMs, however, do not reliably select the strongest mode for each task, although few-shot examples improve selection in some benchmark--model combinations. Overall, reliable agent planning requires both effective mode selection and faithful execution: routing substantially closes the execution gap, while task-specific mode selection remains open.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents

    Jun 3, 2026Haoyu Sun, Wenxuan Wang, Mingyang Song +5Large Language Model PlanningAgentic Benchmarks

  2. Does The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web Agents

    May 28, 2026Alejandra Zambrano, Sara Vera Marjanovic, Imene Kerboua +2Large Language Model PlanningWeb Agents