Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution
Organizations: ADIA Lab, Abu Dhabi, United Arab Emirates · University of Granada, Granada, Spain · Luxembourg Institute of Science and Technology, Luxembourg · Cornell University, Ithaca, USA · Lawrence Berkeley National Laboratory, Berkeley, CA
Abstract
Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner--executor systems can fail at either stage, while final task success alone cannot distinguish selection from execution failures. We therefore study the Plan Declaration--Execution Gap and introduce Planning-as-Routing, where an LLM declares one of four planning modes: Predefined, Sequential, Hierarchical, or Search, and a deterministic router dispatches the task to the corresponding pattern-specific executor. Across four benchmarks and three LLMs, we find three consistent patterns. First, generic Plan+ReAct often fails to preserve declared planning structure, especially for longer plans: across three benchmarks, only (22)--(45%) of trajectories preserve it, whereas pattern-specific executors enforce the intended structure. Second, planning-mode effectiveness varies across environments and models: Search performs best on ALFWorld, Hierarchical on SWE-bench, and the strongest pattern can vary across models within the same benchmark. Third, the largest gains come from execution: pattern-specific executors improve task success from (0.48) to (0.92) on ALFWorld and from (0.36) to (0.44) on SWE-bench Verified over Plan+ReAct. Current LLMs, however, do not reliably select the strongest mode for each task, although few-shot examples improve selection in some benchmark--model combinations. Overall, reliable agent planning requires both effective mode selection and faithful execution: routing substantially closes the execution gap, while task-specific mode selection remains open.
Figures & tables
| Benchmark | Model | Overall | Avg. steps | Structure maintained by declared plan length | ||||
|---|---|---|---|---|---|---|---|---|
| 4–5 | 6–7 | 8–10 | ||||||
| ALFWorld | DeepSeek-V4 | |||||||
| Qwen3.6-35B | ||||||||
| Gemma-4-26B | ||||||||
| Mind2Web | DeepSeek-V4 | |||||||
| Qwen3.6-35B | ||||||||
| Pattern | Model | Planned / Task | Completed / Task | Adherence | Full Plan Completed (%) |
|---|---|---|---|---|---|
| Predefined | DeepSeek-V4 | ||||
| Qwen3.6-35B | |||||
| Gemma-4-26B | |||||
| Sequential | DeepSeek-V4 | ||||
| Qwen3.6-35B | |||||
| Gemma-4-26B |
| Bench | Model | SEQ | PRED | HIER | SEARCH | DSR max | DSR max w/o Search | w/o Search | ||
|---|---|---|---|---|---|---|---|---|---|---|
| ALFWorld | DeepSeek-V4 | 134 | ||||||||
| Qwen3.6-35B | 134 | |||||||||
| Gemma-4-26B | 134 | |||||||||
| Mind2Web | DeepSeek-V4 | 1341 | ||||||||
| Qwen3.6-35B | 1341 | |||||||||
| Gemma-4-26B | 1341 |
| Mode | Model | PQ | TSR | PQ–success |
|---|---|---|---|---|
| Predefined | DeepSeek-V4 | |||
| Qwen3.6-35B | ||||
| Sequential | DeepSeek-V4 | |||
| Qwen3.6-35B | ||||
| Hierarchical | DeepSeek-V4 | |||
| Qwen3.6-35B |
| Benchmark | Model | Metric | Flat ReAct | Plan+ReAct | Best Fixed | Routing@1 | Oracle |
|---|---|---|---|---|---|---|---|
| ALFWorld | DeepSeek-V4 | TSR | (SEARCH) | ||||
| Qwen3.6-35B | TSR | (SEARCH) | |||||
| Gemma-4-26B | TSR | (HIER) | |||||
| Mind2Web | DeepSeek-V4 | TSR | (SEARCH) | ||||
| SSR | (SEARCH) | ||||||
| Qwen3.6-35B | TSR | (SEARCH) |
| Benchmark | Model | Metric | Think | @1 | @2 | @3 | Few-shot@1 | Oracle | |
|---|---|---|---|---|---|---|---|---|---|
| ALFWorld | DeepSeek-V4 | TSR | Off | ||||||
| TSR | On | ||||||||
| Qwen3.6 | TSR | Off | |||||||
| TSR | On | ||||||||
| Gemma-4-26B | TSR | Off | |||||||
| TSR | On |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | Environment | Task type | Evaluation | |
|---|---|---|---|---|
| Mind2Web | Web interaction | 1,341 | Multi-step website interaction | Benchmark task success |
| WebArena | Web interaction | 204 | Navigation, retrieval, and website interaction | Benchmark task success |
| SWE-bench Verified | Software engineering | 500 | GitHub issue resolution and code modification | Patch-based evaluation |
| ALFWorld | Embodied environment | 134 | Multi-step household tasks | Goal completion |
| Model | Role | Thinking |
|---|---|---|
| Qwen3.6-35B-A3B | Declaration and execution | Enabled |
| DeepSeek-V4-Flash | Declaration and execution | Enabled |
| Gemma-4-26B-A4B-it | Declaration and execution | Enabled |
| Benchmark | Group | Tasks | Auxiliary statistic | Additional information |
| ALFWorld | Pick & Place | 24 | Mean expert steps: 4.58 | |
| Examine in Light | 18 | Mean expert steps: 3.78 | ||
| Clean & Place | 31 | Mean expert steps: 6.32 | ||
| Heat & Place | 23 | Mean expert steps: 6.04 | ||
| Cool & Place | 21 | Mean expert steps: 6.10 | ||
| Pick Two & Place | 17 | Mean expert steps: 8.65 |
| Model | Temp. | Top- | Top- | Pres. Pen. | Rep. Pen. | Thinking |
|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B | 1.0 | 0.95 | 20 | 1.5 | 1.0 | Yes |
| Gemma4-26B-A4B-it | 1.0 | 0.95 | 64 | 0.0 | 1.0 | Yes |
| DeepSeek-V4-Flash | 1.0 | 0.95 | Disabled | 0.0 | 1.0 | Yes |
| Model | Model Budget | Default Profile | Hard Profile | Tool Choice |
|---|---|---|---|---|
| Qwen3.6-35B-A3B | 81,920 | 32,768 | 81,920 | Required |
| Gemma-4-26B-A4B-it | 131,072 | 32,768 | 81,920 | Required |
| DeepSeek-V4-Flash | 81,920 | 32,768 | 81,920 | Required |
| Benchmark | Model | Episode mean | Across seeds | Range |
| ALFWorld | DeepSeek-V4 | 3–4 | ||
| Qwen3.6-35B | 0–4 | |||
| Gemma-4-26B | 2–4 | |||
| Mind2Web | DeepSeek-V4 | 2–4 | ||
| Qwen3.6-35B | 0–4 | |||
| Gemma-4-26B | 2–3 |
| Benchmark | Max env. steps | max_steps | max_branch | max_depth |
|---|---|---|---|---|
| ALFWorld | 8 | 4 | 3 | |
| Mind2Web | 8 | 4 | 3 | |
| WebArena | 8 | 4 | 3 | |
| SWE-bench | 8 | 4 | 3 |
| Benchmark | Model | Declared mode | Structure maintained | Steps declared | |
| ALFWorld | DeepSeek-V4 | Sequential | /seed | ||
| DeepSeek-V4 | Predefined ‡ | ||||
| DeepSeek-V4 | Hierarchical ‡ | /seed | |||
| DeepSeek-V4 | Search ‡ | /seed | |||
| Qwen3.6-35B | Sequential | /seed | |||
| Qwen3.6-35B | Hierarchical | /seed |
| Model | Planning mode | Planned/task | Completed/task | Adherence | Full plan |
|---|---|---|---|---|---|
| DeepSeek-V4 | Predefined | ||||
| Sequential | |||||
| Hierarchical | |||||
| Search | |||||
| Qwen3.6-35B | Predefined | ||||
| Sequential |
| Model | Planning mode | Planned/task | Completed/task | Adherence | Full plan |
|---|---|---|---|---|---|
| DeepSeek-V4 | Predefined | ||||
| Sequential § | |||||
| Hierarchical | |||||
| Search | |||||
| Qwen3.6-35B | Predefined | ||||
| Sequential § |
| Benchmark | Model | Metric | Flat ReAct | Flat ReAct + Judge | Flat ReAct Pass@3 | Search |
|---|---|---|---|---|---|---|
| ALFWorld | DeepSeek-V4 | TSR | ||||
| Qwen3.6-35B | TSR | |||||
| Mind2Web | DeepSeek-V4 | TSR | ||||
| SSR | ||||||
| Qwen3.6-35B | TSR | |||||
| SSR |
| Benchmark | Model | Metric | (Pattern) | pass@2 | pass@3 | ||||
|---|---|---|---|---|---|---|---|---|---|
| ALFWorld | DeepSeek-V4 | TSR | 134 | (SEARCH) | |||||
| Qwen3.6-35B | TSR | 134 | (SEARCH) | ||||||
| Gemma-4-26B | TSR | 134 | (HIER) | ||||||
| Mind2Web | DeepSeek-V4 | TSR | 1341 | (SEARCH) | |||||
| SSR | 1341 | (SEARCH) | |||||||
| Qwen3.6-35B | TSR | 1341 | (SEARCH) |
| Benchmark | Model | 95% null interval | ( ) | (two-sided) | Holm | ||||
|---|---|---|---|---|---|---|---|---|---|
| ALFWorld | DeepSeek-V4 | 134 | |||||||
| Qwen3.6-35B | 134 | ||||||||
| Gemma-4-26B | 134 | ||||||||
| Mind2Web | DeepSeek-V4 | 1341 | |||||||
| Qwen3.6-35B | 1341 | ||||||||
| Gemma-4-26B | 1341 |