Plan-and-Patch: Diffusion Language Models for Agentic Planning
Organizations: University of Texas at Austin · Amazon
Abstract
Planning is increasingly important for long-horizon agents, where successful execution requires coordinating subgoals, tool use, and intermediate outcomes over many steps. Yet assumptions made during planning may be invalidated by the environment, tools may return unexpected results, or actions may fail. Effective agents must therefore not only generate plans, but also revise them. Such revisions often affect only part of a plan, leaving the preceding and subsequent structure intact. Rather than regenerate the entire plan and risk unnecessary changes, repair can regenerate the affected region conditioned on the preserved prefix and suffix. We introduce Plan-and-Patch, a plan-and-act framework in which a diffusion language model (dLLM) generates a structured, program-like plan through parallel unmasking and repairs it by filling in selected regions while keeping the surrounding steps fixed. We compare DreamReasoner-8B and Qwen3-8B as diffusion and autoregressive (AR) planners. On Natural Plan without task-specific training, diffusion (53.7%) achieves nearly twice the plan repair success rate of AR (27.0%). After task-specific training on agentic benchmarks, ALFWorld and TextCraft, the planners achieve similar observed success in plan generation, while diffusion reduces mean plan-generation latency by 39-46% relative to AR. Our results show that Plan-and-Patch provides a framework for faster plan generation and effective plan repair in long-horizon agents.
Figures & tables
| Environment | Model | Plan generation | Reference-assisted repair | History-based repair | End-to-end |
|---|---|---|---|---|---|
| ALFWorld | ReAct | ||||
| Diff-Plan | |||||
| AR-Plan | |||||
| TextCraft | ReAct | ||||
| Diff-Plan | |||||
| AR-Plan |
| Task | Model | Validator | GPT-5 | Sonnet 5 |
|---|---|---|---|---|
| Generation | Diff-Plan | |||
| AR-Plan | ||||
| Repair | Diff-Plan | |||
| AR-Plan |
| Environment (tasks) | Both succeed | Both fail | Only diffusion fails | Only AR fails |
|---|---|---|---|---|
| ALFWorld ( ) | ||||
| TextCraft ( ) |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Task split | Repair-training examples | ||||||
|---|---|---|---|---|---|---|---|
| Environment | train-gen | train-rep | dev | test | real | synthetic | total |
| ALFWorld | |||||||
| TextCraft | |||||||
| Environment | Split | Candidates | Success | Mean | Median | Max |
|---|---|---|---|---|---|---|
| ALFWorld | train | |||||
| dev | ||||||
| test | ||||||
| all | ||||||
| TextCraft | train | |||||
| dev |
| Plan generation | Reference-assisted repair | |||||||
| Environment | both | Diff-Plan only | AR-Plan only | neither | both | Diff-Plan only | AR-Plan only | neither |
| ALFWorld | ||||||||
| TextCraft | ||||||||
| Task family | Tasks | Both fail | Only diffusion fails | Only AR fails |
|---|---|---|---|---|
| Light inspection | ||||
| Simple placement | ||||
| Clean then place | ||||
| Cool then place | ||||
| Heat then place | ||||
| Two-object placement |
| ALFWorld | TextCraft | |||
|---|---|---|---|---|
| Recorded stopping reason | Diffusion | AR | Diffusion | AR |
| Failure at a leaf step | ||||
| No-progress limit | ||||
| Plan completed without terminal success | ||||
| Invalid-action limit | ||||
| Total | ||||
| Diagnostic | ALFWorld | TextCraft |
|---|---|---|
| All failed executions | ||
| Identified failing leaf step | ||
| Unidentified failing leaf step | ||
| Action at the identified failing leaf step | ||
| take | ||
| fetch | ||
| Qwen3-8B executor | Sonnet 5 executor | ||||
|---|---|---|---|---|---|
| Environment | Planner | Delimited | Prose | Delimited | Prose |
| ALFWorld | Diff-Plan | ||||
| AR-Plan | |||||
| TextCraft | Diff-Plan | ||||
| AR-Plan | |||||
| Environment | Executor | Failures | Eligible | Diff-Plan repair | AR-Plan repair |
|---|---|---|---|---|---|
| ALFWorld | Qwen3-8B | ||||
| TextCraft | Qwen3-8B | ||||
| ALFWorld | Sonnet 5 | ||||
| TextCraft | Sonnet 5 |
| Diff-Plan | AR-Plan | |||
|---|---|---|---|---|
| Environment | Reference-assisted | History-based | Reference-assisted | History-based |
| ALFWorld ( ) | ||||
| TextCraft ( ) | ||||
| Pooled ( ) | ||||
| Domain | Diff-Plan | AR-Plan | |
|---|---|---|---|
| Trip planning | |||
| Meeting scheduling | |||
| Calendar scheduling | – | – |
| Setting | Plan generation | Plan repair |
|---|---|---|
| Training endpoint | Convergence of development-set performance | |
| Learning rate | ||
| Batch gradient accumulation | ||
| Learning-rate warmup | of optimization steps | |
| Maximum sequence length | tokens at training and inference | |
| Diff-Plan block sizes | ||