Consistent Plan-Act for Long-Horizon Agentic Tasks
Organizations: School of Artificial Intelligence, Nanjing University · National Key Laboratory for Novel Software Technology, Nanjing University · LongCat Team, Meituan · University of Science and Technology of China
Abstract
Long-horizon agentic tasks demand strong reasoning and efficient execution across successive interactions with dynamic environments. A common approach decouples high-level planning from low-level execution through separate planner and actor roles. To investigate coordination failures in these tasks, we prompt both agents for structured state assertions and compare their reports programmatically to detect explicit contradictions. Our analyses reveal systematic disagreement about the same task-relevant state facts, a phenomenon we term planner-actor state mismatch. We further find that providing agents with task-relevant state information reduces mismatch and improves coordination and task performance. Based on the systematic analysis of the state mismatch, we propose Consistent Plan-Act (ConPAct), which feeds detected contradictions back to both agents to form consistent state interpretations and fine-tunes them on curated consistent interactions for better coordination. ConPAct improves performance across various environments and model configurations, e.g., increasing MiniGrid success rate from 38.6% to 54.4% with GPT-5.6-sol/terra as planner and actor respectively, demonstrating that state consistency can guide both inference-time correction and coordination training.
Figures & tables
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Game | Target for target_visible | Threat for hazard_nearby |
| bigfish | At least one visibly smaller fish | A larger fish |
| chaser | At least one uncollected green orb | A non-vulnerable enemy |
| coinrun | The goal coin | A saw, enemy, or deadly gap |
| jumper | The carrot | Spikes |
| maze | The cheese | Not included in the schema |
| miner | At least one diamond | A boulder or diamond that could fall onto the player |
| Environment | Source instances | Evaluation instances | Evaluation rollouts | ||
| Sokoban | 100 | 200 | 1,600 | 60 | 60 |
| Crafter | 100 | 200 | 1,600 | 160 | 160 |
| Procgen | 100 | 140 | 1,120 | 160 | 160 |
| MiniGrid / BabyAI | 108 | 150 | 1,200 | 60–200 | 60–200 |
| Game | Source levels | Evaluation levels | ||
| bigfish | 14 | 20 | 1 | 40 |
| chaser | 14 | 20 | 0.5 | 13 |
| coinrun | 15 | 20 | 5 | 10 |
| jumper | 14 | 20 | 3 | 10 |
| maze | 14 | 20 | 5 | 10 |
| miner | 15 | 20 | 1.5 | 13 |
| Task ID | Instances | Seed offsets | |
| BabyAI-BlockedUnlockPickup-v0 | 5 | 0, 1, 2, 3, 4 | 136 |
| BabyAI-GoTo-v0 | 5 | 0, 1, 2, 3, 4 | 200 |
| BabyAI-GoToDoor-v0 | 5 | 0, 1, 2, 3, 4 | 84 |
| BabyAI-GoToLocal-v0 | 5 | 0, 1, 2, 3, 4 | 60 |
| BabyAI-GoToObj-v0 | 5 | 0, 1, 2, 3, 4 | 60 |
| BabyAI-GoToRedBall-v0 | 5 | 0, 1, 2, 3, 4 | 60 |
| Setting | Value |
| Trainable parameters | Language-model backbone; vision encoder/projector frozen |
| Reference schedule | 3 epochs of ReAct SFT; token-matched across methods |
| Training seed | 42 |
| Observation window | 3 frames, including the current observation |
| Optimizer | AdamW (fused) |
| Learning rate |
| Method | Supervised targets |
| ReAct / Plan-Act / ECoT | Complete valid responses from retained demonstrations or progress prefixes. |
| PAA | The planner’s annotated plan and the actor’s action-only responses. |
| HSL | Expert responses and goal-relevant actions from hindsight examples; irrelevant actions are excluded. |
| WebSTAR | Original responses for actions with judge scores greater than 5; other responses are excluded. |
| ConPAct-S | Valid responses from consistent segments and successful corrections; conflicting replies remain context only. |
| ConPAct-R | Validated responses under the same consistency rule; hypothetical errors and recorded nonoptimal actions are excluded from targets. |
| Method | Planner input | Actor input |
| ReAct / HSL / WebSTAR | No separate planner | Up to 3 recent frames |
| Plan-Act / Plan-Act SFT | Current frame plus up to 2 sub-goal-start frames | Up to 3 recent frames |
| PAA | Initial frame; called once | Up to 3 recent frames |
| ECoT | Current frame plus up to 2 sub-goal-start frames | Up to 3 recent frames |
| HiPlan | Initial frame for the global plan; up to 3 recent frames for hints | Up to 3 recent frames |
| TAPE | Up to 3 recent frames for visual calls | Up to 3 recent frames |
| Model | Method | Chrome | GIMP | Calc | Impress | Writer | Multi- apps | OS | Thunder- bird | VLC | VS Code | Overall |
| Opus-5 | ReAct | 71.0/69.3 | 87.2/87.2 | 80.6/80.6 | 83.9/82.3 | 89.2/85.6 | 81.4/75.0 | 82.7/82.7 | 83.4/83.4 | 94.1/90.2 | 82.0/82.0 | 82.0/79.5 |
| Plan-Act | 70.3/68.5 | 88.0/88.0 | 81.9/81.9 | 84.8/83.3 | 88.7/84.9 | 81.9/75.7 | 81.7/81.7 | 84.1/84.1 | 93.4/89.1 | 80.0/80.0 | 82.2/79.7 | |
| [2.5pt][2.5pt] | ConPAct-I | 71.7/70.1 | 86.5/86.5 | 82.3/82.3 | 85.3/83.8 | 90.3/87.1 | 82.6/76.7 | 81.5/81.5 | 82.7/82.7 | 94.6/91.1 | 84.4/84.4 | 82.9/80.6 |
| GPT-5.6-sol | ReAct | 82.6/82.6 | 80.8/80.8 | 78.7/78.7 | 83.0/83.0 | 78.2/69.6 | 83.5/78.5 | 87.5/87.5 | 86.7/86.7 | 86.8/ 82.4 | 91.3/91.3 | 83.2/81.2 |
| Plan-Act | 84.8/84.8 | 76.9/76.9 | 78.7/78.7 | 85.1/85.1 | 77.3/69.6 | 83.6/78.5 | 83.3/83.3 | 86.7/86.7 | 87.1/ 82.4 | 87.0/87.0 | 82.9/80.9 | |
| [2.5pt][2.5pt] | ConPAct-I | 82.6/82.6 | 76.9/76.9 | 80.9/80.9 | 85.1/85.1 | 81.7/73.9 | 85.9/80.7 | 83.3/83.3 | 86.7/86.7 | 88.2/82.4 | 91.3/91.3 | 84.1/82.0 |
| Configuration | Correction Trigger | ✗ | ✓ | ✓ | ✓ |
| Conflict FeedBack | ✗ | ✗ | ✓ | ✓ | |
| Planner Participation | ✗ | ✓ | ✗ | ✓ | |
| GPT-5.6-sol GPT-5.6-sol | |||||
| Sokoban | avg@8 | 85.0 | 90.8 | 89.8 | 92.1 |
| pass@8 | 96.0 | 97.5 | 97.0 | 98.0 | |
| pass 8 | 71.0 | 81.5 | 78.0 | 82.5 | |
| Method | Sokoban | MiniGrid | ||||
| avg@8 | pass@8 | pass 8 | SR | Return | Steps | |
| GPT-5.6-sol | ||||||
| ReAct | 83.8 | 97.0 | 59.5 | 54.5 | 0.50 | 10.9 |
| Self-Refine | 86.5 | 93.5 | 73.5 | 51.2 | 0.49 | 7.2 |
| Plan-Act | 79.4 | 91.5 | 52.5 | 53.1 | 0.46 | 13.1 |
| Assertion | 82.8 | 92.5 | 55.5 | 57.7 | 0.51 | 13.3 |
| Environment | Metric | Baseline | No Sup. | No Replay | Full |
| Qwen3-VL-8B Qwen3-VL-8B | |||||
| Sokoban | 40.7 | 44.4 | 43.5 | 49.1 | |
| 87.5 | 89.0 | 91.0 | 92.5 | ||
| 6.5 | 9.0 | 5.5 | 11.0 | ||
| State mismatch | 46.7 | 34.7 | 39.8 | 12.9 | |
| MiniGrid | SR | 18.0 | 20.0 | 11.6 | 24.8 |
| Environment | Metric | Baseline | ConPAct-R | |
| Qwen3-VL-8B Qwen3-VL-8B | ||||
| Sokoban | Initial D | 47.4 | 20.8 | |
| Planner correctness | 45.5 | 83.0 | ||
| Actor correctness | 46.6 | 83.1 | ||
| MiniGrid | Initial D | 51.3 | 29.9 | |
| Planner correctness | 50.9 | 81.5 | ||
| Training data | Sokoban | MiniGrid | |||||
| Sup. | Replay | Discuss. | Init. D | Resid. D | Discuss. | Init. D | Resid. D |
| Qwen3-VL-8B | |||||||
| ✗ | ✓ | 8.3 | 39.1 | 92.4 | 7.2 | 42.7 | 93.2 |
| ✓ | ✗ | 21.7 | 45.9 | 90.1 | 19.4 | 42.2 | 87.2 |
| [][]✓ | ✓ | 58.4 | 20.8 | 62.0 | 54.1 | 29.9 | 78.9 |
| Qwen3-VL-32B | |||||||
| Training data | Sokoban | MiniGrid | |||||
| Sup. | Replay | CC | CW | D | CC | CW | D |
| Qwen3-VL-8B | |||||||
| ✗ | ✓ | ||||||
| ✓ | ✗ | ||||||
| [][]✓ | ✓ | ||||||
| Qwen3-VL-32B | |||||||
| Method | Calls / episode | Output tokens / episode | Calls / action | Usage coverage (%) |
| ReAct (Planner) | 14.02 | 3,095.37 | 1.04 | 71.5 |
| ReAct (Actor) | 36.29 | 8,703.19 | 1.00 | 71.5 |
| Plan-Act | 35.36 | 5,285.03 | 1.06 | 100.0 |
| HiPlan | 51.86 | 7,992.10 | 2.10 | 100.0 |
| TAPE | 37.84 | 18,900.17 | 6.29 | 99.9 |
| [][] ConPAct-I | 33.02 | 7,199.86 | 1.43 | 83.3 |