Organizations: School of Artificial Intelligence, Nanjing University · National Key Laboratory for Novel Software Technology, Nanjing University · LongCat Team, Meituan · University of Science and Technology of China
Long-horizon agentic tasks demand strong reasoning and efficient execution across successive interactions with dynamic environments. A common approach decouples high-level planning from low-level execution through separate planner and actor roles. To investigate coordination failures in these tasks, we prompt both agents for structured state assertions and compare their reports programmatically to detect explicit contradictions. Our analyses reveal systematic disagreement about the same task-relevant state facts, a phenomenon we term planner-actor state mismatch. We further find that providing agents with task-relevant state information reduces mismatch and improves coordination and task performance. Based on the systematic analysis of the state mismatch, we propose Consistent Plan-Act (ConPAct), which feeds detected contradictions back to both agents to form consistent state interpretations and fine-tunes them on curated consistent interactions for better coordination. ConPAct improves performance across various environments and model configurations, e.g., increasing MiniGrid success rate from 38.6% to 54.4% with GPT-5.6-sol/terra as planner and actor respectively, demonstrating that state consistency can guide both inference-time correction and coordination training.
Figures & tables
Figure 1: Planner-actor state mismatch . (a) Misreading the upper box as one column farther right, the actor pushes Right twice, causing deadlock; the correct sequence is Right → Up , with consistent interpretation as the planner. (b) Mismatch occurs in 85.0% of failed trajectories, far exceeding non-adherence (31.1%) and low planning/execution capability (8.4%).
Figure 2: Prevalence of planner-actor state mismatch. (a) Mismatch occurs across various model configurations, with even homogeneous configurations have notable state mismatch; (b) State mismatch coexists with high sub-goal adherence, with marker size indicating trajectory frequency. Model sources are provided in Appendix C.2 .
Figure 3: Planner-actor state mismatch is systematic and prevalent in failed trajectories. (a) Mismatch persists under both visual and textual observations; (b) mean cross-role disagreement exceeds within-role sampling variability; (c) mismatch-only failures dominate across task configurations.
Figure 4: State information and consistent plan-act coordination. (a) Task-relevant state facts improve the actor’s decisions. (b) Injecting reliable state assertions reduces mismatch and improves task success. (c) ConPAct feeds detected contradictions back for coordination reconciliation.
Figure 5: OSWorld performance and Sokoban inference-time comparisons. (a) OSWorld Score and SR relative to the stronger of ReAct and Plan-Act within each model, measured in percentage points; point labels give absolute percentages and purple annotations give ConPAct-I gains. (b) Task success (bars, left axis) and disagreement repair (lines, right axis) under different correction settings. (c) Task success with assertion prompting, resampling, Self-Refine, and ConPAct-I.
Figure 6: Supervised-training ablations and state analysis on Sokoban. (a) Task success and mismatch under ConPAct-R data ablations; Full includes consistency filtering, reconciliation supervision, and reconciliation replay. (b) Initial mismatch and assertion correctness; annotations show changes in percentage points from the baseline. (c) Discussion and repair rates for ConPAct-R and its data ablations. The components together contribute to the effectiveness of ConPAct.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Game
Target for target_visible
Threat for hazard_nearby
bigfish
At least one visibly smaller fish
A larger fish
chaser
At least one uncollected green orb
A non-vulnerable enemy
coinrun
The goal coin
A saw, enemy, or deadly gap
jumper
The carrot
Spikes
maze
The cheese
Not included in the schema
miner
At least one diamond
A boulder or diamond that could fall onto the player
Appendix
Table 3: Game-specific referents for Procgen assertions. Maze has no hazard_nearby field.
Environment
Source instances
Evaluation instances
Evaluation rollouts
A
C
Sokoban
100
200
1,600
60
60
Crafter
100
200
1,600
160
160
Procgen
100
140
1,120
160
160
MiniGrid / BabyAI
108
150
1,200
60–200
60–200
Appendix
Table 4: Environment splits and evaluation budgets. A limits executed actions and C limits total logical model calls per rollout. Each instance has eight evaluation rollouts. MiniGrid budgets are specified per task in Table 6 .
Game
Source levels
Evaluation levels
Rmin
Rmax
bigfish
14
20
1
40
chaser
14
20
0.5
13
coinrun
15
20
5
10
jumper
14
20
3
10
maze
14
20
5
10
miner
15
20
1.5
13
Appendix
Table 5: Procgen task allocation and easy-mode normalization bounds from Cobbe et al. [2020] . Each game has 20 evaluation levels with IDs 100–119. Source levels are selected from IDs 0–49.
Task ID
Instances
Seed offsets
A=C
BabyAI-BlockedUnlockPickup-v0
5
0, 1, 2, 3, 4
136
BabyAI-GoTo-v0
5
0, 1, 2, 3, 4
200
BabyAI-GoToDoor-v0
5
0, 1, 2, 3, 4
84
BabyAI-GoToLocal-v0
5
0, 1, 2, 3, 4
60
BabyAI-GoToObj-v0
5
0, 1, 2, 3, 4
60
BabyAI-GoToRedBall-v0
5
0, 1, 2, 3, 4
60
Appendix
Table 6: MiniGrid and BabyAI evaluation tasks. Seed offsets are added to 1000; each resulting instance has eight rollouts. The last column gives the common action and model-call limit for each rollout.
3 epochs of ReAct SFT; token-matched across methods
Training seed
42
Observation window
3 frames, including the current observation
Optimizer
AdamW (fused)
Learning rate
1×10−5
Appendix
Table 7: Supervised fine-tuning settings shared by the training methods. Token budgets are matched to the three-epoch ReAct SFT reference for each environment and model size.
Method
Supervised targets
ReAct / Plan-Act / ECoT
Complete valid responses from retained demonstrations or progress prefixes.
PAA
The planner’s annotated plan and the actor’s action-only responses.
HSL
Expert responses and goal-relevant actions from hindsight examples; irrelevant actions are excluded.
WebSTAR
Original responses for actions with judge scores greater than 5; other responses are excluded.
ConPAct-S
Valid responses from consistent segments and successful corrections; conflicting replies remain context only.
ConPAct-R
Validated responses under the same consistency rule; hypothetical errors and recorded nonoptimal actions are excluded from targets.
Appendix
Table 8: Response selection for supervised training. Loss is applied only to the retained target response.
Method
Planner input
Actor input
ReAct / HSL / WebSTAR
No separate planner
Up to 3 recent frames
Plan-Act / Plan-Act SFT
Current frame plus up to 2 sub-goal-start frames
Up to 3 recent frames
PAA
Initial frame; called once
Up to 3 recent frames
ECoT
Current frame plus up to 2 sub-goal-start frames
Up to 3 recent frames
HiPlan
Initial frame for the global plan; up to 3 recent frames for hints
Up to 3 recent frames
TAPE
Up to 3 recent frames for visual calls
Up to 3 recent frames
Appendix
Table 9: The common three-frame observation limit. A planner called only at initialization receives the single available frame. Textual task instructions and active plans are retained separately.
Figure 7: Paired Sokoban observations for the initial state of level Stage_1_5_2387 . The archived screenshot and the text generated from the recorded board show the same two boxes, two targets, player position, and wall layout. Neither box is on a target.
Model
Method
Chrome
GIMP
Calc
Impress
Writer
Multi- apps
OS
Thunder- bird
VLC
VS Code
Overall
Opus-5
ReAct
71.0/69.3
87.2/87.2
80.6/80.6
83.9/82.3
89.2/85.6
81.4/75.0
82.7/82.7
83.4/83.4
94.1/90.2
82.0/82.0
82.0/79.5
Plan-Act
70.3/68.5
88.0/88.0
81.9/81.9
84.8/83.3
88.7/84.9
81.9/75.7
81.7/81.7
84.1/84.1
93.4/89.1
80.0/80.0
82.2/79.7
[2.5pt][2.5pt]
ConPAct-I
71.7/70.1
86.5/86.5
82.3/82.3
85.3/83.8
90.3/87.1
82.6/76.7
81.5/81.5
82.7/82.7
94.6/91.1
84.4/84.4
82.9/80.6
GPT-5.6-sol
ReAct
82.6/82.6
80.8/80.8
78.7/78.7
83.0/83.0
78.2/69.6
83.5/78.5
87.5/87.5
86.7/86.7
86.8/ 82.4
91.3/91.3
83.2/81.2
Plan-Act
84.8/84.8
76.9/76.9
78.7/78.7
85.1/85.1
77.3/69.6
83.6/78.5
83.3/83.3
86.7/86.7
87.1/ 82.4
87.0/87.0
82.9/80.9
[2.5pt][2.5pt]
ConPAct-I
82.6/82.6
76.9/76.9
80.9/80.9
85.1/85.1
81.7/73.9
85.9/80.7
83.3/83.3
86.7/86.7
88.2/82.4
91.3/91.3
84.1/82.0
Appendix
Table 17: Full OSWorld results, averaged over three rollouts per task. Each entry is Score/SR (%; ↑ ), rounded to one decimal place. Bold marks the best value for each metric within each model and category, including ties, using the values before rounding.
Configuration
Correction Trigger
✗
✓
✓
✓
Conflict FeedBack
✗
✗
✓
✓
Planner Participation
✗
✓
✗
✓
GPT-5.6-sol → GPT-5.6-sol
Sokoban
avg@8
85.0
90.8
89.8
92.1
pass@8
96.0
97.5
97.0
98.0
pass 8
71.0
81.5
78.0
82.5
Appendix
Table 18: Effects of online consistency ablations on task performance and planner–actor state mismatch.
Method
Sokoban
MiniGrid
avg@8
pass@8
pass 8
SR
Return
Steps
GPT-5.6-sol
ReAct
83.8
97.0
59.5
54.5
0.50
10.9
Self-Refine
86.5
93.5
73.5
51.2
0.49
7.2
Plan-Act
79.4
91.5
52.5
53.1
0.46
13.1
Assertion
82.8
92.5
55.5
57.7
0.51
13.3
Appendix
Table 19: Inference controls on Sokoban and MiniGrid.
Environment
Metric
Baseline
No Sup.
No Replay
Full
Qwen3-VL-8B → Qwen3-VL-8B
Sokoban
avg@8
40.7
44.4
43.5
49.1
pass@8
87.5
89.0
91.0
92.5
pass8
6.5
9.0
5.5
11.0
State mismatch
46.7
34.7
39.8
12.9
MiniGrid
SR
18.0
20.0
11.6
24.8
Appendix
Table 20: Training-data ablations on Sokoban and MiniGrid across the three planner–actor configurations. No Sup. and No Replay retain consistency filtering; Full uses all three components.
Environment
Metric
Baseline
ConPAct-R
Δ
Qwen3-VL-8B → Qwen3-VL-8B
Sokoban
Initial D
47.4
20.8
−26.6
Planner correctness
45.5
83.0
+37.5
Actor correctness
46.6
83.1
+36.5
MiniGrid
Initial D
51.3
29.9
−21.4
Planner correctness
50.9
81.5
+30.6
Appendix
Table 21: Initial state mismatch and assertion correctness before discussion. Rates are percentages and changes are in percentage points.
Training data
Sokoban
MiniGrid
Sup.
Replay
Discuss.
Init. D
Resid. D
Discuss.
Init. D
Resid. D
Qwen3-VL-8B
✗
✓
8.3
39.1
92.4
7.2
42.7
93.2
✓
✗
21.7
45.9
90.1
19.4
42.2
87.2
[][]✓
✓
58.4
20.8
62.0
54.1
29.9
78.9
Qwen3-VL-32B
Appendix
Table 22: Discussion responses and planner–actor state mismatch (%).
Training data
Sokoban
MiniGrid
Sup.
Replay
CC ±Δ
CW ±Δ
D ±Δ
CC ±Δ
CW ±Δ
D ±Δ
Qwen3-VL-8B
✗
✓
40.7+1.5
19.8+1.5
39.5−3.0
38.4+1.6
17.9+1.5
43.7−3.1
✓
✗
28.4+2.4
28.0+2.0
43.6−4.4
36.2+3.3
17.1+2.7
46.7−6.0
[][]✓
✓
72.6+5.1
6.9+2.8
20.5−7.9
67.0+4.1
4.0+2.2
29.0−6.3
Qwen3-VL-32B
Appendix
Table 23: State-assertion consistency and correctness: initial values and signed changes after discussion.
Figure 8: Sokoban: state mismatch and ConPAct coordination. Condensed recorded outputs for the same positioning sub-goal.
Figure 9: MiniGrid: state mismatch and ConPAct coordination. Condensed recorded outputs for the same gap-alignment sub-goal.
Figure 10: Crafter: state mismatch and ConPAct coordination. Condensed recorded outputs for the same second-wood sub-goal.
Figure 11: Procgen Miner: state mismatch and ConPAct coordination. Condensed recorded outputs for the same one-cell positioning sub-goal.
Method
Calls / episode
Output tokens / episode
Calls / action
Usage coverage (%)
ReAct (Planner)
14.02
3,095.37
1.04
71.5
ReAct (Actor)
36.29
8,703.19
1.00
71.5
Plan-Act
35.36
5,285.03
1.06
100.0
HiPlan
51.86
7,992.10
2.10
100.0
TAPE
37.84
18,900.17
6.29
99.9
[][] ConPAct-I
33.02
7,199.86
1.43
83.3
Appendix
Table 24: Inference costs averaged across the four environments and the evaluated inference-time configurations. Usage coverage is the fraction of recorded rollouts with output-token usage.
Department of Electronic Engineering, Tsinghua University · Qiuzhen College, Tsinghua University · Institute of Microelectronics, Chinese Academy of Sciences +3