Humans solve complex problems by constructing plans and mentally simulating their outcomes with an internal model of the world. Machine learning has made substantial progress in learning world models that predict the consequences of action sequences, yet the procedures used to plan with these models remain largely hand-designed. Most planners rely on fixed search or optimization rules; approaches that learn aspects of search typically imitate a predefined optimizer or use planning to inform an amortized policy, rather improving multi-step plans. We introduce \textbf{Reinforced Planning}, a method that learns the plan-update itself by reinforcing update rules that produce better plans, using gradients propagated through a differentiable world model. We instantiate Reinforced Planning in RP1, which learns a critic over imagined outcomes via temporal-difference learning and a neural plan-improvement operator trained via imagined rollouts with a pretrained world model. RP1 can be trained fully offline without environment interaction; environment episodes are used only for checkpoint selection. Across visual navigation, arm reaching, and robotic manipulation on two world-model backbones, RP1 matches or exceeds existing planners, achieving near-perfect success in several settings while using 1,000× fewer world-model rollouts than the strongest alternative (CEM) and planning up to 67× faster under concurrent planners inference.
Figures & tables
Figure 2: Experiment environments. In Reacher (left) the agent moves a two-link arm to a goal configuration, here shaded. In OGBench Cube (middle) a robot arm picks up a cube and moves it to a goal position, also shaded. In TwoRoom (right) an agent navigates to a goal position, (here the star).
Figure 3: RP1 in TwoRoom. (a) The learned critic better captures temporal cost-to-go than latent L2 distance. (b) RP1 iteratively refines its plan, with later updates focusing on fine corrections to the final actions. Additional visualizations are provided in Appendix D.2 .
LeWM
PLDM
planner
rollouts
25 steps
100 steps
25 steps
100 steps
latent L2 (value critic)
CEM
9000
84.0 ( 100.0 )
13.3 (73.0)
93.3 ( 100.0 )
52.0 (49.3)
MPPI
9000
70.7 (95.3)
20.0 (68.0)
64.0 (95.7)
33.3 (53.3)
Adam
3000
94.7 (96.3)
24.0 (68.0)
90.7 (94.0)
42.0 (53.0)
DMPO
256
73.3
69.3
76.7
44.0
TwoRoom
LeWM
PLDM
planner
rollouts
τ=.1
τ=.05
τ=.1
τ=.05
latent L2 (value critic)
CEM
9000
98.7 (88.0)
80.3 (63.3)
96.7 (75.3)
80.0 (56.3)
MPPI
9000
90.0 (89.0)
69.3 (70.3)
89.3 (86.0)
63.3 (65.3)
Adam
3000
94.0 (92.7)
66.0 (71.7)
94.3 (92.7)
66.0 (68.7)
DMPO
256
40.7
20.7
36.0
19.3
Reacher
LeWM
PLDM
25 steps
100 steps
25 steps
100 steps
planner (roll.)
easy
hard
easy
hard
easy
hard
easy
hard
latent L2 (value-critic)
CEM 9000
74.0 (80.0)
40.9 (54.5)
58.0 (77.7)
23.2 (59.2)
62.7 (75.3)
15.2 (43.9)
58.7 (69.3)
24.5 (43.8)
MPPI 9000
56.7 (67.0)
1.6 (25.0)
46.0 (58.3)
1.2 (23.7)
60.0 (65.3)
9.1 (21.1)
47.3 (53.0)
3.6 (14.0)
Adam 3000
74.0 (77.3)
40.9 (48.4)
57.3 (70.7)
21.9 (46.4)
63.3 (66.0)
16.6 (22.7)
55.3 (57.0)
18.2 (21.3)
OGBench Cube.
LeWM
PLDM
easy
hard
easy
hard
RP1 (pretrained)
92.7
83.4
85.3
66.6
RP1 Dyna (finetuned)
95.0
88.6
93.3
84.8
Table 4: Effect of Dyna at h25 (%). Rollouts are collected at both goal horizons by the deployed planners and the locked configuration is retrained on the finetuned model; we run two rounds and report the one chosen on validation (round 1 for LeWM, round 2 for PLDM)..
Figure 4: End-to-end planning latency. On OGBench Cube with LeWM (one NVIDIA H200, fp32), with RP1 13× faster than CEM for one planner and 67× faster for 50 concurrent planners.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
25 steps
100 steps
domain
base
critic
planner
critic
planner
TwoRoom
LeWM
15
2
6
8
PLDM
3
2
3
10
Reacher
LeWM
18
2
—
—
PLDM
18
2
—
—
Cube
LeWM
3
12
9
2
Appendix
Table 5: Selected stopping points . The offline critic is snapshotted every 3 k to its 18 k cap and the planner every 2 k to its 12 k cap; one of the 6×6 pairs is fixed per cell as described above. Reacher is evaluated at h25 only. A planner entry of 12 is the step cap, i.e. the final iterate.
Hyperparameter
Value
Hyperparameter
Value
Plan refiner
Co-trained critic
refinement iterations K
8
γ ( h25 / h100 )
0.98 / 0.99
plan horizon H (chunks)
5
n -step
1
replan every
5 chunks
expectile (annealed)
0.1→0.03
clip range amax
2.5
critic LR (annealed)
10−3→10−4
refiner MLP (width × hidden layers)
512×2
live fraction of actor steps
0.8
Appendix
Table 6: The RP1 configuration.
Figure 5: Latent-Distance vs. learned Cost-to-go. 3 randomly sampled tasks from seeds (42,43,44) and their corresponding cost landscapes.
Figure 6: Plan refinement in TwoRoom. Each panel shows the planner’s best-scoring candidate plan at refinement iteration k on the same task, with all planners scoring plans using the learned value critic; “final” is k=30 for Adam and CEM and k=8 for RP1. Iterations differ greatly in cost: one CEM iteration evaluates 300 sampled plans ( 9,000 rollouts in total), one Adam iteration takes a gradient step on 100 plans in parallel ( 3,000 rollouts), whereas one RP1 iteration is a single forward pass of the learned refiner costing one rollout ( 9 in total, including the initial evaluation).
Planning with a learned latent world model is a promising route to control from raw pixels, but a strong world model alone is not enough. We show this experimentally: even with a perfect world model (operationalized by replacing the learned forward predictor with an idealized rollout of the true environment dynamics), a finite-budget sample-based planner still fails on some tasks, indicating that the bottleneck can lie in search rather than in world-model accuracy. Motivated by this gap, we propose IMWM (Intuition Model + World Model), which pairs the world model with an intuition model trained from demonstrations to recognize promising actions. The two models collaborate through three lightweight components: (i) Retrieval Initialization, which initializes the planner's action proposal from a retrieved demonstration; (ii) Hybrid Cost, which combines the intuition score with the world-model rollout cost; and (iii) a Reliability Gate, which adjusts how much the planner trusts intuition in each setting. Across four pixel-based goal-reaching tasks (Two-Room, Reacher, Push-T, and OGBench-Cube), IMWM has higher mean success than the world-model-only planner on all four, with the largest gains on Two-Room (99.2%, +11.5 percentage points) and OGBench-Cube (94.7%, +28.5 percentage points).
Baoqi Gao, Ruize Han, Miao Wang +1
Beihang University · Shenzhen University of Advanced Technology
Neural world models coupled with model predictive control (MPC) replan at every environment step to bound accumulated prediction error, but this incurs substantial computational overhead. Reusing a cached plan reduces this overhead, yet its effectiveness depends on how prediction mismatch propagates through the local dynamics. We analyze this trade-off with a perturbation-based dynamic-regret framework and show that stale-plan penalties scale with the reuse tolerance, the accumulated mismatch since the last replanning step, and the local dynamics sensitivity. Based on this structure, we propose AdaReP, a training-free wrapper that adapts the replanning tolerance online using the current deviation from the cached rollout and a local sensitivity estimate, without modifying the learned world model or planner. Across image-space planning, latent-space control, and real-world robotic manipulation, AdaReP substantially reduces planner-side computation while maintaining comparable task performance, including over 80% fewer queries on a 50-trial physical robot study.
Yutian Cheng, Xiaojian Ma, Xianhao Wang +6
Shanghai Jiao Tong University · Beijing Institute for General Artificial Intelligence (BIGAI) · University of Science and Technology of China
Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models. For visual planning, however, LeWM evaluates candidate action sequences by repeatedly applying a local one-step latent transition model. This autoregressive rollout makes planning computationally expensive and exposes the predicted trajectory to accumulated latent errors as the horizon grows. We propose Fast LeWorldModel (Fast-LeWM), a fast latent world model that replaces repeated local rollout with action-prefix prediction. Given the current latent and a candidate action sequence, Fast-LeWM encodes its prefixes and predicts the future latents reached after executing those prefixes in parallel. By making action prefixes the basic prediction unit, Fast-LeWM directly models action effects accumulated to different extents over multiple horizons. This prefix-level supervision forces the model to learn how states continuously evolve under different action prefixes, rather than only fitting one-step state transitions. During planning, the predictor can use the prefix token from the encoded action sequence to evaluate the corresponding future latent without explicitly rolling through each intermediate imagined state. Across multiple tasks, Fast-LeWM improves average success over LeWM while substantially reducing planning time, achieving lower open-loop latent loss whose growth becomes significantly slower as the rollout horizon increases.