cs.LGAug 19, 2026

Reinforced Planning with Latent World Models

Authors: Armin Sommer, Jannik Schilling

Organizations: Pantheon Industries

Abstract

Humans solve complex problems by constructing plans and mentally simulating their outcomes with an internal model of the world. Machine learning has made substantial progress in learning world models that predict the consequences of action sequences, yet the procedures used to plan with these models remain largely hand-designed. Most planners rely on fixed search or optimization rules; approaches that learn aspects of search typically imitate a predefined optimizer or use planning to inform an amortized policy, rather improving multi-step plans. We introduce \textbf{Reinforced Planning}, a method that learns the plan-update itself by reinforcing update rules that produce better plans, using gradients propagated through a differentiable world model. We instantiate Reinforced Planning in RP1, which learns a critic over imagined outcomes via temporal-difference learning and a neural plan-improvement operator trained via imagined rollouts with a pretrained world model. RP1 can be trained fully offline without environment interaction; environment episodes are used only for checkpoint selection. Across visual navigation, arm reaching, and robotic manipulation on two world-model backbones, RP1 matches or exceeds existing planners, achieving near-perfect success in several settings while using 1,000×1{,}000\times fewer world-model rollouts than the strongest alternative (CEM) and planning up to 67×67\times faster under concurrent planners inference.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. IMWM: Intuition Models Complement World Models for Latent Planning

    Jun 1, 2026Baoqi Gao, Ruize Han, Miao Wang +1World Model PlanningLatent World Models

  2. AdaReP:Adaptive Re-Planning under Model Mismatch for Neural World-Model Predictive Control

    Jun 22, 2026Yutian Cheng, Xiaojian Ma, Xianhao Wang +6Model Predictive ControlMismatch

  3. Fast LeWorldModel

    Jun 24, 2026Yuntian Gao, Xiangyu XuLatent World ModelsWorld Models