Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without policy-improvement training? Our random-walk analysis explains what temporal relations contain: short horizons reveal geodesic geometry in the diffusion limit, while longer horizons reveal connectivity between regions before mixing removes these distinctions. We learn these relations with a conditional energy-based model that estimates temporal log-density ratios through horizon-conditioned embeddings. The model is trained on observation pairs by noise-contrastive estimation, without action or reward labels. The planner queries these learned relations at different horizons as it moves toward the goal. At test time, a separate local dynamics model predicts candidate action outcomes, and the temporal model evaluates their progress toward the goal by selecting or aggregating estimated improvements across horizons. The agent executes one action and replans with both models fixed. Experiments demonstrate long-range maze planning from random exploration using states and images. Learned score fields, embedding probes, and planned routes exhibit properties of a multiscale cognitive map. We further demonstrate egocentric navigation from random exploration and manipulation planning from suboptimal data.
Figures & tables
Figure 1: Emergent cognitive maps from random exploration. (A) Score Gθ(x,yg,τ) over positions x in Large, with goal yg fixed (star). Short horizons ( τ=16 ) peak near the goal, while long horizons ( τ=512 ) spread through connecting passages. (B) t-SNE of hθ(x,τ) in Giant at τ=1 and 4096 . At τ=1 , embeddings preserve local geodesic geometry, while at τ=4096 , they reflect connectivity between distant regions. 3D views appear in Fig. 5 . (C) GP and PAP routes through Large and Giant. Colour denotes the selected horizon for GP and the dominant score-contributing horizon for PAP. Matching markers pair starts ( ∘ ) and goals ( ⋆ ).
Official tasks
Random pairs
Navigate
Random Data
Random Data
Large
Giant
Large
Giant
Large
Giant
Method
SR ↑
SR ↑
SR ↑
SPL ↑
SR ↑
SPL ↑
SR ↑
SPL ↑
SR ↑
SPL ↑
Random policy (floor)
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
Shortest-path oracle
1.00
1.00
1.00
0.94
1.00
0.98
1.00
0.95
1.00
0.95
GCIQL ( Kostrikov et al., 2022 )
0.34
0.00
0.00
0.00
0.00
0.00
0.15
0.15
0.07
0.06
Table 1: PointMaze goal reaching from states (A) and images (B) on official tasks and random start–goal pairs. The GCIQL, QRL, and HIQL Navigate results are quoted from Park et al. (2025) .
Figure 2: Official-task success versus maximum planning horizon. Color distinguishes Large and Giant and line style distinguishes GP and PAP. Lines show three-seed means and bands one standard deviation. Planning uses every integer horizon from 1 to τmax .
Table 4
Figure 3: An egocentric plan in Habitat Apartment 1 by the model that plans from the image and the position (horizons capped at 512, no constraint on consecutive headings). Left: the route, coloured by the selected horizon, from the start ( ∘ ) to the goal ( ⋆ ). Right: the agent’s view at ten decisions and the horizon selected there.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Learned embeddings h(x,τ) on the Large (top) and Giant (bottom) state mazes; the left column colours each cell by its position and the other columns keep those colours. Every panel is rotated and reflected onto the maze by Procrustes, since t-SNE leaves the orientation undetermined.
τ
1
16
256
4096
near
0.95
0.91
0.58
0.57
far
− 0.01
0.88
0.97
0.70
Appendix
Table 4: Spearman correlation between embedding distance and geodesic distance for near ( ≤8 steps) and far pairs, Large maze.
Figure 5: 3D t-SNE of state embeddings on Large (top) and Giant (bottom). At τ=1 , 3D views reveal connectivity obscured in 2D. Colours match the maze regions at left; axes are not physical coordinates.
Figure 6: Three held-out cube-single episodes planned from pixels, one per row, each showing eight evenly spaced executed steps with the goal image last. The cube travels 0.455, 0.277, and 0.325 m, respectively; in every row the planner reaches the cube, grasps it, and carries it to the goal.
Latent world models plan by rolling a frozen predictor forward under candidate action sequences and ranking the candidates by the latent distance between their imagined end state and the goal. However, this ranking breaks down when the goal lies several plans away, because the latent distance measures how closely an end state resembles the goal rather than how far it remains from reaching it. To address this, we propose TEMPO, a temporal-distance planning objective that leaves the world model untouched, learns only from the demonstrations already used to train it, and adds negligible cost to the planner's search. TEMPO learns a small map of the frozen latent in which the distance between two states of an episode reflects the number of environment steps between them, and blends this distance into the planner's cost. It requires no rewards, policies or success labels and, being a cost rather than a model, applies to frozen world models with one latent vector per state that plan by a latent distance. We evaluate TEMPO on eleven simulated environments (e.g., maze navigation, tabletop pushing, robotic arm control and three-dimensional manipulation) with the LeWM and PLDM planners. With a small MLP that adds at most 0.3% to a plan's arithmetic, TEMPO improves both planners at every goal distance, including the one-plan setting of their evaluations, raises LeWM from 36% to 99% on TwoRoom three plans from the goal, and remains competitive on a broad range of 2D and 3D navigation, reaching and manipulation tasks.
Compositional diffusion models offer a promising route to long-horizon planning by denoising multiple overlapping sub-trajectories while ensuring that together they constitute a global solution. However, enforcing local behavior over long chains is often insufficient for a coherent global structure to emerge. Recent works tackle this limitation through intrinsic search, which explores multiple paths during the denoising process. While intrinsic search improves global coherence, it comes at the cost of repeated evaluations of an already compute-heavy model. In this work, we argue that extrinsic search, performed outside the denoising process, offers a more effective mode of exploration for long-horizon planning while naturally enabling the use of classical algorithms to solve unseen combinatorial tasks at test time. Our eXtrinsic search-guided Diffuser (XDiffuser) first computes a plan over a state-space graph -- serving as a lightweight local connectivity oracle for the diffusion model. The plan is then used to guide denoising for a single trajectory, effectively offloading the burden of exploration. XDiffuser outperforms diffusion-based baselines on long-horizon tasks, with particularly large gains in the low-quality data regime and on unseen tasks beyond goal-reaching, including multi-agent coordination and TSP-style reasoning. Project website: https://yanivhass.github.io/XDiffuser-site/
Humans solve complex problems by constructing plans and mentally simulating their outcomes with an internal model of the world. Machine learning has made substantial progress in learning world models that predict the consequences of action sequences, yet the procedures used to plan with these models remain largely hand-designed. Most planners rely on fixed search or optimization rules; approaches that learn aspects of search typically imitate a predefined optimizer or use planning to inform an amortized policy, rather improving multi-step plans. We introduce \textbf{Reinforced Planning}, a method that learns the plan-update itself by reinforcing update rules that produce better plans, using gradients propagated through a differentiable world model. We instantiate Reinforced Planning in RP1, which learns a critic over imagined outcomes via temporal-difference learning and a neural plan-improvement operator trained via imagined rollouts with a pretrained world model. RP1 can be trained fully offline without environment interaction; environment episodes are used only for checkpoint selection. Across visual navigation, arm reaching, and robotic manipulation on two world-model backbones, RP1 matches or exceeds existing planners, achieving near-perfect success in several settings while using 1,000× fewer world-model rollouts than the strongest alternative (CEM) and planning up to 67× faster under concurrent planners inference.