Organizations: Nanjing University · Shenzhen University of Advanced Technology · Shanghai Jiao Tong University · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
Latent world models often struggle with long-horizon planning despite accurate short-term predictions. Recursive rollouts accumulate errors, while distance concentration in high-dimensional latent spaces can weaken goal discrimination. We introduce the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning through distinct state representations and dynamics models. The low-level model predicts action-conditioned transitions, while the high-level model uses learned macro-actions to plan over longer temporal spans. We also propose Long-Horizon Representation Learning with Weighted Rollout (LoRe), which supervises self-generated predictions at both levels. An analysis of recursive error propagation motivates exponential horizon weights with separate decay rates for the two temporal scales. During planning, the high-level model generates latent subgoals that the low-level model refines into actions for precise execution. We evaluate from-scratch Dual-WM on five goal-conditioned visual control tasks against the task-wise strongest baselines without actor-guided proposals. At goal offsets of 50 and 100 environment steps, mean success increases from 75.9% to 84.4% and from 61.4% to 69.5%, respectively. At offset 100, Dual-WM outperforms these baselines on all five tasks and improves mean success over LeWM by 30.8 percentage points. Ablations and supporting analyses provide evidence of more informative representations for goal evaluation and greater consistency under recursive prediction. These results highlight the value of separating temporal roles and training across multiple horizons for reliable latent planning. Our core implementation is available at https://github.com/DeLin1001/Dual-WM-Official.
Figures & tables
Figure 1: Overview of the long-horizon planning problem and our dual-latent solution. (a) Two failure modes in a single latent space: goal costs collapse for distant states, and recursive rollouts drift from the ground-truth future. (b) Latent-space diagnostics: the high-level space preserves task-distance contrast, while weighted rollout training reduces multi-step prediction error relative to one-step teacher forcing. (c) Dual-WM assigns coarse route planning to a high-level latent and precise local execution to a low-level latent.
Figure 2: Training the dual-timescale latent world model. The low-level encoder EL and predictor PL model fine-scale transitions, while EH encodes length- k low-level latent windows and PH models temporally extended transitions. The action encoder EA produces macro-actions, MAPS regularizes their posterior, and LoRe supervises self-generated predictions at both levels.
Figure 3: Three-stage planning with the trained dual-latent world model. Stage 1 plans a high-level route and produces latent subgoals; Stage 2 refines subgoals with low-level actions while comparing latent windows through EH ; Stage 3 performs direct low-level goal convergence.
Offset 50
Offset 100
Task
LeWM
Best other
Dual-WM
LeWM
Best other
Dual-WM
TwoRoom
61.50
94.83
(Gemini)
99.33
27.00
90.00
(Gemini)
92.50
Reacher
85.33
92.33
(INTACT)
98.67
75.67
84.83
(Gemini)
93.30
Sokoban-Long
41.50
52.00
(Gemini)
75.17
16.33
35.00
(Gemini)
52.33
PushT
56.00
78.00
(HWM)
86.17
20.00
31.00
(HWM)
41.33
Cube-Single
51.33
62.17
(DINO-WM)
62.67
54.17
66.33
(Fast-LeWM)
67.83
Table 1: Success (%) at goal offsets 50 and 100 environment steps. Dual-WM uses the from-scratch variant throughout. Best other selects the highest non-Dual-WM mean without actor-guided proposals separately for each task and offset, with the selected method shown beside each value. The external controller is included. Bold marks the largest displayed mean per task/offset. Full results and uncertainty are in Tables 5 – 7 .
Figure 5
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Method
State spaces
Dynamics trained
Prediction training
LeWM
Single
Primitive
One-step transition
HWM
Shared across scales
Primitive + macro
Uniform-weight rollout
Hi-LeWM
Shared; low frozen
Macro only
One-step waypoint
Dual-WM
Separate across scales
Primitive + macro
Scale-specific decayed rollout
Appendix
Table 2: LeWM and hierarchical latent world models with learned macro-actions.
Setting
TwoRoom
Reacher
Sokoban-Long
PushT
Cube-Single
Episodes
10,000
10,000
2,500
18,685
10,000
Environment steps
921,000
2,010,000
395,000
2,337,000
2,010,000
f
5
5
1
5
5
k
3
3
2
2
2
N/M
5/6
3/6
10/8
3/6
6/6
αL/αH
0.5/0.5
0.2/0.5
0.1/0.5
0.2/0.5
0.05/0.5
Appendix
Table 3: Task-specific training configurations. Paired entries are low/high level. Environment-step counts are approximate dataset totals; N,M count model transitions at their respective scales.
TwoRoom
Reacher
Sokoban-Long
PushT
Cube-Single
HL
8
5
12
5
5
HH
6
6
8
6
6
Appendix
Table 4: Planning horizons. HL counts low-level model steps and HH counts high-level macro-steps.
Figure 6: The evaluation environments: TwoRoom, Reacher, Sokoban-Long, PushT and Cube-Single (left to right).
Method
TwoRoom
Reacher
25
50
100
25
50
100
LeWM
87.0
61.50±1.22†
27.00±0.82†
86.0
85.33±0.62†
75.67±0.62†
DINO-WM
100.0
60.50±0.71†
32.17±0.94†
79.0
67.00±0.71†
34.17±1.03†
HWM
100 †
73.67±1.03†
48.33±0.94†
–
–
–
PLDM
97
–
–
78
–
–
Fast-LeWM
98
67.17±1.25†
38.67±0.62†
90.0
81.33±0.94†
79.33±0.62†
Appendix
Table 5: Goal-reaching success rates (%) on TwoRoom (left) and Reacher (right); columns are goal offsets measured in original environment steps. † denotes results reproduced by us.
Method
25
50
100
LeWM
68.17±0.62†
41.50±0.71†
16.33±0.47†
Gemini
69.67±0.85†
52.00±0.82†
35.00±0.41†
Dual-WM
From scratch
87.83±0.62
75.17±0.85
52.33±0.94
VFM adaptor
83.67±0.24
68.83±0.85
48.00±0.41
Low-level only
87.00±0.00
70.17±0.24
45.83±0.85
Appendix
Table 6: Goal-reaching success rates (%) on Sokoban-Long; columns are goal offsets measured in original environment steps. † denotes results reproduced by us.
Method
PushT
Cube-Single
25
50
75
100
25
50
100
LeWM
96.0
56.00±0.71†
45.67±0.62†
20.00±0.82†
74.0
51.33±1.03†
54.17±0.62†
DINO-WM
74.0
55.33±0.62†
38.17±0.85†
12.00±0.71†
86.0
62.17±0.24†
62.33±0.47†
PLDM
78
–
–
–
65
–
–
HWM
89
78
61
31.00±1.41†
–
–
–
Hi-LeWM
90.7±6.2
42.0±6.8
15.3±4.1
–
–
–
–
Appendix
Table 7: Goal-reaching success rates (%) on PushT (left) and Cube-Single (right); columns are goal offsets measured in original environment steps. PushT carries the offset- 75 column because HWM reports its long-horizon result at offset 75 . On Cube-Single, Dual-WM (with INTACT actor) adds the INTACT actor as an action prior for low-level planning; all other Dual-WM rows use plain CEM. † denotes results reproduced by us.
Variant
25
50
100
INTACT
Pure CEM
68.44±0.77
53.50±1.08
61.17±1.31
Actor+CEM
96.89±0.19
57.17±1.03
80.67±0.24
Dual-WM
Pure CEM
74.00±0.41
62.67±0.62
67.83±0.47
With INTACT actor
97.33±0.24
58.50±0.41
82.50±0.71
Appendix
Table 8: Cube-Single success (%) grouped by method, with and without actor-guided proposals. Columns are goal offsets in environment steps. INTACT’s offset- 25 results are quoted; the remaining entries are our evaluations. Bold marks a Dual-WM variant only when it attains the highest mean in the column.
Figure 7: Success–time tradeoff on TwoRoom at goal offset 100 and an execution budget of 125 environment steps. Each method has four measured planning-budget settings; lines connect successive settings. Error bars show ± one standard deviation across three planning seeds.
Method / level
Tier 1
Tier 2
Tier 3
Tier 4
Dual-WM high
30/10/7
40/10/8
80/10/16
100/10/20
Dual-WM low
60/30/7
80/30/8
160/30/16
192/30/20
Low-level only
60/30/7
96/30/9
192/30/20
256/30/25
LeWM
60/30/7
96/30/9
192/30/20
256/30/25
HWM high
30/10/7
40/10/8
80/10/12
80/10/12
HWM low
30/30/3
35/30/3
40/30/3
60/30/3
Appendix
Table 9: CEM settings for the TwoRoom planning-budget sweep. Each entry lists candidates / elites / iterations. Tiers correspond to successive points on each method’s curve in Figure E.2 .
Figure 8: TwoRoom execution. Dual-WM navigates through the doorway and approaches the goal in the other room; LeWM remains on the starting side of the wall.
Figure 9: PushT execution. Dual-WM changes the block configuration toward the sampled goal, while LeWM leaves it near its initial configuration. The green shape is a fixed renderer reference; the task goal is the rightmost image.
Figure 10: Reacher execution. Dual-WM brings the arm toward the goal configuration; LeWM moves the arm but retains a different configuration.
Figure 11: Cube-Single execution. Dual-WM brings the gripper and object toward the goal arrangement, whereas LeWM changes the gripper pose without completing the required object movement.
Figure 12: Sokoban-Long execution. Dual-WM navigates to a box and performs a sequence of aligned pushes toward the goal, while LeWM does not complete the required box movement. This supplemental held-out source pair lies outside the standard 200 -task sample.
Figure 13: Latent-distance concentration on held-out state pairs. We plot dimension-normalized Euclidean latent distance against A* endpoint path distance in TwoRoom (left) and wrapped joint-angle distance in Reacher (right). Curves are equal-width-bin means and shaded regions are 95% episode-bootstrap confidence intervals. The displayed task-relevant ranges are 0–150 px and 0–3 rad, respectively; annotated rank correlations use all held-out pairs. Only bins with at least 100 valid pairs are displayed. The dotted line marks the independent isotropic-Gaussian reference.
Figure 14: TwoRoom distance fields for three prespecified members of a geometry-only nine-anchor maximin set; stars mark the query states. Each row contains one A* reference followed by LeWM, Dual-WM-L and Dual-WM-H latent distances normalized by 2D . All A* panels share one path-length scale, and all model panels share one latent-distance scale. The wall-safe Gaussian bandwidth ( σ=10 px) is selected by blocked spatial cross-validation over all three models and all nine anchors, without using A* similarity or visual appearance.
Figure 16: PushT goal-geometry improvements relative to LeWM on 128 held-out episodes. Dots are paired differences for Dual-WM-L and Dual-WM-H; error bars are 95% episode-cluster bootstrap confidence intervals (2,000 draws). Positive values favor Dual-WM. Correlations use physical goal distance, while ordering accuracy measures which state in a same-episode pair is closer to the goal; exact physical ties are excluded.
Figure 17: Physical state decoded from recursively predicted low-level latents. A single horizon-pooled ridge probe is fitted separately for each checkpoint and physical target, selected on disjoint validation episodes, and evaluated on disjoint test episodes. Curves report mean task-native error and shaded regions show 95% episode-cluster bootstrap confidence intervals. The rollout-trained checkpoint has lower point estimates at all horizons; PushT orientation has the widest uncertainty.
Figure 18: Goal ordering on TwoRoom. Accuracy compares latent and A* goal-distance rankings; bins use mean goal distance. Dual-WM-H and Window-Concat use identical observation windows, with learned encoding and direct concatenation, respectively. Dashed curves are contextual references. Corresponding planning success is shown in Figure 5 .
Figure 19: Low-level training-horizon sensitivity on TwoRoom. Left: mean success using low-level-only planning at goal offsets 25 , 50 , and 100 environment steps. Right: mean position-probe error over recursive low-level predictions. Both panels compare N∈{1,3,5,8,10} ; the right horizontal axis is evaluation rollout length, not training length.
Figure 20: Rollout-weighting comparison on TwoRoom at training horizon N=5 . Bars show mean success across low-level-only planning runs. The difference between profiles grows as the goal offset increases.
Figure 21: Macro-action prior shaping on TwoRoom. Left: mean success at goal offset 100 . Middle and right: mean one-step and six-step latent MSE. Det. denotes deterministic macro-action encoding; numeric labels denote stochastic encoding with the indicated MAPS coefficient, including β=0 without the KL penalty. All panels show the same six configurations. Raw MSE is measured in each learned latent space.
Figure 22: Additional TwoRoom distance fields for the six remaining query states. Each row contains one A* reference followed by LeWM, Dual-WM-L and Dual-WM-H. The A* column and the three model columns use the same respective shared scales as Figure 15 .
State Key Lab of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University · Amazon AGI SF Lab · Institute for Artificial Intelligence, Peking University