World models are becoming central to robotic planning and control by predicting future state transitions. Existing approaches mainly rely on visual, latent, or language prediction, which can be difficult to ground in executable robot actions and prone to compounding errors over long horizons. In contrast, traditional robotic task and motion planning enables structured long-horizon reasoning through compact symbolic representations of world transitions, but typically lacks synchronized visual prediction. We propose Hierarchical World Model (H-WM), which jointly predicts logical and visual state transitions by combining a high-level logical world model with a low-level visual world model. The predicted logical actions and latent visual state transitions are jointly incorporated into Vision-Language-Action (VLA) models as intermediate state guidance for long-horizon task execution. Experiments on three long-horizon benchmarks and real robots show that H-WM consistently improves VLA's performance by stabilizing long-horizon execution and mitigating error accumulation. We also construct LIBERO-Logic, a frame-level aligned dataset that pairs visual observations and continuous robot states with logical actions and predicate-based logical states.
Figures & tables
Fig. 1: Overview of H-WM and guided VLA execution. H-WM jointly models logical and visual transitions at the subtask level m , while the VLA performs low-level control at the higher-frequency time step t . The logical world model searches and evaluates candidate transitions to predict the next logical action am+1 and logical state Xm+1 . Conditioned on the current observation obsm , robot configuration qm , am+1 , and Xm+1 , the visual world model uses an understanding expert and a prediction expert to generate the latent visual subgoal fpredm+1 . During execution, the understanding expert encodes obst and am+1 , while the goal expert processes fpredm+1 . The action expert takes qt and noisy action tokens, attends to both representations, and generates the low-level action chunk α^mt:t+k , grounding long-horizon logical guidance into visually conditioned robot execution.
Methods
Tasks Q-score / Task Success Rate
Task 1
Task 2
Task 3
Task 4
Task 5
Average
H-WM- π0.5 (ours)
98.0/94.0
86.7/60.0
74.0/46.0
70.7 / 42.0
95.0/82.0
84.9/64.8
H-WM-Stable-Diffusion- π0.5
92.7/82.0
80.0/46.0
65.3/ 30.0
74.0 / 34.0
93.5/80.0
81.1/54.4
Logic-guided π0.5
95.3/86.0
84.7/58.0
54.7/16.0
39.3/4.0
92.0/78.0
73.2/48.4
LLM-guided π0.5
84.7/54.0
80.7/42.0
68.0 /24.0
41.3/4.0
59.5/10.0
66.8/26.8
π0.5
66.0/4.0
73.3/24.0
54.7/4.0
44.7/0.0
38.0/0.0
55.3/6.4
TABLE I: Performance on the LIBERO-LoHo benchmark. For each task, we report Q-Score and Success Rate (in %). The best results are shown in bold , and the second-best results are underlined .
Fig. 2: (a) Evaluation results of H-WM-guided π0.5 and various VLA baselines on the LIBERO-10 benchmark. (b) Evaluation results of H-WM-guided π0.5 and baseline methods on the RoboCerebra benchmark.
Fig. 3: (a) Real-world experiment with UR5e robot evaluating H-WM guided π0.5 on long-horizon manipulation task. (b) Step-wise success rates in real-robot experiments.