World models are becoming central to robotic planning and control by predicting future state transitions. Existing approaches mainly rely on visual, latent, or language prediction, which can be difficult to ground in executable robot actions and prone to compounding errors over long horizons. In contrast, traditional robotic task and motion planning enables structured long-horizon reasoning through compact symbolic representations of world transitions, but typically lacks synchronized visual prediction. We propose Hierarchical World Model (H-WM), which jointly predicts logical and visual state transitions by combining a high-level logical world model with a low-level visual world model. The predicted logical actions and latent visual state transitions are jointly incorporated into Vision-Language-Action (VLA) models as intermediate state guidance for long-horizon task execution. Experiments on three long-horizon benchmarks and real robots show that H-WM consistently improves VLA's performance by stabilizing long-horizon execution and mitigating error accumulation. We also construct LIBERO-Logic, a frame-level aligned dataset that pairs visual observations and continuous robot states with logical actions and predicate-based logical states.
Figures & tables
Fig. 1: Overview of H-WM and guided VLA execution. H-WM jointly models logical and visual transitions at the subtask level m , while the VLA performs low-level control at the higher-frequency time step t . The logical world model searches and evaluates candidate transitions to predict the next logical action am+1 and logical state Xm+1 . Conditioned on the current observation obsm , robot configuration qm , am+1 , and Xm+1 , the visual world model uses an understanding expert and a prediction expert to generate the latent visual subgoal fpredm+1 . During execution, the understanding expert encodes obst and am+1 , while the goal expert processes fpredm+1 . The action expert takes qt and noisy action tokens, attends to both representations, and generates the low-level action chunk α^mt:t+k , grounding long-horizon logical guidance into visually conditioned robot execution.
Methods
Tasks Q-score / Task Success Rate
Task 1
Task 2
Task 3
Task 4
Task 5
Average
H-WM- π0.5 (ours)
98.0/94.0
86.7/60.0
74.0/46.0
70.7 / 42.0
95.0/82.0
84.9/64.8
H-WM-Stable-Diffusion- π0.5
92.7/82.0
80.0/46.0
65.3/ 30.0
74.0 / 34.0
93.5/80.0
81.1/54.4
Logic-guided π0.5
95.3/86.0
84.7/58.0
54.7/16.0
39.3/4.0
92.0/78.0
73.2/48.4
LLM-guided π0.5
84.7/54.0
80.7/42.0
68.0 /24.0
41.3/4.0
59.5/10.0
66.8/26.8
π0.5
66.0/4.0
73.3/24.0
54.7/4.0
44.7/0.0
38.0/0.0
55.3/6.4
TABLE I: Performance on the LIBERO-LoHo benchmark. For each task, we report Q-Score and Success Rate (in %). The best results are shown in bold , and the second-best results are underlined .
Fig. 2: (a) Evaluation results of H-WM-guided π0.5 and various VLA baselines on the LIBERO-10 benchmark. (b) Evaluation results of H-WM-guided π0.5 and baseline methods on the RoboCerebra benchmark.
Fig. 3: (a) Real-world experiment with UR5e robot evaluating H-WM guided π0.5 on long-horizon manipulation task. (b) Step-wise success rates in real-robot experiments.
We propose world-language-action (WLA) models as a new class of embodied foundation models. WLA takes textual instructions, images, and robot states as inputs to jointly predict textual subtasks, subgoal images, and robot actions, conjoining the \emph{world modeling interface} to learn from extensive egocentric videos as in the world-action model (WAM) and the \emph{language reasoning} capacities to solve complex long-horizon tasks as in vision-language-action (VLA) models. At the core of WLA lies an \emph{autoregressive (AR)} Transformer backbone, instead of a bidirectional diffusion Transformer as in WAMs, to predict the \emph{next state}, comprising the \emph{semantic-level} textual intention and complementary \emph{fine-grained} physical dynamics. The physical dynamics are supervised by the world modeling objective based on a dedicated World Expert, and are leveraged to ease the characterization of the state-action correlation for the Action Expert. WLA leverages meta-queries to make the world prediction \emph{implicitly} impact the action generation so that the former can be disabled during inference. The world prediction can also be activated to enable test-time scaling for improved robot control. Our WLA-0 prototype, with 2B active parameters, achieves 40 ms per inference on an NVIDIA RTX 5090. Evaluations across simulated and real-world environments demonstrate that WLA-0 achieves state-of-the-art multi-task and long-horizon learning abilities, e.g., 92.94% success rate on RoboTwin2.0 Clean and 56.5% success rate on RMBench. WLA-0 also holds the promise to learn novel tasks directly from \emph{cross-embodiment robot videos} without action annotations.
Latent world models often struggle with long-horizon planning despite accurate short-term predictions. Recursive rollouts accumulate errors, while distance concentration in high-dimensional latent spaces can weaken goal discrimination. We introduce the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning through distinct state representations and dynamics models. The low-level model predicts action-conditioned transitions, while the high-level model uses learned macro-actions to plan over longer temporal spans. We also propose Long-Horizon Representation Learning with Weighted Rollout (LoRe), which supervises self-generated predictions at both levels. An analysis of recursive error propagation motivates exponential horizon weights with separate decay rates for the two temporal scales. During planning, the high-level model generates latent subgoals that the low-level model refines into actions for precise execution. We evaluate from-scratch Dual-WM on five goal-conditioned visual control tasks against the task-wise strongest baselines without actor-guided proposals. At goal offsets of 50 and 100 environment steps, mean success increases from 75.9% to 84.4% and from 61.4% to 69.5%, respectively. At offset 100, Dual-WM outperforms these baselines on all five tasks and improves mean success over LeWM by 30.8 percentage points. Ablations and supporting analyses provide evidence of more informative representations for goal evaluation and greater consistency under recursive prediction. These results highlight the value of separating temporal roles and training across multiple horizons for reliable latent planning. Our core implementation is available at https://github.com/DeLin1001/Dual-WM-Official.
Delin Zhao, Zhengrong Yue, Shaobin Zhuang +6
Nanjing University · Shenzhen University of Advanced Technology · Shanghai Jiao Tong University +1
Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change the scene. World-Action Models (WAMs) address this limitation by conditioning policies on predicted futures, yet existing approaches typically rely on computationally expensive video generation with substantial pixel-level redundancy. We present LaWAM, a Latent World Action Model that exposes predictive dynamics to robot policies through compact latent visual subgoals instead of reconstructed future video. At the core of LaWAM is a latent-action-conditioned Latent World Model (LaWM). We obtain LaWM by training a latent action model in the latent space of a pretrained vision foundation model and repurposing its forward decoder to predict future observation features for scene evolution. LaWAM then conditions action generation on these predicted latent visual subgoals to enable dynamics-aware robot control. LaWAM achieves state-of-the-art or competitive success rates (SRs) across LIBERO (98.6% SR), RoboTwin (91.22% SR), and real-world manipulation tasks while retaining low-latency inference. LaWAM runs in 187 ms per action-chunk prediction and achieves up to 24x lower wall-clock latency than pixel-space WAMs.
Jialei Chen, Kai Wang, Kang Chen +9
Jilin University · Zhongguancun Academy · Nankai University +4