FLEX-WAM: Flexible Block-Causal World-Action Models for Long-Horizon Imagination and Planning
Authors: R. Khorrambakht, Joseph Amigo, Félix Lebel, Leon Seetoo, Jean Ponce, Zhenzhen Li, Ludovic Righetti
Organizations: Center for Robotics and Embodied Intelligence (CREO), New York University · Courant Institute of Mathematical Sciences and Center for Data Science, New York University · Ecole normale supérieure - PSL · NVIDIA · Artificial and Natural Intelligence Toulouse Institute (ANITI)
World--action models (WAMs) promise a unified model that predicts action-conditioned futures, generates feasible actions, and supports planning in imagination. However, existing joint video--action models often use computationally heavy, fixed-horizon backbones ill-suited to streaming inference and stable long-horizon open-loop rollouts. We introduce FLEX-WAM, a Flexible and Efficient Block-Causal World--Action Model for unified simulation and policy inference. FLEX-WAM supports variable-length contexts and non-causal prediction horizons, as well as infinite autoregressive generation frame by frame or block by block. Its block-causal, KV-cacheable architecture combines axial attention and blockwise diffusion forcing to enable efficient real-time rollout and deployment-time latency--throughput tradeoffs without retraining. Joint training can nevertheless produce plausible futures that weakly respond to commanded actions. We address this failure mode by balancing state and action flow-matching gradient contributions across the state--action diffusion-noise grid and regulating world-model sampling using Forward-Dynamics (FD) elasticity, an efficient training-time proxy for action responsiveness. Across simulated and real-world datasets, FLEX-WAM achieves superior multi-step prediction quality and latency while producing stable joint state--action rollouts for thousands of steps. As a joint action proposer and simulator within MCTS, it solves long-horizon PushT and all five OGBench Puzzle-4x4 tasks entirely in imagination. On a bimanual OpenArm-based robot, a single checkpoint jointly serves as a play policy and expected-outcome predictor, enabling real-time identification and collection of model--reality mismatches for future self-improvement.
Figures & tables
Fig. 1: WAM. (a) The tokenizer encodes every camera view into one shared set of latents per frame, and is trained to reconstruct masked patches; views never attend to each other, so the latents are the only path between them. (b) Both networks hold their tokens on a frame \mathchar8706 token grid and attend along one axis per layer. (c) The dynamics model denoises one \mathchar28722\mathchar28726\mathchar28723 -token set per frame — latents, registers, the two noise levels, and the action — into clean latents and action velocities. (d) Its attention in time is block-causal: a frame attends to its own block in both directions and to every earlier block, never to a later one.
Domain
Dataset / Platform
Evaluation Purpose
Model Size
Primary Evidence
Simulation
LIBERO
Generation quality and inference latency
690M
SSIM, PSNR, and latency
Real-world play
PushT
Action sensitivity and planning
400M
DINO response and imagined planning
Puzzle, Cube, Scene
OGBench
Long-horizon planning
512M
Imagined task success
Humanoid
Unitree G1
Long-horizon rollout
690M
Qualitative rollout stability
Deployment
OpenArm
Model-reality mismatch detection
690M
Closed-loop case study
TABLE I: Experimental domains and their roles in our evaluation.
Fig. 2: WAM response to synthetic counterfactual stop command along a random set of PushT dataset trajectories with and without training-time feedback regulation.
Fig. 3: Bi-directional attention along the horizon facilitates temporal consistency and smoothness by allowing past frames to attend to the future ones during denoise process. Snapshots show decoded frames from FLEX-WAM with no seed history frames under causal and fully bidirectional attention showing broken temporal consistency in fully causal mode. Plots quantify this by measuring cross frame visual and action distances and dataset-model FVD distance.
Fig. 4: Inference-time scaling of FLEX-WAM. (a) Per-step latency as a function of model size under fully causal inference. (b) Block-generation rate and effective frame throughput as the inference block size varies. All block-size results use the same checkpoint; increasing \mathchar28994 trades causal latency for greater frame-level throughput.
TABLE II: Inference time (latency) of our model vs. baselines.
Fig. 5: (a) FLEX-WAM’s KV-cached block-causal architecture and blockwise diffusion forcing training enables very long stable autoregressive generation and simulation for thousands of frames. Above shows sample 1024-step autoregressive rollouts of FLEX-WAM trained on two teleoperated real-world datasets (human play) and a simulated puzzle task with \mathchar28722\mathchar28721\mathchar28726 states from the OGBench [ 13 ] benchmark. (b) Visualization of two long-horizon plans solved entirely in FLEX-WAM imagination with MCTS. On the left, the search trees: node colours are the task rewards (in \delimiter67482370\mathchar28720\mathchar24891\mathchar28721\delimiter84267779 , light to dark), hollow rings are children created but never simulated, highlighted path is the returned plan. On the right, the decoded imagined states along the returned plan.
Fig. 6: Autoregressive rollout in world model mode with actions provided from the evaluation dataset and decoded states visualized.
Fig. 7: Open-loop rollout (4.5 seconds) of DreamZero and our model both trained under identical WM mode ( \mathchar28956\mathchar29025\mathchar12349\mathchar28721 ) vs ground truth.
TABLE IV: Open-loop performances of our WAM model with MCTS on OGBench visual-puzzle-4x4 evaluation set.
Fig. 8: The joint state-action prediction capability and real-time speed of FLEX-WAM allows closed loop deployment and near-future prediction in image-space on the fly, facilitating the identification of counterfactual training samples for continual self-improvement and early warning for human intervention based training. Above shows the model trained on a small dataset hallucinating a box to pick next to the real camera feed.
World-Action Models (WAMs) have emerged as a promising paradigm for embodied control by coupling future visual prediction with action generation. However, most existing WAMs rely on photorealistic future prediction, which incurs high inference latency and makes real-time robot deployment difficult. This motivates a more efficient WAM design that preserves the control benefits of future visual prediction while reducing its inference cost. We introduce Efficient-WAM, a World-Action Model that reduces the cost of future imagination while preserving its control benefit. Efficient-WAM improves inference efficiency via a compact video expert transferred from WAN-2.2-5B, token-sparse video latents, and asymmetric video-action denoising that allocates fewer sampling steps to video than to actions. Instead of optimizing the future branch for visual fidelity, Efficient-WAM treats future video prediction as a compact guidance signal for action generation. Comprehensive experiments on RoboTwin 2.0 and real-world manipulation tasks show that Efficient-WAM maintains strong action performance despite visibly coarse future predictions. While maintaining competitive control capabilities, our 1B-parameter model can reduce per-chunk latency to around 100 ms during physical deployment, achieving a 30x speedup over existing WAMs.
Jiajun Li, Tiecheng Guo, Yifan Ye +9
1The University of Hong Kong · 2Peking University · 3Muka Robotics +2
World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observations at test time incurs substantial latency. Fast-WAM removes this process for efficiency; however, our matched implementations show lower generalization for Fast-WAM than for future-aware alternatives, especially with scarce robot demonstrations and in out-of-distribution scenarios. To bridge this gap, we introduce LAWA, a WAM architecture that uses compact latent actions as an operational representation of future intentions, enabling efficient test-time future imagination without generating future observations. Specifically, a discrete tokenizer enhanced by action-free pre-training produces manipulation-centric codebook targets. LAWA jointly denoises a continuous latent state anchored to these targets with executable action chunks while omitting the future-video branch at inference. On RoboCasa, LAWA achieves state-of-the-art average success rates of 65.6% and 80.8% in the few-shot and full data settings, improving over the matched Fast-WAM baseline by 9.6 and 4.5 points, respectively. It also preserves the performance level of the matched Joint-WAM variant while requiring 42.9% lower inference latency. LAWA also demonstrates competitive zero-shot robustness on LIBERO-Plus and superior performance on real-world tasks. These results show that future imagination need not be discarded: retaining it with compact latent actions yields an effective trade-off among performance, generalization, and latency. Code and models will be released.
World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense multi-frame future tokens make inference costly, full video prediction spends capacity on action-irrelevant temporal and appearance details, and long-horizon future imagination may introduce errors that mislead action prediction. These issues raise a simple question: Does world action model really need video generation? We propose ImageWAM, a simple WAM framework that repurposes pretrained image editing models for robot action prediction. In contrast to video generation, image editing provides a better-matched prior: it only needs to model a target-frame transformation, focuses on action-relevant current-to-target visual differences, and grounds task instructions to localized visual changes through edit pretraining. In practice, ImageWAM does not decode the target frame at inference time; instead, it conditions a flow-matching action expert on the KV caches produced by image-editing denoising, using them as a compact world-action context. ImageWAM outperforms standard VLA baselines and matching competitive WAMs without additional policy pretraining across different simulator and real-world experiments. It also reduces FLOPs to 1/6 and latency to 1/4 of video-based WAMs. Attention analysis further shows that editing caches focus on task-relevant change regions, supporting image editing as an effective alternative to video-based world-action modeling.
Yuyang Zhang, Wenyao Zhang, Zekun Qi +7
Shanghai Jiao Tong University · Tsinghua University · Tencent Robotics X +1