FLEX-WAM: Flexible Block-Causal World-Action Models for Long-Horizon Imagination and Planning
Authors: R. Khorrambakht, Joseph Amigo, Félix Lebel, Leon Seetoo, Jean Ponce, Zhenzhen Li, Ludovic Righetti
Organizations: Center for Robotics and Embodied Intelligence (CREO), New York University · Courant Institute of Mathematical Sciences and Center for Data Science, New York University · Ecole normale supérieure - PSL · NVIDIA · Artificial and Natural Intelligence Toulouse Institute (ANITI)
World--action models (WAMs) promise a unified model that predicts action-conditioned futures, generates feasible actions, and supports planning in imagination. However, existing joint video--action models often use computationally heavy, fixed-horizon backbones ill-suited to streaming inference and stable long-horizon open-loop rollouts. We introduce FLEX-WAM, a Flexible and Efficient Block-Causal World--Action Model for unified simulation and policy inference. FLEX-WAM supports variable-length contexts and non-causal prediction horizons, as well as infinite autoregressive generation frame by frame or block by block. Its block-causal, KV-cacheable architecture combines axial attention and blockwise diffusion forcing to enable efficient real-time rollout and deployment-time latency--throughput tradeoffs without retraining. Joint training can nevertheless produce plausible futures that weakly respond to commanded actions. We address this failure mode by balancing state and action flow-matching gradient contributions across the state--action diffusion-noise grid and regulating world-model sampling using Forward-Dynamics (FD) elasticity, an efficient training-time proxy for action responsiveness. Across simulated and real-world datasets, FLEX-WAM achieves superior multi-step prediction quality and latency while producing stable joint state--action rollouts for thousands of steps. As a joint action proposer and simulator within MCTS, it solves long-horizon PushT and all five OGBench Puzzle-4x4 tasks entirely in imagination. On a bimanual OpenArm-based robot, a single checkpoint jointly serves as a play policy and expected-outcome predictor, enabling real-time identification and collection of model--reality mismatches for future self-improvement.
Figures & tables
Fig. 1: WAM. (a) The tokenizer encodes every camera view into one shared set of latents per frame, and is trained to reconstruct masked patches; views never attend to each other, so the latents are the only path between them. (b) Both networks hold their tokens on a frame \mathchar8706 token grid and attend along one axis per layer. (c) The dynamics model denoises one \mathchar28722\mathchar28726\mathchar28723 -token set per frame — latents, registers, the two noise levels, and the action — into clean latents and action velocities. (d) Its attention in time is block-causal: a frame attends to its own block in both directions and to every earlier block, never to a later one.
Domain
Dataset / Platform
Evaluation Purpose
Model Size
Primary Evidence
Simulation
LIBERO
Generation quality and inference latency
690M
SSIM, PSNR, and latency
Real-world play
PushT
Action sensitivity and planning
400M
DINO response and imagined planning
Puzzle, Cube, Scene
OGBench
Long-horizon planning
512M
Imagined task success
Humanoid
Unitree G1
Long-horizon rollout
690M
Qualitative rollout stability
Deployment
OpenArm
Model-reality mismatch detection
690M
Closed-loop case study
TABLE I: Experimental domains and their roles in our evaluation.
Fig. 2: WAM response to synthetic counterfactual stop command along a random set of PushT dataset trajectories with and without training-time feedback regulation.
Fig. 3: Bi-directional attention along the horizon facilitates temporal consistency and smoothness by allowing past frames to attend to the future ones during denoise process. Snapshots show decoded frames from FLEX-WAM with no seed history frames under causal and fully bidirectional attention showing broken temporal consistency in fully causal mode. Plots quantify this by measuring cross frame visual and action distances and dataset-model FVD distance.
Fig. 4: Inference-time scaling of FLEX-WAM. (a) Per-step latency as a function of model size under fully causal inference. (b) Block-generation rate and effective frame throughput as the inference block size varies. All block-size results use the same checkpoint; increasing \mathchar28994 trades causal latency for greater frame-level throughput.
TABLE II: Inference time (latency) of our model vs. baselines.
Fig. 5: (a) FLEX-WAM’s KV-cached block-causal architecture and blockwise diffusion forcing training enables very long stable autoregressive generation and simulation for thousands of frames. Above shows sample 1024-step autoregressive rollouts of FLEX-WAM trained on two teleoperated real-world datasets (human play) and a simulated puzzle task with \mathchar28722\mathchar28721\mathchar28726 states from the OGBench [ 13 ] benchmark. (b) Visualization of two long-horizon plans solved entirely in FLEX-WAM imagination with MCTS. On the left, the search trees: node colours are the task rewards (in \delimiter67482370\mathchar28720\mathchar24891\mathchar28721\delimiter84267779 , light to dark), hollow rings are children created but never simulated, highlighted path is the returned plan. On the right, the decoded imagined states along the returned plan.
Fig. 6: Autoregressive rollout in world model mode with actions provided from the evaluation dataset and decoded states visualized.
Fig. 7: Open-loop rollout (4.5 seconds) of DreamZero and our model both trained under identical WM mode ( \mathchar28956\mathchar29025\mathchar12349\mathchar28721 ) vs ground truth.
TABLE IV: Open-loop performances of our WAM model with MCTS on OGBench visual-puzzle-4x4 evaluation set.
Fig. 8: The joint state-action prediction capability and real-time speed of FLEX-WAM allows closed loop deployment and near-future prediction in image-space on the fly, facilitating the identification of counterfactual training samples for continual self-improvement and early warning for human intervention based training. Above shows the model trained on a small dataset hallucinating a box to pick next to the real camera feed.