cs.ROOct 4, 2026

FLEX-WAM: Flexible Block-Causal World-Action Models for Long-Horizon Imagination and Planning

Authors: R. Khorrambakht, Joseph Amigo, Félix Lebel, Leon Seetoo, Jean Ponce, Zhenzhen Li, Ludovic Righetti

Organizations: Center for Robotics and Embodied Intelligence (CREO), New York University · Courant Institute of Mathematical Sciences and Center for Data Science, New York University · Ecole normale supérieure - PSL · NVIDIA · Artificial and Natural Intelligence Toulouse Institute (ANITI)

Abstract

World--action models (WAMs) promise a unified model that predicts action-conditioned futures, generates feasible actions, and supports planning in imagination. However, existing joint video--action models often use computationally heavy, fixed-horizon backbones ill-suited to streaming inference and stable long-horizon open-loop rollouts. We introduce FLEX-WAM, a Flexible and Efficient Block-Causal World--Action Model for unified simulation and policy inference. FLEX-WAM supports variable-length contexts and non-causal prediction horizons, as well as infinite autoregressive generation frame by frame or block by block. Its block-causal, KV-cacheable architecture combines axial attention and blockwise diffusion forcing to enable efficient real-time rollout and deployment-time latency--throughput tradeoffs without retraining. Joint training can nevertheless produce plausible futures that weakly respond to commanded actions. We address this failure mode by balancing state and action flow-matching gradient contributions across the state--action diffusion-noise grid and regulating world-model sampling using Forward-Dynamics (FD) elasticity, an efficient training-time proxy for action responsiveness. Across simulated and real-world datasets, FLEX-WAM achieves superior multi-step prediction quality and latency while producing stable joint state--action rollouts for thousands of steps. As a joint action proposer and simulator within MCTS, it solves long-horizon PushT and all five OGBench Puzzle-4x4 tasks entirely in imagination. On a bimanual OpenArm-based robot, a single checkpoint jointly serves as a play policy and expected-outcome predictor, enabling real-time identification and collection of model--reality mismatches for future self-improvement.

Figures & tables

Explore similar work

CardsList
  1. Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination

    Jun 8, 2026Jiajun Li, Tiecheng Guo, Yifan Ye +9Efficient World-Action ModelFaster-Wam

  2. Latent Action as Intention Enables Efficient Future Imagination for World Action Models

    Aug 25, 2026Xiang Li, Yupeng Zheng, Songen Gu +11Efficient World-Action ModelLatent Actions

  3. ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

    Jun 17, 2026Yuyang Zhang, Wenyao Zhang, Zekun Qi +7World ModelsVideo Generation