cs.CVOct 8, 2026

WorldGuide: Goal-Directed Video World Model for Procedural Task Execution

Authors: Ankan Deria, Komal Kumar, Hisham Cholakkal, Fahad Shahbaz Khan, Salman Khan

Organizations: Mohamed bin Zayed University of Artificial Intelligence

Abstract

Video generators and video-based world models can synthesize plausible visual trajectories, but long-horizon procedural tasks require generation to adapt to what has actually been produced. A model must determine the next action from its generated state, execute that action, and recognize when the task is complete. Open-loop generation cannot adapt to execution outcomes, while existing closed-loop systems often rely on pretrained executors or indirect verification. This leaves a gap between deciding an action and successfully realizing it. We formulate procedural video generation as \emph{closed-loop task execution in visual world space} and introduce \textbf{WorldGuide}. Given only an initial image and a task goal, WorldGuide predicts an atomic action, generates its corresponding video clip, and uses the generated result to select the next action or terminate. The Planner and Executor are trained on the same step-level procedural demonstrations: the Planner learns to predict the next atomic action or task completion from visual progress, while the Executor is directly trained to realize the predicted actions. Hierarchical visual memory maintains state across long-horizon execution with bounded history token cost. Due to the lack of step-level action-video supervision for joint planner-executor training, we introduce \textbf{WorldGuide Bench}: approximately 59K step-annotated videos across 245 tasks and 27 procedural categories. WorldGuide achieves a 33.33% Task Success on \textbf{WorldGuide-Bench}, compared with 29.90% for the strong recent video model MiniMax-H3, even though MiniMax-H3 receives reference action plans, and achieves 47.69% on \textbf{VideoCraft-Bench} compared with 32.73% for MiniMax-H3 under goal-only conditioning. These results demonstrate the importance of coupling planning with learned execution for goal-directed procedural video generation.

Figures & tables

Appendix figures & tables19 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. World Model Self-Distillation: Training World Models to Solve General Tasks

    Jun 10, 2026Sebastian Stapf, Pablo Acuaviva Huertos, Aram Davtyan +1Video Diffusion ModelsRL for Video Generation

  2. RECIPE: Procedural Planning via Grounding in Instructional Video

    May 19, 2026Luigi Seminara, Antonino Furnari, Lorenzo TorresaniVideo Understanding

  3. Astronex-World 1.0: Real-Time Interactive World Model Foundation

    Sep 17, 2026Xin Zhou, Cong MiaoCamera-Conditioned Video GenerationVideo Diffusion Models