WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
Organizations: Mohamed bin Zayed University of Artificial Intelligence
Abstract
Video generators and video-based world models can synthesize plausible visual trajectories, but long-horizon procedural tasks require generation to adapt to what has actually been produced. A model must determine the next action from its generated state, execute that action, and recognize when the task is complete. Open-loop generation cannot adapt to execution outcomes, while existing closed-loop systems often rely on pretrained executors or indirect verification. This leaves a gap between deciding an action and successfully realizing it. We formulate procedural video generation as \emph{closed-loop task execution in visual world space} and introduce \textbf{WorldGuide}. Given only an initial image and a task goal, WorldGuide predicts an atomic action, generates its corresponding video clip, and uses the generated result to select the next action or terminate. The Planner and Executor are trained on the same step-level procedural demonstrations: the Planner learns to predict the next atomic action or task completion from visual progress, while the Executor is directly trained to realize the predicted actions. Hierarchical visual memory maintains state across long-horizon execution with bounded history token cost. Due to the lack of step-level action-video supervision for joint planner-executor training, we introduce \textbf{WorldGuide Bench}: approximately 59K step-annotated videos across 245 tasks and 27 procedural categories. WorldGuide achieves a 33.33% Task Success on \textbf{WorldGuide-Bench}, compared with 29.90% for the strong recent video model MiniMax-H3, even though MiniMax-H3 receives reference action plans, and achieves 47.69% on \textbf{VideoCraft-Bench} compared with 32.73% for MiniMax-H3 under goal-only conditioning. These results demonstrate the importance of coupling planning with learned execution for goal-directed procedural video generation.
Figures & tables
| Dataset | Size | Tasks/Dom. | Step-Clip | Atomic Act. | Coverage Target | Task Goal | Completion | Type | Public |
| EgoPlan-IT | 50K QA | –/1 | ✗ | ✓ | ✗ | ✓ | ✗ | Planning QA | ✓ |
| Ego4D Goal-Step | 48K seg. | 86/1 | ✓ | ✗ | ✗ | ✓ | ✗ | Localization | ✓ |
| COIN | 11.8K vid. | 180/12 | Fixed-vocab | Partial | ✗ | ✓ | ✗ | Localization | ✓ |
| CrossTask (primary) | 2.75K vid. | 18/4 | ✓ | Partial | ✗ | ✓ | ✗ | Weak localization | ✓ |
| YouCook2 | 2K vid. | 89/1 | ✓ | ✗ | ✓ | ✓ | ✗ | Segmentation | ✓ |
| HT-Step | 19.7K vid. | 433/1 | ✓ | Partial | ✗ | ✓ | ✗ | Grounding | ✓ |
| Model | Video Quality Check | Video Planning Check | ||||||||
| Aesth. | Imag. | Dynamic | Smooth. | Consist. | Align. | Plan Acc. | Order | Repeat/Skip | Task Succ. | |
| Cosmos-Predict2.5-2B | 38.77 | 62.11 | 82.52 | 96.32 | 17.46 | 63.75 | 38.05 | 84.08 | 3.98 | 11.94 |
| Open-Sora 2.0-11B | 33.42 | 44.91 | 60.19 | 95.06 | 17.42 | 62.43 | 31.27 | 79.56 | 6.53 | 8.54 |
| HunyuanVideo-1.5-8.3B | 42.43 | 65.57 | 88.84 | 96.00 | 18.91 | 62.28 | 43.53 | 80.00 | 3.00 | 12.00 |
| Wan2.2-14B | 41.46 | 58.87 | 16.02 | 93.67 | 17.52 | 67.33 | 4.37 | 21.65 | 2.06 | 2.58 |
| CogVideoX-5B | 38.76 | 54.99 | 64.08 | 96.58 | 18.09 | 68.30 | 34.33 | 75.74 | 3.96 | 9.90 |
| Model | Video Quality Check | Video Planning Check | ||||||||
| Aesth. | Imag. | Dynamic | Smooth. | Consist. | Align. | Plan Acc. | Order | Repeat/Skip | Task Succ. | |
| Cosmos-Predict2.5-2B | 35.83 | 63.66 | 95.92 | 98.61 | 19.49 | 63.54 | 5.66 | 32.42 | 5.46 | 2.05 |
| Open-Sora 2.0 | 33.86 | 30.26 | 47.62 | 99.04 | 21.61 | 52.97 | 22.03 | 82.96 | 2.05 | 14.32 |
| HunyuanVideo-1.5 | 45.30 | 67.62 | 99.66 | 98.91 | 21.91 | 63.94 | 14.66 | 64.97 | 3.06 | 24.15 |
| Wan2.2 | 46.05 | 72.47 | 100.00 | 98.31 | 22.28 | 71.98 | 2.98 | 18.22 | 4.11 | 2.40 |
| CogVideoX | 40.60 | 63.49 | 89.46 | 98.61 | 21.52 | 75.13 | 15.04 | 69.73 | 8.16 | 2.38 |
| Variant | Video Quality Check | Video Planning Check | ||||||||
| Aesth. | Imag. | Dynamic | Smooth. | Consist. | Align. | Plan Acc. | Order | Repeat/Skip | Task Succ. | |
| WorldGuide Bench | ||||||||||
| WorldGuide (open-loop) | 36.76 | 59.38 | 73.30 | 98.42 | 18.86 | 67.40 | 22.87 | 53.81 | 0.49 | 11.71 |
| WorldGuide (w/o visual) | 36.77 | 58.74 | 83.92 | 98.25 | 18.67 | 68.11 | 29.29 | 68.70 | 0.08 | 14.72 |
| WorldGuide (w/o Mem) | 38.70 | 55.35 | 88.35 | 98.14 | 19.70 | 64.74 | 55.36 | 91.12 | 4.39 | 27.56 |
| WorldGuide (oracle ContextPlanner) | 52.44 | 66.72 | 97.68 | 99.49 | 33.37 | 77.14 | 65.42 | 94.28 | 2.04 | 38.82 |
| Setting | FT | Plan Acc. | Order | Repeat/Skip | Success | |
| Qwen2.5-VL baseline | 3 | 1.04 | 0.00 | 8.82 | 0.00 | |
| ContextPlanner | 1 | 56.48 | 42.35 | 21.16 | 39.17 | |
| ContextPlanner | 2 | 75.83 | 71.28 | 15.56 | 44.66 | |
| ContextPlanner | 3 | 79.66 | 72.56 | 14.71 | 47.04 | |
| WorldGuide (w/o visual) | 3 | 63.84 | 77.46 | 20.19 | 86.41 | |
| WorldGuide planner-executor | 3 | 64.03 | 79.45 | 15.73 | 91.75 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Information | Contents |
| Task context | Category, task label, and task goal. |
| Action annotation | Step index, action caption, and start/end timestamps in the original video. |
| Target clip | Video segment corresponding to the annotated action. |
| Preceding history | Previous clips and their action captions. The ContextPlanner uses up to previous steps, while the Executor uses visual memory. |
| Progress | Step index normalized by the total number of annotated steps. |
| Completion | Flag marking the final step. After it, one ContextPlanner-only record with target <|Task Completed|> and no target clip is appended. |
| Quantity | Count |
| Top-level categories | 27 |
| Task topics | 245 |
| Merged raw video metadata records | 133,360 |
| Validated (decodable) video records | 132,911 |
| Videos with a non-null Gemini annotation | 119,798 |
| Videos with planning score 5 (retained) | 58,679 |
| Step | Time | Action caption |
| 1 | 00:04–00:05 | Folds the bottom right corner of the paper to the top left. |
| 2 | 00:05–00:06 | Unfolds the paper. |
| 3 | 00:06–00:07 | Folds the bottom left corner of the paper to the top right. |
| 4 | 00:07–00:08 | Unfolds the paper. |
| 5 | 00:08–00:09 | Flips the paper over. |
| 6 | 00:09–00:10 | Folds the bottom edge of the paper upwards. |
| Criterion | Level | Annot. 1 | Annot. 2 | Annot. 3 | Majority |
| Caption correctness | Step | 82.4 | 82.6 | 81.9 | 83.0 |
| Action atomicity | Step | 82.2 | 81.5 | 82.3 | 82.9 |
| Task-label correctness | Video | 82.9 | 82.4 | 83.2 | 83.7 |
| Essential-action coverage | Video | 83.2 | 84.4 | 83.9 | 84.5 |
| Temporal order | Video | 82.1 | 81.0 | 81.6 | 81.6 |
| Demonstration completeness | Video | 81.4 | 81.7 | 82.6 | 82.9 |
| Segment | Latent frames | Rate | Cost upper bound |
| Present | 1 | 3.00 | |
| Near present | 2 | 0.50 | |
| Recent past | 4 | 1.00 | |
| Mid past | 8 | 1.00 | |
| Far past | 16 | 1.00 | |
| Distant past | 64 | 0.25 |
| Component | Configuration |
| Base model | Qwen2.5-VL-7B-Instruct |
| Training objective | Autoregressive next-step prediction |
| Trainable parameters | Full-parameter fine-tuning |
| Precision | BF16 |
| Distributed training | DeepSpeed ZeRO-3 |
| Hardware | 8 nodes 8 AMD Instinct MI210 (64 GB) |
| Component | Configuration |
| Base model | HunyuanVideo-1.5 DiT |
| Training objective | Masked MSE flow matching |
| Input mode | Image-to-video |
| Target resolution | |
| Target FPS | 24 |
| Text conditioning | Action-caption embeddings from the frozen ContextPlanner, and ByT5 [ 43 ] features |
| Video judge | Text judge | |
| Input | Rollout, goal, reference plan | Predicted plan, goal, reference plan |
| Plan Accuracy | Visibly achieved sub-goals | Semantic correctness and coverage |
| Order Score | Executed prerequisite order | Planned prerequisite order |
| Repeat/Skip | Unnecessary executions; omissions reduce Plan Accuracy | Redundant actions and missing sub-goals |
| Success | Final visible task completion | Plan would succeed if executed correctly |
| Dimension | Predictor | Measures |
| Aesthetic Quality (AQ) | LAION CLIP aesthetic | Mean frame-level aesthetic score |
| Imaging Quality (IQ) | MUSIQ | Technical image quality; penalizes blur, noise, and distortion |
| Dynamic Degree (DD) | RAFT optical flow | Percentage of videos classified as containing sufficient motion |
| Motion Smoothness (MS) | Frame interpolation | Agreement between observed and interpolated intermediate frames |
| Overall Consistency (OC) | ViCLIP | Video–text semantic consistency with the task description |
| Model | Avg. score / 5 |
| Bernini* | 1.583 |
| Cosmos-Predict2.5 | 1.000 |
| HunyuanVideo-1.5 | 1.000 |
| LTX-2.5 | 2.472 |
| MiniMax-H3 | 2.542 |
| PhysAgent* | 1.000 |