Video generators and video-based world models can synthesize plausible visual trajectories, but long-horizon procedural tasks require generation to adapt to what has actually been produced. A model must determine the next action from its generated state, execute that action, and recognize when the task is complete. Open-loop generation cannot adapt to execution outcomes, while existing closed-loop systems often rely on pretrained executors or indirect verification. This leaves a gap between deciding an action and successfully realizing it. We formulate procedural video generation as \emph{closed-loop task execution in visual world space} and introduce \textbf{WorldGuide}. Given only an initial image and a task goal, WorldGuide predicts an atomic action, generates its corresponding video clip, and uses the generated result to select the next action or terminate. The Planner and Executor are trained on the same step-level procedural demonstrations: the Planner learns to predict the next atomic action or task completion from visual progress, while the Executor is directly trained to realize the predicted actions. Hierarchical visual memory maintains state across long-horizon execution with bounded history token cost. Due to the lack of step-level action-video supervision for joint planner-executor training, we introduce \textbf{WorldGuide Bench}: approximately 59K step-annotated videos across 245 tasks and 27 procedural categories. WorldGuide achieves a 33.33% Task Success on \textbf{WorldGuide-Bench}, compared with 29.90% for the strong recent video model MiniMax-H3, even though MiniMax-H3 receives reference action plans, and achieves 47.69% on \textbf{VideoCraft-Bench} compared with 32.73% for MiniMax-H3 under goal-only conditioning. These results demonstrate the importance of coupling planning with learned execution for goal-directed procedural video generation.
Figures & tables
Figure 1: From open-loop synthesis to learned procedural execution. (a) Open-loop generation follows a fixed prompt with no intermediate decisions, so it can drift into a wrong step and keep generating after the task is done. (b) A planner can inspect outcomes and replan, but with a frozen pretrained executor, a correct instruction can still fail to execute. (c) WorldGuide trains the ContextPlanner and Executor on the same demonstrations, the Executor conditioned on the planner’s action embeddings. Thus, generated outcomes guide the next action and the decision to stop.
Figure 2: Overview of WorldGuide. Given the initial image x0 , task goal g , and the latest K=3 clip–action pairs ht , the ContextPlanner πθ predicts the next atomic action at or DONE . The Executor fϕ renders at into the next clip xt+1 from the current visual state xt and hierarchical memory Mt , which progressively compresses older latent frames. Each new clip xt+1 becomes the next visual state, is paired with at in the history ht+1 , and updates the memory Mt+1 , and the loop continues until DONE is predicted or the step budget is reached. Notation follows Eq. 1 .
Figure 3: Longest-history branch ( 343≤F≤1366 ). Age 1 is newest; age F is the anchor. Rates r denote spatial reduction relative to base patch embedding, and costs are nominal latent-frame equivalents before spatial padding, with the separately embedded reference adding cost 1 at r=1 . Shorter branches appear in Appendix C.3 .
Figure 4: Task taxonomy of WorldGuide Bench. The inner ring shows procedural categories and their dataset shares; the outer ring lists representative tasks within each category.
Dataset
Size
Tasks/Dom.
Step-Clip
Atomic Act.
Coverage Target
Task Goal
Completion
Type
Public
EgoPlan-IT
50K QA
–/1
✗
✓
✗
✓
✗
Planning QA
✓
Ego4D Goal-Step
48K seg.
86/1
✓
✗
✗
✓
✗
Localization
✓
COIN
11.8K vid.
180/12
Fixed-vocab
Partial
✗
✓
✗
Localization
✓
CrossTask (primary)
2.75K vid.
18/4
✓
Partial
✗
✓
✗
Weak localization
✓
YouCook2
2K vid.
89/1
✓
✗
✓
✓
✗
Segmentation
✓
HT-Step
19.7K vid.
433/1
✓
Partial
✗
✓
✗
Grounding
✓
Table 1: Procedural video datasets and their supervision for planning and generation. Sizes retain their original units: videos (vid.), segments (seg.), clips. Tasks/Dom.: tasks/domains; Step-Clip: temporally localized step annotations; Atomic Act.: single-action granularity; Coverage Target: annotation protocol targets all visible task-relevant steps, excluding background intervals, rather than a verified coverage rate; Task Goal: explicit task-level goal; Completion: explicit task-completion supervision. Gen., Eval., and Recog. denote generation, evaluation, and recognition.
Model
Video Quality Check
Video Planning Check
Aesth.
Imag.
Dynamic
Smooth.
Consist.
Align.
Plan Acc.
Order
Repeat/Skip ↓
Task Succ.
Cosmos-Predict2.5-2B
38.77
62.11
82.52
96.32
17.46
63.75
38.05
84.08
3.98
11.94
Open-Sora 2.0-11B
33.42
44.91
60.19
95.06
17.42
62.43
31.27
79.56
6.53
8.54
HunyuanVideo-1.5-8.3B
42.43
65.57
88.84
96.00
18.91
62.28
43.53
80.00
3.00
12.00
Wan2.2-14B
41.46
58.87
16.02
93.67
17.52
67.33
4.37
21.65
2.06
2.58
CogVideoX-5B
38.76
54.99
64.08
96.58
18.09
68.30
34.33
75.74
3.96
9.90
Table 2: WorldGuide Bench evaluation. Unstarred baselines use reference actions; ∗ denotes methods with their own planning-execution loop. WorldGuide predicts actions from the task goal and generated history. Aesth. : Aesthetic Quality; Imag. : Imaging Quality; Dynamic : Dynamic Degree; Smooth. : Motion Smoothness; Consist. : Overall Consistency; Align. : Alignment; Plan Acc. : Plan Accuracy; Task Succ. : Task Success Rate. Higher is better except Repeat/Skip (Appendix E ).
Model
Video Quality Check
Video Planning Check
Aesth.
Imag.
Dynamic
Smooth.
Consist.
Align.
Plan Acc.
Order
Repeat/Skip ↓
Task Succ.
Cosmos-Predict2.5-2B
35.83
63.66
95.92
98.61
19.49
63.54
5.66
32.42
5.46
2.05
Open-Sora 2.0
33.86
30.26
47.62
99.04
21.61
52.97
22.03
82.96
2.05
14.32
HunyuanVideo-1.5
45.30
67.62
99.66
98.91
21.91
63.94
14.66
64.97
3.06
24.15
Wan2.2
46.05
72.47
100.00
98.31
22.28
71.98
2.98
18.22
4.11
2.40
CogVideoX
40.60
63.49
89.46
98.61
21.52
75.13
15.04
69.73
8.16
2.38
Table 3: Video-CraftBench: goal-conditioned generation. All models receive the initial image and task goal without reference actions. ∗ denotes methods with their own planning–execution loops. Higher is better except Repeat/Skip.
Table 5: ContextPlanner evaluation on WorldGuide Bench. Comparison of ground-truth-history planning against autonomous rollout with generated history; visual feedback improves plan quality during autonomous execution. FT: supervised fine-tuning; K : history length.
Figure 5: Qualitative comparison for “step-by-step fried rice tutorial” . MiniMax-H3, LTX-2.5, HunyuanVideo, and Cosmos-Predict2.5 receive reference action sequences; Bernini and TempAct use their own planning–execution loops. WorldGuide predicts actions from the task goal and generated progress. Frames are uniformly sampled from each trajectory.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Per-video prompt supplied to Gemini 2.5 Flash for dataset captioning. The video, together with the retrieval query, task, category, and description, is passed alongside this prompt, prefixed by a fixed system instruction: “You are a professional technical writer. Your task is to provide a non-repetitive, high-accuracy timeline of physical actions from video.” Only videos assigned a Video Planning Score of 5 are retained.
Information
Contents
Task context
Category, task label, and task goal.
Action annotation
Step index, action caption, and start/end timestamps in the original video.
Target clip
Video segment corresponding to the annotated action.
Preceding history
Previous clips and their action captions. The ContextPlanner uses up to K=3 previous steps, while the Executor uses visual memory.
Progress
Step index normalized by the total number of annotated steps.
Completion
Flag marking the final step. After it, one ContextPlanner-only record with target <|Task Completed|> and no target clip is appended.
Appendix
Table 6: WorldGuide Bench sample format. Each record contains an annotated action, its corresponding video clip, and the preceding procedural context. The completion flag identifies the final step.
Quantity
Count
Top-level categories
27
Task topics
245
Merged raw video metadata records
133,360
Validated (decodable) video records
132,911
Videos with a non-null Gemini annotation
119,798
Videos with planning score = 5 (retained)
58,679
Appendix
Table 7: WorldGuide Bench construction statistics. Counts are reported after each filtering stage. Only videos receiving a Gemini Video Planning Score of 5/5 are retained. The video-disjoint test split holds out four videos per task.
Step
Time
Action caption
1
00:04–00:05
Folds the bottom right corner of the paper to the top left.
2
00:05–00:06
Unfolds the paper.
3
00:06–00:07
Folds the bottom left corner of the paper to the top right.
4
00:07–00:08
Unfolds the paper.
5
00:08–00:09
Flips the paper over.
6
00:09–00:10
Folds the bottom edge of the paper upwards.
Appendix
Table 8: Complete annotation of one WorldGuide Bench video. Task: origami, airplane_dart . Work subject: folding a paper airplane (dart style). Each row is one clip; Time gives its start and end in the source video (mm:ss).
Criterion
Level
Annot. 1
Annot. 2
Annot. 3
Majority
Caption correctness
Step
82.4
82.6
81.9
83.0
Action atomicity
Step
82.2
81.5
82.3
82.9
Task-label correctness
Video
82.9
82.4
83.2
83.7
Essential-action coverage
Video
83.2
84.4
83.9
84.5
Temporal order
Video
82.1
81.0
81.6
81.6
Demonstration completeness
Video
81.4
81.7
82.6
82.9
Appendix
Table 9: Human audit of WorldGuide Bench annotations (% judged correct) on 245 task-balanced videos with 3,168 action steps. Majority aggregates the three annotators by item-level vote.
Segment
Latent frames
Rate
Cost upper bound
Present
3
1
3.00
Near present
2
2
0.50
Recent past
16
4
1.00
Mid past
64
8
1.00
Far past
256
16
1.00
Distant past
F−342≤1024
64
0.25
Appendix
Table 10: Nominal token cost for 343≤F≤1366 history latent frames. One full-resolution latent-frame equivalent is HW/p2 tokens; the reference is counted separately.
Component
Configuration
Base model
Qwen2.5-VL-7B-Instruct
Training objective
Autoregressive next-step prediction
Trainable parameters
Full-parameter fine-tuning
Precision
BF16
Distributed training
DeepSpeed ZeRO-3
Hardware
8 nodes × 8 AMD Instinct MI210 (64 GB)
Appendix
Table 11: ContextPlanner training configuration. Full-parameter fine-tuning of Qwen2.5-VL-7B-Instruct for next-step action prediction from goal, visual state, and a history of K=3 clip–action pairs.
Component
Configuration
Base model
HunyuanVideo-1.5 DiT
Training objective
Masked MSE flow matching
Input mode
Image-to-video
Target resolution
480×832
Target FPS
24
Text conditioning
Action-caption embeddings from the frozen ContextPlanner, and ByT5 [ 43 ] features
Appendix
Table 12: Executor training configuration. HunyuanVideo-1.5 DiT fine-tuned for language-conditioned image-to-video generation on step-level clips under the masked flow-matching objective (Eq. 2 ).
Video judge
Text judge
Input
Rollout, goal, reference plan
Predicted plan, goal, reference plan
Plan Accuracy
Visibly achieved sub-goals
Semantic correctness and coverage
Order Score
Executed prerequisite order
Planned prerequisite order
Repeat/Skip
Unnecessary executions; omissions reduce Plan Accuracy
Redundant actions and missing sub-goals
Success
Final visible task completion
Plan would succeed if executed correctly
Appendix
Table 13: Metric meanings on WorldGuide Bench. Video scores assess rendered execution; text scores assess the predicted procedure. Video-CraftBench uses goal-based video judging without a reference plan.
Figure 7: Prompt structure for procedural video evaluation with the Gemini 3.5 Flash judge (temperature 0.0 ), used to produce the video-judge columns of Table 2 . The judge receives the generated rollout, task description, and ground-truth plan with the <|Task Completed|> token stripped, and scores Plan Accuracy, Order Score, Repeat/Skip, and Task Success from visible execution alone; the ground-truth video is never shown.
Figure 8: Prompt structure for text-only planner evaluation with a GPT-5.2 judge, used to produce Table 5 . The judge receives the task description, predicted plan, and ground-truth plan, with no video, and scores the procedural validity of the predicted text. The ground truth serves as a reference rather than a unique target, so semantically equivalent actions, orderings, and granularities are accepted.
Dimension
Predictor
Measures
Aesthetic Quality (AQ)
LAION CLIP aesthetic
Mean frame-level aesthetic score
Imaging Quality (IQ)
MUSIQ
Technical image quality; penalizes blur, noise, and distortion
Dynamic Degree (DD)
RAFT optical flow
Percentage of videos classified as containing sufficient motion
Motion Smoothness (MS)
Frame interpolation
Agreement between observed and interpolated intermediate frames
Overall Consistency (OC)
ViCLIP
Video–text semantic consistency with the task description
Appendix
Table 14: The five VBench dimensions reported in Tables 2 and 3 , with the predictor used by the official implementation. All are perceptual or temporal quality measures; none assesses procedural correctness.
Model
Avg. score / 5
Bernini*
1.583
Cosmos-Predict2.5
1.000
HunyuanVideo-1.5
1.000
LTX-2.5
2.472
MiniMax-H3
2.542
PhysAgent*
1.000
Appendix
Table 15: Human evaluation of procedural video generation. Average task-completion scores on a 1–5 scale. Higher scores indicate better task execution. Model labels follow Table 2 .
Figure 9: Building a tower from colorful wooden blocks. WorldGuide progressively adds upright pieces and a top to the arch base. TempAct substitutes different blocks, Bernini shows floating pieces and scene changes, and PhysAgent retains several small stacks.
Figure 10: Burrito assembly and folding. WorldGuide adds fillings before folding the tortilla. TempAct folds a tortilla without visible fillings, PhysAgent shifts to stacking round objects, and Bernini replaces the split-screen demonstration with a different scene.
Figure 11: Layered dessert preparation. WorldGuide shows mixing, cream addition, and toppings. HunyuanVideo-1.5 and Cosmos-Predict2.5 also show preparation stages, while Helios develops severe distortion and SpMem changes the vessel and contents.
Figure 12: Video-CraftBench: a second block-horse example. WorldGuide adds a yellow bar and orange cylinder to the starting arch. MiniMax-H3 introduces shaped toy pieces, while LTX-2.5 substitutes different blocks.
Figure 13: Video-CraftBench: folding a blue paper airplane. WorldGuide progresses through several folds, with visible late texture distortion. PhysAgent repeatedly handles a broad fold, while TempAct changes the blue sheet into a white model aircraft.
Figure 14: Video-CraftBench: folding an orange paper boat. WorldGuide shows successive folds with late changes in paper color and scene texture. TempAct substitutes a white boat-shaped object, while PhysAgent remains near a rectangular fold.