World-Action Models (WAMs) utilize future visual prediction as an intermediate reasoning process to guide action generation. However, existing WAMs typically structure visual imagination according to predefined temporal intervals, without explicitly accounting for the different roles of task-critical interactions and connecting transitions. We argue that effective visual foresight should align directly with task-relevant interactions and their corresponding reasoning demands. To this end, we introduce an event-aligned visual action reasoning framework that organizes visual-action prediction around interaction events. Through event-aligned visual-action supervision, WAM learns to generate event-aligned visual context in each imagined rollout, placing greater emphasis on critical state changes that inform action generation. This shapes the visual reasoning granularity according to the underlying interaction dynamics, with detailed reasoning around task-critical events and coarser progression through connecting transitions. Furthermore, we introduce an execution validity head that identifies the valid portion of each predicted action sequence, avoiding redundant actions during chunked inference. Experiments demonstrate a 10.26 percentage point improvement in DOMINO success rate over baseline and competitive performance on RoboTwin 2.0. It also transfers from DOMINO Level 1 to Levels 2 and 3 without target-level adaptation.
Figures & tables
Figure 1: Interaction-dependent temporal granularity of visual reasoning. (a) Our event-aligned visual action reasoning framework organizes visual-action targets around interaction events and transitions. (b) DOMINO success rates under a matched sequence budget of training visual-action slots.
Figure 2: Interaction-Structured Visual Reasoning framework. Event-aligned chunks allocate visual reasoning according to interaction structure, while execution validity prediction identifies which generated actions to execute.
Methods
DOMINO
RoboTwin 2.0
SR ↑
MS ↑
Clean SR ↑
Rand. SR ↑
Avg. SR ↑
π0 ( Black et al., 2026 )
8.17
23.96
65.92
58.40
62.16
π0.5 ( Black et al., 2025 )
9.63
26.17
82.74
76.76
79.75
PUMA ( Fang et al., 2026 )
17.20
34.97
–
–
–
DynamicWAM † † footnotemark: ( Lou et al., 2026 )
38.20
53.20
–
–
–
ABot-M0 ( Yang et al., 2026 )
–
–
81.20
80.40
80.80
Table 1: Evaluation on DOMINO and RoboTwin 2.0. For LingBot-VA on RoboTwin, the upper row reports published results, while the lower row ( † ) reports our re-evaluation of the released checkpoint (3 steps for video tokens to s=0.6 , 10 steps for action tokens to s=1.0 ).
Table 4
Sampling
Granularity
Placement
Slots
SR ↑
MS ↑
LingBot-VA (reference)
Episode
Uniform
291,296
32.57
45.55
Episode Uniform
Episode
Uniform
211,664
29.17
41.32
Event Uniform
Event
Uniform
211,664
39.60
52.26
Reduced Evidence
Event
Transition-Protected
268,592
36.17
48.36
Reduced Context
Event
Evidence-protected
173,216
37.83
52.99
Event-aligned (Ours)
Event
Evidence-protected
211,664
42.83
55.92
Table 4: Ablations of temporal sampling scheme and execution validity prediction on DOMINO. Action slots count all sampled positions, including repetitions, across all retained demonstrations. All variants except the backbone reference and the no-head ablation use the execution validity head during training and inference. Full sampling statistics are reported in Table C5 .
Figure 3: Visual predictions and attention on a Scan Object demonstration episode. (a) Video attention to the latest observation. (b) Episode outcomes. (c) The last predicted frame of each displayed chunk, overlaid with action attention. Red boxes highlight interaction regions. Both methods start from the same initial observation. Later chunks follow each policy’s own rollout, so frames in the same column can differ.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
DOMINO
RoboTwin 2.0
Tasks
35
50
Source demonstrations
1,750 clean Level 1
2,500 clean + 25,000 randomized
Retained demonstrations
1,693
2,465 clean + 22,379 randomized
Optimizer updates
10,000
50,000
Learning rate
10−5
10−5
Effective batch size
32
64
Appendix
Table B1: Benchmark settings and training schedules for our model.
Method
Updates
Global batch
Sampling unit
Action targets (M)
LingBot-VA
10,000
32
Episode segment
≈55.21
Fast-WAM
15,000
128
32-action window
≈56.19
ImageWAM
10,000
384
16-action window
≈58.90
Appendix
Table B2: DOMINO baseline training budgets. The last column reports the estimated total number of non-padding training action target occurrences during training, in millions.
Table C3: Zero-shot transfer to unseen DOMINO dynamic levels.
Method
Initialization
SR (%) ↑
MS ↑
ImageWAM
FLUX.2-klein-base-4B
18.86
35.47
Fast-WAM
Wan2.2-TI2V-5B
19.09
35.33
LingBot-VA
Wan2.2-TI2V-5B
20.83
36.57
Ours
Wan2.2-TI2V-5B
37.43
52.07
Appendix
Table C4: Comparison on DOMINO without robot pretraining. Ours and LingBot-VA are initialized directly from Wan2.2-TI2V-5B, without using LingBot-VA-Base.
Sampling
Granularity
Slots
Inside evidence phases
Outside evidence phases
LingBot-VA (reference)
Episode
291,296
94,233 (32.35%)
197,063 (67.65%)
Episode Uniform
Episode
211,664
80,953 (38.25%)
130,711 (61.75%)
Event Uniform
Event
211,664
100,653 (47.55%)
111,011 (52.45%)
Reduced Evidence
Event
268,592
61,913 (23.05%)
206,679 (76.95%)
Reduced Context
Event
173,216
116,114 (67.03%)
57,102 (32.97%)
Event-aligned (Ours)
Event
211,664
116,114 (54.86%)
95,550 (45.14%)
Appendix
Table C5: Slot composition of sampling configurations in Tab. 4 .
Task
Paired n
Nexecbase
Nexecours
Reduction (%)
Nrollbase
Nrollours
adjust_bottle
42
114.40
86.02
24.81
4.83
3.50
beat_block_hammer
27
111.44
77.26
30.67
4.19
3.26
click_alarmclock
6
68.33
75.17
-10.00
3.00
3.17
click_bell
1
216.00
77.00
64.35
8.00
3.00
dump_bin_bigbin
13
196.85
123.38
37.32
7.15
4.85
grab_roller
45
110.64
62.60
43.42
4.69
3.09
Appendix
Table C6: Execution and visual-rollout requests on common successes.
Figure C2: Visual prediction and attention during task execution on stamp seal and scan object .
Figure C3: Visual prediction and attention during task execution on put bottle dustbin and place shoe .
Figure C4: Visual prediction and attention during task execution on hanging mug and shake bottle horizontally .
Input
Provenance
Use in annotation
Physical states and contacts
Recorded robot values and replayed simulator state
Identify gripper changes, object contact and motion, relative-pose stability, and articulation changes.
Task-success predicates
Native check_success() in benchmark task code
Localize task-specific state changes using the original predicate and our temporal confirmation rule.
Expert execution intervals
Recorded execution of benchmark expert routines
Localize designated task motions, such as the DOMINO shaking routines.
Task specifications and detection rules
Our annotation protocol, with LLM-assisted specifications reviewed task by task
Select actors, contact pairs, detector mappings, and the rules used to localize and verify interactions.
Appendix
Table D1: Inputs to the event annotation pipeline. Reviewed task specifications select the entities and detection rules. Measurements and recorded execution determine the times for each demonstration.
Detector source
Temporal evidence retained
Stable closed contact
From onset through the saved post-evidence end, including the complete qualifying stability window.
Merged bimanual grasp
From onset through the latest of the merged confirmation row and both component post-evidence windows.
Handover receive; object release / handover give
[ton,tconf+1) , including the detector’s stability or release-confirmation observations.
Task-specific body-pair contact
From contact onset through confirmation and the associated task-success or object-release effect, as specified by the detector.
Handle hold and drawer motion
The linked drawer record supplies the held-run window for handle verification; drawer-motion evidence spans the saved baseline-to-peak-search interval.
Native success or expert motion
[ton,tconf+1) , using the predicate-confirmation interval or the designated expert-method record.
Appendix
Table D2: Detector evidence windows in DOMINO. Windows retain the source observations used for verification. Their relation to the method’s event and phase annotations is described in Section D.3 .
Task(s)
Configured detector subtype(s) and interpretation
adjust_bottle
S: high_at_side Bottle reaches the required side and height.
beat_block_hammer
F: hit Hammer satisfies the task-specific hit condition.
click_alarmclock , click_bell , press_stapler
F: press Gripper satisfies the task-specific press condition.
dump_bin_bigbin
S: contents_height_band Bin reaches the required height; waste enters the height band.
grab_roller
S: lifted_bimanually ; G: bimanual_grasp Two-arm acquisition and the roller-lift predicate.
handover_block
S: placed_at_target ; G: handover_receive Handover followed by placement at the target.
Appendix
Table D3: DOMINO task-to-detector-subtype mapping. Prefixes S, F, and G identify detector categories.
Task(s)
Configured detector subtype(s)
adjust_bottle
S: high_at_instructed_side
beat_block_hammer
F: hit ; F: hit_contact
blocks_ranking_rgb
S: ordered_red_green_blue
blocks_ranking_size
S: ordered_large_medium_small
click_alarmclock , click_bell , press_stapler
F: press
dump_bin_bigbin
S: contents_height_band
Appendix
Table D4: RoboTwin 2.0 task-to-detector-subtype mapping (1/2). Prefixes S, F, and G identify detector categories.
World Action Models (WAMs) offer a promising route for robot manipulation by using video generation models to model future scene evolution before producing control actions. However, our empirical observations reveal a phenomenon: generating plausible visual futures does not always guarantee the extraction of accurate actions. To diagnose this failure, we conduct action-head attention analysis and causal interventions. We find that the action decoder fails to focus on task-relevant interaction regions and remains sensitive to perturbations in task-irrelevant areas. This reveals a representation mismatch: hidden states optimized for visual reconstruction are not inherently organized in a form useful for low-level action control. In this paper, we propose AGRA, an Action-Grounded Representation Alignment objective that regularizes the world-action interface by aligning intermediate video diffusion features with spatially coherent semantic representations from a foundation visual encoder. We evaluate AGRA on real-world manipulation tasks. Experiments show that AGRA makes world model representations more action-grounded: by focusing the action decoder on the correct interaction regions, it improves object localization accuracy and affordance understanding, and makes the policy more robust to perturbations in task-irrelevant regions. As a result, AGRA consistently improves both in-distribution performance and out-of-distribution generalization over the baseline world action model.
World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense multi-frame future tokens make inference costly, full video prediction spends capacity on action-irrelevant temporal and appearance details, and long-horizon future imagination may introduce errors that mislead action prediction. These issues raise a simple question: Does world action model really need video generation? We propose ImageWAM, a simple WAM framework that repurposes pretrained image editing models for robot action prediction. In contrast to video generation, image editing provides a better-matched prior: it only needs to model a target-frame transformation, focuses on action-relevant current-to-target visual differences, and grounds task instructions to localized visual changes through edit pretraining. In practice, ImageWAM does not decode the target frame at inference time; instead, it conditions a flow-matching action expert on the KV caches produced by image-editing denoising, using them as a compact world-action context. ImageWAM outperforms standard VLA baselines and matching competitive WAMs without additional policy pretraining across different simulator and real-world experiments. It also reduces FLOPs to 1/6 and latency to 1/4 of video-based WAMs. Attention analysis further shows that editing caches focus on task-relevant change regions, supporting image editing as an effective alternative to video-based world-action modeling.
Yuyang Zhang, Wenyao Zhang, Zekun Qi +7
Shanghai Jiao Tong University · Tsinghua University · Tencent Robotics X +1
World Action Models (WAMs) predict visual futures and robot actions, yet they remain susceptible to task-incomplete imagination, where plausible, action-consistent predictions omit the transition needed for task completion. In this paper, we show that this failure is not inherent to the world model backbone, but emerges when adapted for short-chunk control, which can repeatedly favor plausible local continuations over task-completing transitions. To address this, we introduce Completion Aware Guidance (CAG), a training-free sampling method that guides generation toward task completion. Across representative WAMs, CAG improves success from 64% to 70% on a RoboTwin 2.0 subset and from 69% to 75% in zero-shot simulation, while reducing task-incomplete imagination from 79% to 40%.