cs.CVOct 7, 2026

Event-Aligned Visual Action Reasoning for World Action Models

Authors: Xiaomeng Yang, Yushu Wu, Yi Gao, Yuhao Lei, Xuan Zhang, Pu Zhao, Yanzhi Wang

Organizations: Northeastern University

Abstract

World-Action Models (WAMs) utilize future visual prediction as an intermediate reasoning process to guide action generation. However, existing WAMs typically structure visual imagination according to predefined temporal intervals, without explicitly accounting for the different roles of task-critical interactions and connecting transitions. We argue that effective visual foresight should align directly with task-relevant interactions and their corresponding reasoning demands. To this end, we introduce an event-aligned visual action reasoning framework that organizes visual-action prediction around interaction events. Through event-aligned visual-action supervision, WAM learns to generate event-aligned visual context in each imagined rollout, placing greater emphasis on critical state changes that inform action generation. This shapes the visual reasoning granularity according to the underlying interaction dynamics, with detailed reasoning around task-critical events and coarser progression through connecting transitions. Furthermore, we introduce an execution validity head that identifies the valid portion of each predicted action sequence, avoiding redundant actions during chunked inference. Experiments demonstrate a 10.26 percentage point improvement in DOMINO success rate over baseline and competitive performance on RoboTwin 2.0. It also transfers from DOMINO Level 1 to Levels 2 and 3 without target-level adaptation.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Making Foresight Actionable: Repurposing Representation Alignment in World Action Models

    Jun 10, 2026Lu Qiu, Yizhuo Li, Yi Chen +3Efficient World-Action ModelWorld Models

  2. ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

    Jun 17, 2026Yuyang Zhang, Wenyao Zhang, Zekun Qi +7World ModelsVideo Generation

  3. Completion Aware Guidance for World Action Models

    Oct 1, 2026Seungyeon Kim, Junhoo Lee, Baekseung Kim +2Efficient World-Action ModelWorld Models