cs.CVOct 8, 2026

WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation

Authors: Jing He, Kaixin Ding, Xingye Tian, Guibao Shen, Wenhang Ge, Xin Tao, Pengfei Wan, Ying-Cong Chen

Organizations: The Hong Kong University of Science and Technology (Guangzhou) · KlingAI · The University of Hong Kong · The Hong Kong University of Science and Technology

Abstract

Faithful visual world simulation requires generated videos to maintain 4D world consistency, encompassing both static and dynamic consistency. Static consistency requires coherent 3D structure in static environments across viewpoints, while dynamic consistency requires plausible subject motion and consistent appearance over time. Geometry-aware post-training offers a promising way to improve world consistency. However, existing methods often rely on a static-scene assumption. Even those that accommodate dynamic scenes struggle to provide reliable static-consistency feedback, while dynamic consistency is often overlooked or inadequately assessed. To address these limitations, we introduce WorldAlign, a decoupled 4D reward framework that semantically separates static regions and dynamic subjects and provides feedback by aligning each with a world prior suited to its assumptions. For static regions, WorldAlign aligns static geometry with a geometric world prior through semantically guided masked reprojection, enabling more reliable static-consistency evaluation; an auxiliary camera-motion reward discourages nearly static solutions. For dynamic subjects, WorldAlign uses a strong vision-language model (VLM) as a dynamic world prior and constructs a VLM-as-a-judge reward based on sample-specific checklists that assess dynamicity, physical plausibility, shape, and texture consistency. This decoupled design enables more effective online post-training without requiring human preference annotations. Across two pretrained image-to-video generators, Wan2.1 and Wan2.2, WorldAlign jointly improves static and dynamic consistency over existing methods without suppressing overall or subject motion. These results support decoupled world-prior alignment for more faithful visual world simulation. Project page: https://worldalign.github.io/.

Explore similar work

CardsList
  1. GeoFlow: Enforcing Implicit Geometric Consistency in Video Generation

    May 18, 2026Jan Ackermann, Shengqu Cai, Boyang Deng +3Physical Consistency in Video GenerationVideo Diffusion Models

  2. World-R1: Reinforcing 3D Constraints for Text-to-Video Generation

    Apr 27, 2026Weijie Wang, Xiaoxuan He, Youping Gu +9Physical Consistency in Video GenerationText-to-Video Generation