cs.CVOct 8, 2026

IntactWorld: Joint World Modeling with Intact Features

Authors: Boming Tan, Xiangdong Zhang, Yan Xia, Qi Zhu, Deyi Ji, Xue Yang, Shaofeng Zhang

Organizations: University of Science and Technology of China · Shanghai Jiao Tong University · KOKONI 3D, Moxin Technology

Abstract

While recent video generation models synthesize highly realistic visuals, they lack a genuine understanding of intrinsic real-world logic. Existing methods attempt to understand the world by internalizing diverse world knowledge, yet constrained by computational overhead or dimensionality alignment, their learning processes inevitably compress features, causing a severe loss of structural information. To address this, we propose \textbf{IntactWorld}, a \textbf{Joint World Modeling Architecture} utilizing uncompressed \textbf{Intact Features}. Since data naturally reside on a low-dimensional manifold within a high-dimensional space, predicting the flow velocity vv within this uncompressed high-dimensional space induces a severe manifold gap. To successfully eliminate this optimization bottleneck, our framework instead predicts the clean feature x0x_0 at intermediate layers. Furthermore, to mitigate the computational overhead of incorporating complete world knowledge, we introduce a \textit{Full-to-Compact Training Paradigm}. By replacing raw full features with highly refined CLS tokens, this paradigm enables efficient single-branch guidance, reducing spatial memory consumption by 11.4% and cutting inference latency by 43.8%. Extensive evaluations demonstrate the effectiveness of IntactWorld, outperforming established baselines by 2.46 points on the VBench 2.0 benchmark.

Figures & tables

Explore similar work

CardsList
  1. Render, Don't Decode: Weight-Space World Models with Latent Structural Disentanglement

    May 7, 2026Roussel Desmond Nzoyem, Mauro ComiDisentangled Representation LearningImplicit Neural Representations

  2. Learning Visual Feature-Based World Models via Residual Latent Action

    May 8, 2026Xinyu Zhang, Zhengtong Xu, Yutian Tao +3Latent Action LearningWorld Model Learning

  3. Latent Spatial Memory for Video World Models

    Jun 8, 2026Weijie Wang, Haoyu Zhao, Yifan Yang +7Video Diffusion ModelsMemory-Augmented Video Generation