While recent video generation models synthesize highly realistic visuals, they lack a genuine understanding of intrinsic real-world logic. Existing methods attempt to understand the world by internalizing diverse world knowledge, yet constrained by computational overhead or dimensionality alignment, their learning processes inevitably compress features, causing a severe loss of structural information. To address this, we propose \textbf{IntactWorld}, a \textbf{Joint World Modeling Architecture} utilizing uncompressed \textbf{Intact Features}. Since data naturally reside on a low-dimensional manifold within a high-dimensional space, predicting the flow velocity v within this uncompressed high-dimensional space induces a severe manifold gap. To successfully eliminate this optimization bottleneck, our framework instead predicts the clean feature x0 at intermediate layers. Furthermore, to mitigate the computational overhead of incorporating complete world knowledge, we introduce a \textit{Full-to-Compact Training Paradigm}. By replacing raw full features with highly refined CLS tokens, this paradigm enables efficient single-branch guidance, reducing spatial memory consumption by 11.4% and cutting inference latency by 43.8%. Extensive evaluations demonstrate the effectiveness of IntactWorld, outperforming established baselines by 2.46 points on the VBench 2.0 benchmark.
Figures & tables
Figure 1: Overview of the IntactWorld Framework. (a) Training: We employ a Full-to-Compact Training Paradigm. Stage I utilizes lossless full features, predicting the clean target x0 to compute the World Loss in the v -space. Stage II trains exclusively on compact CLS tokens for semantic abstraction. (b) Inference: These condensed tokens enable Compact Inner-Guidance, achieving efficient generation via a single additional branch atop standard CFG.
Method
Temporal
Semantic
Spatial
Summary
Overall Score
Subject Consistency
Background Consistency
Dynamic Degree
Object Class
Human Action
Scene
Spatial Relationship
Quality Score
Semantic Score
Wan2.1-T2V-1.3B
91.83
94.71
65.00
76.09
74.60
20.03
62.37
79.81
65.43
76.93
Baseline
93.59
95.81
54.08
79.90
78.98
28.55
63.31
81.26
68.47
78.71
DreamWorld
93.62
94.95
79.16
81.32
81.20
29.71
70.47
83.49
70.89
80.97
IntactWorld (Ours)
93.60
97.44
78.94
82.63
83.40
31.34
69.78
83.76
72.19
81.45
Table 1: Quantitative comparison on VBench. Bold and underline indicate the best and second best results, respectively. IntactWorld achieves the best performance, clearly outperforming existing methods.
Table 3
Figure 2: Qualitative comparison. Comparison of IntactWorld with Baseline and DreamWorld where red boxes denote structural errors and artifacts. Notably, IntactWorld achieves superior visual quality with fewer artifacts.
Figure 5
Method
Quality
Semantic
Total
w/o full feature
83.60
71.80
81.24
w/o cls tokens
83.64
72.02
81.31
pred_v
83.16
70.50
80.63
IntactWorld
83.76
72.19
81.45
Table 4: Ablation Study. Quantitative evaluation on VBench for full feature training, CLS token abstraction, and the prediction target of x0 versus flow velocity v .
Training world models on vast quantities of unlabelled videos is a critical step toward fully autonomous intelligence. However, the prevailing paradigm of encoding raw pixels into opaque latent spaces and relying on heavy decoders for reconstruction leaves these models computationally expensive and uninterpretable. We address this problem by introducing NOVA, a world modelling framework that represents the system state as the weights and biases of an auxiliary coordinate-based implicit neural representation (INR). This structured representation is analytically rendered, which eliminates the decoder bottleneck while conferring compactness, portability, and zero-shot super-resolution. Furthermore, like most latent action models, NOVA can be distilled into a context-dependent video generator via an action-matching objective. Surprisingly, without resorting to auxiliary losses or adversarial objectives, NOVA can disentangle structural scene components such as background, foreground, and inter-frame motion, enabling users to edit either content or dynamics without compromising the other. We validate our framework on several challenging datasets, achieving strong controllable forecasting while operating on a single consumer GPU at ∼40M parameters. Ultimately, structured representations like INRs not only enhance our understanding of latent dynamics but also pave the way for immersive and customisable virtual experiences.
Roussel Desmond Nzoyem, Mauro Comi
Department of Computer Science University of Manchester Manchester, M13 9PL · Department of Computer Science University of Bristol Bristol, BS8 1QU
World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-based world models, on the other hand, predict future visual features instead of raw video pixels, offering a promising alternative that is more efficient and less prone to hallucination. However, current feature-based approaches rely on direct regression, which leads to blurry or collapsed predictions in complex interactions, while generative modeling in high-dimensional feature spaces still remains challenging. In this work, we discover that a new type of latent action representation, which we refer to as Residual Latent Action (RLA), can be easily learned from DINO residuals. We also show that RLA is predictive, generalizable, and encodes temporal progression. Building on RLA, we propose RLA World Model (RLA-WM), which predicts RLA values via flow matching. RLA-WM outperforms both state-of-the-art feature-based and video-diffusion world models on simulation and real-world datasets, while being orders of magnitude faster than video diffusion. Furthermore, we develop two robot learning techniques that use RLA-WM to improve policy learning. The first one is a minimalist world action model with RLA that learns from actionless videos, and improves VLA on LIBERO and real robot. The second one is a visual RL framework trained entirely inside a world model learned from offline videos only, using a video-aligned reward and no online interactions. Project page: https://mlzxy.github.io/rla-wm
Xinyu Zhang, Zhengtong Xu, Yutian Tao +3
Rutgers University · Purdue University · University of Wisconsin-Madison
Video world models that maintain 3D spatial consistency across generated frames typically rely on explicit point cloud memory constructed in RGB space. This design is both computationally expensive, requiring repeated rendering and VAE encoding, and inherently lossy, as the round trip through pixel space discards rich features of the learned latent representation. In this paper, we introduce \emph{latent spatial memory} for video world models, a persistent 3D cache that stores scene information directly in the diffusion latent space, avoiding pixel-space reconstruction. Building on this, we propose Mirage, a latent-space spatial memory framework that constructs the memory by lifting latent tokens into 3D via depth-guided back-projection and queries it by synthesizing novel views through direct latent-space warping. This unified formulation eliminates both the information loss of pixel-space reconstruction and the computational burden of repeated encoding and rendering. Experiments show that latent spatial memory achieves up to \textbf{10.57}× faster end-to-end video generation and \textbf{55}× reduction in memory footprint relative to explicit 3D baselines. Leveraging the geometric prior of the diffusion model, Mirage attains state-of-the-art performance on WorldScore and strong reconstruction quality on RealEstate10K.
Weijie Wang, Haoyu Zhao, Yifan Yang +7
Zhejiang University · Microsoft Research · Adelaide University +1