Authors: Haotian Zhang, Fengyuan Yu, Dezhi Luo, Haoran Sun, Zehong Zhao, Qingying Gao, Yihan Li, Siyuan An, +23 more
Organizations: University of Southern California · Carnegie Mellon University · University of Michigan · Johns Hopkins University · University of California, San Diego · University of California, Los Angeles · Columbia University · University of Toronto · University of Bristol · University of California, Berkeley · University of Waterloo · Friedrich-Alexander-Universität Erlangen · University of Oxford · New York University · Stanford University · Harvard University
Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models, a paradigmatic class of current world models, have begun to show emerged reasoning abilities, making them ideal candidates for building human-like physical intelligence. Do video models have emerged object permanence in them? If not, could we train them with a core-cognition inspired dataset? We introduce WROP (World Reasoning with Object Permanence), a data infrastructure of 150 hand-designed cognitive science inspired tasks, divided into six cognitive categories. We build Blender generators that randomize speed, lighting, camera angle, and other nuisance parameters while preserving each task's cognitive structure, yielding 10,000+ samples per task. We release a 1.5M-sample training corpus and a 300-question exam. On this exam we evaluate 14 video models: 3 reference-to-video, 7 edit, and 4 continuation, among which PWM-WROP, our 16B world model. In a blind pairwise Elo study, PWM-WROP ranks first among continuation models and third overall, behind only a statistical tie between two reference-to-video models. We release the data, exam, model answers, scores, weights, and PWM, our native-PyTorch training stack on AWS Trainium2.
Figures & tables
Figure 1 : WROP: World Reasoning with Object Permanence. We introduce a benchmark and training resource for object permanence and solidity in video generation models, built on 150 hand-designed Blender generators across six task families. Each generator randomises lighting, camera angle, speed, and other nuisance parameters while preserving the task’s core cognitive-scientific structure, yielding 10,000 samples per task. We release this 1.5M-sample training corpus, a 300-question evaluation exam with human Elo ratings across 14 video generation models, and PWM-WROP, a fine-tuned continuation model trained on WROP data achieving state-of-the-art performance in its category.
Figure 2 : Taxonomy of the six WROP task families. The top row evaluates object permanence through dynamic occlusion, static-scene occlusion, and container-based concealment and movement; the bottom row evaluates object solidity through obstruction, support removal, and collision. Each panel shows representative frames from a sample drawn from one generator in the corresponding task family: an early input frame, the final input frame at the split boundary, and a target frame depicting the expected physical outcome.
Figure 3 : Generator-level composition of WROP’s 150 generators. The inner ring shows 90 OP and 60 OS generators; the outer ring shows the six task families.
Figure 4 : Generator-design example from Marked_Boxes_Swap using the current 60-frame input and 60-frame target clips. The sequence shows the objects in marked boxes, the closed boxes after crossing at the split, completion of the position exchange, and the identity-preserving reveal. The task requires the model to maintain the association between each hidden object and its marked container across the input–target boundary rather than infer identity from final screen position.
Model
Access
Interface class
Native output (res / fps / frames)
Ours
PWM-WROP
Trainium2 48XL
True continuation
320×192 / 24 / 60 predicted †
Open-weight
MAGI-1 24B
Local, multi-GPU
True continuation
1280×720 / 24 / 77
LTX-2.3 Dev
Local, multi-GPU
Edit / transfer (IC-LoRA)
1536×1024 / 24 / 121
Wan-VACE 14B
Local, multi-GPU
Edit / transfer (repaint)
832×480 / 16 / 57
Table 1: The fourteen evaluated models, grouped by provenance, with access route, interface class and native output geometry. † Our model emits a 117-frame clip that replays its 57 conditioning frames before the 60 predicted frames; the replay is trimmed before evaluation. ‡ Stitched or padded outputs whose source prefix is trimmed before evaluation (Section 4.3 ).
Rank
Model
Class
Elo
95% CI
Score rate
Games
1
Wan 3.0 Prime
Reference-to-video
1723.6
[1629.2, 1864.9]
77.9%
52
2
MiniMax H3
Reference-to-video
1723.6
[1644.5, 1837.6]
77.9%
52
3
PWM-WROP (ours)
True continuation
1679.5
[1603.5, 1781.5]
73.1%
52
4
Seedance 2.5
Reference-to-video
1649.6
[1554.2, 1751.5]
69.6%
51
5
Runway Aleph 2
Edit / transfer
1518.3
[1425.1, 1621.6]
52.9%
51
6
Wan-VACE 14B
Edit / transfer
1506.7
[1429.9, 1590.8]
51.0%
52
Table 2 : Human preference leaderboard of the 14 evaluated models on the 300-question WROP exam. Twenty raters made 361 blind pairwise judgments between the fourteen models (50–52 per model; ties count 0.5 for each side); strengths are Bradley–Terry maximum-likelihood estimates, with one virtual draw added per model and distributed evenly across its opponents, rescaled to an Elo scale with mean 1500. 95% confidence intervals come from 1,000 rater-clustered bootstrap resamples; overlapping intervals should not be read as significant rank differences. Score rate is the raw win rate (win =1 , tie =0.5 ). Wan 3.0 Prime and MiniMax H3 have identical records and tie for first.
Figure 5 : Human-preference Elo of the 14 models. Points are Bradley–Terry strengths on an Elo scale (mean 1500, dashed line), bars are 95% rater-clustered bootstrap intervals, colour is interface class, and PWM-WROP (filled) is the top true-continuation model.
Figure 6 : Within-family human-preference ranks across the six WROP task families. Each cell shows the Bradley–Terry rank within that family (1 = best, 14 = worst), fitted separately on the 20-rater pairwise judgments for each family (36–96 games per family; 361 total). Color encodes rank from green (top) to red (bottom). The outlined row marks PWM-WROP; overall Elo (right) is from the joint fit over all judgments. Row order follows overall Elo rank; column groups correspond to the task families in Figure 2 . Interface class is indicated by the label color, matching Figure 5 .
Figure 7 : Qualitative comparison on three object-permanence task families (Figure 2 , top row). Each row shows the same ground-truth input (two key frames, blue), the PWM-WROP target output (two frames, green), and the full baseline output (three frames spanning the generation, red), with ✓ / × badges indicating physical correctness. G43 (OP-1): Gemini Omni Flash collapses all three balls into a tight cluster at the tunnel exit, violating their lane identities and count, while PWM-WROP keeps each ball in its correct lane with continuous motion. G19 (OP-2): Seedance 2.5 generates an oversized occluder over an empty region of the scene while the original objects remain visible, then makes all three objects suddenly appear when the occluder moves away, with no clear continuity from their previous positions; PWM-WROP correctly preserves the scene and reveals the same unchanged configuration. G66 (OP-3): Seedance 2.5 correctly animates the 180° turntable rotation but then lifts the cup at the original front position–a different physical cup–rather than tracking the cup containing the ball to its new location; PWM-WROP correctly tracks the ball with the rotating cup and lifts the correct cup.
Model
LPIPS ↓
MS-SSIM ↑
SSIM ↑
PSNR ↑
MSE ( ×10−3 ) ↓
CLIP ↑
FID ↓
Wan 3.0 Prime
0.115
0.861
0.942
27.15
4.48
0.948
15.1
MiniMax H3
0.105
0.877
0.938
27.51
9.09
0.962
14.8
PWM-WROP (ours)
0.081
0.921
0.917
26.45
2.97
0.956
20.4
Seedance 2.5
0.282
0.616
0.770
18.37
28.55
0.928
24.6
Runway Aleph 2 †
0.181
0.789
0.918
24.98
9.71
0.918
19.7
Wan-VACE 14B
0.349
0.744
0.842
17.88
24.21
0.872
33.0
Table 3 : Automatic full-reference metrics against the reference continuation, averaged over the 300 exam questions and ordered by human Elo ( Table 2 ). For each model, the predicted span is extracted using its per-model trim rule, uniformly resampled to the 60 target frames, and compared at 320 × 192; CLIP (ViT-B/32, 224 × 224) and FID (299 × 299, all frames at stride 2, ≈ 9,030 frames per side) are computed at their respective input resolutions. The best result in each column is shown in bold and the second-best is underlined. These metrics measure similarity to the target rather than object-permanence capability: edit/transfer models can repaint the static scene and score well on structural metrics without depicting the required reappearance.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
ID
Name
Description
OP-1: Baillargeonian_Occlusion (26 tasks)
G02
ramp_tunnel
An object rolls down a ramp, passes through an opaque tunnel, reappears …
G04
high_low_cover
Objects pass behind high/low covers on crossing or straight paths; identity is preserved.
G14
u_tube_three_lanes
Three coloured balls travel parallel U-tube lanes, preserve identity …
G17
moving_tray
A moving tray carries objects behind a fixed screen (one or two lanes).
G28
open_ended_tunnel
An object travels through an open-ended tunnel, hidden in the middle.
Appendix
Table 4 : Task inventory of all 150 WROP generators across the six task families. Each entry names the generator and its core physical scenario.
+ multi-tensor AdamW, one sync per step; FSDP2 prefetch depth 2
5.71
12,454
Appendix
Table 5: Training step time on one trn2.48xlarge (36-layer Cosmos3-Nano, tp4 × fsdp16, batch 16, 288×512×30 latent frames; medians over 10 steps after 3 warm-up steps). Loss trajectories are identical step for step across all rows.
Model
Class
Overall
Object Permanence
Object Solidity
Elo
OP-1
OP-2
OP-3
OS-1
OS-2
OS-3
Wan 3.0 Prime
Ref.-to-video
1723.6
1670 2
1559 6
1730 1
1558 7
1749 3
1915 2
MiniMax H3
Ref.-to-video
1723.6
1614 4
1689 2
1707 2
1625 4
1868 1
1925 1
PWM-WROP (ours)
True cont.
1679.5
1616 3
1785 1
1692 3
1604 5
1825 2
1394 8
Seedance 2.5
Ref.-to-video
1649.6
1710 1
1644 3
1656 4
1687 1
1489 7
1714 4
Runway Aleph 2
Edit/transfer
1518.3
1448 10
1552 7
1596 5
1641 3
1480 8
1331 12
Appendix
Table 6 : Per-family Bradley–Terry Elo for all 14 models across the six WROP task families. Superscripts give within-family rank (1 = best). Underline marks the top-ranked model in each family; bold marks PWM-WROP (ours). Overall Elo is from the joint fit over all 361 judgments; models are ordered by overall Elo rank. Bootstrap CIs are wide especially in OS families (36–44 total games each); ranks indicate tendency rather than significant differences.
Figure 9 : Per-family human-preference Elo leaderboards for all six WROP task families. Points are Bradley–Terry MLE strengths on an Elo scale (mean 1500, dashed line); bars are 95% rater-clustered bootstrap intervals. Colour encodes interface class as in Figure 5 : filled blue circle is PWM-WROP (ours); open circles are other true-continuation models (blue), reference-to-video models (orange), and edit/transfer models (green). Each panel is sorted independently by within-family rank. Intervals are especially wide in OS families (36–44 total games each); adjacent ranks are rarely distinguishable. Numerical values and within-family ranks are reported in Table 6 .
Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving into "world simulators." Yet the community still lacks a benchmark that directly tests whether a model can reason about how an observed world should evolve over time. We introduce WorldReasonBench, which reframes video generation evaluation as world-state prediction: given an initial state and an action, can a model generate a future video whose state evolution remains physically, socially, logically, and informationally consistent? WorldReasonBench contains 436 curated test cases with structured ground-truth QA annotations spanning four reasoning dimensions and 22 subcategories. We evaluate generated videos with a human-aligned two-part methodology: Process-aware Reasoning Verification uses structured QA and reasoning-phase diagnostics to detect temporal and causal failures, while Multi-dimensional Quality Assessment scores reasoning quality, temporal consistency, and visual aesthetics for ranking and reward modeling. We further introduce WorldRewardBench, a preference benchmark with approximately 6K expert-annotated pairs over 1.4K videos, supporting pair-wise and point-wise reward-model evaluation. Across modern video generators, our results expose a persistent gap between visual plausibility and world reasoning: videos can look convincing while failing dynamics, causality, or information preservation. We will release our benchmarks and evaluation toolkit to support community research on genuinely world-aware video generation at https://github.com/UniX-AI-Lab/WorldReasonBench/.
Keming Wu, Yijing Cui, Wenhan Xue +11
Tsinghua University · Nanyang Technological University · Hong Kong University of Science and Technology (Guangzhou) +2
Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions. Existing vision-language models (VLMs) often struggle to interpret these dynamics and reason reliably about future and counterfactual outcomes. We introduce PhysMind, a training-free agentic framework that constructs one reusable, question-agnostic executable world per video. PhysMind recovers a temporally consistent dynamic scene through object segmentation, mesh reconstruction, and 6D pose tracking, then fits analytic continuous-time dynamics and latent physical parameters without unrolling a time-stepped simulator. Given a question, it inspects, continues, or edits the world and answers from the resulting trajectories and interactions. Relative to direct chain-of-thought (CoT) reasoning with the same VLM, PhysMind improves accuracy by 38.23 points on CLEVRER and 8.08 points on Physion++. On counterfactual questions, it exceeds the strongest evaluated VLM baseline, GPT-5.5, by 19.25 points.
Chen Yang, Shenxiang Zeng, Haoyang Zhao +6
1Tsinghua University · 2The University of Hong Kong
We present Astronex-World 1.0, an open controllable video world-model foundation. Given a text prompt (text-to-video) or an initial observation (image-to-video), the model predicts future visual states under frame-aligned camera trajectories, continuous actions, and an embodiment identifier, and accepts text events inserted at a specified position of a rollout. The family provides a bidirectional model for full-context generation and a causal model with block-causal attention and cross-block KV caching for persistent generation, both built on the Wan2.2-TI2V-5B prior. PRoPE injects camera intrinsics and extrinsics, while a 64-dimensional action stream modulates every Transformer layer. A five-stage training path develops bidirectional camera and action control, converts the backbone to block-causal generation, distills a few-step student, restores mixed-domain dynamics, and applies asymmetric DMD/DMD2 distribution matching. The causal model generates 832x480 video at 24 fps. All five training stages run on two NVIDIA L20 48 GB GPUs, and the causal model streams in real time on one. It scores 73.5 on WBench Navi and 70.0 on WBench Full. On Full, this 5B model is above the 13.6B LongCat-Video and the 14B Helios, within one point of the 22B LTX-2.3, and above YUME 1.5, which is post-trained from the same 5B prior on NVIDIA A100 GPUs. The reserved action input and output interfaces allow post-training for embodied intelligence and autonomous driving.
Xin Zhou, Cong Miao
Astronex Robotics · Nanjing University of Information Science and Technology