Video world models track objects they can see; we ask what they keep of objects they cannot. We hide an object from a frozen V-JEPA 2 predictor and compare its prediction for the hidden region with the encoder's representation of two worlds that differ only inside that region. The predictor's decision keeps a stationary object in part and one carried inside a container not at all, and loses a moving one within 0.3 s (0.5 s under V-JEPA's own tube mask; ViT-H keeps it to 1.1 s at pretraining's 90% masking ratio); in projection a trace remains, below the midpoint, at 14-60% of what a baseline copying the last view retains. The information is there: the encoder reads the object's presence at 1.00 and keeps a closed container's contents decodable for 3.5 s, while the predictor's output, read with the encoder's own probe, contains the ball in 2% of scenes once the box has been closed for half a second. On rendered scenes, permanence is missing on the predictor's side, and training installs it cheaply as a prior: three thousand predictor-only steps on synthetic containers take this belief from 0.05 to 1.00 against two matched controls. They also raise IntPhys-2019 from 84.2% to 93.3%, but so does a curriculum without containers, and which training habit the benchmark credits changes with its scoring rule. Continued training with tube masks produces 1.1-1.6 s of moving-object carry-over on manipulation and internet-style video, so the deficit is not intrinsic to latent prediction. VideoMAE keeps almost nothing, and Cosmos's next-token prediction keeps a stationary hidden object but not one carried inside a moving container.
Figures & tables
Figure 1: The belief probe reads object permanence out of a frozen predictor. Once the object is hidden, its region is dropped from the encoder input and the frozen predictor P infers the missing tokens; the prediction z^ is compared with the frozen encoder’s representation of two worlds that are pixel-identical except inside that region. The fraction of decisions for world A is P(present) : one value per 0.27 s tubelet, chance 0.5.
Figure 2: Scale does not buy permanence: on a moving object under the main query all three released models cross to absent after one tubelet, and a predictor trained with temporally-partial masks keeps a moving object for 1.1 s and a stationary one to the end of the window. Decision that the hidden object is still there, by seconds since it was last visible. Grey: the three released scales (thick: ViT-H); red dashed: the never-rendered control for ViT-H (the floor for no memory); dotted: chance. Shaded: binomial 95% intervals over 60 scenes.
Figure 3: The released predictor fills in only from its immediate neighbours: revealing the object in the last four frames moves its prediction only in the adjacent tubelet, while the trained predictor carries the object forward and revises it from the future. Decision in the gap when the reveal shows the object (solid) or its absence (dashed), against no reveal (dotted).
Figure 4: Continued training with tube masks produces moving-object carry-over on either video source but does not keep a stationary object; with the mask mixture the object freezes instead. Belief for a hidden moving (solid) and stationary (dashed) object, ViT-L, one panel per training condition; the two lines in the second panel are two seeds. Shaded: binomial 95% intervals.
Figure 5: VideoMAE keeps almost nothing; Cosmos keeps a stationary hidden object and not a carried one. (a) Belief curves for the textured ball under the pixel readout (VideoMAE) next to the latent readout (V-JEPA 2, released and after (i)). (b) Cosmos-Predict1-4B: fraction of 30 scenes expecting the ball at the reveal, against two controls; only a stationary box survives the exit control.
Figure 6: Training installs what the released model lacks. (a) Three thousand predictor-only steps take a ball inside a moving box from at most 0.05 to 1.00 at every hidden duration, with the never-shown-box and exit controls at 0.00. (b) Training the context encoder on real hands, intervention (v), separates coin-in, removed and never on held-out clips and on two participants never seen in training (bars: mean of two seeds; dots: each seed).
Table 1: IntPhys-2019 dev set, pairwise accuracy (%) under the protocol of Section 3.5 . Two context frames; L1 readout: surprise from the released L1 loss instead of the trained scale head; seeds and other scoring rules in Table 10 ( † RTX 5090 machines, Appendix C ; –: not run).
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Intervention
Data
Trained
Installs
(i) temporally-partial masks
natural video
predictor, 20k
object kept where last seen
(ii) container masks
synthetic
predictor, 3k
containment prior; IntPhys +9.1
(iii) single mask kind
synthetic
predictor, 3k
carry-over only
(iv) tube masks, EMA
video of (i) / internet
enc. + pred., 16k
1.1–1.6 s carry-over
(v) three-world coin
real hands
ctx. enc. + pred., 8k
coin belief, unseen people
Appendix
Table 2: The five training interventions of Section 3.5 , all starting from the released encoder and, except (v), the released predictor; the IntPhys-2019 gain is for ViT-H with two context frames; (v) starts from a predictor already trained through (i), a container curriculum and hand-anchored masks (Appendix C ).
Object
Model
0
0.27
0.53
0.80
1.07
1.33
1.60
moving
ViT-L released
0.58
0.00
0.03
0.02
0.02
0.00
0.02
moving
ViT-H released
0.92
0.00
0.18
0.00
0.03
0.00
0.00
moving
ViT-g released
0.98
0.17
0.10
0.02
0.03
0.02
0.10
moving
ViT-H, never rendered (control)
0.30
0.00
0.02
0.00
0.00
0.00
0.00
moving
ViT-H + temporally-partial (v2)
1.00
0.95
0.95
0.63
0.57
0.37
0.42
moving
v2, never visible (control)
0.18
0.00
0.00
0.00
0.00
0.00
–
Appendix
Table 3: Belief that a hidden object is still there, P(present) , by seconds since it was last visible (60 decisions per cell; chance 0.5; the matched control is the floor for no memory). The released models drop to at most 0.17 at 0.27 s and at most 0.18 afterwards on a moving object; the predictor trained with temporally-partial masks (v2) keeps a moving object for 1.1 s and a stationary one for the whole window (bold: its values at or above 0.5). –: not evaluated.
Training masks
0
0.27
0.53
0.80
1.07
1.33
1.60
IntPhys-2019
released ViT-H (none)
0.92
0.00
0.18
0.00
0.03
0.00
0.00
84.2
tube only (released mask kind)
1.00
0.98
0.97
0.77
0.77
0.58
0.53
91.1
temporally-partial only
1.00
0.97
0.97
0.68
0.75
0.67
0.50
87.8
causal only
1.00
1.00
0.97
0.72
0.72
0.58
0.42
91.9
Appendix
Table 4: Mask kind does not decide carry-over on synthetic rigid-motion scenes; belief by seconds since last visible. Predictors trained on the same rendered scenes with a single mask kind each (3,000 steps, no container masks) acquire the same moving-object belief, keep a stationary object (0.77–1.00 at every step; last step 0.82 / 0.95 / 0.92 for causal / partial / tube) and none acquire containment (ball in a visible box at 0.53 s hidden: 0.05 / 0.18 / 0.18, same order); IntPhys-2019 (%, two context frames) rises under all three; in these scenes an object that enters a hidden block keeps moving, so predicting the block rewards continuation under each of these masks. These runs have no never-visible control of their own (Appendix D ).
ViT-L, 16k
moving, P(present) at
stat.
3 worlds at 0.80 s
IntPhys-
0.27 s
0.80 s
1.33 s
1.07 s
moving / static / gone
2019
released
0.00
0.02
0.00
0.12
0.00 / 0.03 / 0.97
61.4
+ tube masks, seed 1
0.97
0.90
0.52
0.28
0.45 / 0.37 / 0.18
62.8
+ tube masks, seed 2
1.00
0.82
0.48
0.37
0.27 / 0.57 / 0.17
64.2
+ tube masks, internet
0.93
0.90
0.80
0.15
0.72 / 0.17 / 0.12
62.5
+ mask mixture
0.98
0.32
0.15
1.00
0.10 / 0.82 / 0.08
53.1
Appendix
Table 5: Continued training with tube masks for 16k steps produces carry-over on the video of (i) (two seeds) and on internet-style video; with the mask mixture the model freezes the object instead. ViT-L with an EMA target; IntPhys-2019 (%) with two context frames; bold: P(present)≥0.5 after training, and the best IntPhys-2019 score. Never-visible controls in Appendix D .
Model
Condition
0.53
1.07
2.00
2.93
ViT-H released
moving box
0.05
0.02
0.00
0.00
ViT-H released
never-shown box (control)
0.00
0.00
0.00
0.00
ViT-H released
stationary box
0.05
0.03
0.17
0.08
ViT-H released
stationary box, controls
0.02 / 0.03
0.00 / 0.00
0.03 / 0.07
0.00 / 0.00
v2
moving box
0.30
0.32
0.15
0.05
v2
stationary box
0.23
0.38
0.53
0.47
Appendix
Table 6: A ball inside a visible box: belief at the moment the box opens, by seconds hidden. The released predictor loses it within 0.5 s (controls: never-shown box / exit); 3,000 predictor-only steps on synthetic containers install it at 1.00 with both controls at 0.00 (stationary-box controls, not shown: 0.00–0.02).
Model
Held-out set
coin-in
removed
never
released ViT-H
MagicCATs hands
0.36
0.42
0.33
predictor only, 2 worlds
EgoDex
0.75
1.00
0.00
encoder trained, 2 worlds
EgoDex
1.00
1.00
0.00
encoder trained, 3 worlds
EgoDex
1.00
0.33
0.00
encoder trained, 3 worlds + containers
EgoDex
1.00
0.00
0.00
clean recipe, s1 / s2
EgoDex
1.00 / 1.00
0.00 / 0.00
0.00 / 0.00
Appendix
Table 7: A three-world belief about closed real hands transfers to participants never seen in training. Belief that a coin is in the re-opened hand, hands closed for at most 0.6 s; the coin-in world should read high, the other two 0.00. Clean recipe: intervention (v) as specified in Section 3.5 , two seeds (s1 / s2); the rows above it are its development stages. Events: 8 (EgoDex), 4 (EPIC-KITCHENS) and 39 (MagicCATs), of which 3, 2 and 24 have a removed world, since removal needs the coin in view 0.8 s before closing; bold: all three worlds separated. Event counts in Appendix G .
vs. never-shown box
vs. exit control
Condition
Hidden
nats
P(>0)
nats
P(>0)
box moving
1.07 s
+20.1
0.90
−2.7
0.53
box moving
2.13 s
+18.2
0.93
−10.3
0.23
box stationary
1.07 s
+35.8
0.93
+20.7
0.97
box stationary
2.13 s
+38.7
1.00
+18.7
0.87
visible continuation
–
+153.4
1.00
–
–
Appendix
Table 8: An autoregressive model keeps a stationary hidden object and not a carried one. Cosmos-Predict1-4B, teacher-forced log-likelihood of the reveal chunk; memory is the per-scene difference between the main condition and a control (30 scenes per row); visible continuation: the ball stays in view and the chunk after the event is scored; –: not applicable.
Figure 7: Where the hidden moving object goes: nearest of three worlds over the hidden band. The released predictor says gone within a second; the predictor trained with temporally-partial masks (v2) puts it where it was last seen from the second tubelet on; gone is its most frequent answer only at 2.7–3.2 s.
moving / stationary, time since last visible (s)
container, time hidden (s)
readout
condition
0
0.27
0.53
1.07
2.93
0.53
1.07
2.00
2.93
L1 rule (paper)
main
0.92 / 1.00
0.00 / 0.83
0.18 / 0.63
0.03 / 0.52
0.00 / 0.25
0.05
0.02
0.00
0.00
never rendered
0.30 / 0.67
0.00 / 0.23
0.02 / 0.07
0.00 / 0.00
0.00 / 0.02
0.00
0.00
0.00
0.00
exit
0.30 / –
0.00 / –
0.02 / –
0.00 / –
0.00 / –
0.00
0.00
0.00
0.00
copy-last-view
main
0.98 / 0.92
1.00 / 0.80
0.92 / 0.98
0.58 / 0.98
0.08 / 1.00
0.37
0.35
0.30
0.28
encoder probe on z^
main
0.68 / 1.00
0.22 / 0.50
0.08 / 0.02
0.00 / 0.58
0.00 / 0.18
0.02
0.00
0.00
0.00
Appendix
Table 9: Symmetric readouts of the released ViT-H predictor’s inference. Fraction of the 60 scenes read as present (chance 0.5 for the probe on z^ , which reports held-out accuracy of main against a control). Moving and stationary: seconds since the object was last visible; container: seconds hidden before the box opens, read at the opening slot. Copy-last-view: the encoder’s tokens of the last visible slot, read with the L1 rule; –: no stationary exit control.
fixed C=2
best fixed C
min over C
ViT-H
released
84.2
84.2 ( C=2 )
70.0
+ container curriculum
93.3
93.3 ( C=2 )
91.4
+ no container masks (original run)
93.9
94.2 ( C=6 )
93.6
moving / static / non-persist. ∗ (3 seeds)
94.3 / 93.8 / 88.7
94.7 / 93.9 / 90.6
93.3 / 90.6 / 91.0
+ real-hand clean recipe (2 seeds)
27.2 / 30.6
42.8 / 47.8 ( C=10 )
38.6 / 43.9
ViT-L
released
61.4
61.4 ( C=2 )
45.0
Appendix
Table 10: IntPhys-2019 pairwise accuracy (%) under three aggregation rules: the fixed two-frame context of the main text, the best single context length in 2–10 frames chosen per model, and the minimum surprise over those context lengths per start frame (our implementation of the rule of Garrido et al., 2025 ); average surprise throughout. ∗ Evaluated on the single-card RTX 5090 machines (Appendix C ), mean of three seeds trained there; the original run is a separate seed trained and scored on the cluster; the two-seed rows list each seed.
Self-supervised video models are increasingly framed as world models, yet they are still evaluated almost entirely on clean video and reported as a final task score, obscuring how their representations behave under the degraded and ambiguous conditions a deployed world model must handle. We present the first systematic study of this component, analyzing four matched-capacity frontier self-supervised learning models that are strong candidates for world-model encoders -- V-JEPA 2.1, V-JEPA 2, VideoPrism, and VideoMAEv2 -- across five robustness axes relevant to their deployment as video world models: feature discriminability, corruption robustness, fine-grained discrimination, occlusion robustness, and sensitivity to temporal direction. We run the study on Something-Something-v2 and repeat every axis that ports on the egocentric EGTEA Gaze+, where the model ordering is unchanged. Our results reveal a distinct and consistent profile for latent-prediction models across all five axes. They degrade more gracefully under pixel corruption, preserve usable class structure rather than mere geometric stability under occlusion, capture fine-grained physical contact cues without reconstructing pixels, and uniquely encode the arrow of time. We finally test whether these representation-level differences matter for downstream world modeling, pairing each frozen encoder with an identical action-conditioned predictor and planner in a simulated manipulation environment. Only latent prediction yields near-complete task success, remains effective under degraded observations, and transfers to a manipulation task for which the predictor was never trained. Our results provide concrete new evidence that latent prediction is a promising foundation for robust world modeling.
Ali J Alrasheed, Aryan Yazdan Parast, Basim Azam +2
Video world models are increasingly used to provide predictive visual representations, yet it remains unclear which pretraining signals induce action-relevant structure in their latent spaces. We study this question through a unified probe-based evaluation across diverse encoder families, including image-only self-supervision, video pretraining with and without latent prediction, reconstruction-based autoencoders, diffusion models, and shortcut-forcing dynamics models. Using a common inverse-dynamics probing objective, we find that action-relevant structure is driven primarily by temporal video pretraining rather than pixel reconstruction fidelity: models with strong pixel decoding quality can exhibit near-zero action recoverability, while video-pretrained self-supervised encoders consistently achieve the best Pareto trade-off between visual fidelity and action prediction. Comparing V-JEPA and VideoMAE further shows that most gains arise from natural-video temporal context, with feature-level latent prediction providing a smaller additional benefit. These trends transfer across robotic benchmarks, though CALVIN reveals that static-environment tasks can partially mask the importance of temporal structure by allowing strong image priors to suffice. Finally, inverse-dynamics supervision substantially improves robustness to visual corruption, suggesting that action-aware objectives regularize latent geometry beyond clean-setting performance. Our results identify temporal predictive structure -- not reconstruction fidelity -- as the primary ingredient underlying action-relevant video representations.
Jewon Yeom, Hanseul Kim, Jeongjae Park +3
Graduate School of Data Science, Seoul National University
Predictive video models have emerged as promising world models by learning latent visual dynamics from large-scale video. Yet these models remain challenged by physical events under occlusion, where later predictions may depend on object evidence that is no longer available in the current view. Addressing this challenge requires historical evidence not only to be preserved but also to remain accessible when it becomes relevant to a subsequent prediction. Existing approaches mainly enlarge the temporal context, cache generic video features, or impose explicit object-centric states, thereby improving the capacity or structure of retained history. However, they do not directly address how relevant historical evidence can be selectively retrieved and integrated into a pretrained predictor without interfering with its native latent workspace. Accordingly, we introduce HERA (Historical Evidence Routing Adapter), a framework for routing retained historical evidence into a frozen latent predictor, and instantiate it with Register-Routed Patch Memory (RRPM), a lightweight adapter comprising a Structured Memory Bank, Memory Registers, and Workspace Registers. On the IntPhys2 Main split, HERA with RRPM improves the pairwise AvgSurprise accuracy of V-JEPA 2-G from 52.57% to 54.35%. Subgroup analysis shows particularly strong improvements on fixed-camera continuity, from 46.15% to 57.69%, and fixed-camera immutability, from 46.15% to 63.46%. These results support historical evidence routing as a practical adaptation strategy for physical prediction in latent world models.
Yuanruyi, Yue Cao, Haojia Gao +7
Chongqing University · Everwise-Tech Co., Ltd. · Research Institute of Tsinghua University in Shenzhen +2