Tracking Is Not Permanence: What Video World Models Keep of a Hidden Object
Organizations: TUM School of Computation, Information and Technology Technical University of Munich, Germany
Abstract
Video world models track objects they can see; we ask what they keep of objects they cannot. We hide an object from a frozen V-JEPA 2 predictor and compare its prediction for the hidden region with the encoder's representation of two worlds that differ only inside that region. The predictor's decision keeps a stationary object in part and one carried inside a container not at all, and loses a moving one within 0.3 s (0.5 s under V-JEPA's own tube mask; ViT-H keeps it to 1.1 s at pretraining's 90% masking ratio); in projection a trace remains, below the midpoint, at 14-60% of what a baseline copying the last view retains. The information is there: the encoder reads the object's presence at 1.00 and keeps a closed container's contents decodable for 3.5 s, while the predictor's output, read with the encoder's own probe, contains the ball in 2% of scenes once the box has been closed for half a second. On rendered scenes, permanence is missing on the predictor's side, and training installs it cheaply as a prior: three thousand predictor-only steps on synthetic containers take this belief from 0.05 to 1.00 against two matched controls. They also raise IntPhys-2019 from 84.2% to 93.3%, but so does a curriculum without containers, and which training habit the benchmark credits changes with its scoring rule. Continued training with tube masks produces 1.1-1.6 s of moving-object carry-over on manipulation and internet-style video, so the deficit is not intrinsic to latent prediction. VideoMAE keeps almost nothing, and Cosmos's next-token prediction keeps a stationary hidden object but not one carried inside a moving container.
Figures & tables
| ViT-L | ViT-H | ViT-g | |
|---|---|---|---|
| released (baseline) | 61.4 | 84.2 | 54.2 |
| + container curriculum, 3,000 predictor-only steps | 64.4 | 93.3 | 75.0 |
| two further seeds / readout | – / 64.4 | 93.1, 93.3 / 92.8 | – / 73.9 |
| no container masks (3 seeds for ViT-H) † | 67.8 | 94.3 | 77.5 |
| static / non-persistent † | – / 66.7 | 93.8 / 88.7 | – / 70.0 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Intervention | Data | Trained | Installs |
|---|---|---|---|
| (i) temporally-partial masks | natural video | predictor, 20k | object kept where last seen |
| (ii) container masks | synthetic | predictor, 3k | containment prior; IntPhys |
| (iii) single mask kind | synthetic | predictor, 3k | carry-over only |
| (iv) tube masks, EMA | video of (i) / internet | enc. + pred., 16k | 1.1–1.6 s carry-over |
| (v) three-world coin | real hands | ctx. enc. + pred., 8k | coin belief, unseen people |
| Object | Model | 0 | 0.27 | 0.53 | 0.80 | 1.07 | 1.33 | 1.60 |
|---|---|---|---|---|---|---|---|---|
| moving | ViT-L released | 0.58 | 0.00 | 0.03 | 0.02 | 0.02 | 0.00 | 0.02 |
| moving | ViT-H released | 0.92 | 0.00 | 0.18 | 0.00 | 0.03 | 0.00 | 0.00 |
| moving | ViT-g released | 0.98 | 0.17 | 0.10 | 0.02 | 0.03 | 0.02 | 0.10 |
| moving | ViT-H, never rendered (control) | 0.30 | 0.00 | 0.02 | 0.00 | 0.00 | 0.00 | 0.00 |
| moving | ViT-H + temporally-partial (v2) | 1.00 | 0.95 | 0.95 | 0.63 | 0.57 | 0.37 | 0.42 |
| moving | v2, never visible (control) | 0.18 | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | – |
| Training masks | 0 | 0.27 | 0.53 | 0.80 | 1.07 | 1.33 | 1.60 | IntPhys-2019 |
|---|---|---|---|---|---|---|---|---|
| released ViT-H (none) | 0.92 | 0.00 | 0.18 | 0.00 | 0.03 | 0.00 | 0.00 | 84.2 |
| tube only (released mask kind) | 1.00 | 0.98 | 0.97 | 0.77 | 0.77 | 0.58 | 0.53 | 91.1 |
| temporally-partial only | 1.00 | 0.97 | 0.97 | 0.68 | 0.75 | 0.67 | 0.50 | 87.8 |
| causal only | 1.00 | 1.00 | 0.97 | 0.72 | 0.72 | 0.58 | 0.42 | 91.9 |
| ViT-L, 16k | moving, at | stat. | 3 worlds at 0.80 s | IntPhys- | ||
|---|---|---|---|---|---|---|
| 0.27 s | 0.80 s | 1.33 s | 1.07 s | moving / static / gone | 2019 | |
| released | 0.00 | 0.02 | 0.00 | 0.12 | 0.00 / 0.03 / 0.97 | 61.4 |
| + tube masks, seed 1 | 0.97 | 0.90 | 0.52 | 0.28 | 0.45 / 0.37 / 0.18 | 62.8 |
| + tube masks, seed 2 | 1.00 | 0.82 | 0.48 | 0.37 | 0.27 / 0.57 / 0.17 | 64.2 |
| + tube masks, internet | 0.93 | 0.90 | 0.80 | 0.15 | 0.72 / 0.17 / 0.12 | 62.5 |
| + mask mixture | 0.98 | 0.32 | 0.15 | 1.00 | 0.10 / 0.82 / 0.08 | 53.1 |
| Model | Condition | 0.53 | 1.07 | 2.00 | 2.93 |
|---|---|---|---|---|---|
| ViT-H released | moving box | 0.05 | 0.02 | 0.00 | 0.00 |
| ViT-H released | never-shown box (control) | 0.00 | 0.00 | 0.00 | 0.00 |
| ViT-H released | stationary box | 0.05 | 0.03 | 0.17 | 0.08 |
| ViT-H released | stationary box, controls | 0.02 / 0.03 | 0.00 / 0.00 | 0.03 / 0.07 | 0.00 / 0.00 |
| v2 | moving box | 0.30 | 0.32 | 0.15 | 0.05 |
| v2 | stationary box | 0.23 | 0.38 | 0.53 | 0.47 |
| Model | Held-out set | coin-in | removed | never |
|---|---|---|---|---|
| released ViT-H | MagicCATs hands | 0.36 | 0.42 | 0.33 |
| predictor only, 2 worlds | EgoDex | 0.75 | 1.00 | 0.00 |
| encoder trained, 2 worlds | EgoDex | 1.00 | 1.00 | 0.00 |
| encoder trained, 3 worlds | EgoDex | 1.00 | 0.33 | 0.00 |
| encoder trained, 3 worlds + containers | EgoDex | 1.00 | 0.00 | 0.00 |
| clean recipe, s1 / s2 | EgoDex | 1.00 / 1.00 | 0.00 / 0.00 | 0.00 / 0.00 |
| vs. never-shown box | vs. exit control | ||||
|---|---|---|---|---|---|
| Condition | Hidden | nats | nats | ||
| box moving | 1.07 s | 0.90 | 0.53 | ||
| box moving | 2.13 s | 0.93 | 0.23 | ||
| box stationary | 1.07 s | 0.93 | 0.97 | ||
| box stationary | 2.13 s | 1.00 | 0.87 | ||
| visible continuation | – | 1.00 | – | – | |
| moving / stationary, time since last visible (s) | container, time hidden (s) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| readout | condition | 0 | 0.27 | 0.53 | 1.07 | 2.93 | 0.53 | 1.07 | 2.00 | 2.93 |
| rule (paper) | main | 0.92 / 1.00 | 0.00 / 0.83 | 0.18 / 0.63 | 0.03 / 0.52 | 0.00 / 0.25 | 0.05 | 0.02 | 0.00 | 0.00 |
| never rendered | 0.30 / 0.67 | 0.00 / 0.23 | 0.02 / 0.07 | 0.00 / 0.00 | 0.00 / 0.02 | 0.00 | 0.00 | 0.00 | 0.00 | |
| exit | 0.30 / – | 0.00 / – | 0.02 / – | 0.00 / – | 0.00 / – | 0.00 | 0.00 | 0.00 | 0.00 | |
| copy-last-view | main | 0.98 / 0.92 | 1.00 / 0.80 | 0.92 / 0.98 | 0.58 / 0.98 | 0.08 / 1.00 | 0.37 | 0.35 | 0.30 | 0.28 |
| encoder probe on | main | 0.68 / 1.00 | 0.22 / 0.50 | 0.08 / 0.02 | 0.00 / 0.58 | 0.00 / 0.18 | 0.02 | 0.00 | 0.00 | 0.00 |
| fixed | best fixed | min over | ||
| ViT-H | released | 84.2 | 84.2 ( ) | 70.0 |
| + container curriculum | 93.3 | 93.3 ( ) | 91.4 | |
| + no container masks (original run) | 93.9 | 94.2 ( ) | 93.6 | |
| moving / static / non-persist. ∗ (3 seeds) | 94.3 / 93.8 / 88.7 | 94.7 / 93.9 / 90.6 | 93.3 / 90.6 / 91.0 | |
| + real-hand clean recipe (2 seeds) | 27.2 / 30.6 | 42.8 / 47.8 ( ) | 38.6 / 43.9 | |
| ViT-L | released | 61.4 | 61.4 ( ) | 45.0 |