World models learned from high-dimensional visual observations allow agents to make decisions and plan directly in latent space, avoiding pixel-level reconstruction. However, recent latent predictive architectures (JEPAs), including the DINO world model (DINO-WM), display a degradation in test time robustness due to their sensitivity to ``slow features". These include visual variations such as background changes and distractors that are irrelevant to the task being solved. We address this limitation by augmenting the predictive objective with a bisimulation encoder that enforces control-relevant state equivalence, mapping states with similar transition dynamics to nearby latent states while limiting contributions from slow features. We evaluate our model on a navigation task (PointMaze) and on a manipulation task (PushT) under different test-time background changes and visual distractors. Across all benchmarks, our model consistently improves robustness to slow features while operating in a reduced latent space, up to 10× smaller than that of DINO-WM. Moreover, our model is agnostic to the choice of pre-trained visual encoder and maintains robustness when paired with DINOv2, SimDINOv2, and iBOT features.
Figures & tables
Figure 1 : Visually distinct observations that differ only in background (checkerboard and gradient) are first mapped to latent embeddings Z and Z′ by a pretrained encoder (initial state in green, goal in red). A bisimulation encoder then projects these into lower-dimensional representations W and W′ , which are equivalent under on-policy transition dynamics.
Figure 2 : Rollouts for PointMaze navigation under a background change at test time. For clarity, the background change is not depicted in the figure. See Section 4.2 for details on background changes. While DINO-WM fails to reach the goal due to background change our model succeeds.
Figure 3 : Model architecture and training objectives.
Figure 4 : Bisimulation encoder with patch-based design, with 196 of patches and 32 output patch dimension.
Figure 5 : Visualization of the first principal component (PC1) of latent embeddings produced by different visual feature encoders for a PointMaze observation. In all cases, PC1 predominantly encodes background and layout information rather than control-relevant features, motivating our PCA-based VC-Reg in the bisimulation encoder.
Figure 6 : Visual distribution shifts used to evaluate robustness on PointMaze (top) and PushT (bottom). From left to right, the conditions are: NC : No Change, SC : Slight Change, C : Color Gradient, LC : Large Color Change, LCG : Large Color Gradient, and D : moving Distractors.
Model
NC (Abs.)
SC (Rel.%)
C (Rel.%)
LC (Rel. %)
LCG (Rel.%)
D (Rel.%)
Avg (Rel.%)
Comparison with DR
DINO-WM
0.8
10.0
25.0
30.0
40.0
2.5
21.5
DINO-WM w/DR
0.82
0.0
0.0
17.1
22.0
0.0
7.8
Ours (DINOv2)
0.78
−2.6
2.6
−10.3
0.0
−5.1
{\color[rgb]{0,0,1}\mathbf{-3.1}}
Encoder ablation
End-to-End
0.68
35.3
−2.9
61.8
47.1
5.9
29.4
Table 1 : PointMaze. Absolute success rate under NC ( ↑ ) and relative degradation (%, ↓ ) under shifts. Negative values indicate improvement. Avg.: mean degradation over the five background shifts.
Model
NC (Abs.)
SC (Rel.%)
C (Rel.%)
LC (Rel.%)
LCG (Rel.%)
D (Rel.%)
Avg. (Rel.%)
DINO-WM
0.48
21.0
21.0
71.0
50.0
75.0
47.6
DINO-Bisim
0.36
0.0
0.0
11.0
11.0
17.0
7.8
Table 2 : PushT under data scarcity. Both models are trained on 1,000 samples. We report absolute success rate under NC ( ↑ ) and relative degradation (%, ↓ ) under five visual shifts. Avg. denotes mean relative degradation across the shifted conditions.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7 : Overview of the DINO-Bisim architecture.
Table 10
Figure 8 : First two principal components (PCs) of three distinct states (colors) observed under three different backgrounds (marker shapes) for the PointMaze task. The goal of the bisimulation encoder is to map identical states across backgrounds to nearby latent representations. Left. DINOv2 embeddings zt=fθ(ot) show strong separation along the first PC for identical states under different backgrounds, indicating that background variation dominates the leading principal direction. Middle. Bisimulation embeddings wt=hη(zt) after 50 epochs with standard VICReg still exhibit separation along the first PC, preserving background-dependent variation. Right. Bisimulation embeddings after 90 epochs with PCA-based VICReg (activated at epoch 50) show improved clustering of identical states across backgrounds, consistent with suppression of task-irrelevant variation.
Figure 9 : First twp principal components (PC1 and PC2) of latent embeddings produced by DINOv2, SimDINOv2, and iBOT for a PointMaze observation. In all cases, PC1 captures a large percentage of the total variance that predominantly encodes background and layout information rather than control-relevant features.
Figure 10 : PointMaze task, from leftmost the initial state to the rightmost the goal state
Figure 11 : Push T Task, from leftmost the start state to the rightmost the goal state
Figure 12 : PointMaze planning failure (DINO-WM). Six MPC steps (rows) under six visual conditions (columns: NC, SC, C, LC, LCG, D). Markers: start (orange), goal (green); red line: rollout path. Same init/goal as Fig. 13 . DINO-WM fails to reach the goal, especially under LC/LCG.
Figure 13 : PointMaze planning success (DINO-Bisim). Same layout and shared episode as Fig. 12 . Bisimulation-aligned representations yield a goal-reaching trajectory across all six appearance conditions.
Figure 14 : PushT planning failure (DINO-WM). Six MPC steps × six visual conditions; start/goal markers and executed path overlaid. Matched init/goal with Fig. 15 . WM misaligns the block with the target pose under shifted backgrounds.
Figure 15 : PushT planning success (DINO-Bisim). Same episode as Fig. 14 . DINO-Bisim completes the push despite lighting/color perturbations at test time.
Figure 16 : Examples of domain-randomized visual observations for the PointMaze task. During training, background appearance is randomized across trajectories using color shifts, gradient backgrounds, and fixed distractors (i.e., the dot outside the maze) while the underlying maze layout and transition dynamics remain unchanged.
Figure 17 : Training losses value versus the epochs. Ground-truth reward distance is shown only as a diagnostic and is not used during training.
Figure 18 : Moving distractor setting in PointMaze. Yellow and magenta dots move continuously within the maze, introducing dynamic but task-irrelevant visual variation while preserving the underlying transition dynamics.
Horizon
NC
SC
C
5
0.78
0.80
0.76
10
0.88
0.88
0.88
15
0.94
0.88
0.92
20
0.88
0.92
0.90
25
0.94
0.94
0.88
50
0.84
0.90
0.92
Appendix
Table 5 : PointMaze. Planning-horizon ablation for DINO-Bisim. Absolute success rates under the native background (NC) and background changes (SC and C). Higher is better.
Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations. We propose Delta-JEPA, an end-to-end reconstruction-free world model that augments latent forward prediction with a Latent Difference Action Decoder (LDAD). Unlike inverse decoders that infer actions from concatenated endpoint embeddings, LDAD reconstructs the executed action from the latent displacement between consecutive observations. This displacement-level supervision directly regularizes transition geometry: adjacent embeddings cannot collapse without losing action information, and different actions are encouraged to induce distinguishable latent changes for rollout-based planning. Delta-JEPA uses only latent prediction and action reconstruction, avoiding pixel reconstruction and distribution-matching regularizers. Across four visual continuous-control tasks, Delta-JEPA improves planning over JEPA-based and representation-learning world model baselines. Ablations show that displacement-based action decoding is consistently more effective than endpoint concatenation, and action-sensitivity analyses show clearer action-conditioned latent responses. These results indicate that supervising latent differences is a simple and effective mechanism for collapse-resistant and action-sensitive world model learning.
Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan +11
School of Computer Science and Technology, University of Chinese Academy of Sciences, Beijing · Institute of Information Engineering, Chinese Academy of Sciences, Beijing · School of Computer Science and Technology, Harbin Institute of Technology, Weihai +2
Joint-Embedding Predictive Architectures (JEPAs) have emerged as a principled framework for self-supervised learning of world models in compact latent spaces, yet existing methods are fragmented: some predict masked parts of a single image in latent space (I-JEPA), others learn to predict global photometric transformations (Image World Models), while video-scale JEPAs predict future temporal states and are post-trained for action-conditioned planning (V-JEPA~2, DINO-World, DINO-WM). These objectives are treated as distinct recipes with separate encoders, predictors, and anti-collapse regularizers, hindering a single model from unifying image-level and video-level world modeling. We present UniJEPA, a unified JEPA that jointly learns photometric prediction (image-level transformations) and temporal prediction (video-level next-state dynamics) in one shared latent space. A single end-to-end objective, composed of a next-embedding prediction loss and a Gaussian regularizer, yields a provably anti-collapse encoder-predictor pair trainable from raw pixels without EMA, stop-gradient, or pre-trained encoders. We show that the same latent space supports controllable abstraction: photometric prediction learns invariant structure while temporal prediction learns equivariant dynamics. After action-conditioned post-training on offline trajectories, UniJEPA enables zero-shot planning by treating goal features as prediction targets. On image, video, and control benchmarks, UniJEPA matches or surpasses task-specific JEPAs while requiring a single loss hyperparameter, and plans up to tens of times faster than generative world models at comparable accuracy.
An Lanji, Dawei Liu, Jin Li +3
University of Electronic Science and Technology of China, Chengdu, China.
Compact JEPA world models enable efficient latent-space planning, but low-dimensional representation trained under reward-free self-supervision must encode both action-conditioned dynamics and predictable visual context. This competition can entangle controllable state with high-rank nuisance appearance and degrade planning as scenes become more complex. We introduce LRC-JEPA, a lightweight end-to-end world model that routes information into a compact predictive latent z and learned-query residual-context embeddings u. Only z is propagated by the dynamics model and used for planning, while u captures temporally persistent information for cross-attention reconstruction; a differentiable residual connection encourages the latent to retain complementary dynamic content. Under explicit assumptions, we show that the resulting representation is sufficient, minimal, nuisance-invariant, and disentangled. Across four simulated control environments, LRC-JEPA improves average planning success over a parameter-matched JEPA baseline by 9 percentage points and matches or exceeds substantially larger pretrained models. On the real-world Bridge-v2 set, its 5.5M-parameter active encoder outperforms DINO-WM (22.1M) and V-JEPA2 (303.9M) encoders while also enabling faster planning. Physical-state probes, reconstruction interventions, and ablations confirm the effectiveness of LRC-JEPA's representation disentanglement.