cs.LGFeb 20, 2026

JEPA-Bisim: Learning Robust Visual Representations for Planning with Joint-Embedding Predictive World Models

Authors: Leonardo F. Toso, Davit Shadunts, Yunyang Lu, Gloria Geng, Nihal Sharma, Donglin Zhan, Nam H. Nguyen, James Anderson

Organizations: Department of Electrical Engineering, Columbia University, New York, USA. · Capital One, USA.

Abstract

World models learned from high-dimensional visual observations allow agents to make decisions and plan directly in latent space, avoiding pixel-level reconstruction. However, recent latent predictive architectures (JEPAs), including the DINO world model (DINO-WM), display a degradation in test time robustness due to their sensitivity to ``slow features". These include visual variations such as background changes and distractors that are irrelevant to the task being solved. We address this limitation by augmenting the predictive objective with a bisimulation encoder that enforces control-relevant state equivalence, mapping states with similar transition dynamics to nearby latent states while limiting contributions from slow features. We evaluate our model on a navigation task (PointMaze) and on a manipulation task (PushT) under different test-time background changes and visual distractors. Across all benchmarks, our model consistently improves robustness to slow features while operating in a reduced latent space, up to 10×10\times smaller than that of DINO-WM. Moreover, our model is agnostic to the choice of pre-trained visual encoder and maintains robustness when paired with DINOv2, SimDINOv2, and iBOT features.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding

    Jun 30, 2026Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan +11Latent World ModelsVideo Joint Embedding Predictive Architecture