cs.CVSep 29, 2026

Do-JEPA: From Masking to Intervention in Latent World Models

Authors: Hossein Resani, Javen Qinfeng Shi

Organizations: Australian Institute for Machine Learning, Adelaide University · Responsible AI Research Centre

Abstract

Latent world models are trained to predict what happens next, so nothing in their objective separates what an action caused from what merely co-occurred with it. Object-masking models such as C-JEPA intervene on what the predictor can see; we intervene on what physically happens. From one saved simulator state we run the dynamics under an action aa and under a reference action a∅a_{\varnothing}, and train the model to predict the difference Δz=za−za∅Δz=z^{a}-z^{a_{\varnothing}} between the two latent futures. The resulting objective, Do-JEPA, has an effect loss, a support loss (where the action enters), a propagation loss (where its effect travels) and invariance losses (what must not change). In a synthetic system with object-aligned variables, support supervision finds the directly intervened object in 99.95% of test cases, where a sparse action mask sends the action to a nuisance slot in every case, and response-onset supervision recovers the ring-shaped propagation graph (edge AUROC 0.975 vs. 0.624). From pixels, the effect loss beats a control trained on exactly the same data: it lowers latent effect error by 28.4% on an end-to-end LeWM model and physical effect error by 13.5% when trained and tested on natural action sequences, and on three independently generated CausalWorld benchmarks it lowers responsive effect error by about 20% under physics shifts and the latent context sensitivity of predicted effects by 66%. Trained from scratch it costs factual accuracy; fine-tuning an existing model with it removes this cost. Together, these results show that intervening on the world, rather than on what the model sees, helps latent world models predict what their actions cause.

Figures & tables

Appendix figures & tables24 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models

    Aug 6, 2026Xi Zeng, Haojie Ren, Ziying Song +2Latent World ModelsJoint-Embedding Predictive Architectures

  2. D-JEPA: A Decision-Aligned Latent World Model

    Sep 21, 2026Shuaijun Liu, Chengyu Wu, Qifu Wen +5Latent World ModelsWorld Models

  3. Is Forward Prediction Enough? Physical State Grounding for JEPA World Models

    Aug 7, 2026Haodong Yan, Jiaguan Zhu, Mingyuan Jia +12Latent World ModelsWorld Models