cs.ROAug 7, 2026

Is Forward Prediction Enough? Physical State Grounding for JEPA World Models

Authors: Haodong YanJiaguan ZhuMingyuan JiaRuiqing YinJunjie HeZhide ZhongJunfeng LiJinxuan Lu+7 more

Organizations: 1The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China · 2COCO Matrix

Abstract

Learning structured and control-relevant latent representations remains a key challenge for world models. Recent JEPA-based world models learn action-conditioned predictive latent dynamics from observation sequences. However, their forward-prediction objectives do not explicitly enforce reliable identifiability of robot-centric physical state from individual latents or state changes from latent pairs, which can limit downstream planning and policy performance. We propose PSG-JEPA, a physically grounded JEPA world model that shapes its latent space with two complementary grounding objectives beyond forward prediction: grounding individual latents in robot proprioceptive state, and grounding latent pairs in multi-horizon joint-angle changes. Both objectives are applied only during training, leaving the inference architecture and computational cost unchanged. To comprehensively evaluate PSG-JEPA, we conduct experiments at three levels: (1) latent identifiability via probing, (2) goal-conditioned planning on frozen latents, and (3) policy learning in simulation and on a real robot. Experiments demonstrate that our PSG-JEPA consistently outperforms state-of-the-art latent world-model baselines at all three levels.

Explore similar work

CardsList
  1. D-JEPA: A Decision-Aligned Latent World Model

    Sep 21, 2026Shuaijun Liu, Chengyu Wu, Qifu Wen +5Latent World Models