cs.LGAug 29, 2026

Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations and Goals

Authors: Fabio F. Oberweger, Michael Schwingshackl, Markus Murschitz

Organizations: AIT Austrian Institute of Technology Assistive & Autonomous Systems

Abstract

Latent action world models let agents plan new behaviors at test time by predicting how actions change the environment, and joint-embedding predictive architectures (JEPAs) do so by forecasting future latent states rather than pixels. Yet nearly all such models see the world through a camera, even though robotic manipulation is fundamentally geometric: in robotics goals for manipulation are traditionally specified by target object poses, not by images of the object once placed. We ask whether latent planning survives a shift from appearance to geometry, on the observation side as well as on the goal specifications side. To answer this, we extend the stable-worldmodel evaluation platform with simulated LiDAR-style raycast point clouds as a new sensor modality, and adapt three JEPA designs to point clouds: a frozen-encoder model built on Utonia features, a distribution-prior model based on LeWM, and an action-sensitive model based on Delta-JEPA. We further introduce a goal-encoding mechanism that constructs the goal latent from the current latent and a 3D target pose, removing the need for goal images or goal point clouds. A comparative evaluation of the different anti-collapse mechanisms shows that point-cloud world models can match their image-based counterparts, demonstrating that the modality shift from appearance to geometry is achievable. All models are released as open weights with open-source training and inference code, to make world-model planning accessible for LiDAR-driven and pose-directed robotic tasks.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning

    Sep 3, 2026Muyuan Liu, Yue Huang, Zheng Liang +1Inverse DynamicsRobot Planning

  2. Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding

    Jun 30, 2026Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan +11Latent World ModelsVideo Joint Embedding Predictive Architecture

  3. D-JEPA: A Decision-Aligned Latent World Model

    Sep 21, 2026Shuaijun Liu, Chengyu Wu, Qifu Wen +5Latent World ModelsWorld Models