Latent action world models let agents plan new behaviors at test time by predicting how actions change the environment, and joint-embedding predictive architectures (JEPAs) do so by forecasting future latent states rather than pixels. Yet nearly all such models see the world through a camera, even though robotic manipulation is fundamentally geometric: in robotics goals for manipulation are traditionally specified by target object poses, not by images of the object once placed. We ask whether latent planning survives a shift from appearance to geometry, on the observation side as well as on the goal specifications side. To answer this, we extend the stable-worldmodel evaluation platform with simulated LiDAR-style raycast point clouds as a new sensor modality, and adapt three JEPA designs to point clouds: a frozen-encoder model built on Utonia features, a distribution-prior model based on LeWM, and an action-sensitive model based on Delta-JEPA. We further introduce a goal-encoding mechanism that constructs the goal latent from the current latent and a 3D target pose, removing the need for goal images or goal point clouds. A comparative evaluation of the different anti-collapse mechanisms shows that point-cloud world models can match their image-based counterparts, demonstrating that the modality shift from appearance to geometry is achievable. All models are released as open weights with open-source training and inference code, to make world-model planning accessible for LiDAR-driven and pose-directed robotic tasks.
Figures & tables
Figure 1: Encoder of Point-LeWM and Point-Delta-JEPA : Furthest point sampling (FPS) + radius query (only a few queries are shown), PointNet ( Qi et al., 2017a ) tokenizer, and ViT-Tiny whose class token (CLS) yields zt . The next latent zt+1 , from the same encoder, is matched to z^t+1 via MSE. SIGReg (Point-LeWM) and the Action-Reconstruction-Loss (Point-Delta-JEPA) prevent collapse.
Figure 2: Planning toward 3D goal poses. The target-to-latent module enables specifying a goal directly as a 3D pose. This pose is encoded into a goal latent, which the planner then optimises against, removing the need for a goal observation.
Figure 3: From image environments to point cloud environments. Top: the 3D scene. Bottom: the generated point cloud of that same scene, which replaces the image as the observation.
OGB-Cube
Two-Room
Push-T
Reacher
Ground share
∼ 79%
∼ 80%
∼ 98%
∼ 97%
Moving returns
∼ 15.0%
∼ 0.3%
∼ 0.5%
∼ 2.0%
Table 1: The four re-sensed datasets. Ground share and moving returns are the fractions of returns on the ground plane and on non-static entities.
Figure 4: Utonia-WM : frozen Utonia ( Zhang et al., 2026a ) encodes the point cloud into super-point features, which a fixed canonical-frame voxel grid bins into ≈256 tokens (only a few cells shown). Per cell, features are pooled and compressed by a fixed random orthogonal projection to the predictor width. Only the predictor trains, with a token-wise MSE against the frozen targets zt+1 .
Family
Obs.
Method
Two-Room
Reacher
Push-T
OGB-Cube
Random
25.2 ±4.1
10.8 ±4.0
2.4 ±1.3
46.0 ±8.5
SIGReg
Image
LeWM
85.6 ±5.7
82.0 ±5.6
86.4 ±5.2
69.8 ±9.3
Points
Point-LeWM
87.0 ±4.4
80.4 ±4.6
83.6 ±3.4
66.0 ±5.9
Act.-reconst.
Image
Delta-JEPA
99.8 ±0.6
86.0 ±3.3
94.2 ±2.7
80.2 ±6.0
Points
Point-Delta-JEPA
100.0 ±0.0
77.6 ±4.4
70.8 ±7.7
83.4 ±3.8
Frozen enc.
Image
DINO-WM
99.8 ±0.6
77.0 ±4.8
76.0 ±9.5
78.6 ±6.7
Table 2: Planning success rate (%) . Mean ±sd over 10 evaluation seeds of 50 episodes each (sd across seeds, here and throughout), with CEM behind the SWM solver ( Maes et al., 2026a ) on identical fixed start–goal pairs. Bold marks the best model within each family.
Two-Room
Reacher
Push-T
OGB-Cube
Encoder
zt
P-LeWM
PD-JEPA
P-LeWM
PD-JEPA
P-LeWM
PD-JEPA
P-LeWM
PD-JEPA
MLP
✗
87.4 ± 4.8
100.0 ± 0.0
77.4 ± 5.3
80.0 ± 5.5
85.0 ± 3.6
71.8 ± 7.1
65.2 ± 6.5
72.0 ± 5.2
Shortcut
✗
87.6 ± 4.8
100.0 ± 0.0
73.6 ± 6.7
72.8 ± 4.5
83.2 ± 4.4
71.4 ± 6.8
61.6 ± 5.7
78.8 ± 5.3
MLP
✓
88.4 ± 4.3
100.0 ± 0.0
77.0 ± 4.2
77.8 ± 4.9
83.8 ± 5.0
70.8 ± 7.4
66.0 ± 7.8
82.4 ± 5.9
Shortcut
✓
88.6 ± 3.9
100.0 ± 0.0
73.6 ± 5.2
71.2 ± 4.9
84.8 ± 2.3
73.6 ± 5.8
68.8 ± 8.3
83.2 ± 5.0
Goal cloud
87.2 ± 4.1
100.0 ± 0.0
78.8 ± 4.6
76.8 ± 2.3
84.8 ± 4.1
71.2 ± 5.0
65.4 ± 7.5
84.2 ± 4.2
Table 3: Planning success rate (%) toward 3D targets. Mean ±sd over 10 seeds (50 episodes each). Target encoders (P-LeWM/PD-JEPA abbreviate Point-LeWM/Point-Delta-JEPA) map the 3D target to a goal latent. zt marks whether the current latent is an additional input. Goal cloud encodes the stored goal cloud (reference). Best per column in bold, separately with and without zt .
Two-Room
OGB-Cube
P-LeWM
PD-JEPA
P-LeWM
PD-JEPA
Perturbation
Base
+MLP z
Base
+MLP z
Base
+MLP z
Base
+MLP z
Noise 0.25%
61.6 ±5.7
80.2 ±7.0
31.0 ±5.1
41.0 ±5.6
97.0 ±12.0
100.6 ±10.5
93.5 ±6.8
98.3 ±7.6
Noise 0.5%
37.9 ±6.0
53.6 ±8.9
35.8 ±6.8
36.4 ±4.3
89.7 ±15.9
97.6 ±11.2
77.2 ±9.4
90.9 ±8.3
Dropout 25%
96.1 ±5.9
99.3 ±5.4
60.4 ±6.4
72.0 ±6.8
99.1 ±12.3
100.0 ±10.0
98.8 ±4.1
98.6 ±7.4
Dropout 50%
92.2 ±4.6
98.4 ±5.7
50.2 ±5.5
63.0 ±4.2
101.5 ±8.5
100.0 ±13.0
99.5 ±4.4
101.0 ±6.6
Table 4: Target-to-latent module under sensor degradation. Success rate relative to each base model’s clean performance (100 = no degradation, clean rates as in Table 2 ), mean ±sd over 10 seeds (50 episodes each). Base encodes the degraded goal cloud, +MLP z receives the clean 3D target through the zt -conditioned MLP. Better of the pair in bold.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Source
Two-Room
Reacher
Push-T
OGB-Cube
Delta-JEPA
published ( Zhang et al., 2026b , Tab. 1)
100.0
81.3
89.1
79.3
ours (retrained)
99.8
86.0
94.2
80.2
DINO-WM
published ( Maes et al., 2026b , Fig. 6)
100
79
74
86
ours (retrained)
99.8
77.0
76.0
78.6
LeWM
published ( Maes et al., 2026b , Fig. 6)
87
86
96
74
retrained by Zhang et al. (2026b, Tab. 1)
74.9
79.9
84.5
64.1
Appendix
Table 5: Image baselines versus their published success rates (%). Ours are the Table 2 values under our protocol. Published DINO-WM excludes proprioception, as in our runs. Two-Room protocols differ (see text).
OGB-Cube
Two-Room
Push-T
Reacher
Sensor pose
scene camera
derived oblique
derived oblique
derived oblique
Ground share
∼ 79%
∼ 80%
∼ 98%
∼ 97%
Moving returns
∼ 15.0%
∼ 0.3%
∼ 0.5%
∼ 2.0%
Episodes
10,000
10,000
18,685
10,000
Frames
2,010,000
920,809
2,336,736
2,010,000
Episode length
201
31–101
49–246
201
Appendix
Table 6: The four re-sensed datasets. Ground share and moving returns are the fractions of returns on the ground plane and on non-static entities. Both are constant within an environment.
Linear
MLP
Property
Method
MSE ↓
R2↑
MSE ↓
R2↑
Agent Location
Point-LeWM
0.008
0.955
0.001
0.997
Point-Delta-JEPA
0.007
0.957
0.002
0.989
Block Location
Point-LeWM
0.004
0.958
0.000
0.998
Point-Delta-JEPA
0.005
0.946
0.001
0.992
Block Angle
Point-LeWM
0.027
0.891
0.002
0.992
Appendix
Table 7: Physical latent probing results on Push-T. All targets are 3-dimensional: locations are coordinates in the sensor frame, and the block angle is encoded as a heading vector in the sensor frame. Lower MSE and higher R2 indicate better representation quality.
Linear
MLP
Property
Method
MSE ↓
R2↑
MSE ↓
R2↑
Finger Position
Point-LeWM
0.000
1.000
0.000
1.000
Point-Delta-JEPA
0.000
0.999
0.000
1.000
Elbow Position
Point-LeWM
0.000
1.000
0.000
1.000
Point-Delta-JEPA
0.000
0.999
0.000
1.000
Joint Angles
Point-LeWM
0.000
1.000
0.000
1.000
Appendix
Table 8: Physical latent probing results on Reacher. Positions are 3D coordinates in the sensor frame, and joint angles are encoded as per-link unit heading vectors in the sensor frame. Lower MSE and higher R2 indicate better representation quality.
Linear
MLP
Property
Method
MSE ↓
R2↑
MSE ↓
R2↑
Joint Positions
Point-LeWM
0.000
0.992
0.000
0.999
Point-Delta-JEPA
0.000
0.998
0.000
1.000
Joint Velocity
Point-LeWM
0.142
0.021
0.130
0.102
Point-Delta-JEPA
0.071
0.507
0.054
0.626
End-Effector Position
Point-LeWM
0.000
0.992
0.000
0.999
Appendix
Table 9: Physical latent probing results on OGB-Cube. Positions are 3D coordinates in the sensor frame, stacked across the five arm joints for Joint Positions . Yaw is encoded as a heading vector in the sensor frame, and joint velocity is the raw six-dimensional vector in rad/s. Lower MSE and higher R2 indicate better representation quality. The near-zero yaw rows reflect label unidentifiability (a cube is rotationally symmetric, and the wrist returns few points), not representation failure, as discussed in the text.
Two-Room
Reacher
Push-T
OGB-Cube
Perturbation
P-LeWM
PD-JEPA
P-LeWM
PD-JEPA
P-LeWM
PD-JEPA
P-LeWM
PD-JEPA
None
87.0 ±4.4
100.0 ±0.0
80.4 ±4.6
77.6 ±4.4
83.6 ±3.4
70.8 ±7.7
66.0 ±5.9
83.4 ±3.8
Observation degraded, goal clean
Noise 0.25%
80.2 ±9.4
42.8 ±6.1
96.3 ±8.2
89.4 ±6.9
1.9 ±1.9
3.7 ±2.7
103.0 ±10.8
99.8 ±6.2
Noise 0.5%
52.2 ±11.5
36.0 ±6.7
90.8 ±5.9
62.2 ±6.1
1.2 ±1.7
4.0 ±2.0
97.6 ±11.4
98.3 ±6.8
Dropout 25%
99.8 ±5.4
72.2 ±6.4
97.5 ±6.1
90.7 ±6.5
90.7 ±6.8
94.6 ±8.8
101.5 ±12.7
99.3 ±4.8
Appendix
Table 10: Point cloud degradation. Planning success rate over 10 seeds with 50 episodes each, reported relative to each model’s own clean performance (100 = no degradation). The top row gives the absolute clean success rates used as denominators. The two point-cloud world models are shown side by side under each condition, better of the pair in bold. Due to Utonia-WM’s long computation times, it is not part of this ablation table. Noise σ is a fraction of the frame’s scan extent, and dropout is the fraction of returns lost.
P-LeWM
PD-JEPA
Env.
Perturbation
Base
+MLP
+MLP z
Base
+MLP
+MLP z
Two-Room
Noise 0.25%
61.6 ±5.7
80.0 ±6.6
80.2 ±7.0
31.0 ±5.1
44.4 ±5.1
41.0 ±5.6
Noise 0.5%
37.9 ±6.0
53.8 ±10.1
53.6 ±8.9
35.8 ±6.8
35.0 ±5.8
36.4 ±4.3
Dropout 25%
96.1 ±5.9
100.7 ±5.3
99.3 ±5.4
60.4 ±6.4
71.4 ±7.5
72.0 ±6.8
Dropout 50%
92.2 ±4.6
99.5 ±4.5
98.4 ±5.7
50.2 ±5.5
63.4 ±3.9
63.0 ±4.2
Reacher
Noise 0.25%
91.8 ±7.8
90.8 ±7.1
91.0 ±8.2
84.7 ±6.0
86.9 ±7.6
89.2 ±10.1
Appendix
Table 11: Effect of the target-to-latent module under joint observation-and-goal degradation. Planning success rate over 10 seeds with 50 episodes each, reported relative to each base model’s clean performance (100 = no degradation; clean rates as in Table 10 ). Base repeats the observation-and-goal-degraded rows of Table 10 ; +MLP and +MLP z add the unconditioned and the zt -conditioned MLP target encoder of Section 3.2 on top of the same backbone, so the goal enters as a clean 3D target instead of a degraded goal cloud. Best of the three variants per model in bold. Noise σ is a fraction of the frame’s scan extent, and dropout is the fraction of returns lost.
Figure 5: Specified cube goals. 201 cube positions at five heights ( z=2,8,16,24,32 cm), each approached from five grasped start poses with the frozen Point-Delta-JEPA point-cloud world model. Target-to-latent maps the position, together with the latent of the current observation, to the goal latent. Planner and goal latents operate on LiDAR point clouds; the rendered scene is shown only for legibility. Colour is the median over the five starts of the smallest cube-to-goal distance, from 0 cm (green) to the 4 cm success threshold (red). Blue cubes mark the four corner start poses, and the red cube in the gripper marks the center start. Only z=2 cm occurs as a goal in training.
Median error (cm)
Success (%)
Goal height
n
T2L
Goal cloud
Δ [95 % CI]
T2L
Goal cloud
8 cm
165
0.76
0.68
+0.05 [ −0.06 , +0.14 ]
96
95
16 cm
210
0.58
0.51
+0.12 [ +0.03 , +0.22 ]
97
95
24 cm
210
0.63
0.68
−0.07 [ −0.15 , +0.01 ]
98
94
32 cm
210
0.85
0.85
−0.02 [ −0.18 , +0.09 ]
96
91
All
795
0.68
0.67
+0.02 [ −0.05 , +0.07 ]
97
94
Appendix
Table 12: Accuracy on unseen goals. Cube goals at heights the expert never specified. Each 3-D position is approached from five grasped start poses with the same Point-Delta-JEPA point-cloud world model and CEM planner, so only the goal latent differs between target-to-latent, which predicts it from the position and the current latent, and the rebuilt goal cloud, which encodes the state rendered and re-scanned in simulation. Error is the cube’s closest approach to the goal, and success uses the environment’s 4 cm threshold. Δ gives the median paired difference between the two methods with a bootstrap 95 % interval, where negative values favour target-to-latent. Over all elevated goals the two are indistinguishable (Wilcoxon signed-rank, p=0.94 ).
Figure 6: Encoder attention. cls -token attention rollout of the Point-LeWM/Point-Delta-JEPA encoder mapped onto the input points. Warmer colours indicate higher attention, and points no group token covers are omitted. Columns are environments, rows the two models. Attention is normalised per point cloud. Utonia-WM is absent because its frozen PTv3 encoder has no cls token and attends only locally within serialised patches, so it admits no comparable attention map.
Table 13: Paired tests within the families of Table 2 . Difference of means Δ (points, first model minus second) with its 95% confidence interval, and Holm–Bonferroni corrected p -values of the paired t -test and the Wilcoxon signed-rank test over the four benchmarks of each comparison ( n=10 seeds). With ten seeds the exact Wilcoxon p cannot fall below 0.002 , or 0.008 after correction.
Action-conditioned JEPA world models enable planning toward visually specified goals without reconstructing future pixels, yet latent prediction alone does not explicitly encourage the learned representations to retain information relevant to robotic control. We introduce an end-to-end JEPA world model that augments latent prediction with inverse dynamics (IDM) and state alignment (SA). While inverse dynamics discourages latent collapse and makes latent transitions informative of the actions that produced them, state alignment grounds consecutive representations in their associated physical configuration and motion. Across four benchmark tasks, our model attains the highest success rates on TwoRoom (100%), PushT (98%), and OGBench-Cube (87%), while performing comparably to LeWorldModel on Reacher. Our ablation further shows that adding state alignment consistently improves planning success over IDM alone across all four tasks. Although LeWorldModel, our primary baseline, attains higher average straightening on OGBench-Cube, transition-subspace analysis shows that its transition energy is concentrated in a substantially lower-dimensional subspace. Our state-aligned model exhibits a higher effective transition dimension than LeWorldModel and improves planning over IDM alone, supporting state alignment as an effective complement to inverse dynamics for robotic planning.
Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations. We propose Delta-JEPA, an end-to-end reconstruction-free world model that augments latent forward prediction with a Latent Difference Action Decoder (LDAD). Unlike inverse decoders that infer actions from concatenated endpoint embeddings, LDAD reconstructs the executed action from the latent displacement between consecutive observations. This displacement-level supervision directly regularizes transition geometry: adjacent embeddings cannot collapse without losing action information, and different actions are encouraged to induce distinguishable latent changes for rollout-based planning. Delta-JEPA uses only latent prediction and action reconstruction, avoiding pixel reconstruction and distribution-matching regularizers. Across four visual continuous-control tasks, Delta-JEPA improves planning over JEPA-based and representation-learning world model baselines. Ablations show that displacement-based action decoding is consistently more effective than endpoint concatenation, and action-sensitivity analyses show clearer action-conditioned latent responses. These results indicate that supervising latent differences is a simple and effective mechanism for collapse-resistant and action-sensitive world model learning.
Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan +11
School of Computer Science and Technology, University of Chinese Academy of Sciences, Beijing · Institute of Information Engineering, Chinese Academy of Sciences, Beijing · School of Computer Science and Technology, Harbin Institute of Technology, Weihai +2
Latent world models predict the consequences of actions, but accurate prediction does not guarantee that latent distance reflects which candidate will execute successfully. We identify a decision-local prediction gap: among the few futures competing for execution, a candidate predicted closer to the goal can produce a worse realized outcome than an available alternative. We introduce D-JEPA, a decision-aligned latent world model that learns decision-relevant relations among candidate futures from executed outcomes. A bounded, permutation-equivariant operator jointly reasons over goal-relative predictive features and ordinal evidence, refining pretrained predictive geometry where action choices are most consequential. Restricted predictor adaptation and a shared ordinal interface extend this alignment across complementary predictive geometries. D-JEPA further realizes the learned decision structure in JEPA-compatible future representations, enabling deployment through native latent-distance planning. Evaluations across latent control, manipulation, pretrained action-producing models, physical robots and autonomous driving demonstrate improved action selection, including 87.89% success on PushT, a 15.04-point average gain on RoboTwin, and a 17-point gain on physical robot tasks. These results establish decision-relevant relational structure as a direct bridge between predictive world modeling and effective control.
Shuaijun Liu, Chengyu Wu, Qifu Wen +5
The Hong Kong University of Science and Technology (Guangzhou) · Boston University · Shanghai Jiao Tong University