Latent action world models let agents plan new behaviors at test time by predicting how actions change the environment, and joint-embedding predictive architectures (JEPAs) do so by forecasting future latent states rather than pixels. Yet nearly all such models see the world through a camera, even though robotic manipulation is fundamentally geometric: in robotics goals for manipulation are traditionally specified by target object poses, not by images of the object once placed. We ask whether latent planning survives a shift from appearance to geometry, on the observation side as well as on the goal specifications side. To answer this, we extend the stable-worldmodel evaluation platform with simulated LiDAR-style raycast point clouds as a new sensor modality, and adapt three JEPA designs to point clouds: a frozen-encoder model built on Utonia features, a distribution-prior model based on LeWM, and an action-sensitive model based on Delta-JEPA. We further introduce a goal-encoding mechanism that constructs the goal latent from the current latent and a 3D target pose, removing the need for goal images or goal point clouds. A comparative evaluation of the different anti-collapse mechanisms shows that point-cloud world models can match their image-based counterparts, demonstrating that the modality shift from appearance to geometry is achievable. All models are released as open weights with open-source training and inference code, to make world-model planning accessible for LiDAR-driven and pose-directed robotic tasks.
Figures & tables
Figure 1: Encoder of Point-LeWM and Point-Delta-JEPA : Furthest point sampling (FPS) + radius query (only a few queries are shown), PointNet ( Qi et al., 2017a ) tokenizer, and ViT-Tiny whose class token (CLS) yields zt . The next latent zt+1 , from the same encoder, is matched to z^t+1 via MSE. SIGReg (Point-LeWM) and the Action-Reconstruction-Loss (Point-Delta-JEPA) prevent collapse.
Figure 2: Planning toward 3D goal poses. The target-to-latent module enables specifying a goal directly as a 3D pose. This pose is encoded into a goal latent, which the planner then optimises against, removing the need for a goal observation.
Figure 3: From image environments to point cloud environments. Top: the 3D scene. Bottom: the generated point cloud of that same scene, which replaces the image as the observation.
OGB-Cube
Two-Room
Push-T
Reacher
Ground share
∼ 79%
∼ 80%
∼ 98%
∼ 97%
Moving returns
∼ 15.0%
∼ 0.3%
∼ 0.5%
∼ 2.0%
Table 1: The four re-sensed datasets. Ground share and moving returns are the fractions of returns on the ground plane and on non-static entities.
Figure 4: Utonia-WM : frozen Utonia ( Zhang et al., 2026a ) encodes the point cloud into super-point features, which a fixed canonical-frame voxel grid bins into ≈256 tokens (only a few cells shown). Per cell, features are pooled and compressed by a fixed random orthogonal projection to the predictor width. Only the predictor trains, with a token-wise MSE against the frozen targets zt+1 .
Family
Obs.
Method
Two-Room
Reacher
Push-T
OGB-Cube
Random
25.2 ±4.1
10.8 ±4.0
2.4 ±1.3
46.0 ±8.5
SIGReg
Image
LeWM
85.6 ±5.7
82.0 ±5.6
86.4 ±5.2
69.8 ±9.3
Points
Point-LeWM
87.0 ±4.4
80.4 ±4.6
83.6 ±3.4
66.0 ±5.9
Act.-reconst.
Image
Delta-JEPA
99.8 ±0.6
86.0 ±3.3
94.2 ±2.7
80.2 ±6.0
Points
Point-Delta-JEPA
100.0 ±0.0
77.6 ±4.4
70.8 ±7.7
83.4 ±3.8
Frozen enc.
Image
DINO-WM
99.8 ±0.6
77.0 ±4.8
76.0 ±9.5
78.6 ±6.7
Table 2: Planning success rate (%) . Mean ±sd over 10 evaluation seeds of 50 episodes each (sd across seeds, here and throughout), with CEM behind the SWM solver ( Maes et al., 2026a ) on identical fixed start–goal pairs. Bold marks the best model within each family.
Two-Room
Reacher
Push-T
OGB-Cube
Encoder
zt
P-LeWM
PD-JEPA
P-LeWM
PD-JEPA
P-LeWM
PD-JEPA
P-LeWM
PD-JEPA
MLP
✗
87.4 ± 4.8
100.0 ± 0.0
77.4 ± 5.3
80.0 ± 5.5
85.0 ± 3.6
71.8 ± 7.1
65.2 ± 6.5
72.0 ± 5.2
Shortcut
✗
87.6 ± 4.8
100.0 ± 0.0
73.6 ± 6.7
72.8 ± 4.5
83.2 ± 4.4
71.4 ± 6.8
61.6 ± 5.7
78.8 ± 5.3
MLP
✓
88.4 ± 4.3
100.0 ± 0.0
77.0 ± 4.2
77.8 ± 4.9
83.8 ± 5.0
70.8 ± 7.4
66.0 ± 7.8
82.4 ± 5.9
Shortcut
✓
88.6 ± 3.9
100.0 ± 0.0
73.6 ± 5.2
71.2 ± 4.9
84.8 ± 2.3
73.6 ± 5.8
68.8 ± 8.3
83.2 ± 5.0
Goal cloud
87.2 ± 4.1
100.0 ± 0.0
78.8 ± 4.6
76.8 ± 2.3
84.8 ± 4.1
71.2 ± 5.0
65.4 ± 7.5
84.2 ± 4.2
Table 3: Planning success rate (%) toward 3D targets. Mean ±sd over 10 seeds (50 episodes each). Target encoders (P-LeWM/PD-JEPA abbreviate Point-LeWM/Point-Delta-JEPA) map the 3D target to a goal latent. zt marks whether the current latent is an additional input. Goal cloud encodes the stored goal cloud (reference). Best per column in bold, separately with and without zt .
Two-Room
OGB-Cube
P-LeWM
PD-JEPA
P-LeWM
PD-JEPA
Perturbation
Base
+MLP z
Base
+MLP z
Base
+MLP z
Base
+MLP z
Noise 0.25%
61.6 ±5.7
80.2 ±7.0
31.0 ±5.1
41.0 ±5.6
97.0 ±12.0
100.6 ±10.5
93.5 ±6.8
98.3 ±7.6
Noise 0.5%
37.9 ±6.0
53.6 ±8.9
35.8 ±6.8
36.4 ±4.3
89.7 ±15.9
97.6 ±11.2
77.2 ±9.4
90.9 ±8.3
Dropout 25%
96.1 ±5.9
99.3 ±5.4
60.4 ±6.4
72.0 ±6.8
99.1 ±12.3
100.0 ±10.0
98.8 ±4.1
98.6 ±7.4
Dropout 50%
92.2 ±4.6
98.4 ±5.7
50.2 ±5.5
63.0 ±4.2
101.5 ±8.5
100.0 ±13.0
99.5 ±4.4
101.0 ±6.6
Table 4: Target-to-latent module under sensor degradation. Success rate relative to each base model’s clean performance (100 = no degradation, clean rates as in Table 2 ), mean ±sd over 10 seeds (50 episodes each). Base encodes the degraded goal cloud, +MLP z receives the clean 3D target through the zt -conditioned MLP. Better of the pair in bold.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Source
Two-Room
Reacher
Push-T
OGB-Cube
Delta-JEPA
published ( Zhang et al., 2026b , Tab. 1)
100.0
81.3
89.1
79.3
ours (retrained)
99.8
86.0
94.2
80.2
DINO-WM
published ( Maes et al., 2026b , Fig. 6)
100
79
74
86
ours (retrained)
99.8
77.0
76.0
78.6
LeWM
published ( Maes et al., 2026b , Fig. 6)
87
86
96
74
retrained by Zhang et al. (2026b, Tab. 1)
74.9
79.9
84.5
64.1
Appendix
Table 5: Image baselines versus their published success rates (%). Ours are the Table 2 values under our protocol. Published DINO-WM excludes proprioception, as in our runs. Two-Room protocols differ (see text).
OGB-Cube
Two-Room
Push-T
Reacher
Sensor pose
scene camera
derived oblique
derived oblique
derived oblique
Ground share
∼ 79%
∼ 80%
∼ 98%
∼ 97%
Moving returns
∼ 15.0%
∼ 0.3%
∼ 0.5%
∼ 2.0%
Episodes
10,000
10,000
18,685
10,000
Frames
2,010,000
920,809
2,336,736
2,010,000
Episode length
201
31–101
49–246
201
Appendix
Table 6: The four re-sensed datasets. Ground share and moving returns are the fractions of returns on the ground plane and on non-static entities. Both are constant within an environment.
Linear
MLP
Property
Method
MSE ↓
R2↑
MSE ↓
R2↑
Agent Location
Point-LeWM
0.008
0.955
0.001
0.997
Point-Delta-JEPA
0.007
0.957
0.002
0.989
Block Location
Point-LeWM
0.004
0.958
0.000
0.998
Point-Delta-JEPA
0.005
0.946
0.001
0.992
Block Angle
Point-LeWM
0.027
0.891
0.002
0.992
Appendix
Table 7: Physical latent probing results on Push-T. All targets are 3-dimensional: locations are coordinates in the sensor frame, and the block angle is encoded as a heading vector in the sensor frame. Lower MSE and higher R2 indicate better representation quality.
Linear
MLP
Property
Method
MSE ↓
R2↑
MSE ↓
R2↑
Finger Position
Point-LeWM
0.000
1.000
0.000
1.000
Point-Delta-JEPA
0.000
0.999
0.000
1.000
Elbow Position
Point-LeWM
0.000
1.000
0.000
1.000
Point-Delta-JEPA
0.000
0.999
0.000
1.000
Joint Angles
Point-LeWM
0.000
1.000
0.000
1.000
Appendix
Table 8: Physical latent probing results on Reacher. Positions are 3D coordinates in the sensor frame, and joint angles are encoded as per-link unit heading vectors in the sensor frame. Lower MSE and higher R2 indicate better representation quality.
Linear
MLP
Property
Method
MSE ↓
R2↑
MSE ↓
R2↑
Joint Positions
Point-LeWM
0.000
0.992
0.000
0.999
Point-Delta-JEPA
0.000
0.998
0.000
1.000
Joint Velocity
Point-LeWM
0.142
0.021
0.130
0.102
Point-Delta-JEPA
0.071
0.507
0.054
0.626
End-Effector Position
Point-LeWM
0.000
0.992
0.000
0.999
Appendix
Table 9: Physical latent probing results on OGB-Cube. Positions are 3D coordinates in the sensor frame, stacked across the five arm joints for Joint Positions . Yaw is encoded as a heading vector in the sensor frame, and joint velocity is the raw six-dimensional vector in rad/s. Lower MSE and higher R2 indicate better representation quality. The near-zero yaw rows reflect label unidentifiability (a cube is rotationally symmetric, and the wrist returns few points), not representation failure, as discussed in the text.
Two-Room
Reacher
Push-T
OGB-Cube
Perturbation
P-LeWM
PD-JEPA
P-LeWM
PD-JEPA
P-LeWM
PD-JEPA
P-LeWM
PD-JEPA
None
87.0 ±4.4
100.0 ±0.0
80.4 ±4.6
77.6 ±4.4
83.6 ±3.4
70.8 ±7.7
66.0 ±5.9
83.4 ±3.8
Observation degraded, goal clean
Noise 0.25%
80.2 ±9.4
42.8 ±6.1
96.3 ±8.2
89.4 ±6.9
1.9 ±1.9
3.7 ±2.7
103.0 ±10.8
99.8 ±6.2
Noise 0.5%
52.2 ±11.5
36.0 ±6.7
90.8 ±5.9
62.2 ±6.1
1.2 ±1.7
4.0 ±2.0
97.6 ±11.4
98.3 ±6.8
Dropout 25%
99.8 ±5.4
72.2 ±6.4
97.5 ±6.1
90.7 ±6.5
90.7 ±6.8
94.6 ±8.8
101.5 ±12.7
99.3 ±4.8
Appendix
Table 10: Point cloud degradation. Planning success rate over 10 seeds with 50 episodes each, reported relative to each model’s own clean performance (100 = no degradation). The top row gives the absolute clean success rates used as denominators. The two point-cloud world models are shown side by side under each condition, better of the pair in bold. Due to Utonia-WM’s long computation times, it is not part of this ablation table. Noise σ is a fraction of the frame’s scan extent, and dropout is the fraction of returns lost.
P-LeWM
PD-JEPA
Env.
Perturbation
Base
+MLP
+MLP z
Base
+MLP
+MLP z
Two-Room
Noise 0.25%
61.6 ±5.7
80.0 ±6.6
80.2 ±7.0
31.0 ±5.1
44.4 ±5.1
41.0 ±5.6
Noise 0.5%
37.9 ±6.0
53.8 ±10.1
53.6 ±8.9
35.8 ±6.8
35.0 ±5.8
36.4 ±4.3
Dropout 25%
96.1 ±5.9
100.7 ±5.3
99.3 ±5.4
60.4 ±6.4
71.4 ±7.5
72.0 ±6.8
Dropout 50%
92.2 ±4.6
99.5 ±4.5
98.4 ±5.7
50.2 ±5.5
63.4 ±3.9
63.0 ±4.2
Reacher
Noise 0.25%
91.8 ±7.8
90.8 ±7.1
91.0 ±8.2
84.7 ±6.0
86.9 ±7.6
89.2 ±10.1
Appendix
Table 11: Effect of the target-to-latent module under joint observation-and-goal degradation. Planning success rate over 10 seeds with 50 episodes each, reported relative to each base model’s clean performance (100 = no degradation; clean rates as in Table 10 ). Base repeats the observation-and-goal-degraded rows of Table 10 ; +MLP and +MLP z add the unconditioned and the zt -conditioned MLP target encoder of Section 3.2 on top of the same backbone, so the goal enters as a clean 3D target instead of a degraded goal cloud. Best of the three variants per model in bold. Noise σ is a fraction of the frame’s scan extent, and dropout is the fraction of returns lost.
Figure 5: Specified cube goals. 201 cube positions at five heights ( z=2,8,16,24,32 cm), each approached from five grasped start poses with the frozen Point-Delta-JEPA point-cloud world model. Target-to-latent maps the position, together with the latent of the current observation, to the goal latent. Planner and goal latents operate on LiDAR point clouds; the rendered scene is shown only for legibility. Colour is the median over the five starts of the smallest cube-to-goal distance, from 0 cm (green) to the 4 cm success threshold (red). Blue cubes mark the four corner start poses, and the red cube in the gripper marks the center start. Only z=2 cm occurs as a goal in training.
Median error (cm)
Success (%)
Goal height
n
T2L
Goal cloud
Δ [95 % CI]
T2L
Goal cloud
8 cm
165
0.76
0.68
+0.05 [ −0.06 , +0.14 ]
96
95
16 cm
210
0.58
0.51
+0.12 [ +0.03 , +0.22 ]
97
95
24 cm
210
0.63
0.68
−0.07 [ −0.15 , +0.01 ]
98
94
32 cm
210
0.85
0.85
−0.02 [ −0.18 , +0.09 ]
96
91
All
795
0.68
0.67
+0.02 [ −0.05 , +0.07 ]
97
94
Appendix
Table 12: Accuracy on unseen goals. Cube goals at heights the expert never specified. Each 3-D position is approached from five grasped start poses with the same Point-Delta-JEPA point-cloud world model and CEM planner, so only the goal latent differs between target-to-latent, which predicts it from the position and the current latent, and the rebuilt goal cloud, which encodes the state rendered and re-scanned in simulation. Error is the cube’s closest approach to the goal, and success uses the environment’s 4 cm threshold. Δ gives the median paired difference between the two methods with a bootstrap 95 % interval, where negative values favour target-to-latent. Over all elevated goals the two are indistinguishable (Wilcoxon signed-rank, p=0.94 ).
Figure 6: Encoder attention. cls -token attention rollout of the Point-LeWM/Point-Delta-JEPA encoder mapped onto the input points. Warmer colours indicate higher attention, and points no group token covers are omitted. Columns are environments, rows the two models. Attention is normalised per point cloud. Utonia-WM is absent because its frozen PTv3 encoder has no cls token and attends only locally within serialised patches, so it admits no comparable attention map.
Table 13: Paired tests within the families of Table 2 . Difference of means Δ (points, first model minus second) with its 95% confidence interval, and Holm–Bonferroni corrected p -values of the paired t -test and the Wilcoxon signed-rank test over the four benchmarks of each comparison ( n=10 seeds). With ten seeds the exact Wilcoxon p cannot fall below 0.002 , or 0.008 after correction.
School of Computer Science and Technology, University of Chinese Academy of Sciences, Beijing · Institute of Information Engineering, Chinese Academy of Sciences, Beijing · School of Computer Science and Technology, Harbin Institute of Technology, Weihai +2