Planning from pixels needs more than a latent space that is stable and predictable. The state the planner scores must also be organized by how actions move the system. Joint-embedding predictive architectures (JEPAs) avoid pixel reconstruction by predicting future representations, but existing action-conditioned JEPAs ask one embedding to serve both perception and control. We introduce H-JEPA, which separates the two. A wide perceptual code is regularized toward a well-scaled isotropic geometry with a Bures-Wasserstein prior, and a fixed orthonormal slice of that code is the control state, which inherits the code's covariance without any objective of its own. The state evolves under phase-conditioned dissipative port-Hamiltonian dynamics whose input port has orthonormal columns. Port-inverse consistency (PIC) reads the executed action back through the transpose of that port. We show that this readout is exactly the rollout error projected onto the port directions, so PIC is a parameter-free reweighting of prediction error and not an auxiliary action decoder. Untying the readout from the port breaks this identity and loses half of the gain. H-JEPA matches or exceeds reconstruction-free baselines, including the action-decoding Delta-JEPA, on four pixel-based control benchmarks after at most 10 training epochs, and its largest gain is on OGB-Cube (91.9 against 79.3 percent). Ablations on PushT and OGB-Cube separate the contributions of the structured predictor, PIC, the prediction horizon, the state rank, and the anti-collapse prior.
Figures & tables
Relaxed choice
Observed failure
Adopted constraint
BatchNorm in the projector, or a standardizer on s
Hides collapse: s keeps unit marginals while h degenerates to rank one
Per-sample affine-free LN0 , no normalizer on s (§ 3.1 , § 3.2 )
LSUV-style whitening of the projector activations at initialization ( Mishkin and Matas, 2016 )
Amplifies projector weights about 104 -fold, point-mass collapse within tens of steps
Isotropy as a loss with bounded per-step influence (§ 3.3 )
Learnable projection U
Co-adapts with the losses and collapses the rank of the state
Fixed seeded orthonormal U (§ 3.2 )
Learned action embedding
The PIC regression target collapses to effective rank about 1.3
Raw normalized actions (§ 3.1 )
SIGReg, or the Euclidean penalty ∥Σ−I∥F2
The restoring force fades as variance shrinks, collapse within about 50 to 100 steps
Bures-Wasserstein prior (§ 3.3 )
Unconstrained port gain
The gain trades off against the encoder scale, so optimization moves scale and not content
Orthonormal port via differentiable QR (§ 3.4 )
Table 1: Design choices we relaxed during development, the failure each one produced in this architecture, and the constraint adopted instead. These observations motivate the constraints of § 3 and complement the controlled ablations of § 4.1 .
Method
Two-Room
Reacher
PushT
OGB-Cube
PLDM
93.73±1.03
64.33±2.14
76.13±1.70
57.27±1.53
LeWM
74.93±0.42
79.87±0.90
84.53±1.50
64.13±1.89
Sub-JEPA
90.60±0.53
81.00±2.40
63.73±0.12
62.67±1.45
Delta-JEPA
100.00±0.00
81.33±0.50
89.07±1.90
79.27±1.81
H-JEPA (ours)
100.00±0.00
86.13±0.23
90.40±0.60
91.93±1.30
Table 2: Planning success rate (%) on the four LeWM benchmarks over 500 test episodes per task, mean ± std over three training seeds. Baselines are evaluated under the Delta-JEPA protocol, in which Delta-JEPA trains for 50 epochs. H-JEPA trains for 10 epochs and is evaluated at the per-task checkpoint epoch selected on a held-out seed (§ 4 ).
Predictor
Params total / dyn.
Success
Representation
LeWM, native embedding
18.0 M / 10.8 M
20%
SIGReg
AdaLN transformer on H-JEPA, r=64
9.6 M / 3.3 M
84%
H-JEPA
Port-Hamiltonian + PIC, r=64
7.2 M / 0.86 M
92%
H-JEPA
Port-Hamiltonian + PIC, r=192
11.6 M / 5.4 M
94%
H-JEPA
Table 3: Predictor class on PushT at a matched representation and horizon ( Kp=5 in all rows). Parameters are total model followed by dynamics module. H-JEPA rows share the ViT-T/14 encoder and 0.79 M projector.
Table 4
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1: H-JEPA. The Bures-Wasserstein prior Liso acts only on the spherical code ht . A fixed orthonormal slice U⊤ yields the control state st . The phase-conditioned port-Hamiltonian predictor Fθ is supervised open-loop ( Lpred ), PIC reads actions back through the port transpose ( LPIC ), and planning scores s alone.
Quantity
Value
Perceptual dimension D
192
Encoder
ViT-T/14 from scratch, MLP projector 192→2048→192
Output normalization
affine-free LayerNorm, ∥h∥=D , per sample
Projection U
fixed random orthonormal, seeded QR
History and phase
2 frames, c=[s,v] with v=s−sprev , left-pad v0=0
Net action dimension da
2 on Two-Room, Reacher, PushT, 5 on OGB-Cube
Appendix
Table 7: Default H-JEPA configuration. Task-specific entries are shown explicitly for rank, action dimension, and checkpoint epoch.
Property
Method
Linear MSE ↓
Linear ρ↑
MLP MSE ↓
MLP ρ↑
Agent location
H-JEPA
0.028±0.003
0.986
0.008±0.003
0.996
LeWM
0.062±0.012
0.968
0.017±0.006
0.992
Sub-JEPA
0.065±0.009
0.966
0.016±0.004
0.992
Block location
H-JEPA
0.152±0.025
0.926
0.049±0.024
0.979
LeWM
0.038±0.011
0.982
0.024±0.014
0.990
Sub-JEPA
0.033±0.009
0.984
0.017±0.007
0.992
Appendix
Table 8: Physical latent probing on PushT ( ρ is the Pearson correlation).
Property
Method
Linear MSE ↓
Linear ρ↑
MLP MSE ↓
MLP ρ↑
Finger position
H-JEPA
0.002±0.000
0.999
0.000±0.000
1.000
LeWM
0.002±0.000
0.999
0.000±0.000
1.000
Sub-JEPA
0.002±0.000
0.999
0.000±0.000
1.000
Joint position
H-JEPA
0.601±0.125
0.588
0.742±0.111
0.562
LeWM
0.636±0.120
0.573
0.738±0.116
0.550
Sub-JEPA
0.639±0.125
0.575
0.726±0.129
0.559
Appendix
Table 9: Physical latent probing on Reacher ( ρ is the Pearson correlation).
Property
Method
Linear MSE ↓
Linear ρ↑
MLP MSE ↓
MLP ρ↑
Joint position
H-JEPA
0.445±0.044
0.670
0.554±0.056
0.685
LeWM
0.579±0.060
0.645
0.775±0.109
0.647
Sub-JEPA
0.492±0.067
0.676
0.557±0.059
0.675
Joint velocity
H-JEPA
0.740±0.027
0.432
0.713±0.039
0.541
LeWM
1.069±0.038
0.088
2.674±0.236
0.050
Sub-JEPA
0.988±0.032
0.203
1.595±0.226
0.101
Appendix
Table 10: Physical latent probing on OGB-Cube ( ρ is the Pearson correlation).
Task
∥J∇H∥∥R∇H∥
∥passive∥∥Gaˉ∥
tr(R)tr(G⊤RG)
corr(trR,∥v∥)
PushT
4.35
0.13
0.001
0.48
Reacher
1.16
0.56
0.008
0.24
OGB-Cube
1.11
1.53
0.184
0.13
Appendix
Table 11: Learned port-Hamiltonian structure on held-out transitions. The relative contribution of dissipation, passive dynamics, and direct control changes markedly across tasks.
Task
Passive →Δy
Gaˉ→Δy
∇H→Δy
s→y
PushT
0.881
0.875
0.899
0.953
Reacher
0.351
0.033
0.291
0.514
OGB-Cube
0.268
0.323
0.312
0.561
Appendix
Table 12: Linear physical alignment of the learned predictor. R2 is measured on held-out transitions using ridge regression.
exceptional
measured eigenvalues of Cov(s)
Model
∥U⊤e∥22
r/D
1−∥U⊤e∥22
min
median
max
MP interval
PushT, r=64
0.356
0.333
0.644
0.236
0.713
2.48
[0.42,1.83]
PushT, r=192
1.000
1.000
0.000
1.3×10−6
0.877
3.31
[0.15,2.60]
OGB-Cube, r=32
0.116
0.167
0.884
0.015
0.625
1.46
[0.56,1.56]
Appendix
Table 13: Measured geometry of the control state on n=512 encoded frames from the selected checkpoints. The Marchenko-Pastur (MP) interval is the range of sample eigenvalues that an exactly isotropic state would show at this n and r ( Marčenko and Pastur, 1967 ) .
Figure 2: Eigenvalues of Cov(s) for the selected checkpoints. The dashed line marks the exceptional eigenvalue 1−∥U⊤e∥22 of the target covariance and the dotted line marks one. At r=D=192 the predicted null direction appears as the last eigenvalue.
Figure 3: Eigenvalue derivatives of the three penalties on synthetic data. The plot shows ∣∂ℓ/∂λ∣ against λ . Bures-Wasserstein rises as λ−1/2 when λ→0 , while the Euclidean penalty and SIGReg flatten to constants. By equation 20 , a flat curve means a force on the standard deviation that vanishes like λ , and the λ−1/2 rise means a constant force. Vertical offsets reflect the native normalization of each statistic and are not comparable across curves. The comparison is between slopes.
Figure 4: H-JEPA training on PushT. Training and validation losses across the 10 -epoch run for (a) open-loop latent prediction, (b) port-inverse consistency (PIC), and (c) the Bures-Wasserstein isotropy objective. All objectives decrease smoothly, with closely tracking training and validation curves.
Figure 5: H-JEPA training on OGB-Cube. Training and validation losses for (a) open-loop latent prediction, (b) PIC, and (c) isotropy. All three objectives fall rapidly during the first epochs and subsequently improve smoothly, with little separation between training and validation curves. The benchmark model uses the epoch- 3 OGB-Cube checkpoint selected in § 4 . The extended trace shows that lower representation-level losses after convergence are not required for the reported planning result.
School of Mechanical and Aerospace Engineering, Nanyang Technological University, Singapore · The University of Sheffield, Sheffield, United Kingdom · School of Artificial Intelligence (School of Software), Yanshan University, Qinhuangdao, China