Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State
Organizations: The Blavatnik School of Computer Science, Tel Aviv University
Abstract
Planning from pixels needs more than a latent space that is stable and predictable. The state the planner scores must also be organized by how actions move the system. Joint-embedding predictive architectures (JEPAs) avoid pixel reconstruction by predicting future representations, but existing action-conditioned JEPAs ask one embedding to serve both perception and control. We introduce H-JEPA, which separates the two. A wide perceptual code is regularized toward a well-scaled isotropic geometry with a Bures-Wasserstein prior, and a fixed orthonormal slice of that code is the control state, which inherits the code's covariance without any objective of its own. The state evolves under phase-conditioned dissipative port-Hamiltonian dynamics whose input port has orthonormal columns. Port-inverse consistency (PIC) reads the executed action back through the transpose of that port. We show that this readout is exactly the rollout error projected onto the port directions, so PIC is a parameter-free reweighting of prediction error and not an auxiliary action decoder. Untying the readout from the port breaks this identity and loses half of the gain. H-JEPA matches or exceeds reconstruction-free baselines, including the action-decoding Delta-JEPA, on four pixel-based control benchmarks after at most training epochs, and its largest gain is on OGB-Cube ( against percent). Ablations on PushT and OGB-Cube separate the contributions of the structured predictor, PIC, the prediction horizon, the state rank, and the anti-collapse prior.
Figures & tables
| Relaxed choice | Observed failure | Adopted constraint |
|---|---|---|
| BatchNorm in the projector, or a standardizer on | Hides collapse: keeps unit marginals while degenerates to rank one | Per-sample affine-free , no normalizer on (§ 3.1 , § 3.2 ) |
| LSUV-style whitening of the projector activations at initialization ( Mishkin and Matas, 2016 ) | Amplifies projector weights about -fold, point-mass collapse within tens of steps | Isotropy as a loss with bounded per-step influence (§ 3.3 ) |
| Learnable projection | Co-adapts with the losses and collapses the rank of the state | Fixed seeded orthonormal (§ 3.2 ) |
| Learned action embedding | The PIC regression target collapses to effective rank about | Raw normalized actions (§ 3.1 ) |
| SIGReg, or the Euclidean penalty | The restoring force fades as variance shrinks, collapse within about to steps | Bures-Wasserstein prior (§ 3.3 ) |
| Unconstrained port gain | The gain trades off against the encoder scale, so optimization moves scale and not content | Orthonormal port via differentiable QR (§ 3.4 ) |
| Method | Two-Room | Reacher | PushT | OGB-Cube |
|---|---|---|---|---|
| PLDM | ||||
| LeWM | ||||
| Sub-JEPA | ||||
| Delta-JEPA | ||||
| H-JEPA (ours) |
| Predictor | Params total / dyn. | Success | Representation |
|---|---|---|---|
| LeWM, native embedding | M / M | SIGReg | |
| AdaLN transformer on H-JEPA, | M / M | H-JEPA | |
| Port-Hamiltonian PIC, | M / M | H-JEPA | |
| Port-Hamiltonian PIC, | M / M | H-JEPA |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Quantity | Value |
|---|---|
| Perceptual dimension | |
| Encoder | ViT-T/14 from scratch, MLP projector |
| Output normalization | affine-free LayerNorm, , per sample |
| Projection | fixed random orthonormal, seeded QR |
| History and phase | frames, with , left-pad |
| Net action dimension | on Two-Room, Reacher, PushT, on OGB-Cube |
| Property | Method | Linear MSE | Linear | MLP MSE | MLP |
|---|---|---|---|---|---|
| Agent location | H-JEPA | ||||
| LeWM | |||||
| Sub-JEPA | |||||
| Block location | H-JEPA | ||||
| LeWM | |||||
| Sub-JEPA |
| Property | Method | Linear MSE | Linear | MLP MSE | MLP |
|---|---|---|---|---|---|
| Finger position | H-JEPA | ||||
| LeWM | |||||
| Sub-JEPA | |||||
| Joint position | H-JEPA | ||||
| LeWM | |||||
| Sub-JEPA |
| Property | Method | Linear MSE | Linear | MLP MSE | MLP |
|---|---|---|---|---|---|
| Joint position | H-JEPA | ||||
| LeWM | |||||
| Sub-JEPA | |||||
| Joint velocity | H-JEPA | ||||
| LeWM | |||||
| Sub-JEPA |
| Task | ||||
|---|---|---|---|---|
| PushT | 4.35 | 0.13 | 0.001 | 0.48 |
| Reacher | 1.16 | 0.56 | 0.008 | 0.24 |
| OGB-Cube | 1.11 | 1.53 | 0.184 | 0.13 |
| Task | Passive | |||
|---|---|---|---|---|
| PushT | 0.881 | 0.875 | 0.899 | 0.953 |
| Reacher | 0.351 | 0.033 | 0.291 | 0.514 |
| OGB-Cube | 0.268 | 0.323 | 0.312 | 0.561 |
| exceptional | measured eigenvalues of | ||||||
|---|---|---|---|---|---|---|---|
| Model | min | median | max | MP interval | |||
| PushT, | |||||||
| PushT, | |||||||
| OGB-Cube, | |||||||