cs.LGSep 27, 2026

Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State

Authors: Tamim Zoabi, Ameen Ali, Lior Wolf

Organizations: The Blavatnik School of Computer Science, Tel Aviv University

Abstract

Planning from pixels needs more than a latent space that is stable and predictable. The state the planner scores must also be organized by how actions move the system. Joint-embedding predictive architectures (JEPAs) avoid pixel reconstruction by predicting future representations, but existing action-conditioned JEPAs ask one embedding to serve both perception and control. We introduce H-JEPA, which separates the two. A wide perceptual code is regularized toward a well-scaled isotropic geometry with a Bures-Wasserstein prior, and a fixed orthonormal slice of that code is the control state, which inherits the code's covariance without any objective of its own. The state evolves under phase-conditioned dissipative port-Hamiltonian dynamics whose input port has orthonormal columns. Port-inverse consistency (PIC) reads the executed action back through the transpose of that port. We show that this readout is exactly the rollout error projected onto the port directions, so PIC is a parameter-free reweighting of prediction error and not an auxiliary action decoder. Untying the readout from the port breaks this identity and loses half of the gain. H-JEPA matches or exceeds reconstruction-free baselines, including the action-decoding Delta-JEPA, on four pixel-based control benchmarks after at most 1010 training epochs, and its largest gain is on OGB-Cube (91.991.9 against 79.379.3 percent). Ablations on PushT and OGB-Cube separate the contributions of the structured predictor, PIC, the prediction horizon, the state rank, and the anti-collapse prior.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Aug 13, 2026cs.LG

Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency

Joint-embedding predictive architectures (JEPAs) learn world models that predict in a compact latent space rather than in pixels, reducing the pressure to model nuisance appearance. Yet this provides no guarantee against visual perturbations: they can still alter the encoded representation and affect subsequent action-conditioned predictions. Bisimulation captures this requirement precisely: two observations should be treated as the same state only when their action-conditioned consequences agree. Guided by this criterion, we introduce Action-Conditioned Predictive Consistency (ACPC), a diagnostic that measures how far a clean history and a visually perturbed view of it diverge after being rolled forward under the same action sequence. We prove that this divergence bounds the perturbation-induced change in multi-step prediction error and planner cost. Building on pairwise ACPC, we define two complementary measures: the Invariance Radius (IR) summarizes clean-perturbed rollout spread, while the Separation Rate (SR) checks whether different states remain distinguishable after rollout. Experiments on four visual control tasks show that pairwise ACPC predicts perturbation-induced prediction and cost changes. On LeWM, the IR-SR screen transfers across tasks, and the joint diagnostic remains informative under blur and resize. PLDM exhibits similar diagnostic trends under a different architecture.
Aug 20, 2026cs.LG

Orthogonal JEPA: Factorized Predictive States for Latent World Models

World models construct latent states that support prediction, planning, and reasoning about an underlying system. Joint-embedding predictive architectures (JEPAs) offer a direct way to learn such states by predicting targets in representation space instead of reconstructing every detail of the observation. Standard JEPAs, however, organize all predictable content through one target embedding and one prediction pathway. In complex systems, this monolithic state can allocate redundant capacity to dominant signals while providing weak or conflicting gradients to less dominant predictive structure. We introduce \method, a latent world-modeling framework based on orthogonal predictive factorization. Learned basis matrices analyze each target state into multiple components, and a dedicated prediction branch estimates each component from a shared context representation. Predictive regression preserves the factor magnitudes required for state synthesis, an orthogonality objective discourages repeated directions, factor-activity regularization maintains variation in projected targets, and online variance regularization discourages coordinate-wise encoder collapse. Predicted components are synthesized into a complete latent state that can be used by a readout, decoder, planner, or autoregressive rollout. The same predictive-state mechanism applies when the target is temporally future, spatially hidden, or another partial observation of the same system. Experiments on controlled vision, single-cell transcriptomics, longitudinal health records, continuous control, and molecular dynamics evaluate representation quality, forecasting, planning, and long-horizon stability.
Aug 6, 2026cs.CV

PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models

We propose PhyLatent, a dynamics-relevant training objective for Joint-Embedding Predictive Architecture (JEPA) world models. Our key observation is that preventing global latent collapse does not necessarily ensure that the learned representation preserves physically meaningful state and action relationships. We identify three failure modes: sensitivity to appearance changes that leave the physical state unchanged, insufficient separation of distinct physical states, and insufficient separation of different action-conditioned futures. We refer to these as Physical Invariance Collapse, Physical Distinguishability Collapse, and Counterfactual Dynamics Collapse, respectively. PhyLatent targets these failures through three coordinated training pathways, implemented with static visual invariance, physical state grounding, future representation alignment, counterfactual branch separation, and latent denoising. On OGBench-Cube, PhyLatent reduces the three collapse rates by 43.9%, 27.9%, and 47.1%, respectively, while improving model predictive control (MPC) success by 12.0 percentage points (17.2% relative). Across four visual-control tasks, average planning success increases by 6.62 percentage points (8.3% relative). These results show that global non-collapse alone is insufficient for learning a reliable JEPA world-model state space, and that explicitly preserving dynamics-relevant structure can improve closed-loop planning.