cs.ROSep 30, 2026

CF-JEPA: Improving Robustness of JEPA World Models via Controllability Factorization

Authors: Morgan Byrd, Robert Wright, Sehoon Ha

Organizations: Georgia Institute of Technology, Atlanta, GA, 30308, USA · Georgia Tech Research Institute, Atlanta, GA, 30308, USA

Abstract

Controlling an agent with vision requires being able to separate useful information from irrelevant background information. JEPA-style latent world models seem like a natural approach for this, as they do not perform pixel-level reconstruction; however, they are still sensitive to these distractor signals and experience latent collapse. In this work, we introduce Controllability Factorized JEPA (CF-JEPA), a JEPA-style world model which splits the latent space into controllable and uncontrollable subspaces. This factorization allows us to capture all the distractor information into the uncontrollable region, while we use the control-relevant latent information for our task. With this, we show comparable performance across 2D and 3D control tasks under nominal conditions and improved performance under distracted conditions, where CF-JEPA is the only model that does not experience latent collapse. We also validate our model under distracted conditions for a simulated robot task, highlighting the practical application of such a scheme.

Figures & tables

Explore similar work

Aug 6, 2026cs.CV

PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models

We propose PhyLatent, a dynamics-relevant training objective for Joint-Embedding Predictive Architecture (JEPA) world models. Our key observation is that preventing global latent collapse does not necessarily ensure that the learned representation preserves physically meaningful state and action relationships. We identify three failure modes: sensitivity to appearance changes that leave the physical state unchanged, insufficient separation of distinct physical states, and insufficient separation of different action-conditioned futures. We refer to these as Physical Invariance Collapse, Physical Distinguishability Collapse, and Counterfactual Dynamics Collapse, respectively. PhyLatent targets these failures through three coordinated training pathways, implemented with static visual invariance, physical state grounding, future representation alignment, counterfactual branch separation, and latent denoising. On OGBench-Cube, PhyLatent reduces the three collapse rates by 43.9%, 27.9%, and 47.1%, respectively, while improving model predictive control (MPC) success by 12.0 percentage points (17.2% relative). Across four visual-control tasks, average planning success increases by 6.62 percentage points (8.3% relative). These results show that global non-collapse alone is insufficient for learning a reliable JEPA world-model state space, and that explicitly preserving dynamics-relevant structure can improve closed-loop planning.
Sep 21, 2026cs.RO

D-JEPA: A Decision-Aligned Latent World Model

Latent world models predict the consequences of actions, but accurate prediction does not guarantee that latent distance reflects which candidate will execute successfully. We identify a decision-local prediction gap: among the few futures competing for execution, a candidate predicted closer to the goal can produce a worse realized outcome than an available alternative. We introduce D-JEPA, a decision-aligned latent world model that learns decision-relevant relations among candidate futures from executed outcomes. A bounded, permutation-equivariant operator jointly reasons over goal-relative predictive features and ordinal evidence, refining pretrained predictive geometry where action choices are most consequential. Restricted predictor adaptation and a shared ordinal interface extend this alignment across complementary predictive geometries. D-JEPA further realizes the learned decision structure in JEPA-compatible future representations, enabling deployment through native latent-distance planning. Evaluations across latent control, manipulation, pretrained action-producing models, physical robots and autonomous driving demonstrate improved action selection, including 87.89% success on PushT, a 15.04-point average gain on RoboTwin, and a 17-point gain on physical robot tasks. These results establish decision-relevant relational structure as a direct bridge between predictive world modeling and effective control.
May 29, 2026cs.LG

Subspace-Decomposed JEPAs: Disentangling Progression and Content in Latent World Models

Joint-Embedding Predictive Architectures (JEPAs) learn compact latent world models by predicting future embeddings, but no single coordinate of the latent is designated to encode task progression. We carve the JEPA latent into two orthogonal subspaces with disjoint roles: a low-dimensional progression subspace shaped by a cosine-margin triplet loss, and a high-dimensional content subspace regularised by the existing SIGReg objective of LeWM. We prove that the two anti-collapse forces act on disjoint coordinates, so they compose additively rather than competing on the same dimensions. Our method, SD-JEPA improves over the LeWM baseline on the majority of its control benchmarks at matched compute, and outperforms the strongest non-LeWM JEPA baseline on Push-T; a subspace-ablation falsifier confirms the split is the load-bearing ingredient. Beyond planning, the resulting 1-D angular progression coordinate functions as a scene-aware compass on the latent. It advances with task progress, regresses when the agent backtracks, and under controlled perturbations both spikes and relocalises to a semantically appropriate new task-phase sector, separating the moment of surprise from its meaning in a way that prediction-error scalars cannot.