cs.LGAug 29, 2026

Flow-JEPA: Robust Latent Dynamics for JEPA World Models via Flow Matching

Authors: Yanchen Huo, Ziying Song, Yadan Luo

Organizations: Nanyang Technological University · The University of Queensland

Abstract

Joint-Embedding Predictive Architectures (JEPAs) provide a powerful framework for latent world modeling and planning in a reconstruction-free manner. Although numerous JEPA-based approaches have been proposed to mitigate representation collapse, our experiments on localized, out-of-distribution visual noise reveal that performance degradation remains pronounced and unresolved. We propose Flow-JEPA (F-JEPA), a flow-based latent dynamics model that jointly generates a sequence of future latent states conditioned on the current observation and actions. A Gaussian distribution serves as the flow source, exposing the vector field to perturbed latent trajectories as it learns to transport them toward clean future representations. This formulation retains the reconstruction-free JEPA framework while switching from pointwise transition regression to stochastic trajectory-level prediction. F-JEPA raises mean success from 86%86\% to 92%92\% under clean observations and from 67%67\% to 86%86\% under noisy conditions. Further evaluations over varying perturbation severity and inference settings show that the performance advantage persists across a broad range of conditions. These results suggest that conditional flow matching provides a promising alternative to deterministic autoregressive prediction as a dynamics formulation in JEPA world models.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Aug 7, 2026cs.CV

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling

Joint-Embedding Predictive Architectures (JEPAs) have emerged as a principled framework for self-supervised learning of world models in compact latent spaces, yet existing methods are fragmented: some predict masked parts of a single image in latent space (I-JEPA), others learn to predict global photometric transformations (Image World Models), while video-scale JEPAs predict future temporal states and are post-trained for action-conditioned planning (V-JEPA~2, DINO-World, DINO-WM). These objectives are treated as distinct recipes with separate encoders, predictors, and anti-collapse regularizers, hindering a single model from unifying image-level and video-level world modeling. We present UniJEPA, a unified JEPA that jointly learns photometric prediction (image-level transformations) and temporal prediction (video-level next-state dynamics) in one shared latent space. A single end-to-end objective, composed of a next-embedding prediction loss and a Gaussian regularizer, yields a provably anti-collapse encoder-predictor pair trainable from raw pixels without EMA, stop-gradient, or pre-trained encoders. We show that the same latent space supports controllable abstraction: photometric prediction learns invariant structure while temporal prediction learns equivariant dynamics. After action-conditioned post-training on offline trajectories, UniJEPA enables zero-shot planning by treating goal features as prediction targets. On image, video, and control benchmarks, UniJEPA matches or surpasses task-specific JEPAs while requiring a single loss hyperparameter, and plans up to tens of times faster than generative world models at comparable accuracy.
Aug 6, 2026cs.CV

PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models

We propose PhyLatent, a dynamics-relevant training objective for Joint-Embedding Predictive Architecture (JEPA) world models. Our key observation is that preventing global latent collapse does not necessarily ensure that the learned representation preserves physically meaningful state and action relationships. We identify three failure modes: sensitivity to appearance changes that leave the physical state unchanged, insufficient separation of distinct physical states, and insufficient separation of different action-conditioned futures. We refer to these as Physical Invariance Collapse, Physical Distinguishability Collapse, and Counterfactual Dynamics Collapse, respectively. PhyLatent targets these failures through three coordinated training pathways, implemented with static visual invariance, physical state grounding, future representation alignment, counterfactual branch separation, and latent denoising. On OGBench-Cube, PhyLatent reduces the three collapse rates by 43.9%, 27.9%, and 47.1%, respectively, while improving model predictive control (MPC) success by 12.0 percentage points (17.2% relative). Across four visual-control tasks, average planning success increases by 6.62 percentage points (8.3% relative). These results show that global non-collapse alone is insufficient for learning a reliable JEPA world-model state space, and that explicitly preserving dynamics-relevant structure can improve closed-loop planning.
Jun 25, 2026cs.LG

A Generalization Theory for JEPA-Based World Models

Joint Embedding Predictive Architectures (JEPAs) have recently emerged as a promising paradigm for world modeling by learning predictive dynamics in a latent space rather than generating future observations at the input level. Despite their empirical success, the theoretical understanding of JEPA-based world models remains limited. In this paper, we develop the first generalization theory for JEPA-based world models. We formulate JEPA pretraining as a conditional spectral graph learning problem and show that the JEPA objective is equivalent to a low-rank factorization of an action-conditioned co-occurrence matrix. Building on this characterization, we establish a connection between JEPA pretraining error and downstream planning regret, leading to a finite-sample generalization bound for JEPA-based world models. Our analysis reveals an inherent trade-off between approximation and sample errors with respect to the latent dimension, providing theoretical insights into the advantages and limitations of latent predictive models compared with input-level predictive approaches.