We introduce Adaptive LeWorldModel (ALeWM), a world model based on a joint-embedding predictive architecture (JEPA) that learns to concentrate predictive information in compact prefixes of a wide latent representation. To encourage this ordering, ALeWM learns a sequence-conditioned distribution over prefix lengths and trains the predictor to estimate the full next embedding from a sampled input prefix. As standard anti-collapse objectives encourage variation across latent coordinates and do not organize them by predictive importance, we also introduce MixSIGReg. MixSIGReg regularizes the masked embeddings against a prior-weighted mixture with Gaussian active prefixes and zeros in the remaining coordinates. As a result, the ALeWM objective encourages early coordinates to retain information useful for prediction and recursive planning. Our analysis shows that the mixture distribution used by MixSIGReg assigns higher variance to earlier coordinate blocks and lower variance to later ones. In addition, we show that, under specified assumptions, prediction error is minimized by placing the information most useful for prediction in earlier blocks. Empirically, we study the behavior of ALeWM in a controlled dynamical system with known state variables and in goal-conditioned visual control. We show that ALeWM consistently achieves higher mean success rates than tuned fixed-width LeWM, with lower planning capacity on average.
Figures & tables
Figure 1: ALeWM overview. Left: A shared encoder maps consecutive observations to latent states, and a learned network selects the active prefix capacity k for masking the state st . The predictor estimates the full next latent state from the masked embedding and action, while the prediction loss and MixSIGReg objective jointly train the representation. In orange , our modifications to the LeWM framework. Right: Gains of ALeWM over fixed-width LeWM using ViT-Tiny models with latent dimension 192 . Moving left and up indicates lower average capacity and higher test success rate.
Figure 2: Four-factor controlled damped oscillator diagnostics. (a) Mean linear state-recovery R2 as prefix size increases; the dashed line marks the true factor count k=4 . For LeWM with d∈{4,6} the curves are continued horizontally beyond their trained dimension only as a visual reference. (b) Raw unmasked latent covariance spectra. (c) ALeWM whitened Procrustes alignment, fitted using training statistics. (d) ALeWM empirical masked-latent variance vs the target prior variance.
Backbone ( dmax )
Method
TwoRoom
PushT
Reacher
OGBench-Cube
SR ( ↑ )
E[Kplan]
SR ( ↑ )
E[Kplan]
SR ( ↑ )
E[Kplan]
SR ( ↑ )
E[Kplan]
ViT-Tiny (96)
LeWM (best d )
95.00±0.00
32.00
89.33±0.44
96.00
83.67±0.73
64.00
72.33±1.17
96.00
LeWM ( d=96 )
92.17±0.17
96.00
89.33±0.44
96.00
82.00±1.76
96.00
72.33±1.17
96.00
ALeWM
99.50±0.00
13.33
94.67±0.83
32.00
85.83±0.60
21.33
77.17±1.17
16.00
ViT-Tiny (192)
LeWM (best d )
95.00±0.00
32.00
92.67±0.33
160.0
84.50±1.15
160.0
72.33±1.17
96.00
LeWM ( d=192 )
88.17±0.33
192.0
92.67±0.33
192.0
83.67±1.30
192.0
68.67±0.17
192.0
Table 1: LeWM and ALeWM mean test success rate ( ± SEM) and mean planning prefix E[Kplan] .
Figure 4
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Rollout target
PushT SR ( ↑ )
Reacher SR ( ↑ )
OGBench-Cube SR ( ↑ )
Full Capacity (dmax)
96.00±0.58
83.50±2.25
75.17±0.44
Max Probability (Kplan)
96.00±0.00
85.83±1.01
79.00±0.76
Appendix
Table 2: Effect of the rollout target capacity size on test success rate (SR, %) with ViT-Tiny ALeWM having dmax=192 . Results present test success rate (SR, %) ± SEM over 3 seeds.
Figure 6: Mean capacity probability over training for the ViT-Tiny ALeWM models with dmax=192 for one seed. Each curve is the capacity probability averaged over the training batch at each logged step.
Figure 7: Dataset-conditioned capacity dynamics for joint DMC–TwoRoom training, averaged over three seeds at fixed steps. Each panel corresponds to one capacity; solid and dashed curves show the mean selector probability for DMC and TwoRoom, respectively. For better visibility, the vertical range is capped at 0.4 , thereby omitting the initial probability 0.837 of capacity k=192 .
Figure 8: Reacher rollout capacity allocation for ViT-Tiny ALeWM with dmax=192 . Left: Euclidean start–goal arm-tip displacement in pixels, grouped by selected capacity; black bars indicate means and parentheses give episode counts. Right: the logged selector log probability ratio versus Euclidean distance between the start and goal encoder CLS tokens. Squares denote K=32 and circles K=64 ; blue and purple indicate successful rollouts, respectively, and orange indicates failures.
Figure 9: Qualitative comparison on two matched test episodes per dataset, selected for ALeWM success and LeWM failure. Each pair places LeWM above ALeWM. Columns show the initial observation, four approximately uniformly spaced intermediate observations, the final recorded observation, and the shared goal. Images are environment observations under model-planned actions.
dmax
Dataset
λ
sq
α
96
TwoRoom
0.15
3
3
96
PushT
0.09
2
−1
96
Reacher
0.30
3
2
96
OGBench-Cube
0.30
3
1
192
TwoRoom
0.30
3
3
192
PushT
0.15
2
−0.5
Appendix
Table 3: Selected ALeWM hyperparameters for the main control results. λ is the MixSIGReg coefficient, sq multiplies the base learning rate for qψ , and α is the polynomial-prior degree.
Backbone ( dmax )
Dataset
Seed
Kplan : #rollouts
ViT-Tiny (96)
TwoRoom
3072
16:200
3073
8:200
3074
16:200
PushT
3072
32:200
3073
32:200
3074
32:200
Appendix
Table 4: Per-seed planning-capacity counts for the selected ALeWM configurations in Table 1 . An entry k:n means that Kplan=k for n episodes; capacities that were never selected are omitted. Test contains 200 episodes per seed.
Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models. For visual planning, however, LeWM evaluates candidate action sequences by repeatedly applying a local one-step latent transition model. This autoregressive rollout makes planning computationally expensive and exposes the predicted trajectory to accumulated latent errors as the horizon grows. We propose Fast LeWorldModel (Fast-LeWM), a fast latent world model that replaces repeated local rollout with action-prefix prediction. Given the current latent and a candidate action sequence, Fast-LeWM encodes its prefixes and predicts the future latents reached after executing those prefixes in parallel. By making action prefixes the basic prediction unit, Fast-LeWM directly models action effects accumulated to different extents over multiple horizons. This prefix-level supervision forces the model to learn how states continuously evolve under different action prefixes, rather than only fitting one-step state transitions. During planning, the predictor can use the prefix token from the encoded action sequence to evaluate the corresponding future latent without explicitly rolling through each intermediate imagined state. Across multiple tasks, Fast-LeWM improves average success over LeWM while substantially reducing planning time, achieving lower open-loop latent loss whose growth becomes significantly slower as the rollout horizon increases.
Latent world models enable planning from high-dimensional observations by predicting future states in a compact latent space. However, these models are typically kept frozen at test time: when their predictions become inaccurate, planning can fail, especially under test-time distribution shift. To address this, we propose AdaJEPA, an adaptive latent world model that performs test-time adaptation within the closed loop of model predictive control (MPC). After training, AdaJEPA plans and executes the first action chunk, uses the observed next-state transition as a self-supervised adaptation signal, and replans with the updated model. This closed-loop update continuously recalibrates the world model without additional expert demonstrations. Across a range of goal-reaching tasks, AdaJEPA substantially improves planning success with as few as one gradient step per MPC replanning step.
Latent world models learn action-conditioned dynamics in representation space and often score candidate actions by Euclidean distance to a goal representation. Joint training typically regularizes the representation to prevent collapse, but the resulting representation geometry also determines how terminal errors are weighted during planning. We show that accurate prediction and noncollapsed representations do not guarantee a task-aligned latent planning cost: isotropic Gaussian regularization can induce a geometry that ranks feasible outcomes differently from the task cost. To address this mismatch, we introduce AnisoWM with ΛReg, which replaces the fixed isotropic Gaussian target with a learnable diagonal covariance under fixed-trace and anisotropy constraints. The prediction objective, predictor architecture, and Euclidean planner remain unchanged; the target is used only during training. Our analysis characterizes the prediction-driven allocation of target variance, its dependence on the training distribution, and the conditions under which the induced metric reduces planning regret. Across four visual control environments, AnisoWM improves planning success over LeWorldModel in all four. Its latent planning cost also shows better agreement with task outcomes. Project website: https://rkdrn79.github.io/AnisoWM-page/
Mingu Kang, Yoori Oh, Sookyung Kim +1
Seoul National University · Ewha Womans University