World models enable agents to plan by predicting future states of the environment, but their predictions can become unreliable when test-time dynamics differ from those seen during training. We present JEPA-TTT, which adapts the latent dynamics predictor of a pretrained action-conditioned Joint-Embedding Predictive Architecture world model throughout test time. Self-supervised updates accumulate across episodes, while the visual encoder and reward head remain fixed, preserving the pretrained representation and task objective. Planning requires neither a goal image nor online environment reward. JEPA-TTT uses dense replay, which forms prediction windows at every temporal offset, retains them in a growing buffer, and samples minibatches from that buffer for predictor updates. Across eight dynamics shifts in four continuous-control environments, JEPA-TTT improves planning on every shift. After 500 test-time episodes, it reduces autoregressive latent prediction error by 83% on average and improves planning performance by 153% over the frozen JEPA world model. These results show that persistent self-supervised test-time training can adapt a pretrained latent world model under changed dynamics.
Figures & tables
Figure 1 : A dynamics shift can cause a pretrained world model to fail during planning. The model predicts the outcome of a candidate action based on the dynamics learned during training, but the test environment has different transitions. This mismatch can cause the planner to select wrong actions.
Figure 2 : Overview of JEPA-TTT . Cross Entropy Method (CEM) plans with the current latent dynamics predictor Fθ and frozen reward head Rω⋆ . After executing the actions, the frozen visual encoder maps the new observation ot+1 to a target latent zt+1 . JEPA-TTT stores prediction windows at every temporal offset and samples from the growing buffer to update Fθ with a self-supervised latent prediction objective. The adapted predictor is retained across episodes.
Figure 3 : Test-time training can correct the ranking of candidate plans under dynamics shift. For each environment, Plan A is optimized using the frozen JEPA world model, and Plan B is optimized using JEPA-TTT from the same initial state. The frozen JEPA world model predicts that A will achieve a higher task score, but the ground-truth rollouts show that B is better. After adaptation, JEPA-TTT predicts the correct ordering. Both planners use the same frozen reward head Rω⋆ , so the change in plan ranking comes from the adapted latent dynamics predictor. A pretrained decoder is used only to visualize latent predictions.
Figure 4 : Train and test dynamics illustrated for the eight shifts in our evaluation suite.
Figure 5 : Planning performance under the test-time dynamics shift. Panel (a) reports the best held-out score. Panel (b) reports normalized held-out planning-score AUC from episode 0 to 500. Bars and error bars report means and standard deviations over three deployment runs. JEPA-TTT outperforms other methods including the frozen JEPA world model baseline, PPO test-time training (PPO-TTT), and AdaJEPA.
Figure 6 : Held-out planning performance throughout 500 test-time training episodes. Curves show means over three deployment runs evaluated on the same fixed set of 100 held-out episodes, and shaded regions show standard deviations. JEPA-TTT is very close to, or even slightly surpasses, the offline reference using only 500 episodes of online interaction, and outperforms other methods including the frozen JEPA world model baseline, PPO test-time training (PPO-TTT), and AdaJEPA.
Figure 7 : Five-block autoregressive latent prediction error under the test-time dynamics. The orange bars show JEPA-TTT’s prediction errors drop significantly after 500 test-time episodes. Percent labels show the relative MSE reduction after adaptation, and error bars show standard deviations over the three adapted models.
Shift
Metric
Sparse stream
Compute-matched sparse stream
Dense stream
Dense replay
PushT: action rotation
Best
0.253±0.063
0.236±0.049
0.511±0.009
0.490±0.022
AUC
0.169±0.007
0.169±0.021
0.359±0.044
0.378±0.022
PushT: contact rotation
Best
0.317±0.258
0.521±0.319
0.646±0.044
0.619±0.033
AUC
0.205±0.082
0.320±0.157
0.490±0.041
0.555±0.028
Two-Room: spatial wave
Best
0.505±0.031
0.790±0.002
0.797±0.002
0.799±0.002
AUC
0.427±0.034
0.672±0.020
0.697±0.012
0.728±0.012
Table 1 : Comparisons of different update rules for test-time training. The sparse stream uses its natural number of updates. Entries report means and standard deviations over three runs. The Mean rows summarize the three run-level averages across shifts. Bold marks the highest unrounded mean for each metric and shift. Dense replay has the highest aggregate mean best score and AUC.
Component
Setting
Test-time interaction horizon
500 episodes under the test-time dynamics
Episode horizon
PushT: 300 raw steps, other tasks: 50
Held-out evaluations
episodes 0,50,…,500
Evaluation size
100 held-out episodes per evaluation
CEM
256 candidates, 6 iterations, 32 elites
Planning horizon
5 world-model blocks
Table 2 : Shared test-time planning and training settings. These settings define the common test-time protocol for all eight shifts.
Shift
Metric
Frozen JEPA
PPO-TTT
AdaJEPA
JEPA-TTT
PushT: action rotation
Best
0.080
0.119±0.009
0.121±0.015
0.490±0.022
AUC
0.080
0.110±0.006
0.121±0.015
0.378±0.022
PushT: contact rotation
Best
0.127
0.237±0.029
0.186±0.016
0.619±0.033
AUC
0.127
0.219±0.005
0.186±0.016
0.555±0.028
Two-Room: spatial wave
Best
0.557
0.554±0.334
0.579±0.013
0.799±0.002
AUC
0.557
0.421±0.295
0.579±0.013
0.728±0.012
Table 3 : Exact values for Figure 5 . JEPA-TTT and PPO-TTT’s entries are means and sample standard deviations over three deployment runs. For AdaJEPA, each run averages the task scores from 100 episodes, with the pretrained model reloaded before every episode. Entries report the mean and sample standard deviation across three runs. Because AdaJEPA does not retain adaptation across episodes, the same score is reported for both metrics. JEPA-TTT outperforms all three baselines on every shift under both metrics.
Figure 8 : Full held-out score curves for JEPA-TTT and the three ablations of dense replay. Curves show means and sample standard deviations over three runs. Dashed gray lines show the frozen JEPA world model. Dotted green lines show the data-rich offline reference trained under the test-time dynamics. This reference is neither a deployable baseline nor a mathematical upper bound because it uses a different data-collection distribution and receives 3,000 complete offline episodes from the shifted environment (Table 9 ). Dense replay achieves the highest aggregate best score and AUC.
Shift
pred-last + enc-last (default)
pred-last + enc-frozen
PushT: action rotation
0.1210±0.0155
0.1042±0.0163
PushT: contact rotation
0.1864±0.0156
0.1527±0.0145
Two-Room: spatial wave
0.5791±0.0134
0.5850±0.0125
Two-Room: grid rotation
0.2031±0.0101
0.2470±0.0028
Reacher: joint phase
0.4259±0.0565
0.4266±0.0552
Reacher: harmonic
0.4176±0.0490
0.4183±0.0487
Table 4 : AdaJEPA performance with two choices of trainable components. Each run averages the task scores from 100 episodes, with the pretrained model reloaded before every episode. Entries report the mean and sample standard deviation across three runs. Because adaptation does not carry across episodes, the same performance estimate is used for both best held-out score and normalized held-out planning-score AUC. Bold marks the higher mean within each shift. The difference between the two configurations is only 0.0017. All main comparisons use the default pred-last+enc-last configuration.
Shift
Frozen JEPA MSE
Adapted MSE
Reduction
PushT action rotation
1.82540
0.19988±0.00491
89.05%
PushT contact rotation
0.32117
0.17751±0.00953
44.73%
Two-Room spatial wave
1.74541
0.18697±0.02870
89.29%
Two-Room grid field
1.48501
0.02205±0.00345
98.52%
Reacher joint phase
2.11969
0.48927±0.02162
76.92%
Reacher harmonic
2.00316
0.35488±0.00216
82.28%
Table 5 : Five-block autoregressive latent MSE before and after dense replay adaptation. Adaptation reduces MSE on all eight shifts, with reductions from 45% to 99%.
Figure 9 : Decoded predictions for both PushT and Two-Room shifts. Rows within each subpanel are ground truth, Frozen JEPA prediction, and TTT-trained-world-model prediction; columns are 5, 15, and 25 raw steps ahead. After adaptation, decoded predictions more closely follow the ground truth across prediction horizons.
Figure 10 : Decoded predictions for both Reacher and OGB-Cube shifts, using the same layout as Figure 9 . After adaptation, decoded predictions more closely follow the ground truth across prediction horizons.
Shift
Metric
Persistent
Episodic reset
Difference
PushT action rotation
Best
0.4683±0.0239
0.1049±0.0168
+0.3634
AUC
0.3863±0.0247
0.0934±0.0016
+0.2929
PushT contact rotation
Best
0.6393±0.0292
0.1735±0.0924
+0.4658
AUC
0.5324±0.0352
0.1254±0.0127
+0.4069
Two-Room spatial wave
Best
0.7971±0.0016
0.5972±0.0135
+0.1999
AUC
0.7692±0.0041
0.5614±0.0089
+0.2078
Table 6 : Best held-out score and normalized held-out planning-score AUC for the matched persistence-versus-reset control. Entries are means and sample standard deviations over three runs. Persistence improves both metrics on every shift.
Task
Frozen JEPA best/AUC
Continued training best
Continued training AUC
PushT
0.2682
0.3599±0.0187
0.3247±0.0264
Two-Room
0.8039
0.8013±0.0019
0.7894±0.0084
Reacher
0.9995
0.9520±0.0065
0.9438±0.0042
OGB-Cube
0.6439
0.6492±0.0010
0.6465±0.0003
Table 7 : Best held-out score and normalized held-out planning-score AUC for dense replay under the unchanged training dynamics. Continued training produces no consistent improvement when the dynamics do not change.
Task
Planning score
Strict success
Reward-head MSE
Decoder MSE
PushT
0.2682
0%
2.54×10−4
8.10×10−4
Two-Room
0.8039
100%
1.24×10−6
5.21×10−6
Reacher
0.9995
100%
1.36×10−6
3.85×10−5
OGB-Cube
0.6439
100%
3.34×10−5
5.77×10−4
Table 8 : Planning and held-out reward-head and decoder diagnostics under the training dynamics. All reward heads and decoders have low held-out MSE, while PushT is the only task without strict CEM success.
Shift
Frozen JEPA
JEPA-TTT best
JEPA-TTT AUC
Offline model [95% CI]
PushT action rotation
0.0804
0.490±0.022
0.378±0.022
0.5667[0.5447,0.5886]
PushT contact rotation
0.1268
0.619±0.033
0.555±0.028
0.6321[0.6145,0.6497]
Two-Room spatial wave
0.5572
0.799±0.002
0.728±0.012
0.7993[0.7970,0.8017]
Two-Room grid field
0.2106
0.799±0.003
0.725±0.019
0.8021[0.7993,0.8050]
Reacher joint phase
0.3673
0.758±0.012
0.619±0.019
0.9994[0.9993,0.9994]
Reacher harmonic
0.3687
0.688±0.075
0.518±0.012
0.9994[0.9994,0.9995]
Table 9 : Offline references trained under the test-time dynamics. The best held-out score and normalized held-out planning-score AUC follow the definitions in Section 4.1 . The offline reference is a separately pretrained model evaluated once on 100 held-out episodes. The best JEPA-TTT score is within 0.08 of the offline reference on six shifts, with larger gaps on the two Reacher shifts.
Figure 11 : Online test-time performance for JEPA-TTT and the three ablations of dense replay. Curves show means ± sample standard deviations over three runs, binned into consecutive groups of 50 episodes. Dense replay achieves the highest aggregate online score.