World-model planners typically scale outward by rolling farther, sampling more trajectories, or optimizing longer, while assigning the same computation to every imagined transition. We show that making every transition uniformly deeper wastes computation and can degrade planning because useful refinement is concentrated at a small set of decision-critical events. We introduce DeepJEPA, a weight-tied joint-embedding predictive world model that treats transition depth as an inner test-time scaling axis and learns when another recurrent update is worth computing for each candidate and rollout step. Across five visual-control settings, DeepJEPA improves or matches the strongest fixed-depth planner while averaging only 1.00-1.26 updates per transition. Its additional computation concentrates at contact onset and sustained object interaction, where latent corrections can change which candidates enter the planner's elite set and which action is selected. Representation probes further show that improved planning does not require uniformly better object-state decodability. DeepJEPA therefore reframes world-model scaling as a problem of allocating internal computation where it can change the planner's decision: think deeper at decision-critical transitions instead of making every rollout uniformly deeper or longer.
Figures & tables
Figure 1 : Think before you roll. Conventional planners scale the outer search while assigning every imagined transition the same predictor depth. DeepJEPA scales from within by keeping most transitions shallow and refining only decision-critical ones before the planner selects an action. The task, rollout horizon, candidate budget, and outer planning loop remain unchanged.
Figure 2 : DeepJEPA scales world-model computation from within. (a) At inference each CEM candidate enters a weight-tied transition cell that may update its latent prediction up to four times; filled pips give the depth selected for that transition. (b) A continue head decides depth independently for every candidate–time pair: the dashed ring is the target latent, the dot the current prediction, and the head continues while the predicted marginal gain exceeds η . (c) During training, the marginal reduction in latent error after each update supplies the continue label. Future target latents are used only in training and never at inference.
Method
Reacher
Single
Double
Triple
PushT
LeWM
81.3
72.0
74.7
74.0
10.4
K=1
84.0
79.3
72.0
74.0
10.7
K=2
82.0
78.7
72.0
74.0
10.7
K=3
83.3
78.0
72.7
73.3
10.4
K=4
82.7
78.0
72.0
74.0
9.8
DeepJEPA
85.3
79.3
74.0
77.3
11.6
Table 1: Adaptive depth improves the fixed-depth frontier on four of five settings. Mean success rate in percent over three training/evaluation seed pairs, 50 episodes each (150 episodes per cell). Single, Double and Triple are the Cube tasks. LeWM is a separate non-recurrent checkpoint family, not a depth setting. The last row is in recurrent updates, not percent. Per-seed standard deviations are in Table 4 . Higher is better.
Figure 3 : DeepJEPA improves the fixed-depth frontier while staying at the shallow endpoint. Left: the move from the best fixed depth (hollow dot) to DeepJEPA (arrowhead) in percentage points; the two absolute success rates are printed either side of the track, and the stub on Cube Single is a measured zero, not a missing bar. Centre: mean selected depth Kˉ as a bar inside the full K=1 to K=4 budget the model could have spent. Right: every arm relative to K=1 , so a reader can check that no fixed depth beats the adaptive one; blue is above K=1 , brown below. Means over three training/evaluation seed pairs, 150 episodes per cell; per-seed spreads are in Table 4 . Higher is better.
Task
η=.30
.45
.50
.60
.70
Mean success (%)
Reacher
82.7
85.3
84.0
84.7
82.7
Single
78.0
77.3
78.0
77.3
79.3
Double
73.3
74.0
73.3
72.0
72.7
Triple
77.3
73.3
74.7
73.3
72.7
Mean selected depth Kˉ
Table 2 : Threshold controls allocation, not average depth. Success and selected depth across η . Bold marks the main operating point.
Figure 4 : Adaptive computation is interaction aligned. (a) Fraction of PushT transitions refined beyond K=1 , aligned on contact onset; the band is the 95% cluster-bootstrap interval. (b) Phase effects on refinement probability for Cube Triple, adjusted for rollout-step fixed effects, distance and action norm; dots are point estimates, whiskers 95% cluster-bootstrap intervals, and squares mark negative effects. Blue is more refinement, brown less. 100 roots per task, one checkpoint per task.
Figure 5 : Additional depth acts at CEM decision boundaries. (a) One mark per root, 150 roots per task, ordered so the flips form a seam: pale blue squares succeed at both K=1 and K=4 , pale grey squares fail at both, blue squares are helped by depth and brown discs are harmed. 93–99% of roots never change outcome. (b) Across 450 Cube Triple root–search-seed pairs the net effect is +1.1 points and its 95% interval crosses zero (paired exact p=0.424 ). (c) Re-sampling the search on the ten reference-discordant roots changes the flip direction in five of thirty trials. Together with exact-batch replay this is the signature of a local ranking mechanism: depth matters when a small cost correction can change the elite set.
Figure 6 : Global object-state decodability is not the mechanism. Paired MLP-probe error differences between fixed K=1 and adaptive DeepJEPA on 10,000 held-out transitions per task. Dots are paired means, whiskers 95% intervals. Positive favours adaptive depth. PushT effects stay below 0.011 pixels; Cube Triple object position moves 1.7μ m in the opposite direction. This null rules out a uniform state-semantics explanation and supports the local CEM-ranking mechanism.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Horizon
Action block
CEM samples
Iterations
Elites
η
Reacher
5
5
300
30
30
0.45
Cube Single
5
5
300
30
30
0.70
Cube Double
5
5
300
30
30
0.45
Cube Triple
5
5
300
30
30
0.30
PushT
10
—
—
—
—
0.60
Appendix
Table 3: Main evaluation protocol. Each row uses three training/evaluation seed pairs and 50 episodes per pair. Cube goals are sampled 25 dataset steps after the initial observation. All fixed-depth and learned-depth evaluations within a recurrent checkpoint share episode roots and CEM seeds. PushT uses the planner configuration shipped with the benchmark, identical across all methods, marked —.
Method
Reacher
Single
Double
Triple
PushT
LeWM
4.2
12.0
7.6
8.0
2.3
K=1
5.3
8.1
3.5
4.0
1.2
K=2
2.0
7.6
5.3
0.0
2.4
K=3
1.2
8.7
5.0
5.0
2.0
K=4
1.2
8.7
5.3
6.0
1.7
DeepJEPA
4.2
8.1
3.5
7.6
2.5
Appendix
Table 4: Seed spread for every entry of Table 1 . Standard deviation of success rate in percentage points across the same three training/evaluation seed pairs. Rows and columns match Table 1 exactly.