Latent world models based on Joint-Embedding Predictive Architecture (JEPA) are deterministic by design. While successful in fully observable scenarios, this paradigm breaks down when past observations and actions lead to multiple plausible future possibilities, e.g., due to occlusion. We introduce EpicWorldModel, a framework to train stochastic JEPAs for environments and tasks with inherent uncertainty under partially observability. We jointly train the EpicWorldModel predictor with its latent representation space to directly predict multiple potential future states using a flow-matching objective, when the goal-relevant scene content is absent from the conditioning history. We show that flow predictive variance, motivated by its relation to an upper bound on predictive entropy, serves as a useful exploration guidance for planning. By incorporating this uncertainty signal into Cross-Entropy Method (CEM)-based planning, our approach balances goal-reaching with exploration of uncertain regions where occluded goals are most likely to be located. We demonstrate the effectiveness of EpicWorldModel through a series of latent planning experiments with the best or on-par performance across tasks, showing up to 22% empirical improvement in success rate over LeWorldModel.
Figures & tables
Figure 1 : A robot is asked to navigate to a fridge initially outside its field of view. A deterministic world model memorizes the location of the fridge from training data and hallucinates it at inference, even when the true location differs. A stochastic model instead maintains a distribution over plausible futures, enabling uncertainty-aware planning. Dark gray framed images indicate observation inputs; colored frames indicate predicted futures. We only show four possibilities at t and the results of a single choice at t+1 . In practice observations and live in the jointly learned representation space.
Figure 2 : Loss curves of EpicWorldModel on the visual PointMaze environment. The bold colored lines represents the smoothed loss and the faint colored lines the actual per step loss. Overall we observe minimal variance and smooth convergence of the encoder and stochastic predictor.
Figure 3 : Success rate comparison across benchmark environments. FM and IMF represent EpicWorldModel with the flow matching objectives introduced in Equation 5 (FM) and Equation 6 (iMF). EpicWorldModel achieves the best performance in both the default- and long-horizon settings.
Figure 4 : Evaluation Tasks. In each of the five tasks the most left image represents the initial observation o0 and the right a goal observation. All tasks are POMDP evaluations in existing benchmarks. The first three tasks are available in OGbench Park et al. (2025) , Car Racing is available in Gymnasium Towers et al. (2024) , and LabMaze in DMLab Beattie et al. (2016) .
Figure 5 : Success rate comparison across benchmark environments. EpicWorldModel achieves competitive planning success rate across diverse navigation task under partially observable scenes and is on-par for complex manipulation in the Visual Scene.
Figure 6 : Ablation study on disagreement weight across environments. (a) Visual PointMaze Giant shows clear sensitivity to hyperparameters with optimal performance at β=0.4 and Number Flow Evolutions(NFE)=2. (b) AntMaze Giant demonstrates more stable performance across settings, with modest gains from moderate exploration. (c) Visual Scene shows NFE values (NFE=1) achieve better performance.
Figure 7 : Latent uncertainty contraction during planning on a single rollout. Each panel shows the top- k latent predictions in a shared PCA space, together with the current latent ( z0 or z25 ), the fixed goal latent ( zgoal ), and the top-k covariance ellipse. The red arrow measures ∥zt−zgoal∥
Start Dist.
LeWM
LeWM 3×
iMF ( β=0 )
iMF ( β=.1 )
FM ( β=.1 )
0.50 m
60.0
93.3
100.0
100.0
100.0
0.92 m
46.7
75.0
80.0
80.0
86.7
1.26 m
53.3
40.0
46.7
80.0
80.0
1.46 m
20.0
25.0
60.0
60.0
93.3
Average
45.0
58.3
71.7
80.0
90.0
Table 1 : Controlled difficulty analysis on RoboCasa NavigateKitchen. Success rate (%) across increasing start distances. Best results are shown in bold , and the second-best average is underlined . (iMF NFE=2; FM NFE=5; LeWM 3× n = 8–12)
Figure 8 : Qualitative rollouts of EpicWorldModel on RoboCasa NavigateKitchen. Each row shows an episode in a different kitchen layout from the validation scene split. The initial visibility is < 0.5 for all episodes, only displaying the hardest tasks. All episodes are conditioned on the same instruction, and frames are ordered left to right in time. At the start of each episode the fridge is outside the field of view or occluded by cabinets and counters. Guided by the uncertainty cost U , the robot first turns and moves toward unobserved regions of the kitchen, then commits to the goal once the fridge becomes visible. All shown results are from an iMF model ( β=0.1 , NFE =2 ).
Initial Visibility
LeWM
iMF ( β=0 )
iMF ( β=.1 )
n
<0.5
0/6
1/6
5/6
6
0.5 – 0.8
2/7
5/7
6/7
7
0.8 – 0.95
1/4
2/4
1/4
4
≥0.95
10/22
21/22
19/22
22
Pooled
13/39 (33.3)
29/39 (74.4)
31/39 (79.5)
39
Table 2 : Occlusion/visibility analysis on RoboCasa NavigateKitchen. Successes / episodes (pooled %) stratified by initial goal visibility. Lower visibility indicates stronger partial observability. Best results are shown in bold . (iMF NFE=2; FM NFE=5; LeWM 3× n = 8–12)
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
iMF
FM
β\ NFE
1
2
4
8
10
4
10
0.0
51.0
69.5
76.5
81.5
81.5
78.0
79.5
0.4
51.5
74.5
83.5
86.0
86.0
81.0
85.5
Appendix
Table 3: Effect of disabling the exploration term ( β=0 ) on Visual PointMaze Giant, as a function of the number of function evaluations (NFE) at planning time. Left: iMeanFlow (iMF); right: flow matching (FM). Success rate (%), single training/evaluation run per cell. Best value per column in bold .
CEM candidates
PointMaze
AntMaze
300 (default) †
64.0
26.0
600
68.0
26.5
900
72.0
24.0
1500
74.0
26.5
6000
74.0
25.0
EpicWM (iMF, β=0.4 , NFE=2)
74.5
—
Appendix
Table 4: Compute-matched comparison: LeWM success rate (%) against the number of CEM action candidates (default: 300), on Visual PointMaze Giant and Visual AntMaze Giant. Bottom rows: EpicWorldModel (iMF, β=0.4 ) at NFE = 2 and NFE = 4 for reference (from Tab. 3 ). Best value per column in bold .
Method
β
N=2
N=5∗
N=10
Range
iMF
0.0
78.0
76.5
77.5
1.5
iMF
0.4
80.0
83.5
78.0
5.5
FM
0.0
79.0
82.0
81.0
3.0
FM
0.4
81.5
81.0
81.0
0.5
Appendix
Table 5: Effect of the number of sampled futures N on Visual PointMaze Giant success rate (%), for both predictor families (FM and iMF) at β∈{0,0.4} (NFE fixed at each method’s main-text default; App. B). ∗ marks N=5 , the default value used in the main results. “Range” is max−min over the row. Best value per row in bold .
School of Mechanical and Aerospace Engineering, Nanyang Technological University, Singapore · The University of Sheffield, Sheffield, United Kingdom · School of Artificial Intelligence (School of Software), Yanshan University, Qinhuangdao, China