World models predict future observations from current experience and actions, yet prediction can depend on observations seen far in the past. Episodic memory preserves past observations for later recall; however, as memory accumulates, it raises a fundamental question: which memories are useful for the current prediction, and which available retrieval cues should be trusted to find them? This is challenging because fixed criteria based on recency, pose overlap, or visual similarity can be unreliable across environments and queries. We propose Future-Aware Recall (FAR), a framework that learns episodic recall from future-aware predictive supervision and adaptive multi-cue scoring. During training, FAR measures predictive utility by the conditional log-likelihood of the realized future given recalled context, approximated by negative diffusion prediction loss, and uses it to train a retriever that remains future-blind at inference. The retriever learns cue-specific relevance and automatically determines which available retrieval cues, such as time, pose, vision, and audio, to trust for each query when selecting memories. Across three complementary settings, FAR outperforms hand-designed recall even with the same retrieval cues, automatically adapts which available cues to trust, and recalls the right history as the world changes. Together, these results establish FAR as a flexible, principled approach to episodic memory access in world models.
Figures & tables
Figure 1: The three advantages of Future-Aware Recall (FAR). Top: FAR retrieves context based on its usefulness for predicting the future, beating hand-designed memory recall rules. In the figure, FAR is compared against WorldMem ( Xiao et al., 2025 ) . Middle: FAR adaptively learns query-dependent weights over multiple available retrieval cues. Here, FAR learns to rely more on audio cue to distinguish memories with occluded and clear line of sight. Bottom: FAR recalls memory to support state-correct future predictions. Here, FAR distinguishes between closed and open fridges, in contrast to WorldMem.
Figure 2: Overview of Future-Aware Recall (FAR). FAR scores memories using general retrieval cues Zt−1 , adaptively fuses cue-specific relevance, and recalls a compact Top- K context Ct from episodic memory Mt−1 for world-model prediction. During training, the observed future ot+1 provides predictive utility to supervise the retriever, while recall remains future-blind at inference.
Figure 3: LoopNav rollout quality. We report frame-wise PSNR ↑ , LPIPS ↓ , and DreamSim ↓ between generated and ground-truth return trajectories as a function of normalized trajectory time. Curves show the mean across trajectories, with shaded regions indicating 95% bootstrap confidence intervals. Errors are highest near the middle of the return phase, where relevant observations are typically farthest in time and memory is most sparse.
Figure 4: Illustration of a data trajectory.
Figure 5: SoundSpaces rollout quality. We report PSNR, LPIPS, and DreamSim, where the error is averaged over predicted frames along return trajectory. Thus, the x -axis denotes the length of the trajectory over which the rollout is evaluated, not the rollout horizon . Top: results on the corpus where the agent performs periodic 360∘ scans during exploration. Bottom: results on the corpus where the agent scans at the endpoints, i.e., at the beginning and end of exploration.
Figure 6: Retrieval comparison on SoundSpaces. Given the observation and goal view, we visualize the top-1 context retrieved by each method and the generated goal frame, with regions of large error overlaid in red . The spectrograms show the audio cue associated with each state. The map shows the exploration and return trajectories with the memories selected by each method.
Figure 7: History-dependent counterfactual prediction in AI2-THOR. We compare two trajectories with the same current closed-fridge observation and opening action but different histories: one history contains the event of placing a tomato inside the fridge, while the other does not.
Figure 8: State accuracy under object manipulation.
Figure 9: Off-scene dynamics prediction. Left: trajectory-level accuracy for predicting the future location of the independently moving agent A2 . Right: qualitative example showing that FAR recalls a motion-informative episode containing A2 , and predicts the correct future state with A2 . WorldMem retrieves a spatially relevant but dynamically uninformative memory, failing to render A2 . All A2 appearances are outlined in green for visual aid.
Retrieval
PSNR ↑
LPIPS ↓
DreamSim ↓
Temporal (NWM)
13.229±0.114
0.582±0.007
0.200±0.004
LongLive-RAG
14.847±0.121
0.463±0.006
0.133±0.003
WorldMem
15.372±0.112
0.448±0.006
0.130±0.003
FAR with Vision Cue
Enc. Pretrained
16.248±0.112
0.390±0.006
0.100±0.003
+ MLP Adapter
16.480±0.108
0.374±0.006
0.091±0.002
Table 1: Ablation on LoopNav. We report average loop closure errors for various configurations.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
World Model
External Retriever
Prediction Supervised
Adaptive Cue Fusion
Cue Options
FramePack ( Zhang et al., 2025 )
✗
✓
✗
✗
Time, Vision
LongLive-RAG ( Hu et al., 2026 )
✗
✓
✗
✗
Vision
WorldMem ( Xiao et al., 2025 )
✓
✓
✗
✗
Time, Pose
Context-as-Memory ( Yu et al., 2025 )
✓
✓
✗
✗
Pose
VRAG ( Chen et al., 2025 )
✓
✓
✗
✗
Pose
SPMem ( Wu et al., 2025 )
✓
✓
✗
✗
Pose
Appendix
Table 2: Conceptual comparison of episodic memory recall for video and world models. World Model indicates whether the method is designed for action-conditioned future prediction; External Retriever whether memory selection is performed outside the generator; Prediction Supervised whether retrieval is optimized using the downstream prediction objective; Adaptive Cue Fusion whether the retriever dynamically weights multiple retrieval cues for each query rather than using a fixed cue or predefined combination; and Cue Options lists the information sources used to determine memory relevance in the respective papers.
LoopNav
SoundSpaces v1
SoundSpaces v2
AI2-THOR
AI2-THOR-dyn
Environment
Minecraft
Matterport3D
Matterport3D
iTHOR
iTHOR
Episodes
19,200
14,425
14,425
8,125
1,202
Scenes
120
81
81
119
7
Train/test split
episode-level
scene-disjoint
scene-disjoint
scene-disjoint
episode-level
Train/test scenes
120/120
57/24
57/24
95/24
7/7
Frames
12.62 M
9.28 M
10.95 M
9.23 M
1.85 M
Appendix
Table 3: Dataset statistics. SoundSpaces and AI2-THOR use scene-disjoint evaluation, whereas LoopNav and the reported AI2-THOR-dyn experiments use episode-level splits with the same scenes appearing in training and test.
Setting
Value
DiT depth / hidden / heads
12/768/12
Patch size
2
Training context size
4 (current +3 recalled)
Diffusion steps
1000
Noise schedule
Linear
Prediction target
ϵ -prediction
Appendix
Table 4: Shared model and optimization settings. These settings are used across all corpora unless stated otherwise.
LoopNav
SoundSpaces v1
SoundSpaces v2
AI2-THOR
AI2-THOR-dyn
Training pool size
200
200
200
400
400
Prediction horizon (frames)
128
128
128
32
32/135†
Goals per observation
4
4
4
4
–
Retrieval chunk size
10
10
10
10
30
FAR retrieval cues
Meta+Vision
Meta+Audio
Meta+Audio
Meta+Vision
Meta / Meta+Agent
Metadata MLP layers
2
2
2
3
3
Appendix
Table 5: Dataset-specific training settings. The FAR cue row gives the signals used for retrieval; visual and audio retrieval representations are not passed to the world model as generation conditions. † AI2-THOR-dyn uses a 135 -frame horizon for reveal-anchored pairs and 32 frames otherwise.
LoopNav
SoundSpaces v1
SoundSpaces v2
AI2-THOR
AI2-THOR-dyn
Evaluation checkpoint
700 k
700 k
700 k
700 k
100 k
Evaluation protocol
rollout
rollout
rollout
rollout
both-ends probe
Stride (frames / seconds)
4/0.2
16/1.6
16/1.6
4/0.4
–
Context size
12
12
12
12
4
Retrieval pool size
1000
1000
1000
1000
–
Retrieval pool phase
outbound
outbound
outbound
exploration
pre-reveal history
Appendix
Table 6: Dataset-specific inference settings. The four rollout benchmarks expand the context from 4 frames during training to 12 frames at evaluation and query up to 1000 real memory frames.