Beyond the Remembered World: Predictive 4D Belief for Persistent Navigation in Evolving Worlds
Organizations: Zhejiang University · University of California, San Diego · Deeprobotics · Chinese University of Hong Kong
Abstract
Persistent spatial memory enables embodied agents to navigate familiar environments across repeated visits. However, targets may move while unobserved, including during navigation, making remembered locations unreliable by the time an agent arrives. Despite advances in memory retrieval and state prediction, accounting for continued hidden world evolution and revising beliefs under limited visibility remain challenging. We study Evolving-World Navigation, where agents infer target locations from intermittent observations, predict their states at inspection time, and revise beliefs using visual evidence. We propose EvolvingNav, which constructs a time-indexed belief from timestamped 3D object histories through a structured persistence-relocation model. The belief distinguishes persistence at the last observed location from relocation to alternative locations and retains probability mass outside the known candidate set. An event-driven filter propagates the current belief as time elapses, forecasts target occupancy at candidate inspection times, and incorporates new RGB-D evidence. Negative observations downweight location hypotheses according to calibrated, visibility-conditioned detection probabilities, while evidence tracking prevents repeated use of the same observations. A frozen, zero-shot vision-language controller uses the updated belief to choose actions and replan. We further introduce EvoWorld-Bench, a benchmark grounded in human activity traces, comprising 54 scenes and 803,680 tasks with controlled changes before and during navigation. In simulation and real-robot experiments, EvolvingNav improves navigation success and search efficiency over the evaluated baselines. Paired experiments show the clearest gains under learnable temporal patterns, while ablations demonstrate the value of preserving uncertainty and incorporating visibility-aware evidence.
Figures & tables
| FindingDory | GOAT-Bench | EvoWorld-Bench | ||||||
| Method | HL-SR | HL-SPL | SR | SPL | Repeat-SR | First-Inspection SR | Search SR | SPL |
| Open-vocabulary and lifelong navigation | ||||||||
| VLFM ( Yokoyama et al., 2024 ) | 13.42 | 10.86 | 15.79 | 5.45 | 11.82 | 20.65 | 28.36 | 24.67 |
| ZSON ( Majumdar et al., 2022 ) | 27.89 | 21.81 | 29.64 | 2.45 | 16.37 | 26.84 | 32.13 | 28.55 |
| FindingDory Agent ( Yadav et al., 2025 ) | 52.44 PR | 40.92 | New A | New A | New A | New A | New A | New A |
| LagMemo GLUE ( Zhou et al., 2025 ) | 32.79 | 23.25 | 20.61 | 1.74 | 11.27 | 30.18 | 42.27 | 33.82 |
| Method | Dynamic SR | Online Recovery SR | Excess Distance (m) | Revisit Success |
| SLaTe-PRO ( Patel et al., 2023 ) | 34.7 | 21.5 | 18.9 | 8.7 |
| SGM+NEP ( Kurenkov et al., 2023 ) | 41.3 | 27.8 | 15.6 | 12.4 |
| FlowMaps ( Argenziano et al., 2026 ) | 29.6 | 18.9 | 22.4 | 6.5 |
| PredictiveGraphs ( Saavedra-Ruiz et al., 2026 ) | 46.8 | 36.7 | 13.2 | 18.9 |
| EvolvingNav (ours) | 65.2 | 58.7 | 8.4 | 32.6 |
| Method | First-Inspection SR | Search SR | Recovery SR | Dist. (m) | Time (s) | Inspect. |
| Last Seen + Search | 17.2 | 26.6 | 12.8 | 56.2 | 284 | 2.19 |
| Time Frequency | 21.9 | 31.3 | 14.0 | 52.9 | 277 | 2.05 |
| Retrieval + Reasoning | 26.6 | 37.5 | 17.1 | 50.3 | 281 | 2.03 |
| EvolvingNav | 34.4 | 48.4 | 24.3 | 43.8 | 248 | 1.84 |
| Update rule | Rank | Top-1 | NLL | ECE |
| No Update | 5.82 | 62.98 | 2.838 | 0.172 |
| Hard Removal | 10.09 | 42.69 | 11.514 | 0.204 |
| Bayesian Update | 5.25 | 64.71 | 1.497 | 0.151 |
| Bayesian + Calibration | 4.96 | 66.93 | 1.530 | 0.130 |
| Temporal | LOSO | |||||
| Predictor | Top-1 | MRR | NLL | Top-1 | MRR | NLL |
| Last Seen | 0.00 | 0.1699 | 3.4011 | 50.00 | 0.5850 | 1.8825 |
| Time Frequency | 33.85 | 0.5462 | 1.8400 | 49.53 | 0.6166 | 1.6455 |
| Direct Transformer | 34.16 | 0.5586 | 1.7292 | 51.71 | 0.6884 | 1.3157 |
| Shared semantic prior | 36.02 | 0.5957 | 1.5766 | 59.47 | 0.7462 | 1.0448 |
| 38.24 | 0.5971 | 1.4925 | 61.18 | 0.7523 | 0.9778 | |
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Source | ||||
| CASAS Aruba | 1,602,820 | 20,252 | 428 | 480 |
| ARAS | 5,184,000 | 5,080 | 312 | 360 |
| HD-EPIC | 59,454 | 10,786 | 742 | 864 |
| ParaHome | 212 seq. | 860 | 186 | 216 |
| OPPORTUNITY | 869,387 | 2,551 | 318 | 372 |
| HOMER+ (sim.) | 65 seq. | 1,991 | 356 | 420 |
| Regime | Mechanism | Frozen signal | Role |
| Stable/rare | Infrequent events | Home/return tendency | Persistence control |
| Routine | Activity state machine | Chain-specific locations | Routine prediction |
| Personal | Placed/carried lifecycle | Owner placement habit | Long-term memory |
| Irregular | Affordance sampling | Affordance only | Uncertainty control |
| Task | Decision signal | Primary measure | |
| P | Current-state prediction | Region/receptacle belief | R@ / NLL |
| N1 | Predictive navigation | Event-filtered belief | First-Inspection SR |
| N2 | Belief-guided search | Full belief + cost | Search SR / cost |
| N3 | Evidence-aware replanning | Updated posterior | Recovery SR |
| N4 | Online-dynamic navigation | Arrival-time belief | Online Recovery SR |
| N5 | Cross-scene generalization | Any protocol above | Held-out SR |
| Standard Simulation Worlds | Photorealistic Simulation | Evolving Simulation Worlds | ||||||||
| R2R-CE | HM3D-OVON | PointNav | SAGE-Bench | EvoWorld-Bench | ||||||
| Method | SR | SPL | SR | SPL | SR | SPL | SR | SPL | SR | SPL |
| Navigation-trained or task-adapted | ||||||||||
| MTU3D ( Zhu et al., 2025 ) | 8.8 | 6.5 | 40.8 PR | 12.1 PR | 75.8 | 72.7 | 25.9 | 15.6 | 35.2 | 30.7 |
| NaVid ( Zhang et al., 2024b ) | 37.4 PR | 35.9 PR | 60.3 | 36.0 | 15.3 | 13.4 | 17.4 | 15.1 | 31.0 | 22.6 |
| Uni-NaVid ( Zhang et al., 2024a ) | 47.0 PR | 42.7 PR | 39.5 PR | 19.8 PR | 15.9 | 13.2 | 26.4 | 23.1 | 55.3 | 41.7 |
| Predictor | Temporal Top-1 | MRR | NLL | LOSO Top-1 | MRR | NLL |
| Uniform Random | 9.32 | 0.2607 | 2.4509 | 11.96 | 0.3050 | 2.2980 |
| Last Seen | 0.00 | 0.1699 | 3.4011 | 50.00 | 0.5850 | 1.8825 |
| Object Frequency | 30.75 | 0.5476 | 1.7397 | 50.47 | 0.6441 | 1.8455 |
| Time Frequency | 33.85 | 0.5462 | 1.8400 | 49.53 | 0.6166 | 1.6455 |
| Markov | 18.94 | 0.4435 | 2.3772 | 49.38 | 0.6286 | 1.6982 |
| Activity-Markov | 25.16 | 0.4582 | 2.2436 | 49.84 | 0.6221 | 1.6183 |
| Method | Inspect. | Distance (m) | Recovery SR (%) |
| Navigation and persistent memory | |||
| VLFM ( Yokoyama et al., 2024 ) | 4.85 | 18.72 | 22.41 |
| LagMemo GLUE ( Zhou et al., 2025 ) | 3.92 | 13.24 | 45.62 |
| HOV-SG ( Werby et al., 2024 ) | 3.36 | 12.08 | 51.24 |
| DynaMem ( Liu et al., 2024a ) | 2.51 | 11.59 | 63.93 |
| Predictive world and state modeling | |||
| Different VLMs | Last-seen Agent | EvolvingNav | Search SR | ||||
| First-Inspection SR | Search SR | SPL | First-Inspection SR | Search SR | SPL | ||
| GPT-4o | 31.88 | 48.94 | 0.3215 | 59.72 | 71.86 | 0.6027 | +22.92 |
| GPT-5.5 | 42.36 | 51.31 | 0.3732 | 63.48 | 73.57 | 0.6531 | +22.26 |
| GPT-5.6-Luna | 46.92 | 57.43 | 0.3970 | 63.87 | 83.36 | 0.6639 | +25.93 |
| Qwen2.5-VL-3B | 16.26 | 26.55 | 0.1502 | 38.27 | 45.02 | 0.3727 | +18.47 |
| Qwen2.5-VL-32B | 18.84 | 31.30 | 0.1591 | 42.44 | 46.21 | 0.3713 | +14.91 |
| Group | Variant | First-Inspection SR | Search SR | SPL | Inspect. |
| Temporal encoding | w/o elapsed-time encoding | 48.37 | 80.28 | 64.82 | 4.36 |
| w/o calendar context | 47.75 | 76.11 | 66.20 | 4.05 | |
| w/o both temporal cues | 40.84 | 76.03 | 64.98 | 4.12 | |
| Historical evidence | w/o instance-specific history | 57.57 | 80.77 | 66.61 | 4.30 |
| w/o negative history | 45.86 | 79.53 | 63.83 | 4.44 | |
| 39.55 | 63.77 | 56.00 | 4.57 |
| Predictor | Top-1 | MRR | NLL |
| GRU | 32.61 | 0.5471 | 1.8279 |
| Direct Transformer | 34.16 | 0.5586 | 1.7292 |
| Structured Transformer (ours) |
| Method | Static SR | Routine SR | Random SR | Routine gain over Last Seen |
| Last Seen | 84.38 | 38.49 | 23.19 | – |
| Markov Transition | 80.22 | 32.57 | 20.64 | |
| Direct Transformer | 69.79 | 48.82 | 31.39 | |
| EvolvingNav | 82.19 | 60.85 | 30.12 |
| Platform | Morphology | Standing size | Mass | Endurance / range |
| LYNX M20 | Wheel-legged | mm | 35 kg | 3 h / 15 km unloaded |
| X30 | Quadruped | mm | 56 kg | 2.5–4 h / 10 km |
| Lite3 (LiDAR) | Quadruped | mm | 13.5 kg | 1.5–2 h / 2.7 km |
| Indoor | Outdoor | |||||
| Method | First-Inspection | Search SR | Dist. | First-Inspection | Search SR | Dist. |
| Last Seen + Search | 18.8 | 31.3 | 28.6 | 15.6 | 21.9 | 83.8 |
| Time Frequency | 25.0 | 34.4 | 27.2 | 18.8 | 28.1 | 78.6 |
| Retrieval + Reasoning | 31.3 | 40.6 | 25.1 | 21.9 | 34.4 | 75.5 |
| EvolvingNav | 40.6 | 53.1 | 22.4 | 28.1 | 43.8 | 65.2 |
| Method | Unchanged | Routine- consistent | Weak- routine | Broken- routine |
| Last Seen + Search | 68.8 | 18.8 | 12.5 | 6.3 |
| Time Frequency | 50.0 | 50.0 | 18.8 | 6.3 |
| Retrieval + Reasoning | 62.5 | 43.8 | 25.0 | 18.8 |
| EvolvingNav | 62.5 | 68.8 | 37.5 | 25.0 |
| Platform | Method | First-Inspection SR | Search SR | Dist. (m) |
| LYNX M20 | Last Seen + Search | 18.8 | 25.0 | 55.4 |
| EvolvingNav | 37.5 | 50.0 | 44.6 | |
| X30 | Last Seen + Search | 18.8 | 25.0 | 54.1 |
| EvolvingNav | 31.3 | 43.8 | 46.8 | |
| Lite3 | Last Seen + Search | 18.8 | 31.3 | 51.8 |
| EvolvingNav | 25.0 | 43.8 | 45.9 |