When searching for an object, people choose their next move by considering both likely target locations and places already explored. The current view can cue place-associated memories, bringing target relevance and prior exploration into the same spatial context. In many zero-shot object navigation (ZSON) methods, however, vision-language models (VLMs) infer promising search areas from egocentric images, while exploration history is represented separately, e.g., as text or maps. This separation either requires an additional fusion step or leaves the correspondence between memory and route choices implicit for the VLM to recover. We instead make exploration memory directly visible on visual route choices. We propose MarvisNav, a ZSON framework that maintains a topological graph and projects candidate nodes together with their exploration states onto egocentric views as memory-bearing visual route choices. These states capture local exploration progress beyond binary visitation. By binding exploration state directly to each visual candidate, MarvisNav enables the VLM to jointly evaluate target relevance and exploration state without a separate post-hoc fusion or reranking stage. Without policy training, MarvisNav achieves state-of-the-art performance on HM3D (81.2% SR and 42.5% SPL), while remaining competitive on MP3D. It also outperforms representative VLM-based methods with far fewer VLM calls (e.g., 7.5% of WMNav). Real-robot experiments across diverse scenes further validate its practical deployability. Beyond MarvisNav, our study shows that memory representation shapes VLM decisions and ZSON performance, highlighting that effective memory use depends not only on its availability, but also on how it is represented. Code and project page will be available at https://wangjincheng1998.github.io/MarvisNav/.
Figures & tables
Figure 1: Top: Object search requires a VLM to consider both where the target might be found and which areas have already been searched. Left: With separate representations, memory either modifies the VLM’s preference externally or leaves the VLM to infer its correspondence with the choices implicitly. Right: We instead ground topological exploration memory onto egocentric route choices, enabling joint reasoning over target relevance and exploration state. Junctions such as D are also valid choices. The selected choice directly specifies a navigation goal.
Figure 2: Motivating comparison of memory representations. Gemini-3.5-Flash and Qwen3.5-Plus score each candidate from 0-10 (higher is more worth exploring), with mean and standard deviation over 10 runs. Full prompts and representation-specific changes are detailed in Appendix D .
Figure 3: Overview of MarvisNav. (1) Given RGB-D observations and the agent pose, MarvisNav constructs a state-aware RVG by skeletonizing navigable space and assigning each node a local exploration state. (2) At each decision site, the local RVG topology determines the number and orientation of egocentric views. Visible paths and candidate nodes are projected into these views, where each candidate is annotated with a letter label and a color-coded exploration state. (3) Given the target and annotated views, a VLM jointly evaluates target relevance and exploration state. Its selected letter directly determines the next navigation goal.
HM3D v0.1
HM3D v0.2
MP3D
Reasoning
Method
ZS
SR ↑
SPL ↑
SR ↑
SPL ↑
SR ↑
SPL ↑
Policy
PONI ( Ramakrishnan et al., 2022 )
×
–
–
–
–
31.8
12.1
Uni-NaVid ( Zhang et al., 2025a )
×
–
–
73.7
37.1
–
–
Embedding similarity-based guidance
ZSON ( Majumdar et al., 2022 )
√
25.5
12.6
–
–
15.3
4.8
CoW ( Gadre et al., 2023 )
√
–
–
–
–
7.4
3.7
VLFM ( Yokoyama et al., 2024 )
√
52.5
30.4
63.6
32.5
36.4
17.5
Table 1: Comparison with ObjectNav baselines on HM3D v0.1, HM3D v0.2, and MP3D. ZS denotes zero-shot ObjectNav, without task-specific fine-tuning. Results are taken from the original papers when available; otherwise, we use reproductions reported in ( Zhang et al., 2025c ) : HM3D-v0.2(L3MVN, VLFM, and SG-Nav), and MP3D (L3MVN and OpenFMNav).
Representation
To VLM
SR (%) ↑
SPL (%) ↑
No Memory
–
71.9
35.4
Post-filtering
×
74.3
37.1
Pre-filtering
×
75.0
37.2
Textual History
√
73.6
36.6
Visited Markers
√
76.5
40.6
Choice-aligned
√
81.2
42.5
Table 2: Ablation of memory representations.
Table 6
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7
Figure 5: Real-world deployment of MarvisNav on a Unitree Go2. The robot searches for a potted plant in the reception . The point-cloud overview shows the executed trajectory and decision sites O1–O9. The accompanying panels present the annotated egocentric views, selected route identifiers, and original VLM responses. Route identifiers are local to each decision.
Figure 6: State-Aware RVG and executed trajectory in the representative indoor deployment. Starting from an empty map, MarvisNav incrementally constructs the RVG and updates node exploration states from online RGB-D observations. The final map shows the resulting topology, exploration states, and robot trajectory from start to end.
Figure 7: Scene reconstruction used to create method illustrations. The reconstruction follows the layout and relative spatial proportions of HM3D scene 00877-4ok3usBNeis .
Figure 8: Prompt templates for the motivation study. All variants share the same task and scoring instructions and differ only in how exploration memory is provided.
Figure 9: Shared navigation decision prompt
Figure 10: Failure analysis across HM3D v0.1, HM3D v0.2, and MP3D. Evaluation episodes are decomposed into successes and three failure outcomes: false-positive target detection, maximum step budget reached, and robot stuck. Numbers denote episode counts.
Issue Type
Scene ID
Episode IDs
Number of Episodes
No valid target in reachable region
7MXmsvcQjpJ
84, 89, 96, 109
4
Annotation error / omission
Dd4bFSTQ8gi
196, 197, 222
3
bCPU9suPUw9
508, 511, 528
3
jxsvVRusffK
538, 539, 558
3
p53SfW6mjZe
732
1
a8BtkwhxdRV
492, 499
2
Appendix
Table 4: Benchmark issues identified in the HM3D v0.2 evaluation split. We report the 52 identified episodes grouped by issue type and scene.
Figure 11: Representative benchmark-related failure cases from HM3D v0.2. Each row shows a representative problematic episode, including the egocentric observation, the corresponding depth/geometry view, and a top-down scene map. In the third column, red regions indicate the benchmark ground-truth object region. (a) Incomplete scene reconstruction ( k1cupFYWXJ6 , ep 649, target: sofa ): missing or corrupted scene geometry around the target region makes the episode unreliable for evaluation. (b) Annotation error / omission ( Dd4bFSTQ8gi , ep 197, target: chair ): the agent successfully reaches a visible chair, but the benchmark ground-truth annotation is incomplete or inconsistent, so the episode is incorrectly counted as a false positive. (c) No valid target in the reachable region ( 7MXmsvcQjpJ , ep 96, target: sofa ): the episode asks the agent to search for a sofa in a basement area where no valid target instance exists. These examples illustrate three representative benchmark defects that can cause apparent navigation failures unrelated to policy quality.
Building memory is essential for long-horizon planning in zero-shot embodied navigation. Detector-centric scene graphs often compress observations into sparse nodes, discarding fine-grained visual evidence and accumulating noise, while 3D reconstruction-based methods remain computationally prohibitive. We present EvoMemNav, an efficient, self-evolving, fine-grained memory framework for zero-shot embodied navigation. EvoMemNav constructs a Visual-Semantic Memory Graph (VSMGraph) that keeps raw views as first-class memory and organizes them with lightweight semantic cues and topological relations into a room-view-object hierarchy, preserving fine-grained details for disambiguation and Stop verification. To scale to growing memory, we introduce a budgeted coarse-to-fine policy: a coarse stage compresses the search space into promising regions, and a fine stage invokes a VLM only for targeted verification and decision. Beyond static memories, EvoMemNav performs reflection-driven write-back after each subtask, updating graph-attached priors that encode accumulated environmental knowledge to refine future decisions without retraining. Experiments on GOAT-Bench and HM3D across object, text-description, and image-goal modalities show consistent gains in SR/SPL, with better multi-instance disambiguation, fewer premature stops, and stronger zero-shot generalization.
Zuhao Ge, Xiaosong Jia, Chao Wu +3
Institute of Trustworthy Embodied AI (TEAI), Fudan University · 2Shanghai Key Laboratory of Multimodal Embodied AI
Object navigation requires a robot to search for an unobserved target in an unknown environment by deciding where to explore next under partial observability. Effective search resembles human-like exploration: selectively probing visually promising frontiers while relying on spatial memory to avoid redundant revisits. We propose IntentNav, a spatial-visual imitation framework that learns human-like ObjectNav policies from human demonstrations. To infer high-level search intent from low-level human actions, we introduce Frontier-based Human-Intent Labeling, which looks ahead in human demonstrations and labels the frontier that best explains the demonstrator's future search direction. We construct a spatial-visual candidate space, where BEV memory tracks explored regions, unexplored frontiers, and trajectory history, while egocentric visual memory provides semantic cues for each candidate. A VLM policy is trained to select among these grounded candidates, using Intent-Aligned Objective to encourage consistent and human-like exploration. IntentNav achieves state-of-the-art performance on the MP3D, HM3D-v1 and HM3D-v2 ObjectNav benchmarks. The proposed candidate-level navigation interface transfers zero-shot to wheeled, quadruped, and humanoid robots without further VLM fine-tuning. Project page.
Yuxin Cai, Zongtai Li, Maonan Wang +9
Nanyang Technological University · Carnegie Mellon University · The Chinese University of Hong Kong +1
Zero-shot object-goal navigation (ZSON) requires a robot to find a named object category in a building it has never entered. The prevailing approach scores frontiers with a vision-language value map: every decision is another argmax over the map as it currently stands, and the evidence behind that score is discarded the moment it is taken. Systems that place a large vision-language model inside the perception-action loop typically query it on a fixed schedule from the current view alone; a room the robot walked through minutes earlier is never reconsidered, and a failed call has no defined fallback. We turn what the robot has already seen into the object of deliberation. Our hierarchical fast-slow agent leaves the value-map controller running at every step and writes a coordinate-anchored memory as it moves: a semantic grid of room types and confirmed object instances, together with a bounded store of pose-tagged keyframes. A VLM screens each candidate detection before it is written. A deliberative layer reads this memory in a bounded reason-retrieve-act loop. It wakes on structural events the reactive layer computes, reasons first over text, and recalls a first-person view only for candidates that text alone cannot separate. Per-invocation and per-run caps bound its calls, a call-free first tier resolves the most frequent stall, and any failure returns control to the reactive controller. Our system reaches 68.75% SR on HM3D v1 val and 47.29% on MP3D val, the highest success rate among the zero-shot methods compared here. Choosing among far frontiers by argmax instead of deliberating costs 3.40 SR points in a paired comparison over all 2000 HM3D episodes (95% CI [1.70, 5.05]); deliberating over every frontier does not recover them.
Zhaochen Lan, Zhi Yang, Yuxiang Fu +1
School of Mechanical Engineering and Automation, Beihang University, Beijing, China · School of Automation, Beijing Institute of Technology, Beijing, China