Organizations: School of Mechanical and Aerospace Engineering, Nanyang Technological University, 639798, Singapore · Schaeffler Hub for Advance REsearch(SHARE) at NTU, Singapore
Object navigation requires an embodied agent to find an object in an unseen environment under partial observability and a limited motion budget. Existing methods primarily optimize where the robot should go next by ranking candidate destinations. In contrast to these methods, we present SOR-Nav, a hierarchical navigation system that explicitly arbitrates between continuing to explore the current context and abandoning it for a more promising reachable region. First, an autonomous semantic exploration system is built that accumulates persistent 3D object clusters and organizes reachable frontiers into a cluster decision graph to provide an efficient search abstraction. Then, SOR-Nav uses a context-gated LLM-driven object-search supervisor to evaluate the suitability of the current search context and decide whether to continue exploration or perform cross-region relocation to another reachable frontier cluster. Across the complete, unfiltered validation sets of HM3D-v1, HM3D-v2, and MP3D, SOR-Nav achieves the strongest reported Success Rate (SR) and Success weighted by Path Length (SPL) on all three benchmarks. On MP3D in particular, it more than doubles the previous best SPL from 18.1% to 38.5% while increasing SR from 50.7% to 61.8%. Nested HM3D-v2 ablations validate the proposed decision structure, while a continuous three-target physical deployment demonstrates persistent ObjectNav operation in real-world scenarios.
Figures & tables
Fig. 1: Overview of the SOR-Nav object-search workflow. RGB and depth observations update persistent 3D object clusters and a cluster decision graph in the autonomous semantic exploration system. The LLM-driven supervisor prioritizes confirmed target evidence; otherwise, its context gate determines whether to continue exploration or relocate to a graph-reachable frontier cluster. All outcomes share the same mission executor, and subsequent observations and mission results close the loop.
Method
ZS
HM3D-v1
HM3D-v2
MP3D
SR
SPL
SR
SPL
SR
SPL
L3MVN [ 2 ]
Yes
50.4
23.1
36.3
15.7
34.9
14.5
ESC [ 1 ]
Yes
39.2
22.3
–
–
28.7
11.2
VoroNav [ 6 ]
Yes
42.0
26.0
–
–
–
–
VLFM [ 3 ]
Yes
52.5
30.4
63.6
32.5
36.4
17.5
SG-Nav [ 5 ]
Yes
54.0
24.9
49.6
25.5
40.2
16.0
TABLE I: ObjectNav comparison based on the SysNav simulation table and SOR-Nav evaluation results. SR and SPL are reported in percentages.
Variant
N
SR ↑
SPL ↑
SL ↓
TO ↓
SOR-Nav (full)
1000
87.6
52.7
5.0
7.4
w/o relocation gate
1000
82.2
39.9
9.7
8.1
w/o relocation gate and TSP
1000
76.4
29.2
16.3
7.3
TABLE II: Nested ablation study on HM3D-v2. SR, SPL, step-limit (SL), and timeout (TO) rates are percentages; N is the number of evaluated episodes.
Fig. 4: Paired qualitative results for 18 successful episodes across HM3D-v2, HM3D-v1, and MP3D, with six target categories shown per dataset. Within each episode panel, the goal-facing terminal RGB observation is shown directly above its corresponding explored occupancy map and executed trajectory. The yellow overlay marks the detected target category; the title reports the target and true dataset episode index. Portrait RViz maps with width-to-height ratio below 0.95 are rotated 90∘ clockwise for compact display.
Fig. 5: Representative events from a continuous real-world ObjectNav run in which the next request was issued after each successful target confirmation. Each labeled pair shows the first-person observation (top) and accumulated map, graph, and trajectory (bottom). The sequence highlights the relocation trigger, passage through the office doorway, and terminal observations for television, trash can, and signboard; all map panels use the final viewport. The reported distance is cumulative robot travel from the beginning of the sequence, rather than geodesic distance to a target.
Object navigation requires an agent to locate a target in an unknown environment through visual observations. Existing methods typically rely on open-vocabulary detectors or vision-language models (VLMs) to answer where to search, but often overlook what not to trust - which semantic cues are unreliable. Open-vocabulary perception is prone to systematic misleading evidence: false positives, outdated static priors, and repeated failed exploration due to lack of embodied verification, which contaminates mapping and decision-making. Such errors are rooted in structured object relations in real-world scenes. To address this, we propose DB-Nav, a framework that reshapes the search space via dual relational biases. It factorizes target-centric relations into an Activation Bias (propagates contextual evidence) and an Inhibition Bias (suppresses unreliable regions via perceptual confusion and action-level falsification). These biases are unified into a Relational Activation-Inhibition Exploration Graph that modulates frontier exploration values using online observations and failed accesses. Experiments on ObjectNav benchmarks show that DB-Nav significantly outperforms existing methods in success rate (SR) and Success weighted by Path Length (SPL), offering a lightweight, interpretable, and robust navigation framework without costly online VLM reasoning.
Weitao An, Chenghao Xu, Xu Yang +1
School of Electronic Engineering, Xidian University, Xi’an 710071, China · School of Information Science and Engineering, Hohai University, Nanjing 210098, China
Object navigation requires a robot to search for an unobserved target in an unknown environment by deciding where to explore next under partial observability. Effective search resembles human-like exploration: selectively probing visually promising frontiers while relying on spatial memory to avoid redundant revisits. We propose IntentNav, a spatial-visual imitation framework that learns human-like ObjectNav policies from human demonstrations. To infer high-level search intent from low-level human actions, we introduce Frontier-based Human-Intent Labeling, which looks ahead in human demonstrations and labels the frontier that best explains the demonstrator's future search direction. We construct a spatial-visual candidate space, where BEV memory tracks explored regions, unexplored frontiers, and trajectory history, while egocentric visual memory provides semantic cues for each candidate. A VLM policy is trained to select among these grounded candidates, using Intent-Aligned Objective to encourage consistent and human-like exploration. IntentNav achieves state-of-the-art performance on the MP3D, HM3D-v1 and HM3D-v2 ObjectNav benchmarks. The proposed candidate-level navigation interface transfers zero-shot to wheeled, quadruped, and humanoid robots without further VLM fine-tuning. \href{https://anonymous.4open.science/w/IntentNav/}{Project page}.
Yuxin Cai, Zongtai Li, Maonan Wang +9
Nanyang Technological University · Carnegie Mellon University · The Chinese University of Hong Kong +1
To locate a target object while exploring the unknown environment is a fundamental capability for autonomous agents, with applications ranging from search-and-rescue to field robots. A simplified version of such task is Object Goal Navigation (ObjNav). In ObjNav, successful arrival at the target object provides a basic measure of performance; however, the efficiency of the navigation trajectory is equally important, as it indicates how intelligently the agent explores and how much time remains for subsequent tasks. In unknown environments, the key to efficient navigation lies in deciding where to explore next. While many prior works aim to address this core challenge and achieved promising performance in certain settings, recent training-based models and non-training frameworks still suffer from generalization and efficiency issues respectively, which in the worst cases can lead to excessive exploration of already-visited areas or redundant back-and-forth motion. We evaluate EffiNav on two widely used simulation benchmarks Habitat Matterport 3D (HM3D) and Open-Vocabulary Object goal Navigation (OVON), and further validate its effectiveness on physical robots in real-world settings. We conduct failure analysis on massive simulation episodes. With minimal modification, we also extend EffiNav to a memory-augmented ObjNav task on the GOAT-BENCH dataset, demonstrating its adaptability beyond standard ObjNav settings. Across two standard metrics--Success Rate (SR) and Success weighted by Path Length (SPL), EffiNav matches or outperforms recent baselines, reflecting its efficiency, robustness, and practical applicability. Recognizing the different emphases of the two datasets, the performances reveals this framework is more balanced and generalizable for efficient ObjNav.
Zecheng Yin, Benedict Jun Ma
Systems Hub of Intelligence Transportation HKUST(GZ) Guangzhou, Guangdong