The ability to navigate toward sound sources extends a robot's reach beyond its visual field, enabling response to auditory events in unknown environments. To equip robots with this capability, existing methods couple acoustic and visual information through joint audio-visual learning in acoustic simulators. However, acoustic simulation is both low-fidelity and expensive, producing a domain gap that prevents reliable real-world deployment, while the discrete action spaces inherited from grid-based simulators introduce an additional kinematic gap on physical robots. To alleviate these issues, we propose BCNav, a decoupled framework that separates the acoustic module from the learned navigation policy using direction-of-arrival (DOA) estimation: an estimator provides a scalar bearing to the sound source, so the navigation policy only processes depth images and a bearing angle, two inputs whose domain gaps are well characterized. We collect shortest-path demonstrations with calibrated bearing noise injection and train the policy via imitation learning to output continuous velocity commands directly executable on ground robots. We demonstrate the method in simulation and on a physical robot, navigating unknown environments without any acoustic fine-tuning, prior mapping, or real-world audio data collection. Code is available at https://github.com/york1to/bcnav.
Figures & tables
Fig. 1: System overview. A classical DOA estimator provides a bearing to the sound source. The learned policy maps depth and bearing to continuous velocity commands.
Fig. 2: Policy architecture. Each timestep’s depth image produces 16 spatial tokens via a DeFM ResNet-18. The bearing is encoded as a single token via sinusoidal embedding. Bidirectional cross-attention fuses the two modalities. A temporal transformer aggregates four frames. A bearing shortcut concatenated with the temporal output feeds a regression MLP that predicts 8 velocity pairs.
Fig. 3: Training data in one HM3D scene. (a) Expert demonstrations follow shortest paths between diverse start–goal pairs. (b) DAgger rollouts: the policy (solid) deviates from the expert path, and expert corrections (blue dashed) are computed from the visited states back to the goal, providing on-policy training labels.
Condition
MAE [ \SIUnitSymbolDegree ]
MedAE [ \SIUnitSymbolDegree ]
<15\text{,}\mathrm{\SIUnitSymbolDegree}$$ [%]
LOS (60%)
11.4
5.8
77.7
NLOS (40%)
37.3
26.4
36.6
Overall
21.7
9.0
61.4
TABLE I: SRP-PHAT DOA accuracy in simulation [ 7 ] . MedAE: median absolute error.
Config.
σ
SR
SPL
Closed-loop
9\SIUnitSymbolDegree41\SIUnitSymbolDegree
71.1
0.709
Oracle
15\SIUnitSymbolDegree
96.5
0.964
w/o reflex
15\SIUnitSymbolDegree
64.2
0.641
w/o velocity
15\SIUnitSymbolDegree
21.3
0.231
P-controller
0\SIUnitSymbolDegree
11.8
0.118
TABLE II: Navigation results: (a) simulation ablation on HM3D val and (b) real-world trials (10 each).
Fig. 5: DOA error vs. source distance during real-world navigation. Dashed lines mark the ±15\text{,}\mathrm{\SIUnitSymbolDegree}$$ robust operating range from Sec. IV-B .
Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents. In particular, it is still unclear how to effectively evaluate and model spatial audio understanding in embodied settings. To address this gap, we introduce \textbf{OmniEchoBench}, a unified benchmark for spatial audio-visual perception and audio-vision-language navigation. OmniEchoBench comprises six tasks over 197 real-world spatial audio-visual scenes, 2,972 question-answer pairs, and 900 navigation samples with first-order ambisonics (FOA) audio collected from 30 real-world environments. To enable scalable training supervision, we develop a controllable rendering pipeline for spatial audio. It preserves geometric consistency among sound sources, visual observations, and agent trajectories. Building on this, we propose \textbf{OmniEcho}, a spatially aware omni-modal model. It introduces an FOA spatial encoder alongside a pretrained semantic audio pathway. Extensive experiments show that OmniEcho achieves state-of-the-art performance on spatial audio-visual perception. For our sound-guided navigation, OmniEcho reaches a performance level close to that of traditional vision-language navigation. These results demonstrate that spatial audio can serve as a valuable signal for embodied scene reasoning and navigation, while also highlighting fine-grained spatial localization and distance estimation as important open challenges. Our code and data will be available in https://github.com/PKU-VaLuE-Lab/OmniEcho/tree/main
Ruixun Liu, Yuxuan Wang, Jiacheng Xie +10
School of Intelligence Science and Technology, Peking University · State Key Laboratory of General Artificial Intelligence, Peking University · Alibaba Token Hub, Alibaba Group +1
Object navigation requires a robot to search for an unobserved target in an unknown environment by deciding where to explore next under partial observability. Effective search resembles human-like exploration: selectively probing visually promising frontiers while relying on spatial memory to avoid redundant revisits. We propose IntentNav, a spatial-visual imitation framework that learns human-like ObjectNav policies from human demonstrations. To infer high-level search intent from low-level human actions, we introduce Frontier-based Human-Intent Labeling, which looks ahead in human demonstrations and labels the frontier that best explains the demonstrator's future search direction. We construct a spatial-visual candidate space, where BEV memory tracks explored regions, unexplored frontiers, and trajectory history, while egocentric visual memory provides semantic cues for each candidate. A VLM policy is trained to select among these grounded candidates, using Intent-Aligned Objective to encourage consistent and human-like exploration. IntentNav achieves state-of-the-art performance on the MP3D, HM3D-v1 and HM3D-v2 ObjectNav benchmarks. The proposed candidate-level navigation interface transfers zero-shot to wheeled, quadruped, and humanoid robots without further VLM fine-tuning. Project page.
Yuxin Cai, Zongtai Li, Maonan Wang +9
Nanyang Technological University · Carnegie Mellon University · The Chinese University of Hong Kong +1
Training Deep Reinforcement Learning (DRL) navigation policies for different robot configurations remains time-consuming. We present FlashNav, a GPU-based framework that trains robot-specific navigation policies within tens of seconds. A unified robot specification configures a lightweight simulator for batched motion updates, range sensing, and footprint collision checking over a shared occupancy map. The framework supports nonconvex footprints, different sensor configurations and drive types. Blockwise ray queries and selective observation recomputation after resets reduce simulation overhead, while GPU-resident replay and overlapping experience collection and learner updates support efficient off-policy training. Experiments covered five robot configurations and three computing platforms. With FastDSAC on a single RTX 5090 GPU, FlashNav can train a deployable navigation policy in under 30 seconds. FlashNav achieved the highest success rate and score in the benchmark comparison. The selected policies were deployed on wheeled, quadrupedal, humanoid, and irregularly shaped robots without additional policy training.
Shanze Wang, Yiwei Qian, Xinming Zhang +8
Eastern Institute of Technology, Ningbo · The Hong Kong Polytechnic University · National University of Singapore +2