Language-guided navigation requires connecting partial observations to a persistent spatial reference and learning how actions change that representation. We introduce InsightMap, a framework that uses top-down maps as both explicit spatial memory and action-conditioned prediction targets. Historical views are linked to labeled map locations, and a shared multimodal backbone jointly learns navigation action prediction and post-action map generation. Map prediction provides auxiliary training supervision, while navigation inference decodes actions from the observed spatial context. An aligned RGB-D data pipeline supports a common interface for navigation, visual question answering, situated reasoning, and 3D grounding. On the validation-unseen splits of R2R-CE and RxR-CE, InsightMap achieves success rates (SR) of 56.9% and 54.9%, respectively. Adding map-prediction supervision improves R2R-CE SR by 4.3 and success weighted by path length (SPL) by 3.2 percentage points. On static spatial tasks, InsightMap achieves 103.7 CIDEr on ScanQA, 60.1% exact-match accuracy on SQA3D, and 53.1% grounding accuracy at 0.5 IoU on ScanRefer with detected object proposals. On Unitree Go2, it outperforms NaVid and NaVILA in hallway, lab, and office environments.
Figures & tables
Fig. 1: Overview of InsightMap across vision-and-language navigation, visual question answering, and 3D visual grounding. An observed top-down map provides spatial context for all tasks. Navigation training pairs actions with post-action map targets. Question answering produces answer text, and grounding predicts a 3D box and a target-annotated map.
Fig. 2: Architecture of InsightMap. Labeled historical views, the current observation and map, and the instruction condition a shared understanding and generation backbone. The training targets are action text followed by an action-conditioned post-action map.
(a) R2R-CE
Method
NE ↓
OS ↑
SR ↑
SPL ↑
HPN+DN ∗ [ 29 ]
6.31
40.0
36.0
34.0
CMA ∗ [ 30 ]
6.20
52.0
41.0
36.0
Sim2Sim ∗ [ 31 ]
6.07
52.0
43.0
36.0
GridMM ∗ [ 2 ]
5.11
61.0
49.0
41.0
ScaleVLN ∗ [ 24 ]
4.80
–
55.0
51.0
TABLE I: Navigation results on R2R-CE and RxR-CE val-unseen.
Fig. 4: Recorded navigation trajectory of InsightMap. The eight displayed steps pair top-down maps and current RGB views with history strips (past to current) and recorded actions. Green circles mark the agent, blue points indicate observation history, and yellow lines trace the path. The sequence ends with STOP .
Variant
Gen. module
Map loss
SR ↑
SPL ↑
A1
Frozen
Off
50.3
45.1
A2
Trainable
Off
52.6
47.5
A3 (Full)
Trainable
On
56.9
50.7
TABLE II: Ablation of generation-module trainability and map supervision on R2R-CE val-unseen.
Variant
Map input
Location IDs
SR ↑
SPL ↑
B1
Blank
No
52.6
47.5
B2
Unnumbered map
No
54.1
49.9
B3 (Full)
Full map
Yes
56.9
50.7
TABLE III: Ablation of map input and location IDs on R2R-CE val-unseen.
Variant
Training tasks
DAgger
R2R-CE
VLN
VG
VQA
NE ↓
OS ↑
SR ↑
SPL ↑
C0
✓
–
–
–
6.02
49.8
46.1
41.7
C1
✓
–
–
✓
5.83
53.8
49.2
44.6
C2
✓
✓
–
✓
5.72
58.2
51.7
47.0
C3
✓
–
✓
✓
5.09
61.2
55.1
49.3
C4
✓
✓
✓
✓
4.89
64.7
56.9
50.7
TABLE IV: Task-mixture ablations on R2R-CE val-unseen.
Method
ScanQA
SQA3D
B-4 ↑
M ↑
R-L ↑
C ↑
EM ↑
ScanQA [ 14 ]
10.08
13.14
33.33
64.86
–
3D-LLM [ 17 ]
12.0
14.5
35.7
69.4
–
LL3DA [ 18 ]
13.53
15.88
37.31
76.79
–
Chat-Scene [ 19 ]
14.31
18.00
41.56
87.70
54.57
3D-LLaVA [ 20 ]
17.1
18.4
43.1
92.6
54.5
TABLE V: Question answering on ScanQA validation and SQA3D test.
Method
Acc@0.25 ↑
Acc@0.5 ↑
ScanRefer [ 16 ]
41.19
27.40
Chat-Scene [ 19 ]
55.52
50.23
3D-LLaVA [ 20 ]
51.2
40.6
Video-3D LLM [ 21 ]
57.87
51.18
GPT4Scene [ 22 ]
62.6
57.0
VG-LLM [ 35 ]
57.6 (41.6)
50.9 (14.9)
TABLE VI: ScanRefer validation results. Parentheses denote raw box predictions before proposal refinement.
Fig. 6: Real-world office navigation with InsightMap on Unitree Go2. Third-person frames are ordered left to right, top to bottom.
Does progress on spatial reasoning benchmarks translate into better navigation? Existing benchmarks test isolated inferences from images or videos, with little connection to downstream navigation. Our analysis reveals a gap between benchmark-oriented spatial specialization and navigation performance, and shows how aligning spatial supervision with navigation goals, phases, and decision learning improves navigation. Guided by these findings, we build \textsc{Spatial-Nav-100K} and fine-tune in two stages, \textit{i.e.} first learning a shared spatial-navigation foundation, and then specializing each phase with the abilities it relies on. We further introduce Spatial-NPD, where a teacher conditioned on spatial priors produces grounded action preferences and distills them into a student policy, so no explicit spatial reasoning is needed at inference. With 45 A100 GPU-hours of policy training, our 8B model reaches SR/SPL of 77.4/35.4 on HM3D-v0.2, 60.2/30.5 on HM3D-v0.1, and 47.9/20.6 on train-unseen MP3D. It outperforms several systems that rely on closed-source models or thousands of GPU-hours of training, at 148 ms per action step. All code and datasets will be publicly available at https://github.com/ylwhxht/Spatial-Nav.
Xun Huang, Shijia Zhao, Rongsheng Qu +5
Xiamen University · Zhongguancun Academy · Beihang University +2
Navigating complex, densely packed environments like retail stores, warehouses, and hospitals poses a significant spatial grounding challenge for humans and embodied AI. In these spaces, dense visual features quickly become stale given the quasi-static nature of items, and long-tail semantic distributions challenge traditional computer vision. While Vision-Language Models (VLMs) help assistive systems navigate semantically-rich spaces, they still struggle with spatial grounding in cluttered environments. We present GIST (Grounded Intelligent Semantic Topology), a multimodal knowledge extraction pipeline that transforms a consumer-grade mobile point cloud into a semantically annotated navigation topology. Our architecture distills the scene into a 2D occupancy map, extracts its topological layout, and overlays a lightweight semantic layer via intelligent keyframe and semantic selection. We demonstrate the versatility of this structured spatial knowledge through critical downstream Human-AI interaction tasks: (1) an intent-driven Semantic Search engine that actively infers categorical alternatives and zones when exact matches fail; (2) a one-shot Semantic Localizer achieving a 1.04 m top-5 mean translation error; (3) a Zone Classification module that segments the walkable floor plan into high-level semantic regions; and (4) a Visually-Grounded Instruction Generator that synthesizes optimal paths into egocentric, landmark-rich natural language routing. In multi-criteria LLM evaluations, GIST outperforms sequence-based instruction generation baselines. Finally, an in-situ formative evaluation (N=5) yields an 80% navigation success rate relying solely on verbal cues, validating the system's capacity for universal design.
Shivendra Agrawal, Bradley Hayes
University of Colorado Boulder · Boulder, Colorado, USA
Spatial intelligence is a key frontier for multimodal large language models (MLLMs), enabling them to reason about the physical world from visual experience. Inspired by human spatial cognition, recent approaches construct grid-based cognitive maps from multi-frame visual inputs to maintain coherent spatial representations over time. However, limited context lengths still challenge spatial understanding, while existing methods, such as long-context modeling and external memory, often require architectural changes, memory modules, or finetuning, limiting their applicability to off-the-shelf pretrained MLLMs. This motivates a lightweight, model-agnostic method for preserving spatial information beyond the native context window. To this end, we propose a plug-and-play multi-agent framework that collaboratively constructs cognitive maps as structured spatial memory, enhancing the spatial understanding of arbitrary pretrained MLLMs without architectural modification or additional training. Our framework features local-global agent coordination, cognitive map construction with atomic commits, and cross-agent verification. Extensive experiments demonstrate that our method achieves superior performance on spatial understanding tasks while remaining fully training-free. Code will be released.
Yiming Zhang, Ruoxuan Cao, Zhihang Zhong
2Cornell University · 1Shanghai Jiao Tong University