LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation
Organizations: AIML, Adelaide University · East China Normal University
Abstract
Language-conditioned goal navigation (LGN) requires embodied agents to locate user-specified targets without step-by-step guidance. However, existing benchmarks largely focus on category-level goals or rely on instance descriptions generated by vision-language models, which often contain ambiguities and semantic errors, limiting systematic and reliable evaluation. We introduce HieraNav, an open-vocabulary LGN task with goals at four hierarchical semantic levels: scene, room, region, and instance. We present Language as a Map (LangMap), the first LGN benchmark to enrich real-world indoor 3D scans with human-verified semantic annotations supporting tasks across all four goal levels. Built on HM3D using a contrastive annotation protocol that compares same-scene regions and instances, LangMap provides region labels and discriminative region and instance descriptions covering 414 object categories and contains over 18K tasks. Each target has concise and detailed descriptions, enabling evaluation across instruction styles. Automated and human evaluations validate our annotation quality: our descriptions improve text-to-view matching accuracy over GOAT-Bench's by 23 points on all shared annotated instances, and an independent human audit yields 92.5% unique-and-correct matches. We also propose PlaNaVid, an RGB-only baseline that combines Bounded Diverse Memory with high-level planning to prime a reactive policy for multi-goal navigation, achieving top-tier success rates without depth, 3D scene representations, or object masks. Further analyses reveal that exploration and hierarchical disambiguation failures become more prominent at finer goal levels, while long-tail categories, small objects, distant targets, timely stopping, and multi-goal completion remain challenging. Benchmark and code: https://bo-miao.github.io/LangMap
Figures & tables
| Eval Benchmark | Goal Granularity | Region Annotation | Object Annotation | Small Obj | Tasks | |||||||
| Scene | Room | Region | Instance | Cat | Desc | Words | Cat | Desc | Words | |||
| RoboTHOR [ 65 ] | ✓ | 12 | - | - | ||||||||
| ObjNav-MP3D [ 21 ] | ✓ | 21 | - | - | ||||||||
| ObjNav-HM3D [ 2 ] | ✓ | 6 | - | - | ||||||||
| HM3D-OVON [ 1 ] | ✓ | 178 | 4.2% | 9000 | ||||||||
| LHPR-VLN [ 6 ] | ✓ † | 10 | - | - | 960 | |||||||
| Method | Observation | Mask | 3D Map | Local Nav. | Multi-Goal | Single-Goal | Latency (s/goal) | |||
|---|---|---|---|---|---|---|---|---|---|---|
| SR | SeqSR@2 | SPL | SR | SPL | ||||||
| 3D-Mem-3B [ 68 ] | RGB-D (pano.) | ✓ | ✓ | Shortest | 13.5 | 1.1 | 6.4 | 8.5 | 1.9 | 147.2 |
| 3D-Mem-7B [ 68 ] | RGB-D (pano.) | ✓ | ✓ | Shortest | 30.1 | 5.1 | 17.3 | 13.7 | 6.2 | 95.6 |
| 3D-Mem-3B † [ 68 ] | RGB-D (pano.) | ✓ | ✓ | Shortest | 20.4 | 2.2 | 11.3 | 15.3 | 2.8 | 147.2 |
| 3D-Mem-7B † [ 68 ] | RGB-D (pano.) | ✓ | ✓ | Shortest | 36.8 | 10.0 | 21.2 | 21.2 | 8.7 | 95.6 |
| MTU3D [ 69 ] | RGB-D (pano.) | ✓ | ✓ | Shortest | 41.2 | 11.0 | 24.3 | 29.9 | 15.3 | 52.1 |
| Config | Method | Scene | Room | Region | Region-D | Instance | Instance-D | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SR | SPL | SR | SPL | SR | SPL | SR | SPL | SR | SPL | SR | SPL | ||
| RGB-D,Mask | 3D-Mem-3B | 11.2 | 2.5 | 11.8 | 2.8 | 8.2 | 1.8 | 6.8 | 1.4 | 4.7 | 1.1 | 5.1 | 1.1 |
| 3D-Mem-7B | 16.4 | 7.2 | 12.6 | 6.0 | 14.0 | 6.1 | 15.1 | 5.6 | 13.0 | 5.8 | 16.3 | 7.0 | |
| MTU3D | 31.4 | 15.7 | 32.5 | 16.2 | 33.6 | 16.8 | 33.7 | 16.2 | 23.8 | 12.9 | 28.7 | 15.2 | |
| RGB | PSL | 6.0 | 1.4 | 6.6 | 1.9 | 7.3 | 2.1 | 9.0 | 2.6 | 6.4 | 1.9 | 8.5 | 2.4 |
| SenseAct-M | 10.2 | 5.6 | 8.3 | 4.4 | 8.5 | 4.3 | 9.5 | 4.8 | 7.7 | 4.0 | 7.4 | 4.1 | |
| Method | Category Frequency | Path Length | Object Size | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Head | Long-tail | Short | Medium | Long | Non-small | Small | ||||||||
| SR | SPL | SR | SPL | SR | SPL | SR | SPL | SR | SPL | SR | SPL | SR | SPL | |
| 3D-Mem-7B | 14.4 | 6.6 | 11.1 | 4.6 | 27.3 | 12.8 | 11.6 | 4.9 | 4.4 | 2.0 | 14.7 | 6.7 | 11.0 | 4.6 |
| MTU3D | 32.1 | 16.3 | 22.0 | 11.5 | 46.3 | 21.7 | 28.0 | 14.8 | 17.4 | 9.7 | 32.8 | 16.7 | 20.5 | 10.0 |
| Uni-NaVid | 31.3 | 15.9 | 26.1 | 13.2 | 42.2 | 21.5 | 29.7 | 14.7 | 19.4 | 10.3 | 32.6 | 16.5 | 23.0 | 10.9 |
| PlaNaVid-3B | 32.4 | 15.8 | 27.4 | 13.0 | 43.8 | 21.7 | 30.5 | 14.4 | 20.7 | 10.4 | 34.0 | 16.4 | 22.4 | 10.5 |
| Components | Results | |||||
|---|---|---|---|---|---|---|
| Mem | GUU | SDR | SR | SeqSR | SPL | #Frames |
| 34.4 | 10.4 | 15.0 | - | |||
| ✓ | 40.4 | 13.2 | 16.9 | 656 | ||
| ✓ | ✓ | 40.9 | 13.2 | 17.3 | 50 | |
| ✓ | ✓ | 42.0 | 14.2 | 17.8 | 629 | |
| ✓ | ✓ | ✓ | 42.8 | 14.6 | 18.1 | 50 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | SeqSR@1 | SeqSR@2 | SeqSR@3 | SeqSR@4 | SeqSR@5 |
|---|---|---|---|---|---|
| PSL | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| SenseAct-M | 6.0 | 1.0 | 0.1 | 0.1 | 0.1 |
| 3D-Mem-3B | 6.9 | 1.1 | 0.3 | 0.0 | 0.0 |
| 3D-Mem-7B | 12.4 | 5.1 | 1.9 | 0.8 | 0.1 |
| MTU3D | 25.0 | 11.0 | 5.4 | 2.5 | 1.1 |
| Uni-NaVid | 27.1 | 10.4 | 3.5 | 1.4 | 0.6 |
| SR | SeqSR@2 | SPL | |||
| Latest Memory | 50 | 10 | 36.4 | 11.4 | 14.9 |
| Full Memory | 656 | 10 | 40.4 | 13.2 | 16.9 |
| Ours | 25 | 10 | 41.5 | 13.8 | 17.3 |
| 50 | 5 | 40.0 | 12.2 | 17.0 | |
| 50 | 10 | 42.8 | 14.6 | 18.1 | |
| 75 | 10 | 42.3 | 13.8 | 17.7 |
| Config | Method | SR | SeqSR@2 | SPL |
|---|---|---|---|---|
| RGB-D,Mask | 3D-Mem-3B | 13.5 [11.8, 15.2] | 1.1 [0.4, 1.9] | 6.4 [4.9, 8.1] |
| 3D-Mem-7B | 30.1 [26.6, 33.7] | 5.1 [3.1, 7.4] | 17.3 [15.0, 19.8] | |
| MTU3D | 41.2 [37.8, 44.7] | 11.0 [8.1, 14.0] | 24.3 [21.9, 26.7] | |
| MTU3D ∗ | 38.0 [34.5, 41.4] | 7.9 [5.4, 10.6] | 21.0 [18.8, 23.1] | |
| RGB | Uni-NaVid | 34.4 [31.0, 37.6] | 10.4 [7.9, 12.9] | 15.0 [13.3, 16.7] |
| PlaNaVid-3B | 41.5 [37.9, 44.9] | 14.4 [11.1, 17.9] | 17.5 [15.5, 19.5] |