SuperNav: An Agentic Navigation System for Any Task in Any Scene
Organizations: Zhejiang University · Shenzhen University · Causa Robotics
Abstract
General-purpose service robots need navigation systems that can handle diverse human requests in unfamiliar environments, combining task generality with scene generality. Some existing methods fine-tune multimodal large language models (MLLMs) to predict navigation actions, making their behavior dependent on the coverage of navigation training data and potentially limiting generalization to new requests and environments. Our key insight is to let the MLLM focus on interpreting requests, understanding scenes, and making decisions while preserving its general-purpose capabilities and delegating motion execution to navigation tools. To realize this idea, we introduce SuperNav, which equips a pretrained MLLM with a specialized agent harness without navigation-specific fine-tuning of the MLLM. Our harness supports these decisions with Navigation Skills, agent-oriented Tools for physical interaction, and task-progress and context management. A unified visual-point interface connects decision-making to motion by allowing the model to specify destinations directly in images and revise its decisions from execution feedback. Together, these components support sustained navigation across different task requirements and environments. SuperNav outperforms four evaluated baselines on instance-level, multi-object, and demand-driven tasks. Category-level evaluation on HM3D and deployment on a real quadruped robot further demonstrate its applicability across environments. Project Page: https://zju3dv.github.io/SuperNav/
Figures & tables
| Our instance-navigation benchmark | Demand-driven benchmark | |||||
| Single-object (150) | Multi-object (150) | Demand-driven (200) | ||||
| Method | SR | SPL | SR | SPL | SR | SPL |
| NaVid | 24.67 | 0.1599 | 2.67 | 0.0216 | 17.00 | 0.0920 |
| UniNaVid | 34.00 | 0.1833 | 1.33 | 0.0131 | 25.00 | 0.0943 |
| StreamVLN | 13.33 | 0.1036 | 0.00 | 0.0000 | 25.50 | 0.0775 |
| OmniNav (Action Former) | 27.33 | 0.2133 | 4.00 | 0.0275 | 37.50 | 0.1213 |
| HM3D-OVON val-unseen | HM3Dv2 | |||||||
| 0.25 m | 1 m | 0.2 m | 1 m | |||||
| Method | SR | SPL | SR | SPL | SR | SPL | SR | SPL |
| MTU3D ( Zhu et al., 2025 ) | 40.80 | 0.1210 | — | — | — | — | — | — |
| SoftNav ( Wu et al., 2026 ) | 66.70 | 0.2570 | — | — | — | — | — | — |
| AstraNav-Memory ( Ren et al., 2025 ) | — | — | 62.50 | 0.3490 | — | — | — | — |
| OmniNav (slow + CoT) ( Xue et al., 2026 ) | — | — | 59.20 | 0.3320 | — | — | — | — |
| 0.25 m | 1 m | |||
| Variant | SR | SPL | SR | SPL |
| w/o Navigation Skills | 35.00 | 0.2460 | 54.17 | 0.3894 |
| w/o Agent-oriented Tool Design | 52.50 | 0.2230 | 65.83 | 0.2855 |
| w/o Skills and Tool Design | 31.67 | 0.1870 | 54.17 | 0.3417 |
| Grounding Motion Tool | 55.83 | 0.3201 | 62.50 | 0.3624 |
| SuperNav (GPT-6 Astra medium) | 71.67 | 0.4579 | 79.17 | 0.4974 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Variant | Entry Skill | Tool-usage Skill |
| Geo-based Executor | global-navigation-geo-based-executor | localnav-pointnav-geo-based-executor |
| Learned Executor | global-navigation-learned-executor | localnav-pointnav |
| Multi-goal | multi-global-navigation-learned-executor | localnav-pointnav |
| Multi-object navigation | ||||||||
|---|---|---|---|---|---|---|---|---|
| Method | SR | SR | SR | SR | SR | |||
| Eligible tasks | 150 | 150 | 112 | 75 | 38 | |||
| NaVid | 40.00 | 6.67 | 2.68 | 1.33 | 0.00 | |||
| UniNaVid | 36.67 | 4.00 | 1.79 | 1.33 | 0.00 | |||
| StreamVLN | 22.00 | 3.33 | 0.89 | 0.00 | 0.00 | |||
| OmniNav (Action Former) | 44.00 | 11.33 | 1.79 | 1.33 | 0.00 | |||
| Method | RGB views | Resolution per view | HFOV | Precision |
|---|---|---|---|---|
| NaVid | Front | FP16 | ||
| UniNaVid | Front | FP16 | ||
| StreamVLN | Front | BF16 | ||
| OmniNav (Action Former) | Left, front, right | BF16 |
| Variant | Calls | Navigation | Turns | Path (m) |
|---|---|---|---|---|
| w/o Navigation Skills | 7.43 | 5.33 | 0.10 | 7.56 |
| w/o Agent-oriented Tool Design | 30.28 | 10.72 | 17.56 | 23.60 |
| w/o Skills and Tool Design | 15.37 | 5.61 | 7.72 | 10.47 |
| SuperNav (Full) | 11.17 | 8.76 | 0.41 | 16.44 |