NavHarness: Towards Lifelong Embodied Navigation
Organizations: Adelaide University · Responsible AI Research Centre, AIML · CSIRO Data61
Abstract
Frontier models can now perform well on individual embodied navigation tasks through multi-round multimodal reasoning with simple tools. Across successive tasks, however, an agent must also rely on an evolving map and earlier search records, both of which may be incomplete or conflict with new observations. We present NavHarness, a training-free embodied harness towards lifelong navigation that makes memory processing part of the navigation loop. During navigation, its multi-round agentic session draws on maps, task records, and house knowledge, checking them against observations and recording corrections to guide its actions. NavHarness preserves this experience across fresh conversations for new tasks or recovery attempts, while outcome verification and run-end summaries support its later reuse. On GOAT-Bench, NavHarness improves s-SR over context-only independent sessions by 18.6 points with Astra and 22.6 with Opus 5. Using SLAM-estimated poses, NavHarness with GPT-6 Astra achieves state-of-the-art task success of 83.7 s-SR with 36.9 e-SR on GOAT-Bench and 85.9 s-SR on IR2R-CE. To understand these gains, we examine how experience is carried between sessions and find that structured recovery handovers outperform length-matched summaries. In extended deployments across houses, consolidation improves navigation beyond retaining maps and task records, with case studies showing how agents use earlier experience to interpret new goals, investigate unresolved questions, and resume failed searches. We suggest that progress towards lifelong navigation depends on how successive reasoning sessions build on prior experience, alongside improvements in single-task capability.
Figures & tables
| Method | Base Model | Type | GT pose | s-SR | SPL | e-SR |
|---|---|---|---|---|---|---|
| MemoryExplorer † ( Wang et al., 2026c ) | Qwen2.5-VL-7B | trained | ✓ | 46.4 | 28.0 | – |
| SSMG-Nav ( Niu et al., 2026 ) | Qwen-VL-Plus | zero-shot | ✓ | 46.5 | 34.1 | 8.6 |
| EvoMemNav ( Ge et al., 2026 ) | Qwen3-VL-8B | zero-shot | ✓ | 59.6 | 38.9 | – |
| ReEXplore † ( Zhang et al., 2025a ) | GPT-4o | zero-shot | ✓ | 59.8 | 42.5 | – |
| AstraNav-Memory ( Ren et al., 2025 ) | Qwen2.5-VL-3B | trained | ✓ | 62.7 | 56.9 | – |
| STEGNav ( Chen et al., 2026c ) | GPT-5.4-mini | zero-shot | ✓ | 66.3 | 39.7 | – |
| Method | Type | GT pose | s-SR | SPL | e-SR | t-nDTW |
|---|---|---|---|---|---|---|
| CMA ( Krantz et al., 2023 ) | trained | 19 | 18 | – | 38 | |
| MAP-CMA ( Krantz et al., 2023 ) | trained | ✓ | 35 | 32 | – | 47 |
| ETPNav ( Han et al., 2026 ) | trained | ✓ | 28 | 27 | – | 41 |
| HNR ( Han et al., 2026 ) | trained | ✓ | 30 | 28 | – | 44 |
| OVER-NAV ( Zhao et al., 2024 ) | hybrid | ✓ | 35 | 33 | – | 50 |
| SeqWalker ( Han et al., 2026 ) | trained | ✓ | 36 | 34 | – | 52 |
| Configuration / intervention | s-SR | SPL | s-SR | 95% CI |
|---|---|---|---|---|
| NavHarness (full reference) | – | – | ||
| A. Session policies | ||||
| Independent sessions | ||||
| Single session | ||||
| Independent sessions + pre-stop verification | ||||
| B. Cross-task memory | ||||
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Call | Arguments | Returns or effect |
|---|---|---|
| observe | none | the forward RGB view and remaining step budget |
| step | a list of actions | executes them and reports motion, collision, and budget status |
| get_goal | none | the current goal, as text or as the goal image |
| get_map | floor, zoom | the top-down occupancy map with the markers on it |
| mark | a place name | drops or moves a named marker at the current position |
| preview_path | a map point | the length of a route over known free space; nothing moves |
| Position | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|
| s-SR (%) | 72.2 | 72.2 | 83.3 | 83.3 | 77.8 | 82.1 | 73.1 | 85.7 | 81.2 | 85.7 |
| Path | Writer | Content |
|---|---|---|
| run/ledger.md | Orchestrator | Task-status index |
| run/tasks/ /goal.md | Orchestrator | Goal and modality |
| run/tasks/ /handover.md | Acting session | Task experience |
| run/tasks/ /lambda.md | Acting session | Recovery handover |
| run/tasks/ /verdict.json | Orchestrator | Judge’s assessment |
| house/ | Consolidation session | House notes: index, overview, room notes, skills |
| Scope | State reconsidered |
|---|---|
| continue | No reset |
| attempt | Active conversation |
| place | Conversation and current location |
| room | Conversation, location, and route |
| floor | Conversation, location, route, and floor association |
| house | The same beliefs as floor, with house-level reassessment |
| Benchmark | Scenes | Episodes / tours | Tasks | Tasks per sequence |
|---|---|---|---|---|
| GOAT-Bench | 36 | 360 episodes | 2,669 | 5–10 |
| IR2R-CE | 11 | 36 tours | 1,824 | 3–100 |
| Method | Type | s-SR | SPL |
|---|---|---|---|
| SenseAct-NN Monolithic ( Khanna et al., 2024 ) | trained | 12.3 | 6.8 |
| Modular CLIP on Wheels ( Gadre et al., 2023 ; Khanna et al., 2024 ) | zero-shot | 16.1 | 10.4 |
| VLMnav ( Goetting et al., 2024 ; Ji et al., 2025 ) | zero-shot | 20.1 | 9.6 |
| Modular GOAT ( Khanna et al., 2024 ) | zero-shot | 24.9 | 17.2 |
| DyNaVLM ( Ji et al., 2025 ) | zero-shot | 25.5 | 10.2 |
| SenseAct-NN Skill Chain ( Khanna et al., 2024 ) | trained | 29.5 | 11.3 |
| Method | Type | GT pose | s-SR | SPL | e-SR | t-nDTW |
|---|---|---|---|---|---|---|
| TourCMA | trained | 18 | 17 | – | 36 | |
| PoolCMA | trained | 16 | 15 | – | 36 | |
| PoolEndCMA | trained | 18 | 16 | – | 38 |
| Quantity | Limit or counting rule |
|---|---|
| Movement budget | 500 executed forward or turning actions per task |
| Model-turn budget | 200 navigation-model turns per task |
| Action batching | Each executed action is charged individually |
| Recovery allowance | At most one recovery, sharing the remaining task budgets |
| Fallback trigger | 80 turns, without resetting either budget |
| Observation and memory | No movement-action charge |
| Metric | NavHarness | Fixed access |
|---|---|---|
| s-SR (%) | 81.5 0.7 | 74.5 0.8 |
| SPL | 55.0 0.6 | 49.0 0.7 |
| Mean time / task (s) | 129.5 | 137.2 |
| Paired s-SR (pp) | ||
| 95% CI (pp) | ||
| Navigator | Judge | s-SR (%) | SPL |
|---|---|---|---|
| Qwen3.8-27B | Qwen3.8-27B | 71.7 0.5 | 48.2 0.4 |
| Qwen3.8-27B | Opus 5 | 71.3 0.4 | 48.5 0.5 |
| Opus 5 | Opus 5 | 81.5 0.7 | 55.0 0.6 |
| Opus 5 | Qwen3.8-27B | 81.7 0.6 | 55.2 0.5 |
| Goal | s-SR | SPL |
|---|---|---|
| Object category | 86.5 | 62.4 |
| Language description | 75.2 | 48.8 |
| Image | 82.2 | 53.0 |
| Diagnostic | Full | Independent + pre-stop verification |
|---|---|---|
| Tasks with a refused first stop | 11 | 20 |
| Rejection precision | 73 | 75 |
| Rescue among correctly refused stops | 45 | 35 |
| Rescues / all tasks | 3.6 | 5.0 |
| Harms / all tasks | 0.7 | 1.1 |
| Net local success change |
| Evidence | Certified, true | Certified, false | Rejected, true | Rejected, false | Acc. (accept-all) | |
|---|---|---|---|---|---|---|
| Four stop views | 78.5 | 5.2 | 4.1 | 12.2 | 90.7 (82.6) | 0.67 |
| One forward frame | 74.2 | 10.7 | 6.4 | 8.7 | 82.9 (80.6) | 0.40 |
| Component | s-SR | Absorbs | What shows when it is removed |
|---|---|---|---|
| Map and markers | revisits and displacement | steps , markers written about | |
| Ledger | repeated goals | steps despite retaining the map | |
| Recovery note | a stalled search | recovery success 35% instead of 56% | |
| Recovery | an unproductive attempt | requests come long before the fallback | |
| Four-view evidence | an ambiguous closure | claimed-completion : 0.40 vs. 0.67 in separate runs | |
| Certification | an unsupported claim | 12.2% false and 4.1% true completions rejected (share of claims) |
| Behaviour | Instruction or tool support | Relation |
|---|---|---|
| Report the goal, outcome, and rooms passed | Handover instruction | Requested |
| Report which markers were dropped | Handover instruction | Requested |
| Write down unsuccessful searches | Handover instruction | Requested |
| Move a marker | Marker tool supports re-placement | Tool-supported |
| Question an earlier task account | Warning that accounts may be wrong | Requested |
| Correct a particular marker | General permission to maintain markers | Not scripted |