NavGPT-3: Harnessing Context in a Hierarchical Navigation Runtime
Organizations: AIML, Adelaide University · Metacognition · Roblox · PKU · SJTU · UNC Chapel Hill · UNSW · ANU
Abstract
Language models trained with long-horizon agentic reinforcement learning can generalize knowledge through reasoning, express precise actions, and pursue goals over many steps, raising the ceiling on what an embodied agent can understand and decide. Physical interaction, however, remains the domain of action policies, which provide dense, low-latency control. We present NavGPT-3, a harness that connects the two models, with an OS-like runtime built above it: reasoning, acting, and monitoring run as threads with their own context, tools, and permissions, while the runtime schedules them and decides which thread controls the robot's motion, so that the robot can react to sudden real-world events through interruption and thread switching. Beneath it, our action policy NavGPT VLA, trained on 19.28M examples, allocates visual tokens using codec allocation, in proportion to scene change; its 8B model alone reaches 74.51 SR on R2R-CE and leads RxR-CE with 78.19 SR. With the complete harness, NavGPT-3 sets the state of the art on R2R-CE (81.51 SR) and, for the first time, brings an autonomous agent to human level: on RxR-CE it matches human followers in success (90.43 vs. 90.4 SR) and path fidelity (78.47 vs. 77.7 nDTW) at 1 min 22 s per episode, versus roughly 3 min for a human. We comprehensively ablate the harness design and the interaction between the two models, showing how tools and the action policy shape the path from language-model reasoning to physical control: when NavGPT VLA executes the route, the reasoning loop shortens and the system's minimum reaction time falls from 3-19 s per language-model decision to 0.5-1 s per action-policy step (1-2 Hz). These results show that designing this embodied interface is central to connecting frontier language-model intelligence with low-level physical control. We will release all models, code, and evaluation records.
Figures & tables
| Family | Sources | Sampling | Size (M) |
| Action regression | |||
| VLN: four-view | 7 | 100% | 3.938 |
| VLN: single-view | 5 | 100% | 2.113 |
| VLN: SRDF | 2 | 15% | 0.516 |
| Point-goal navigation | 5 | 100% | 0.984 |
| Object-goal navigation | 1 | 100% | 2.000 |
| Backbone | GPUs | BS | Steps | Length | LR | AH LR |
| Qwen3-VL-4B | 32 | 256 | 75.3k | 8,192 | ||
| Qwen3-VL-8B | 32 | 256 | 75.3k | 8,192 |
| Planner | Tools | Navigation | Resource use | |||||||||
| VLA | Map | Backtrack | NE | OSR | SR | SPL | Tokens (k) | Turns | Cost (\downarrow$ | Infer. (s) | React. (s) | |
| VLA only | ✓ | – | – | 4.23 | 78 | 71 | 66.14 | – | – | – | 44.1 | 0.78 |
| GPT-6 Astra | – | – | – | 4.67 | 74 | 72 | 62.26 | 325.0 | 32.14 | 0.593 | 105.7 | 3.3 |
| GPT-6 Astra | – | ✓ | ✓ | 2.55 | 87 | 83 | 69.04 | 190.8 | 10.71 | 0.469 | 66.8 | 6.2 |
| GPT-6 Astra | ✓ | – | – | 4.58 | 79 | 75 | 67.83 | 227.9 | 3.37 | 0.469 | 84.4 | 0.78 |
| GPT-6 Astra | ✓ | ✓ | – | 4.01 | 79 | 76 | 67.35 | 201.4 | 3.42 | 0.443 | 80.4 | 0.78 |
| Group | Tool | Effect on state | Returned to the Planner |
| Basic | observe_forward() | None | Current forward RGB for fine semantic inspection. |
| observe_panorama() | Updates observed local coverage, not pose | Four labelled views with bearings, clearance, pose displacement, and node information. | |
| navigate_relative() | Executes a validated local turn and translation | Realized motion status, endpoint panorama, and nearby node candidates. | |
| terminate_episode() | Optionally returns to a selected node, then stops | Stop or return status; this call ends the episode and cannot be undone. | |
| VLA | navigate_by_ instruction() | Runs one VLA rollout; the VLA keeps its history | Up to eight sampled route keyframes, endpoint panorama, optional observed map, and execution status. |
| Map | observe_map() | Updates observed coverage, not pose | Observed-only top-down map, map status, and route-ordered node table. |
| Frequency | NE | OSR | nDTW | SR | SPL | Tokens (k) |
| None | 3.82 | 81 | 16.74 | 77 | 67.51 | 179.2 |
| 32 | 3.56 | 81 | 16.66 | 78 | 68.88 | 1,529.9 |
| 16 | 2.01 | 85 | 18.72 | 83 | 77.36 | 2,319.5 |
| Method | Model | R2R-CE Val-Unseen | RxR-CE Val-Unseen | ||||||
| NE | OSR | SR | SPL | NE | nDTW | SR | SPL | ||
| Human † ( Ku et al., 2020 ) | – | – | – | – | – | 1.32 | 77.7 | 90.4 | – |
| NavFoM ( Zhang et al., 2025a ) | 7B | 4.61 | 72.1 | 61.7 | 55.3 | 4.74 | 65.8 | 64.4 | 56.2 |
| ABot-N0 ( AMAP CV Lab, 2026 ) | 4B | 3.78 | 70.8 | 66.4 | 63.9 | 3.83 | – | 69.3 | 60.0 |
| AstraNav-World ( Chen et al., 2026 ) | 3B+5B | 3.86 | 73.9 | 67.9 | 65.4 | 3.82 | – | 72.9 | 61.5 |
| OmniNav ( Xue et al., 2025 ) | 3B | 3.74 | 74.6 | 69.5 | 66.1 | 3.77 | – | 73.6 | 62.0 |
| Scenario | Controlled event | Primary measurements |
| A: Route revision | A detour or endpoint correction becomes necessary during an extended rollout; vary event timing relative to Planner inference. | Completion SR; event-to-authorized-repair time; route efficiency; Planner calls; execution idle time. |
| B: LiDAR protection | A soft obstacle enters a calibrated protected region during motion, including while the Planner is busy. Use matched harmless passages as negatives. | Missed and false stops; trigger-to-halt p50/p95; post-trigger distance; minimum clearance; operator interventions. |
| C: Priority and recovery | Trigger a hazard during repair or near VLA completion; delay a model reply, resend an outdated command to the robot’s command filter, then clear the hazard. | Accepted stale commands; event ordering; halt latency; unauthorized resumption; correct resumption and task completion. |
| A: Revision | B: Protection | C: Recovery | ||||
| Condition | SR | Repair (s) | Missed/ | Halt p95 (ms) | Stale/ | Resume SR |
| VLA-only | 60.0 | – | – | – | – | – |
| Planner + tools | 70.0 | 14.2 | 9/30 | 2150 | 11/30 | 53.3 |
| Sequential + queued | 73.3 | 12.8 | 7/30 | 1840 | 8/30 | 60.0 |
| Sequential + priority | 76.7 | 11.5 | 2/30 | 420 | 3/30 | 70.0 |
| Overlapping + queued | 76.7 | 9.6 | 5/30 | 1310 | 6/30 | 66.7 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Responsibility | Component and role |
| Thread coordination | Runtime: schedule threads, manage permissions, and control motion authority. |
| Navigation decisions | Planner: interpret evidence, choose tools, and decide when the goal is reached. |
| Context construction | Harness: select and render the images and state shown to the models. |
| Persistent navigation state | Harness: keep route references, nodes, and annotations. |
| Tool checking | Harness: validate requests and execute permitted tool calls. |
| Low-level navigation | VLA: predict waypoints from its visual history. |
| Capability | Navigation purpose | What it takes and returns |
| Local observation | Ground landmarks and inspect the current surroundings | Relate forward or panoramic appearance to the current viewing direction and pose. |
| Instruction delegation | Execute an extended route with the learned policy | Pass the requested instruction to the VLA, keep its history across calls, and return images of the executed route with its status. |
| Route inspection | Relate current progress to previously encountered places | Show the observed map and the ordered list of visited places, without marking unobserved space as free. |
| Spatial revision | Revisit a recorded place or correct a local endpoint error | Turn a reference to a visited place, or a relative motion request, into checked movement and report what was executed. |
| Semantic annotation | Retain an interpretation of a visited place | Attach a label to a recorded place without changing its recorded position. |
| Task termination | Accept the final navigation state | Make episode completion explicit and distinct from the end of a delegated rollout. |
| Operation | Planner | VLA | Enforced by |
| Interrupt the running VLA rollout | May request, if permitted | Stops after the current step | Runtime: revokes the permission to move |
| Request an early Planner review | Receives it, if permitted | Keeps executing | Runtime: queues or discards the request |
| Resume or revise the route | Chooses an available tool | Executes a valid delegation | Runtime: grants one operation the permission to move |
| Change physical or route state | Cannot change it directly | Proposes waypoints | Environment: executes motion; harness: updates route records |
| End the episode | Decides explicitly to stop | Its stop ends only the delegation | Runtime: enforces the stop and episode limits |
| Share information between Planner and VLA | Sends structured instructions only | Returns structured status and route evidence only | Runtime: controls visibility; harness: builds context |
| Images differing | Mean (tokens) | |||||
| 3072 | 2 | 1 | 196 | 7.84 | 0 / 2 | 0.0 |
| 3072 | 8 | 1 | 196 | 1.96 | 0 / 8 | 0.0 |
| 3072 | 15 | 1 | 196 | 1.04 | 0 / 15 | 0.0 |
| 4096 | 4 | 4 | 256 | 1.00 | 0 / 16 | 0.0 |
| 3072 | 16 | 1 | 196 | 0.98 | 3 / 16 | 8.0 |
| 3072 | 8 | 4 | 196 | 0.49 | 30 / 32 | 47.3 |
| Setting | Value |
| RGB resolution | 512 pixels |
| Camera order | Front, right, back, left |
| Returned route keyframes | At most 8, uniformly sampled along the rollout |
| Episode / delegation limit | 500 / 200 environment steps |
| Planner-call limit | 200 turns per episode |
| Episode timeout | 2,400 s |
| Statistic | Full | Subset | ||
| Episodes | 1,839 | 100 | – | – |
| Scenes | 11 | 10 | – | – |
| Unique reference trajectories | 613 | 97 | – | – |
| Geodesic distance (m) | 0.033 | 0.072 | ||
| Instruction length (words) | 0.207 | 0.103 | ||
| Directional mentions | 0.034 | 0.040 |
| Dataset | System | SR | SPL | SR vs. full | |
| MP3D | Qwen-RobotNav 8B | 500 | 47.8 | 17.5 | |
| MP3D | NavGPT-3 | 500 | 59.2 | 25.2 | – |
| HM3D v2 | Qwen-RobotNav 8B | 500 | 72.4 | 33.1 | |
| HM3D v2 | NavGPT-3 | 500 | 75.3 | 40.8 | – |
| Judge | Labeled | Agreement (%) |
| GPT-4o-mini | 50 | 72.0 |
| GPT-6 Astra | 50 | 92.0 |
| Astra 4o-mini | +20.0 |