ASENA: Self-evolving Agents for Embodied Navigation
Organizations: NVIDIA · University of California, San Diego
Abstract
We present ASENA, an embodied agent system that connects general-purpose coding agents to robot sensing, computation, supervised execution, and persistent experience. Agents can write and execute programs, inspect recorded outcomes, repair failures, and reuse notes and executable skills while keeping their model weights fixed. We further introduce ASENA-VLN, a 4B monocular navigation policy that serves as an optional tool within this programmable system. ASENA-VLN predicts body-frame trajectories for both extended routes and short-horizon behaviors using a shared vision-language decoder trained on route instructions, visual question answering, and a newly curated dataset of geometry-derived atomic navigation tasks. As a standalone policy, ASENA-VLN achieves state-of-the-art success rates of 68.7% on R2R and 70.2% on RxR. When integrated with a coding agent, learned navigation improves ASENA's success rate by 11 percentage points on both agentic benchmarks while reducing execution time. Through persistent workspace evolution and simulator feedback, ten passes over recurring 100-task subsets further improve success from 72% to 98% on R2R and from 65% to 89% on RxR. On embodied question answering, ASENA achieves state-of-the-art accuracy with fewer interaction steps. Finally, real-world demonstrations on a Unitree G1 combine search, visual inspection, spatial reasoning, and synthesized gestures without a pre-built map, illustrating how online programming extends robot behavior beyond route following and predefined skills.
Figures & tables
| R2R | RxR | |||||||
| Method | NE | OSR | SR | SPL | NE | nDTW | SR | SPL |
| NaVid [ 45 ] | 5.72 | 49.2 | 41.9 | 36.5 | 5.72 | – | 45.7 | 38.2 |
| Uni-NaVid [ 44 ] | 5.58 | 53.3 | 47.0 | 42.7 | 6.24 | – | 48.7 | 40.9 |
| NaVILA [ 7 ] | 5.22 | 62.5 | 54.0 | 49.0 | 6.77 | 58.8 | 49.3 | 44.0 |
| StreamVLN [ 36 ] | 4.98 | 64.2 | 56.9 | 51.9 | 6.22 | 61.9 | 52.9 | 46.0 |
| NaVIDA [ 50 ] | 4.32 | 69.5 | 61.4 | 54.7 | 5.23 | 67.0 | 57.4 | 49.6 |
| HM3D | HM3D-OVON | |||
| Method | v2 val | seen | syn. | unseen |
| VLFM | 63.6/32.5 | 35.2/18.6 | 32.4/17.3 | 35.2/19.6 |
| OpenFMNav ‡ | 52.5/24.1 | – | – | – |
| SG-Nav | 49.6/25.5 | – | – | – |
| TriHelper ‡ | 56.5/25.3 | – | – | – |
| WMNav ‡ | 58.1/31.2 | – | – | – |
| R2R Agentic Split | RxR Agentic Split | |||||||||||
| Method / coding agent | VLN tool | NE | OSR | SR | SPL | Calls/ep. | NE | SR | SPL | Calls/ep. | ||
| Published zero-shot methods | ||||||||||||
| InstructNav [ 20 ] | – | 6.89 | 47.0 | 31.0 | 24.0 | – | – | – | – | – | ||
| Open-Nav [ 24 ] | – | 6.70 | 23.0 | 19.0 | 16.1 | – | – | – | – | – | ||
| CA-Nav [ 5 ] | – | 7.58 | 48.0 | 25.3 | 10.8 | – | 10.4 | 19.0 | 6.0 | – | ||
| GC-VLN [ 40 ] | – | 7.30 | 41.8 | 33.6 | 16.3 | – | 8.80 | 33.8 | 13.8 | – | ||
| HM-EQA | MT-HM3D | |||
| Method | Acc. | Steps | Acc. | Steps |
| Explore-EQA [ 27 ] | 58.4 | .52 | 35.1 | .64 |
| 3D-Mem [ 39 ] | 50.4 | .63 | – | – |
| Fine-EQA [ 16 ] | 56.0 | .54 | – | – |
| GraphEQA [ 28 ] | 63.5 | .20 | 45.6 | .45 |
| MemoryEQA (Qwen2VL-7B) [ 42 ] | – | – | 51.2 | .40 |
| HM-EQA | MT-HM3D | ||||
| Pass | Effort | Acc. | Steps | Acc. | Steps |
| Base | Low | 70.9 | .153 | 64.8 | .101 |
| 4 | Low | 73.4 | .136 | 66.5 | .091 |
| 9 | Low | 75.8 | .147 | 65.5 | .116 |
| 11 | Low | 75.4 | .152 | 65.3 | .114 |
| 17 | Low | 74.8 | .127 | 64.9 | .076 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Stored field | Type | Meaning |
| goal_type goal_value | category; text | Goal category and text. Recaptioned records can store just the destination phrase, while raw instructions and generated atoms retain the complete instruction. Categories include object, area, ego, region and point. |
| route_steps | ordered list | Ordered directions and landmark cues, when separately annotated. |
| constraints | list | Avoidance, traversal, or relative-position requirements. |
| end_pose | anchor list | Target, spatial relation, and distance derived from scene geometry, relative to an object, a room, or the robot’s starting position. |
| quality | 1–5 | Quality annotation or assigned setting used to condition the policy. R2R/RxR evaluation uses 5. |
| Dataset | Supervision | Weighted pool (M) | Environments |
| R2R [ 2 ] | traj, pixel, CoT | 1.502 | MP3D |
| R2R (augmented) | traj, pixel | 3.290 | MP3D |
| FGR2R [ 13 ] | traj, pixel | 0.704 | MP3D |
| RxR [ 17 ] | traj, pixel, CoT | 4.191 | MP3D |
| RxR (augmented) | traj, pixel | 4.836 | MP3D |
| ScaleVLN [ 34 ] | traj, pixel, CoT | 3.120 | HM3D |
| Variant | Export size | Visual | Prefill | Decode (32) | Total | Speedup |
| (GiB) | Latency (ms) | |||||
| PyTorch BF16 | 9.01 | 167.3 | 363.4 | 2,275.5 | 2,806 | 1.00 |
| FP8, BF16 vision | 5.63 | 73.3 | 128.7 | 847.4 | 1,050 | 2.67 |
| FP8, BF16 vision, reduced vocab | 4.91 | 78.3 | 125.5 | 694.1 | 898 | 3.13 |
| FP8, reduced vocab | 4.53 | 64.9 | 129.1 | 683.4 | 877 | 3.20 |
| Variant | First-waypoint agreement | Exact-trajectory agreement |
| FP8 | 90.23% | 75.39% |
| FP8, BF16 vision | 91.41% | 82.81% |
| Navigation instruction | Policy input (JSON) |
| Route following Descend down the stairs. Turn left and stop in the room. | {"goal": "Descend down the stairs. Turn left and stop in the room. ", "quality": 5} |
| Metric motion Could you move forward 1 meters for me? | {"goal": "Could you move forward 1 meters for me?", "quality": 5, "end_pose": [{ "target": "agent_start", "relation": "forward", "distance_m": 1.0 }]} |
| Relative object goal Walk over to the built-in dishwasher off to your left. | {"goal": "Walk over to the built-in dishwasher off to your left.", "quality": 5, "end_pose": [{ "target": "built-in dishwasher", "relation": "near", "distance_m": 0.55 }]} |
| Room goal Stop there once you find a bedroom. | {"goal": "Stop there once you find a bedroom.", "quality": 5, "end_pose": [{ "target": "bedroom", "relation": "inside", "distance_m": 0.0 }]} |
| Floor change Make your way downstairs to the hallway. | {"goal": "Make your way downstairs to the hallway.", "quality": 5, "end_pose": [{ "target": "hallway", "relation": "inside", "distance_m": 0.0 }]} |
| Component | Native input / specification | Configured output |
| Insta360 X5 | USB panorama: ; 30 fps [ 14 ] | 6 Hz JPEG; 180° yaw correction |
| Livox Mid-360 | 200,000 points/s; field of view: 360° H, to V [ 19 ] | SLAM bridge: 10 Hz; at most 20,000 points/batch |
| Body state | 29 joint positions/velocities; IMU, odometry and mode | 20 Hz |
| Parameter | Value | Parameter | Value |
| Command publication | 10 Hz | Path lookahead | 0.40 m |
| Translation speed bounds | 0.12–0.35 m/s | Braking lead time | 0.60 s |
| Near-goal speed ( m) | 0.25 m/s | Velocity-command expiry | 2.0 s |
| Arrival tolerances | 0.10 m; 4° | Active-job heartbeat timeout | 3.0 s |
| Move / turn command limits | 2.0 m; 180° | Path chunk limits | 6.0 m; 16 points |
| Move-job timeout | 25 s | Path-job timeout | 45 s |