We present ASENA, an embodied agent system that connects general-purpose coding agents to robot sensing, computation, supervised execution, and persistent experience. Agents can write and execute programs, inspect recorded outcomes, repair failures, and reuse notes and executable skills while keeping their model weights fixed. We further introduce ASENA-VLN, a 4B monocular navigation policy that serves as an optional tool within this programmable system. ASENA-VLN predicts body-frame trajectories for both extended routes and short-horizon behaviors using a shared vision-language decoder trained on route instructions, visual question answering, and a newly curated dataset of geometry-derived atomic navigation tasks. As a standalone policy, ASENA-VLN achieves state-of-the-art success rates of 68.7% on R2R and 70.2% on RxR. When integrated with a coding agent, learned navigation improves ASENA's success rate by 11 percentage points on both agentic benchmarks while reducing execution time. Through persistent workspace evolution and simulator feedback, ten passes over recurring 100-task subsets further improve success from 72% to 98% on R2R and from 65% to 89% on RxR. On embodied question answering, ASENA achieves state-of-the-art accuracy with fewer interaction steps. Finally, real-world demonstrations on a Unitree G1 combine search, visual inspection, spatial reasoning, and synthesized gestures without a pre-built map, illustrating how online programming extends robot behavior beyond route following and predefined skills.
Figures & tables
Figure 1 : ASENA connects coding agents to physical robot skills. Programs combine sensing, computation, supervised execution, and reusable skills, with learned navigation as an optional tool. The coding example analyzes LiDAR geometry to propose a route through an opening. The inset shows the subsequent observation inside the passage. The bottom panels show a subset of skills evolved by the system. https://asena-bot.github.io .
Figure 2 : ASENA connects coding agents to physical robots. The system integrates robot sensing, a persistent workspace, supervised execution, and an optional learned navigation policy. Each run is stored as an MCAP record [ 10 ] that aligns sensor and robot states with actions, operator feedback, code, and visualizations for inspection and replay. The agent reflects on these records to update its notes and skills while robot motion is disabled.
Figure 3 : Navigation policy architecture and atomic training examples. Left: language and monocular history enter a shared autoregressive decoder. Separate token blocks encode trajectories and pixel goals; visual QA uses text tokens. Motion tokens decode to body-frame waypoints. Architecture glyphs are schematic. Right: recorded training observations with instructions and overlaid trajectories. The overlays do not use metric camera projection. More examples appear in Figure 6 .
R2R
RxR
Method
NE ↓
OSR ↑
SR ↑
SPL ↑
NE ↓
nDTW ↑
SR ↑
SPL ↑
NaVid [ 45 ]
5.72
49.2
41.9
36.5
5.72
–
45.7
38.2
Uni-NaVid [ 44 ]
5.58
53.3
47.0
42.7
6.24
–
48.7
40.9
NaVILA [ 7 ]
5.22
62.5
54.0
49.0
6.77
58.8
49.3
44.0
StreamVLN [ 36 ]
4.98
64.2
56.9
51.9
6.22
61.9
52.9
46.0
NaVIDA [ 50 ]
4.32
69.5
61.4
54.7
5.23
67.0
57.4
49.6
Table 1 : Standalone monocular VLN-CE policy results on R2R and RxR val_unseen. ASENA-VLN-4B achieves a state-of-the-art result compared to prior work, with significantly lower navigation error and higher nDTW.
HM3D
HM3D-OVON
Method
v2 val
seen
syn.
unseen
VLFM
63.6/32.5
35.2/18.6
32.4/17.3
35.2/19.6
OpenFMNav ‡
52.5/24.1
–
–
–
SG-Nav
49.6/25.5
–
–
–
TriHelper ‡
56.5/25.3
–
–
–
WMNav ‡
58.1/31.2
–
–
–
Table 2 : Object-goal navigation: paired SR/SPL (%, both ↑ ). ‡ denotes HM3D v1 result, where emphasis excludes these.
R2R Agentic Split
RxR Agentic Split
Method / coding agent
VLN tool
NE ↓
OSR ↑
SR ↑
SPL ↑
Calls/ep. ↓
NE ↓
SR ↑
SPL ↑
Calls/ep. ↓
Published zero-shot methods
InstructNav [ 20 ]
–
6.89
47.0
31.0
24.0
–
–
–
–
–
Open-Nav [ 24 ]
–
6.70
23.0
19.0
16.1
–
–
–
–
–
CA-Nav [ 5 ]
–
7.58
48.0
25.3
10.8
–
10.4
19.0
6.0
–
GC-VLN [ 40 ]
–
7.30
41.8
33.6
16.3
–
8.80
33.8
13.8
–
Table 3 : Navigation before self-evolution on R2R and RxR Agentic Splits. NE is in meters and OSR/SR/SPL are percentages. Calls/ep. is mean tool calls over 100 episodes. Green percentages show reductions from the same backend without VLN.
Figure 4 : Navigation workspace self-evolution. Success-related metrics (SR, OSR, and SPL) improve across evolution passes, while the cost per pass generally decreases. Access to the VLN tool consistently improves SPL, enabling the agent to reach destinations more efficiently.
HM-EQA
MT-HM3D
Method
Acc. ↑
Steps ↓
Acc. ↑
Steps ↓
Explore-EQA [ 27 ]
58.4
.52
35.1
.64
3D-Mem [ 39 ]
50.4
.63
–
–
Fine-EQA [ 16 ]
56.0
.54
–
–
GraphEQA [ 28 ]
63.5
.20
45.6
.45
MemoryEQA (Qwen2VL-7B) [ 42 ]
–
–
51.2
.40
Table 4 : Embodied question answering. Accuracy (%) and mean normalized steps with differing coverage.
HM-EQA
MT-HM3D
Pass
Effort
Acc. ↑
Steps ↓
Acc. ↑
Steps ↓
Base
Low
70.9
.153
64.8
.101
4
Low
73.4
.136
66.5
.091
9
Low
75.8
.147
65.5
.116
11
Low
75.4
.152
65.3
.114
17
Low
74.8
.127
64.9
.076
Table 5 : Frozen EQA workspaces: accuracy (%) and normalized exploration steps. Shading marks the final workspace.
Figure 5 : Real-world navigation and interaction on the G1 robot. Three supervised tasks pair online maps and recorded trajectories with observations and responses. Blue and dashed orange traces show search and return. Distances are estimated from drive odometry. Prompts and responses are condensed for display.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Stored field
Type
Meaning
goal_type goal_value
category; text
Goal category and text. Recaptioned records can store just the destination phrase, while raw instructions and generated atoms retain the complete instruction. Categories include object, area, ego, region and point.
route_steps
ordered list
Ordered directions and landmark cues, when separately annotated.
constraints
list
Avoidance, traversal, or relative-position requirements.
end_pose
anchor list
Target, spatial relation, and distance derived from scene geometry, relative to an object, a room, or the robot’s starting position.
quality
1–5
Quality annotation or assigned setting used to condition the policy. R2R/RxR evaluation uses 5.
Appendix
Table 6 : Stored navigation annotations. The policy receives goal text, quality, and optional end-pose anchors ( Table 10 ). Coordinate-based point goals use a separate input format.
Dataset
Supervision
Weighted pool (M)
Environments
R2R [ 2 ]
traj, pixel, CoT
1.502
MP3D
R2R (augmented)
traj, pixel
3.290
MP3D
FGR2R [ 13 ]
traj, pixel
0.704
MP3D
RxR [ 17 ]
traj, pixel, CoT
4.191
MP3D
RxR (augmented)
traj, pixel
4.836
MP3D
ScaleVLN [ 34 ]
traj, pixel, CoT
3.120
HM3D
Appendix
Table 7 : Training pools for ASENA-VLN. Pool sizes account for quality and backward-action filtering, alternative views, chain-of-thought (CoT) variants, repeated trajectory-ending windows, and sampling weights. Counts are rounded independently and include repeated examples. They describe the available pools, not the number of samples consumed during training. Sampling details appear in Section A.1 .
Variant
Export size
Visual
Prefill
Decode (32)
Total
Speedup
(GiB)
Latency (ms)
PyTorch BF16
9.01
167.3
363.4
2,275.5
2,806
1.00 ×
FP8, BF16 vision
5.63
73.3
128.7
847.4
1,050
2.67 ×
FP8, BF16 vision, reduced vocab
4.91
78.3
125.5
694.1
898
3.13 ×
FP8, reduced vocab
4.53
64.9
129.1
683.4
877
3.20 ×
Appendix
Table 8 : Policy inference on the Unitree G1’s Jetson AGX Thor. These measurements use an earlier checkpoint, separate from the benchmark policy. All variants use batch size 1, the same eight-frame mixed-resolution history, and 32 greedy decoding steps. Totals sum the three components. Speedups are computed from unrounded sums. Measured BF16 end-to-end latency is 2,735 ms. Export size describes the corresponding ONNX files. The BF16 model occupies 8.28 GiB of allocated memory after loading. This is not a peak runtime measurement. BF16 values are warm medians of five runs: prefill is time to first token minus visual encoding, and decode scales the remaining 31 tokens to 32 steps. TensorRT-Edge-LLM visual and prefill values are 10-run means. Decode is the median of five runs.
Variant
First-waypoint agreement
Exact-trajectory agreement
FP8
90.23%
75.39%
FP8, BF16 vision
91.41%
82.81%
Appendix
Table 9 : Prediction agreement between the earlier quantized model and its BF16 version on 256 held-out inputs. First-waypoint agreement requires the first predicted (x,y,θ) tuple to match. Exact-trajectory agreement requires the entire predicted trajectory to match. These metrics measure output consistency, not navigation success.
Figure 6 : Atomic navigation training examples. Each observation is paired with an instruction and an overlaid trajectory. The overlays do not use metric camera projection. The examples span metric motion, rotation, relative goals, alignment, room goals, doors, corridors, and floor changes.
Navigation instruction
Policy input (JSON)
Route following Descend down the stairs. Turn left and stop in the room.
{"goal": "Descend down the stairs. Turn left and stop in the room. ", "quality": 5}
Metric motion Could you move forward 1 meters for me?
{"goal": "Could you move forward 1 meters for me?", "quality": 5, "end_pose": [{ "target": "agent_start", "relation": "forward", "distance_m": 1.0 }]}
Relative object goal Walk over to the built-in dishwasher off to your left.
{"goal": "Walk over to the built-in dishwasher off to your left.", "quality": 5, "end_pose": [{ "target": "built-in dishwasher", "relation": "near", "distance_m": 0.55 }]}
Room goal Stop there once you find a bedroom.
{"goal": "Stop there once you find a bedroom.", "quality": 5, "end_pose": [{ "target": "bedroom", "relation": "inside", "distance_m": 0.0 }]}
Floor change Make your way downstairs to the hallway.
{"goal": "Make your way downstairs to the hallway.", "quality": 5, "end_pose": [{ "target": "hallway", "relation": "inside", "distance_m": 0.0 }]}
Appendix
Table 10: Navigation instructions and policy inputs. Each JSON input contains the goal text and a quality setting, with an optional end-pose anchor that specifies where the robot should finish.
Figure 7 : G1 sensor mounting. Front and side views showing sensor placement on the robot.
Component
Native input / specification
Configured output
Insta360 X5
USB panorama: 2880×1440 ; 30 fps [ 14 ]
6 Hz JPEG; 180° yaw correction
Livox Mid-360
200,000 points/s; field of view: 360° H, −7° to 52° V [ 19 ]
SLAM bridge: 10 Hz; at most 20,000 points/batch
Body state
29 joint positions/velocities; IMU, odometry and mode
20 Hz
Appendix
Table 11: Sensor and state interfaces. Nominal specifications and configured publication rates. Lidar angles use the manufacturer’s orientation.
Parameter
Value
Parameter
Value
Command publication
10 Hz
Path lookahead
0.40 m
Translation speed bounds
0.12–0.35 m/s
Braking lead time
0.60 s
Near-goal speed ( ≤0.5 m)
0.25 m/s
Velocity-command expiry
2.0 s
Arrival tolerances
0.10 m; 4°
Active-job heartbeat timeout
3.0 s
Move / turn command limits
2.0 m; 180°
Path chunk limits
6.0 m; 16 points
Move-job timeout
25 s
Path-job timeout
45 s
Appendix
Table 12: Navigation settings and safeguards. The slow-walk controller limits the speed, duration, and size of individual commands and path segments.