Pretrained video foundation models encode rich semantic and spatiotemporal priors for embodied navigation, yet converting these priors into UAV motion typically requires expensive future-video synthesis and geometric reconstruction. We investigate whether the motion implicit in future visual prediction can instead be recovered directly from the predictive representations of a frozen video model. To this end, we present DiffWAM, a geometry-conditioned navigation world-action model that directly transforms multi-level predictive features into continuous camera trajectories. Its Grid-Motion module preserves spatial-temporal motion associations, while Latent2Pose grounds them with first-frame geometry to recover metrically meaningful 3D motion. Complete video rollouts and geometric reconstruction are required only for offline supervision, eliminating future-video decoding and multi-frame reconstruction during deployment. We further introduce FastDreamer, which overlaps predictive and geometric computation with ongoing flight and performs timestamp-aware asynchronous trajectory handoff for continuous UAV execution. DiffWAM achieves a trajectory RMSE of 0.3492 m and an endpoint success rate of 74.40% on the 1,000-sample DiffWAM-1000 benchmark, while representative real-world experiments demonstrate complex behaviors including constrained traversal, orbiting, S-shaped flight, and multi-stage navigation. An onboard DiffWAM-Flash implementation further reaches 1.08 s model-pipeline latency on NVIDIA Jetson AGX Thor. These results demonstrate that predictive video representations can be efficiently grounded into continuous 3D motion, providing a direct alternative to generate-then-reconstruct navigation pipelines. Project page: https://zzmmzzm.github.io/diffwam.github.io/.
Figures & tables
Figure 1: Overview of DiffWAM. DiffWAM grounds multi-level predictive representations from a frozen video world model, together with first-frame geometry, into continuous 3D UAV trajectories. Trained with large-scale synthetic navigation data, it supports diverse structured motions, real-world execution, and efficient deployment without online future-video decoding.
Figure 2: DiffWAM’s pipeline. DiffWAM converts selected early predictive features and first-frame geometry into camera-motion trajectories through Grid-Motion and Latent2Pose. Training proceeds through video-backbone preparation (S0), navigation-aware readout pretraining (S1), and geometry-guided predictive distillation (S2), where completed video rollouts and geometric reconstruction provide offline trajectory supervision. Detailed tensor and output dimensions are specified in Fig. 3 .
Figure 3: DiffWAM’s model architecture. Features from DiT blocks 15, 25, and 35 are projected to 512 dimensions, fused, and processed by a frozen factorized spatiotemporal module. Grid-Motion associates predictive features across time and with geometry-conditioned anchors before pooling them into 312 motion-memory tokens. A direct decoder with 38 pose queries and eight transformer blocks predicts 38 future poses from motion memory and first-frame geometry; together with the initial identity pose, these form a 39-pose trajectory. During offline supervision, π3 reconstructs camera motion and depth from completed video rollouts, and MoGe2 provides metric-depth information for translation-scale calibration.
Figure 4: Asynchronous trajectory preparation and scheduled handoff. While the UAV tracks its committed trajectory, prompt rewriting and LiDAR-based geometry preparation proceed in parallel. The predictive backbone, pose readout, planning, and validation produce a candidate update ready at rn . DiffWAM uses four backbone evaluations, whereas DiffWAM-Flash uses one; both terminate the final evaluation after DiT block 35 without future-video decoding. An early candidate waits until the scheduled handoff th,n and is activated only after its validity is rechecked. The remaining budget is Bn=en−un , with a fallback decision required no later than en−Rn if no valid update can be activated. The diagram illustrates an early-ready update; horizontal distances are schematic rather than measured durations.
Figure 5: Distribution of the evaluation tasks. The benchmark contains 18 task types grouped into four families: basic motion, object interaction, spatial navigation, and scene understanding. Percentages denote the fraction of the complete evaluation set.
Task family
Task type
Example of instructions
Basic Motion
Going Forward
Fly forward for 5 seconds
Translation
Move left without turning
Vertical Moving
Move upwards by 3 meters
Rotating
Rotate 45 degrees in place
Forward Turning
Move to the right and face that direction
Object Interaction
Object Navigation
Navigate to the black rock formation
Table 1: Composition of the evaluation benchmark. All task percentages are computed with respect to the complete evaluation set.
Figure 6: Representative trajectory-generation results of DiffWAM. Rows 1–4 illustrate target-directed navigation and approach; Rows 5–6 show constrained gap traversal; and Rows 7–9 show a full orbit and direction-conditioned half-orbits, including the excerpt in Row 8. Row 10 provides a supplemental generated example of descent onto a yellow landing pad. Each visual sequence is accompanied by a top-view trajectory. These selected examples support qualitative analysis rather than aggregate success-rate estimation.
IndoorUAV-VLA
UAV-FLOW-Sim
DiffWAM-1000
Method
Average
Average
BM
OI
SN
SU
Average
WorldVLN [ 32 ]
39.32
80.24
86.00
56.62
50.77
54.00
58.40
ImagineUAV [ 16 ]
33.78
69.65
68.00
39.85
32.67
56.00
43.20
Fast-WAM-UAV [ 30 ]
35.69
71.27
73.00
52.34
38.67
51.00
52.20
WorldFly [ 34 ]
25.71
53.98
64.00
24.08
28.62
41.00
30.50
DiffWAM (ours)
56.77
91.42
92.00
70.46
74.00
84.00
74.40
Table 2: Benchmark success rate (SR, %) on IndoorUAV-VLA, UAV-FLOW-Sim, and DiffWAM-1000. All reported entries use the same endpoint criterion; this endpoint-based SR is distinct from task-specific closed-loop completion. BM: basic motion; OI: object interaction; SN: spatial navigation; SU: scene understanding.
Figure 7: Real-world UAV platforms.
Figure 8: Representative real-world UAV experiments. The image sequences illustrate task-conditioned physical execution, including the transitions between successive tasks in Cases IV and IX.
Archived head
Platform / precision
Model P50/P95 (s)
Peak memory (GiB)
DiffWAM
2 × H20 / BF16
11.199 / 11.273
76.29 / 149.30
DiffWAM-Flash
2 × H20 / BF16
3.182 / 3.242
76.29 / 149.30
DiffWAM
8 × H20 / BF16
3.605 / 3.820
60.33 / 457.67
DiffWAM-Flash
8 × H20 / BF16
0.835 / 0.912
60.33 / 457.67
DiffWAM
AGX-Thor / BF16
3.796 / 3.851
97.60 / 97.60
DiffWAM-Flash
AGX-Thor / BF16
1.080 / 1.101
97.60 / 97.60
Table 3: Measured latency of archived FastDreamer pipeline variants. All configurations use BF16 inference. Memory is reported in GiB as maximum per-GPU/simultaneous aggregate usage.
Figure 9: Stage-wise latency breakdown of the compared inference pipelines. NavDreamer, pure video generation, DiffWAM, and DiffWAM-Flash are measured on eight NVIDIA H20 GPUs; the final row reports DiffWAM-Flash on NVIDIA Jetson AGX Thor. Bar-end values indicate total latency in milliseconds.
Method
RMSE (m)
ADE (m)
FDE (m)
RoE ( ∘ )
Head P50/P95 (ms)
WorldVLN [ 32 ]
0.8280
0.7103
1.2711
57.2106
3.14 / 3.34
Fast-WAM [ 30 ]
0.8818
0.6815
1.5395
6.1389
59.20 / 60.97
Faster-WAM [ 33 ]
0.7011
0.5749
1.0792
4.5870
60.37 / 61.77
MLP [ 18 ]
1.3881
1.0729
2.6189
98.8160
0.22 / 0.25
DiffWAM (ours)
0.3492
0.3151
0.4357
2.9179
37.35 / 38.84
Table 4: Comparison of video-based trajectory-prediction architectures.
Figure 10: Qualitative comparison of WAM trajectory-readout architectures. Each example shows the initial observation and the corresponding 3D trajectories.
Condition
DiT layers
Pose params
RMSE (m)
FDE (m)
RoE ( ∘ )
Native early
3/5/7
72M
0.7250
1.0036
9.2979
Native middle
23/25/27
72M
0.5397
0.7641
2.4314
Native deep
33/35/37
72M
0.3790
0.5562
2.5763
Native mixed
15/25/35
72M
0.3492
0.4357
2.9179
Table 5: Ablation of world-model feature-layer selection. Four native-grid configurations are compared using the same trajectory-decoder architecture and 72M pose-head parameters.
World-model evaluations
RMSE (m)
ADE (m)
FDE (m)
RoE ( ∘ )
1
0.5012
0.4423
0.6865
3.7897
2
0.4107
0.3601
0.5883
3.1867
3
0.3828
0.3347
0.5653
3.7730
4
0.3492
0.3151
0.4357
2.9179
Table 6: Ablation of predictive computation. One step denotes one world-model evaluation, not an iteration of the trajectory decoder.
Training-Data Volume
RMSE (m)
FDE (m)
RoE ( ∘ )
1,000
0.8203
1.4011
9.9409
4,000
0.6519
0.9897
7.2710
16,000
0.5465
0.7739
4.1872
64,000
0.4873
0.6884
2.7894
Table 7: Generated-supervision scaling on DiffWAM-1000. All models use random initialization and 20,000 training updates with a global batch size of 32.
Objective
RMSE (m)
FDE (m)
RoE ( ∘ )
Legacy composite pose loss
0.7288
1.0297
3.5997
Two-term, unnormalized
0.8120
1.1451
4.2803
Two-term, depth-normalized
1.0728
1.5092
3.9662
Selected depth-normalized recipe
0.6873
0.9496
3.8131
Table 8: Matched compact-loss experiments on DiffWAM-1000. The selected recipe uses validation-based hyperparameter tuning.
We present FlowPilot, a compact world-action model for real-time onboard UAV navigation from depth. Unlike map-then-optimize pipelines that require local reconstruction or end-to-end policies that lack explicit scene prediction, FlowPilot jointly denoises future depth observations and executable trajectories with flow matching. A dual-stream mixture-of-transformers couples video and action experts through shared attention, allowing future-scene prediction and trajectory generation to inform each other. At deployment, the model runs action-centrically and outputs only a trajectory. To ensure trackability, actions are parameterized as degree-7 Bernstein polynomials: the current state constrains the initial control points, and the network predicts five free control points, yielding C^2-continuous references with closed-form velocity, acceleration and jerk. FlowPilot is trained on a three-level depth pyramid spanning high-throughput simulation, photorealistic simulation, and real onboard data. In closed-loop simulation, it outperforms learning- and optimization-based baselines under increasing clutter and commanded speeds up to 8m/s. On a physical quadrotor, the full perception-to-action pipeline runs in under 18ms on a Jetson Orin NX and reaches 5.5m/s in cluttered indoor and forest environments using only onboard sensing and computation.
Image-goal visual navigation is a fundamental capability for embodied agents. Existing navigation policies efficiently predict waypoint trajectories but lack visual foresight, while navigation world models can anticipate future observations but often require costly planning rollouts. We present UniNav, a unified world-action model that generates future visual observations and continuous waypoint trajectories through a single diffusion process. Given history frames and a goal image, UniNav jointly denoises visual and waypoint tokens within a single transformer, unifying future prediction and action generation in a shared framework. To improve spatial grounding, we incorporate geometry-aware camera tokens. We also train on both trajectory-labeled navigation data and video-only data, enabling the model to benefit from diverse videos without waypoint annotations. Based on this unified framework, we introduce two variants: UniNav-Full jointly predicts interpretable future observations and their corresponding trajectories, while UniNav-Fast removes future-image tokens at inference for efficient trajectory prediction. Experiments on navigation benchmarks show that UniNav outperforms the strongest baseline in ATE across all datasets. With one-step inference, UniNav-Fast achieves a latency of 0.1s without a substantial accuracy drop. Code will be released.
Changqing Zhou, Yueru Luo, Zeyu Jiang +1
1The Hong Kong University of Science and Technology (Guangzhou) · 2The Chinese University of Hong Kong, Shenzhen
Goal-conditioned visual navigation requires a robot to act under partial observability by anticipating how its motion will change the future egocentric view and whether that change brings it closer to the goal. Navigation world models provide such visual foresight, but they remain prediction modules that require an external planner to convert predicted futures into closed-loop control. We propose Navigation World Action Model (NavWAM), a diffusion-transformer policy that turns navigation world-model prediction into executable action by representing future observations, goal-progress values, and action chunks in a shared latent sequence. By learning future prediction jointly with the action and value targets that determine closed-loop behavior, NavWAM makes visual foresight directly usable for robot control. We build NavWAM through simulation pretraining and real-robot adaptation, and evaluate it on image-goal navigation against planning-based world models and a representative direct navigation policy. Across offline benchmarks and closed-loop real-robot deployment, NavWAM improves over planning-based world-model baselines in our evaluations while using the default policy mode without CEM-style action search. Project page: https://dachii-azm.github.io/navwam/
Daichi Azuma, Taiki Miyanishi, Koya Sakamoto +6
The University of Tokyo · National Institute of Informatics · AIRoA +1