Pretrained video foundation models encode rich semantic and spatiotemporal priors for embodied navigation, yet converting these priors into UAV motion typically requires expensive future-video synthesis and geometric reconstruction. We investigate whether the motion implicit in future visual prediction can instead be recovered directly from the predictive representations of a frozen video model. To this end, we present DiffWAM, a geometry-conditioned navigation world-action model that directly transforms multi-level predictive features into continuous camera trajectories. Its Grid-Motion module preserves spatial-temporal motion associations, while Latent2Pose grounds them with first-frame geometry to recover metrically meaningful 3D motion. Complete video rollouts and geometric reconstruction are required only for offline supervision, eliminating future-video decoding and multi-frame reconstruction during deployment. We further introduce FastDreamer, which overlaps predictive and geometric computation with ongoing flight and performs timestamp-aware asynchronous trajectory handoff for continuous UAV execution. DiffWAM achieves a trajectory RMSE of 0.3492 m and an endpoint success rate of 74.40% on the 1,000-sample DiffWAM-1000 benchmark, while representative real-world experiments demonstrate complex behaviors including constrained traversal, orbiting, S-shaped flight, and multi-stage navigation. An onboard DiffWAM-Flash implementation further reaches 1.08 s model-pipeline latency on NVIDIA Jetson AGX Thor. These results demonstrate that predictive video representations can be efficiently grounded into continuous 3D motion, providing a direct alternative to generate-then-reconstruct navigation pipelines. Project page: https://zzmmzzm.github.io/diffwam.github.io/.
Figures & tables
Figure 1: Overview of DiffWAM. DiffWAM grounds multi-level predictive representations from a frozen video world model, together with first-frame geometry, into continuous 3D UAV trajectories. Trained with large-scale synthetic navigation data, it supports diverse structured motions, real-world execution, and efficient deployment without online future-video decoding.
Figure 2: DiffWAM’s pipeline. DiffWAM converts selected early predictive features and first-frame geometry into camera-motion trajectories through Grid-Motion and Latent2Pose. Training proceeds through video-backbone preparation (S0), navigation-aware readout pretraining (S1), and geometry-guided predictive distillation (S2), where completed video rollouts and geometric reconstruction provide offline trajectory supervision. Detailed tensor and output dimensions are specified in Fig. 3 .
Figure 3: DiffWAM’s model architecture. Features from DiT blocks 15, 25, and 35 are projected to 512 dimensions, fused, and processed by a frozen factorized spatiotemporal module. Grid-Motion associates predictive features across time and with geometry-conditioned anchors before pooling them into 312 motion-memory tokens. A direct decoder with 38 pose queries and eight transformer blocks predicts 38 future poses from motion memory and first-frame geometry; together with the initial identity pose, these form a 39-pose trajectory. During offline supervision, π3 reconstructs camera motion and depth from completed video rollouts, and MoGe2 provides metric-depth information for translation-scale calibration.
Figure 4: Asynchronous trajectory preparation and scheduled handoff. While the UAV tracks its committed trajectory, prompt rewriting and LiDAR-based geometry preparation proceed in parallel. The predictive backbone, pose readout, planning, and validation produce a candidate update ready at rn . DiffWAM uses four backbone evaluations, whereas DiffWAM-Flash uses one; both terminate the final evaluation after DiT block 35 without future-video decoding. An early candidate waits until the scheduled handoff th,n and is activated only after its validity is rechecked. The remaining budget is Bn=en−un , with a fallback decision required no later than en−Rn if no valid update can be activated. The diagram illustrates an early-ready update; horizontal distances are schematic rather than measured durations.
Figure 5: Distribution of the evaluation tasks. The benchmark contains 18 task types grouped into four families: basic motion, object interaction, spatial navigation, and scene understanding. Percentages denote the fraction of the complete evaluation set.
Task family
Task type
Example of instructions
Basic Motion
Going Forward
Fly forward for 5 seconds
Translation
Move left without turning
Vertical Moving
Move upwards by 3 meters
Rotating
Rotate 45 degrees in place
Forward Turning
Move to the right and face that direction
Object Interaction
Object Navigation
Navigate to the black rock formation
Table 1: Composition of the evaluation benchmark. All task percentages are computed with respect to the complete evaluation set.
Figure 6: Representative trajectory-generation results of DiffWAM. Rows 1–4 illustrate target-directed navigation and approach; Rows 5–6 show constrained gap traversal; and Rows 7–9 show a full orbit and direction-conditioned half-orbits, including the excerpt in Row 8. Row 10 provides a supplemental generated example of descent onto a yellow landing pad. Each visual sequence is accompanied by a top-view trajectory. These selected examples support qualitative analysis rather than aggregate success-rate estimation.
IndoorUAV-VLA
UAV-FLOW-Sim
DiffWAM-1000
Method
Average
Average
BM
OI
SN
SU
Average
WorldVLN [ 32 ]
39.32
80.24
86.00
56.62
50.77
54.00
58.40
ImagineUAV [ 16 ]
33.78
69.65
68.00
39.85
32.67
56.00
43.20
Fast-WAM-UAV [ 30 ]
35.69
71.27
73.00
52.34
38.67
51.00
52.20
WorldFly [ 34 ]
25.71
53.98
64.00
24.08
28.62
41.00
30.50
DiffWAM (ours)
56.77
91.42
92.00
70.46
74.00
84.00
74.40
Table 2: Benchmark success rate (SR, %) on IndoorUAV-VLA, UAV-FLOW-Sim, and DiffWAM-1000. All reported entries use the same endpoint criterion; this endpoint-based SR is distinct from task-specific closed-loop completion. BM: basic motion; OI: object interaction; SN: spatial navigation; SU: scene understanding.
Figure 7: Real-world UAV platforms.
Figure 8: Representative real-world UAV experiments. The image sequences illustrate task-conditioned physical execution, including the transitions between successive tasks in Cases IV and IX.
Archived head
Platform / precision
Model P50/P95 (s)
Peak memory (GiB)
DiffWAM
2 × H20 / BF16
11.199 / 11.273
76.29 / 149.30
DiffWAM-Flash
2 × H20 / BF16
3.182 / 3.242
76.29 / 149.30
DiffWAM
8 × H20 / BF16
3.605 / 3.820
60.33 / 457.67
DiffWAM-Flash
8 × H20 / BF16
0.835 / 0.912
60.33 / 457.67
DiffWAM
AGX-Thor / BF16
3.796 / 3.851
97.60 / 97.60
DiffWAM-Flash
AGX-Thor / BF16
1.080 / 1.101
97.60 / 97.60
Table 3: Measured latency of archived FastDreamer pipeline variants. All configurations use BF16 inference. Memory is reported in GiB as maximum per-GPU/simultaneous aggregate usage.
Figure 9: Stage-wise latency breakdown of the compared inference pipelines. NavDreamer, pure video generation, DiffWAM, and DiffWAM-Flash are measured on eight NVIDIA H20 GPUs; the final row reports DiffWAM-Flash on NVIDIA Jetson AGX Thor. Bar-end values indicate total latency in milliseconds.
Method
RMSE (m)
ADE (m)
FDE (m)
RoE ( ∘ )
Head P50/P95 (ms)
WorldVLN [ 32 ]
0.8280
0.7103
1.2711
57.2106
3.14 / 3.34
Fast-WAM [ 30 ]
0.8818
0.6815
1.5395
6.1389
59.20 / 60.97
Faster-WAM [ 33 ]
0.7011
0.5749
1.0792
4.5870
60.37 / 61.77
MLP [ 18 ]
1.3881
1.0729
2.6189
98.8160
0.22 / 0.25
DiffWAM (ours)
0.3492
0.3151
0.4357
2.9179
37.35 / 38.84
Table 4: Comparison of video-based trajectory-prediction architectures.
Figure 10: Qualitative comparison of WAM trajectory-readout architectures. Each example shows the initial observation and the corresponding 3D trajectories.
Condition
DiT layers
Pose params
RMSE (m)
FDE (m)
RoE ( ∘ )
Native early
3/5/7
72M
0.7250
1.0036
9.2979
Native middle
23/25/27
72M
0.5397
0.7641
2.4314
Native deep
33/35/37
72M
0.3790
0.5562
2.5763
Native mixed
15/25/35
72M
0.3492
0.4357
2.9179
Table 5: Ablation of world-model feature-layer selection. Four native-grid configurations are compared using the same trajectory-decoder architecture and 72M pose-head parameters.
World-model evaluations
RMSE (m)
ADE (m)
FDE (m)
RoE ( ∘ )
1
0.5012
0.4423
0.6865
3.7897
2
0.4107
0.3601
0.5883
3.1867
3
0.3828
0.3347
0.5653
3.7730
4
0.3492
0.3151
0.4357
2.9179
Table 6: Ablation of predictive computation. One step denotes one world-model evaluation, not an iteration of the trajectory decoder.
Training-Data Volume
RMSE (m)
FDE (m)
RoE ( ∘ )
1,000
0.8203
1.4011
9.9409
4,000
0.6519
0.9897
7.2710
16,000
0.5465
0.7739
4.1872
64,000
0.4873
0.6884
2.7894
Table 7: Generated-supervision scaling on DiffWAM-1000. All models use random initialization and 20,000 training updates with a global batch size of 32.
Objective
RMSE (m)
FDE (m)
RoE ( ∘ )
Legacy composite pose loss
0.7288
1.0297
3.5997
Two-term, unnormalized
0.8120
1.1451
4.2803
Two-term, depth-normalized
1.0728
1.5092
3.9662
Selected depth-normalized recipe
0.6873
0.9496
3.8131
Table 8: Matched compact-loss experiments on DiffWAM-1000. The selected recipe uses validation-based hyperparameter tuning.