World models let robots imagine possible futures, but exploiting this capability for real-time planning is bottlenecked by a representation misalignment: generative models and planners operate on decoupled manifolds, requiring computationally expensive decoding of every candidate back to the high-dimensional observation space for evaluation. In this paper, we present Hydra, a discrete World Action Model that tackles this by establishing a unified latent manifold over visual states, physical poses, and control actions. By compressing this manifold through modality-specific Vector-Quantized bottlenecks, Hydra yields discrete vocabularies of kinodynamic intents and visual states. This enables Discrete Latent Planning (DLP), where candidates are sampled directly from the shared manifold and ranked by a Kinematic-Perceptual Cost within the discrete latent space. To bridge discrete planning with the continuous commands required for physical actuation, Hydra pairs DLP with conditional Flow Matching to map selected intents to smooth execution trajectories. Evaluated on two physical robotic platforms, Hydra outperforms state-of-the-art navigation world models in goal-directed planning, while matching or exceeding the closed-loop execution capabilities of leading reactive navigation policies.
Figures & tables
Figure 1: Moving the planner into the model’s own latent space. Top: NWM uses an external planner to sample unconstrained trajectories, decoding them to pixels for evaluation ( red ), wasting compute on infeasible candidates. Bottom: Hydra plans directly within its discrete latent manifold, then executes via continuous Flow Matching, achieving >500× faster planning on a robot.
Figure 2: Comparison of Hydra architecture with Visual Navigation Transformer (ViNT), Navigation World Model (NWM), and VertiFormer in terms of input modalities, conditioning, and the model output. Hydra builds on top of VertiFormer and discretizes the output into codebooks to enable the model to search the discrete latent space for a plausible path.
Figure 3: Initial intended future pose/action tokens are generated unconditionally, then Gaussian noise is added to become the exploratory seed ‘intents’. Then these future intents go through VQ codebooks, and we sample the top- k codes. After calculating the KPC cost, the lowest cost is chosen to seed the next iteration of searching. The best trajectory is decoded using flow matching heads to continuous robot poses and actions at the end for robot execution.
Model
FVD ↓
Time To Generate
NWM/B
22.10 ± 10.71
3.5h
VertiFormer
81.03 ± 12.50
2m
Hydra
19.15 ± 5.11
4m
Table 1: Comparison of Video Synthesis Quality. 100 clips of 16-seconds generated at 4 FPS.
Model
Params
Search Space
Iterations
Samples
Len.
Plan Time ↓
Unobs. (SR) ↑
Obs. (SR) ↑
Corner Turn (SR) ↑
NWM/B
200M
Continuous (CEM)
3
18
12
> 500.0s
0
0
0
VertiFormer
27.11M
Continuous (MPPI)
1
18
12
∼ 0.7s
40
0
0
Hydra-NoVQ
141.60M
Continuous (MPPI)
1
18
12
∼ 1.2s
20
10
10
Hydra
143.29M
Discrete (DLP)
3
18
12
∼ 0.9s
100
80
80
Table 2: Planning Performance. While VertiFormer with MPPI runs quickly, it suffers from sample inefficiency and is blind to obstacles (0% SR), as is the continuous variant, Hydra-NoVQ. Hydra ’s DLP achieves high success rates while maintaining real-time execution speeds compared to NWM.
Dense Goal Images (2Hz)
Sparse Waypoints (turns only)
Scenario
Metric
GNM
ViNT
NoMaD
Hydra (CFG)
Hydra (Sampling)
VertiFormer
Hydra (Planning)
Obstructed Goal
Success Rate (%) ↑
90
80
100
90
50
100
90
Traversal Time (s) ↓
36.4 ± 7.9
34.4 ± 6.8
37.4 ± 7.7
45.0 ± 8.0
48.4 ± 18.5
21.0 ± 1.05
37.2 ± 21.0
Interventions ↓
1.1 ± 0.3
1.2 ± 0.6
1.1 ± 0.3
1.0 ± 0.7
0.8 ± 0.8
1.0 ± 0.5
1.2 ± 0.6
Long Distance
Success Rate (%) ↑
0
0
0
100
100
100
40
Traversal Time (s) ↓
-
-
-
138 ± 7.7
107 ± 8.6
109 ± 7.6
119 ± 3.8
Table 3: Local Planning Evaluation across 10 Physical Trials per Scenario. Hydra drastically reduces physical interventions by leveraging its predictive foresight to move around topological traps that ensnare reactive baselines. GNM, ViNT, and NoMaD are conditioned on dense goal-image sequences at 2Hz; Hydra ’s local-following variants (CFG, Sampling) are conditioned on sparse, more widely spaced waypoints.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
\RowStyle Method
Obs. Input
Goal Input
Cond. Input
Output
Nav. Policy
GNM [ 58 ]
I
I
I
P
Reactive
VINT [ 59 ]
I
I
I
P
Reactive
NOMAD [ 62 ]
I
I
I
P
Reactive
VertiFormer [ 46 ]
I, A, P
I, P
A or P
I, A, P
Planning
NWM [ 6 ]
I
I
P
I
Planning
Hydra
I, A, P
I, P
A or P
I, A, P
Planning
Appendix
Table 4: Comparison with visually-grounded robot navigation baselines. I : Image, P : Pose, A : Action.
Model
Params
DreamSim ↓
LPIPS ↓
PSNR ↑
NWM/B
200M
0.2480 ± 0.1265
0.5168 ± 0.1147
10.69
VertiFormer
27.11M
0.3738 ± 0.0560
0.5417 ± 0.0757
18.29
Hydra
143.29M
0.1103 ± 0.0345
0.2809 ± 0.0679
16.20
Appendix
Table 5: Offline single-step image quality on a held-out subset of the RECON dataset. Hydra achieves competitive accuracy with fewer parameters and NFEs.
Model
Params
SCAND
TartanDrive
SACSoN
Mean
ADE ↓
FDE ↓
ADE ↓
FDE ↓
ADE ↓
FDE ↓
ADE ↓
FDE ↓
GNM
8.6M
0.16
0.27
0.67
1.12
0.13
0.21
0.32
0.53
ViNT
29.6M
0.15
0.25
0.80
1.33
0.16
0.26
0.37
0.61
NoMaD
19.0M
0.24
0.43
1.311
2.35
0.22
0.39
0.59
1.06
VertiFormer
27.11M
0.10
0.18
0.43
0.80
0.10
0.19
0.21
0.39
Hydra
143.29M
0.12
0.20
0.48
0.85
0.12
0.21
0.24
0.42
Appendix
Table 6: Offline trajectory forecasting evaluation on a held-out subset across three navigation datasets (SCAND, TartanDrive, and SACSoN). Hydra achieves competitive policy accuracy while retaining the planning capability that reactive baselines lack.
Figure 4: Comparison of a 16-second video generated from a held-out subset of RECON. While NWM’s continuous latents drift from the commanded actions over time, Hydra ’s discrete manifold maintains action adherence over time.
Scenario
SR (%) ↑
Time (s) ↓
Frontal, λimg=1
80
24.7 ± 5.8
Blind corner, λimg=1
20
30.6 ± 12.1
Blind corner, λimg=0
70
34.5 ± 8.9
Appendix
Table 7: Effect of the image-goal cost term ( λimg ) conditioned on goal visibility. The same term that improves performance when the goal is visible degrades it when the goal is not, motivating conditioning Cimg on goal visibility rather than applying it unconditionally.
Spot
Jackal
Configuration
Time (s) ↓
Collisions ↓
Time (s) ↓
Collisions ↓
Original (avoidance costs + Cgeo )
30.6 ± 12.1
0.20 ± 0.42
23.8 ± 0.39
0.20 ± 0.13
λprior=0
33.9 ± 13.0
0.70 ± 0.48
21.2 ± 0.26
0.70 ± 0.59
λvq=0
30.3 ± 10.0
0.50 ± 0.53
20.9 ± 0.34
0.60 ± 0.28
λentropy=0
36.2 ± 18.9
0.40 ± 0.52
26.6 ± 0.45
0.80 ± 0.61
Appendix
Table 8: Ablation of obstacle-avoidance cost terms ( λprior , λvq , λentropy ), across two robot embodiments: a blind corner turn (Spot) and frontal obstacle avoidance (Jackal). Each avoidance term degrades performance when removed on both platforms, indicating the mechanism is embodiment-agnostic.
Figure 5: 1-step spatial coverage of the learned pose codebook Vp . Unlike normalized baselines, the discrete manifold explicitly captures the absolute physical scaling of different robotic platforms, ranging from low-speed indoor maneuvers (dense central clusters) to high-speed off-road dynamics (sparse outer bounds).
Figure 6: Live RViz visualization of Hydra ’s Discrete Latent Planning search on the physical robot in a never-before-seen environment. Left: the robot’s current ego-centric observation. Middle: a subset of Hydra ’s predicted future frames. Right: the corresponding candidate trajectories ( green ) rendered as a branching tree in the robot’s local frame; each branch is one topological maneuver considered before JKPC selects the lowest-cost ( purple ) candidate for execution.
Figure 7: Manifold Rejection and Predictive Entropy During Impending Collision. Top / Bottom insets: Comparison of Ground Truth vs. Hydra ’s imagination at points along a rollout commanded to drive straight toward a physical occlusion. Both the spatially-averaged VQ quantization error ( Cvq , solid blue, left axis) and the predictive entropy of the visual token distribution ( Cent , dashed orange, right axis) rise together as the commanded trajectory approaches the occlusion, and drop together immediately after the model hallucinates a navigable resolution, two independently computed signals converging on the same collision-relevant moments. The smaller early fluctuation (first two insets) is a false positive triggered by visually complex but traversable foliage, not a genuine obstacle.
Figure 8: Vanishing humans in future predictions on SCAND. Top: ground-truth future frames, in which a pedestrian remains visible throughout. Bottom: Hydra ’s predicted frames for the same sequence, in which the pedestrian is smoothed into the background.
Figure 9: Goal-blind seed initialization causes a missed turn. The green trajectory is the path selected by the DLP cost among the sampled candidates at convergence; the seed distribution that generated the candidates was conditioned only on local scene context, not on the goal, delaying goal-aware correction until it is too late to complete the turn.
Figure 10: LPIPS deviation from geometric grounding. Note that these are not generated images.
Figure 11: Physical experiment setup with obstacles.
\RowStyle Parameter
Value
\Block [l]1-2 Training & Optimization
Optimizer
AdamW
Learning Rate
2×10−4
Weight Decay
0.08
Batch Size
8
Gradient Accumulation Steps
32
Appendix
Table 9: Hyperparameters and architectural details of Hydra . The model utilizes a Transformer backbone combined with dimensionally constrained, modality-specific VQ codebooks and Flow Matching decoders.
Figure 12: Early Training Dynamics (2.1K Steps). A comparison between the model’s early-stage generation, both with (middle) and without (bottom) the deterministic anchor head. While the bottom row suffers from spatial smearing, the model utilizing the deterministic head shows a better ability to construct structural details, like the building’s edges against the sky, even after very limited training.
School of Computer Science and Technology, University of Chinese Academy of Sciences, Beijing · Institute of Information Engineering, Chinese Academy of Sciences, Beijing · School of Computer Science and Technology, Harbin Institute of Technology, Weihai +2