CST-WM: A Causally Structured World Model for Embodied Visual Tracking
Organizations: New York University Abu Dhabi
Abstract
Embodied visual tracking requires a robot to choose actions that keep a moving target observable at a suitable distance, and to recover it after occlusion, out-of-view drift, or distractor crossings. We cast the task as planning over future target evidence with an action-conditioned world model. In logged tracking data, however, the behavior policy's actions are correlated with where the target is, so a generic predictor can learn a shortcut: it writes the current action directly into its prediction of target evidence, instead of letting the action affect that evidence only by moving the robot and changing what it observes. We call this failure causal hallucination; the resulting rollouts look plausible but rank candidate actions for the wrong reason. We propose CST-WM, a causally structured world model whose state is split into target-evidence, robot, and observation branches. Its transition removes the same-step edge from action to target evidence but keeps the path through robot motion and the resulting views, so candidate actions are still distinguished by their predicted ego-motion. With rollout-based model-predictive control, a single model handles both steady following and re-acquisition after target loss. On EVT-Bench and Habitat 3.0, covering standard tracking, target-loss recovery, and cross-dataset transfer, CST-WM improves following, distance-range control, safety, and re-acquisition over reactive trackers and world-model baselines, and removing the action mask causes the largest drop in re-acquisition among our ablations. Offline, CST-WM has lower multi-step rollout error, and its ranking of candidate actions agrees better with the simulator's. On a Unitree Go2 quadruped, CST-WM succeeds in 20 of 30 real-world trials under occlusion, distractor crossing, and fast motion, against 14 for TrackVLA.
Figures & tables
| EVT-Bench (STT) | Habitat 3.0, standard tracking | Cross-dataset transfer to Habitat 3.0 | |||||||||||
| Method | SR | TR | CR | Method | F | DRS | CR | ES | Method | F | DRS | CR | ES |
| Uni-NaVid | 25.7 | 39.5 | 41.9 | Habitat 3.0 baseline | 0.29 | 0.47 | 0.48 | 0.40 | Uni-NaVid | 0.40 | 0.56 | 0.39 | 0.45 |
| OA-VAT † | 32.4 | 63.9 | 3.70 | Uni-NaVid | 0.34 | 0.52 | 0.43 | 0.47 | TrackVLA | 0.38 | 0.54 | 0.41 | 0.43 |
| TrackVLA | 85.1 | 78.6 | 1.65 | Adapted NWM | 0.41 | 0.61 | 0.33 | 0.49 | Adapted NWM | 0.43 | 0.58 | 0.35 | 0.47 |
| TrackVLA++ | 86.0 | 81.0 | 2.10 | SDA-S2 | 0.39 | 0.63 | 0.57 | 0.43 | TrackVLA++ | 0.42 | 0.61 | 0.38 | 0.43 |
| Ours | 88.7 | 83.4 | 1.41 | Ours | 0.53 | 0.70 | 0.27 | 0.61 | Ours | 0.48 | 0.65 | 0.30 | 0.54 |
| Short occlusion | Long occlusion | Out-of-view drift | Distractor crossing | |||||
|---|---|---|---|---|---|---|---|---|
| Method | Re-acq. | TTR | Re-acq. | TTR | Re-acq. | TTR | Re-acq. | TTR |
| TrackVLA | 0.71 | 8.9 | 0.48 | 14.6 | 0.52 | 12.8 | 0.43 | 15.2 |
| Adapted NWM | 0.75 | 8.1 | 0.54 | 13.2 | 0.57 | 11.7 | 0.49 | 14.0 |
| Ours | 0.84 | 6.4 | 0.69 | 9.8 | 0.73 | 8.7 | 0.65 | 10.6 |
| Method | Post-F | Rec-ES | Post-F | Rec-ES | Post-F | Rec-ES | Post-F | Rec-ES |
| TrackVLA | 0.58 | 0.47 | 0.41 | 0.29 | 0.45 | 0.33 | 0.37 | 0.24 |
| Target evidence | F | DRS | Re-acq. | CR | Plan. | Time (s) | |
| Privileged | Simulator true distance | 0.54 | 0.72 | 0.78 | 0.26 | – | – |
| Detector on rendered future views | 0.48 | 0.67 | 0.64 | 0.24 | 0.74 | 5.2 | |
| Read-out | Detector on decoded views | 0.38 | 0.46 | 0.46 | 0.40 | 0.58 | 5.8 |
| Separate latent head on | 0.45 | 0.60 | 0.57 | 0.33 | 0.54 | 0.030 | |
| Head, frozen transition | 0.34 | 0.37 | 0.43 | 0.39 | 0.62 | 0.021 | |
| Signal | Raw detector score | 0.46 | 0.60 | 0.66 | 0.32 | – | – |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Item | Details |
|---|---|
| Data | EVT-Bench / Habitat 3.0; scene-disjoint splits. Train scenes: 64 / 72; val: 8 / 9; test: 8 / 9. Trajectories: 48,000 / 61,000. One-step clips: 1.92M / 2.44M. Filtering removes corrupted frames, invalid robot states, reset fragments, terminal-only collision fragments, and action-invalid transitions. |
| Observation and action | Observations sampled at the simulator control frequency; consecutive frames without temporal skipping; actions at the same rate. |
| Model | VAE ( Blattmann et al., 2023 ) (sd-vae-ft-ema; 34.2M-parameter encoder, 49.5M-parameter decoder), views, latent . Target-evidence branch input: the scalar (the evidence head outputs the two components); conditioning on the current step only: one observation latent, no history; robot state: planar position and yaw; action . Transition: 12 causal layers, each with target-evidence, robot (a pose sub-layer and an action sub-layer, giving ), and observation sub-layers (attention + FFN), hidden 512, 8 heads, FFN 2048; 153.6M parameters, 89 GFLOPs per denoising pass. Evidence: GroundingDINO Swin-T (172.8M parameters), prompt “humanoid”, views resized to , on the current observation and for training labels; evidence head (0.26M parameters) on imagined steps. |
| Diffusion training | Linear DDPM with fixed variance, 1000 train steps, -prediction on the next observation latent, 20 DDIM inference steps; AdamW, lr , no weight decay, batch 16, 150 epochs, grad-norm clip 10, EMA 0.9999, bf16 mixed precision; 3 seeds; evidence-head loss weight (joint training from the 12k-step transition); auxiliary distance-aware loss weight . |
| Planning | , , , , 20 DDIM steps; ; ; m; one stochastic rollout per candidate; target evidence read by at every imagined step. |
| Compute | Training: a single GPU per model, bf16 mixed precision, deterministic evaluation seeds. Runtime: measured on one RTX A6000, and on an RTX 6000 Ada for the NWM-protocol trajectory times and the robot-budget planning step; per-component times, memory, and the NWM-protocol trajectory times are in Appendix D.4 and Table 17 . Random seeds per main model: 3. Evaluation episodes per model: 500 starting states 32 candidate sequences for offline diagnostics, plus standard EVT-Bench / Habitat 3.0 splits for downstream metrics. |
| Method | Type | Target evidence | Action selection | Model and compute |
|---|---|---|---|---|
| Uni-NaVid, TrackVLA, TrackVLA++ | reactive VLA | learned, from RGB and instruction | one policy pass per step | own architecture and training data |
| OA-VAT † | reactive tracker | own detector, tracker, and re-identification | PID control law | released weights, no planner |
| Habitat 3.0 baseline, SDA-S2 | reactive | own perception | one policy pass per step | own architecture and training data |
| Adapted NWM | world model, CEM | GroundingDINO on current view, evidence head on imagined views | Eq. ( 5 ), same budget | CDiT-XL/2, 1.01B, 321 GFLOPs per pass; same as ours otherwise |
| Ablations (Table 7 ) | world model, CEM | GroundingDINO on current view, evidence head on imagined views | Eq. ( 5 ), same budget | full model with one component changed |
| CST-WM (ours) | world model, CEM | GroundingDINO on current view, evidence head on imagined views | Eq. ( 5 ) | causal transition, 0.15B, 89 GFLOPs per pass |
| Method | All states | In-view | Out-of-view | Corr. with performance |
|---|---|---|---|---|
| Adapted NWM | 0.112 | 0.096 | 0.131 | |
| Leaky variant B | 0.331 | 0.304 | 0.357 | |
| Leaky variant C | 0.086 | 0.071 | 0.102 | |
| Ours | 0.001 | 0.001 | 0.002 |
| Stage | Default | Fast | Robot | Robot, Ada |
| Budget | (10,128,4,20) | (6,64,2,10) | (4,24,2,10) | (4,24,2,10) |
| GroundingDINO, current view | 0.09 s | 0.09 s | 0.09 s | – |
| VAE encoder, current frame | 7 ms | 7 ms | 7 ms | – |
| Denoising passes (sequential, batch ) | 800 | 120 | 80 | 80 |
| Evidence-head reads (batch ) | 40 | 12 | 8 | 8 |
| Time per action (whole step) | 267 s | 22.5 s | 8.2 s | 1.4 s |
| Model, GPU | Base | +Skip | +6 steps † | +4-bit ∗ |
|---|---|---|---|---|
| NWM, Ada (reported) | 30.3 | 14.7 | 0.4 | 0.1 |
| NWM, A6000 | 166.2 | 84.0 | 2.30 | 0.57 |
| Ours, A6000 | 135.6 | 67.7 | 1.88 | 0.47 |
| Ours, Ada | 23.4 | 12.3 | 0.34 | 0.08 |
| Component | Size | Batch | Time | Peak memory |
|---|---|---|---|---|
| VAE encoder, 1 / 4 frames | 34.2M | 1 | 7 / 23 ms | – |
| Ours, one denoising pass | 153.6M, 89 GFLOPs | 1 / 24 / 64 | 31 / 106 / 198 ms | 1.2 / 1.6 / 2.4 GB |
| 128 / 192 | 362 / 572 ms | 3.6 / 4.8 GB | ||
| NWM CDiT-XL/2, one denoising pass | 1.01B, 321 GFLOPs | 1 / 128 | 39 / 1017 ms | – |
| Evidence head | 0.26M | 24 / 64 / 128 / 192 | 0.21 / 0.20 / 0.20 / 0.27 ms | – |
| VAE decoder, per view | 49.5M | 1 / 64 | 15.6 / 10.6 ms | – |