UltraWorld: Learning Interactive Ultrasound World Models from Untracked Clinical Videos with Acoustic Sampling Map
Organizations: The Chinese University of Hong Kong · The Eighth Affiliated Hospital, Sun Yat-sen University
Abstract
World models can enable autonomous ultrasound scanning by predicting the outcomes of probe motions from local observations. Learning this action--observation relationship typically relies on synchronized video--pose pairs, which are costly to collect at scale and largely unavailable in routine clinical recordings. Reliable action following further requires modeling ultrasound's cross-sectional sampling geometry. We present UltraWorld, a self-distillation recipe that transfers priors from clinical ultrasound videos into interactive world models without real action annotations. Starting from clinical videos, we adapt a video foundation model into an ultrasound generator conditioned on reference images and anatomical masks. Anatomical masks sampled along programmable trajectories through 3D anatomy provide spatial guidance for synthesizing action--video pairs. We then use these synthetic pairs to self-distill the generator into a world model that predicts future observations from local observations and actions, without requiring anatomical masks or other 3D assets at inference time. To further improve action following, we introduce the Acoustic Sampling Map (AsMap), which represents probe poses and imaging settings as pixel-wise 3D sampling positions, beam directions, and depths. Experiments demonstrate improved prediction fidelity and action following. Across nine simulated closed-loop local planning episodes, UltraWorld reduces the mean final distance to the goal and orientation error by 29% and 38%, respectively, compared with visual servoing. Project Page: https://ultraworld-project.github.io/.
Figures & tables
| Conditioning | Prediction fidelity | Action following | |||
|---|---|---|---|---|---|
| a PSNR | LPIPS | Latent | -LPIPS | Motion err. | |
| Fourier Enc. ( Fan et al., 2026 ) | 22.946 | 0.274 | 0.120 | 0.509 | 4.540 |
| MLP Enc. ( Yue et al., 2025 ) | 23.031 | 0.262 | 0.121 | 0.486 | 4.471 |
| PRoPE ( Li et al., 2025 ) | 22.634 | 0.275 | 0.127 | 0.504 | 4.457 |
| AsMap (Ours) | 24.994 | 0.236 | 0.095 | 0.466 | 4.155 |
| Variant | Prediction fidelity | Action following | |||
|---|---|---|---|---|---|
| PSNR | LPIPS | Latent | -LPIPS | Motion err. | |
| Matched MLP conditioning | 23.402 | 0.258 | 0.113 | 0.482 | 4.429 |
| Wan2.1 initialization | 24.640 | 0.235 | 0.099 | 0.472 | 4.414 |
| Full model (Ours) | 24.994 | 0.236 | 0.095 | 0.466 | 4.155 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Encoder | Architecture | Parameters |
|---|---|---|
| Fourier Enc. | Fourier features + MLP | 1.11M |
| MLP Enc. | MLP | 0.79M |
| MLP Enc. (Matched) | Matched-size MLP | 53.49M |
| PRoPE | 30-layer Linear | 70.82M |
| AsMap (Ours) | Convolution + ResBlock | 53.48M |