World Motion Models: Flexible Sequence Modeling of SE(3) Trajectories
Organizations: UC Berkeley · Harvard University
Abstract
Equipping artificial agents with spatial intelligence requires a comprehensive generative prior over the dynamic 3D world. We propose World Motion Models (WMMs) that capture "what was, is, and will be where across time" via sparse SE(3) pose trajectories. WMMs are built on the observation that elements of dynamic scenes can be well approximated by a set of rigid SE(3) trajectories, a minimal yet expressive primitive for 4D modeling. This representation unifies articulated objects, human bodies, hand-object interactions, piecewise-rigid scene dynamics, camera motion, and even robot states and actions into a single shared space. Given this representation, we cast the joint distribution of these entities as a flexible sequence modeling problem, utilizing flow-matching with per-token noise levels. Coupled with a context token mechanism for non-sequential conditioning, this formulation supports any-to-any marginal conditioning across an arbitrary number of entities and time steps. Tasks such as future prediction, motion infilling, model-predictive control, inverse kinematics, cross-embodiment retargeting, and policy learning all reduce to the application of different masks over the same network. Experiments on 6 diverse applications of 3D vision and robotics demonstrate the versatility and flexibility of WMMs with strong performance.
Figures & tables
| Variant | SR | Variant | MAE | MSE | Endp. |
|---|---|---|---|---|---|
| WMM | 0.860 | WMM | 2.167 | 0.329 | 0.580 |
| no AdaLN | 0.800 | no eff. | 2.356 | 0.344 | 0.602 |
| no per-tk | 0.000 | plain DiT | 32.05 | 16.97 | 17.64 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | SR-seq | gerr | lerr |
|---|---|---|---|
| Paper ( ) | 0.8690 | 0.5938 | 0.0420 |
| 0.8645 | 0.6440 | 0.0449 | |
| 0.8030 | 0.7709 | 0.0504 |
| Paper | noFsem | noFgeo | noBoth | |
| Avg Succ Rate | 0.8720 | 0.8560 | 0.1960 | 0.1680 |
| EPIC | DROID | |||||
|---|---|---|---|---|---|---|
| Setting | MAE | MSE | Endpoint | MAE | MSE | Endpoint |
| Paper | 2.2564 | 0.3458 | 0.6146 | 1.1718 | 0.1665 | 0.2253 |
| noFsem | 2.4032 | 0.3627 | 0.6148 | 1.2894 | 0.1876 | 0.2539 |
| Bone length error | Bone direction error | Raw vs. projected MPJPE |
|---|---|---|
| mm | mm |
| Off-hinge-axis rotation | Off-hinge translation | Raw vs. projected MPJPE |
|---|---|---|
| mm | mm |
| Steps | Time, VRAM |
|---|---|
| Calibration : resize and run GeoCalib on key frames to undistort intrinsics and estimate the gravity/vertical direction. | 4.3 s / 5,856 MB |
| Geometry : run chunked slice-and-align DA3 (to keep VRAM bounded on long sequences) to estimate depth and camera parameters; also run DA3-metric on key frames to compute the metric scale. | 34.6 s / 14,137 MB |
| Tracking : run 2D tracking using TAPNext and multiple rounds of motion-aware resampling (to keep VRAM bounded and focus more on moving areas), then lift the 2D tracks to 3D at visible time steps using the camera parameters and depth, with depth-boundary and anti-jitter filtering. | 64.6 s / 8,151 MB |
| [OPT] Full VOS (not used in the paper) : our pipeline also runs DEVA in auto-mode to segment and track anything in the full video. | 27.2 s / 14,079 MB |
| Total cost | 131 s / 14.1 GB |