We present Underwater C3-JEPA (cross-view, control-conditioned, context-extended), an object-centric multi-view predictive world model for near-field heavy-load underwater ROV salvage. Without contact sensors, it predicts in latent space how the task-object state evolves through contact interaction and under the hydrodynamic lag of the vehicle, from synchronized multi-view RGB observations and vehicle control signals. C3-JEPA encodes multi-camera observations into task-object and context tokens, fuses cross-camera evidence through held-out-view attention, and directly predicts future states conditioned on control. Weak binding anchors the target and gripper at low annotation cost, while SIGReg sharpens the geometric representation. Experiments show that the learned representation transfers substantially more task-relevant information to downstream probes than a reconstruction-free latent baseline, while keeping the predictor lightweight. The resulting predictive interface supports model-predictive-control (MPC) candidate evaluation and imagined-rollout behavior-agent training. Validation on real underwater video shows the same architecture recovering a withheld camera's object state and staying ahead of persistence, so the recipe transfers beyond simulation.
Figures & tables
Fig. 1: The heavy-duty underwater salvage ROV used in this study: (a) isometric view; (b) grasping a target object (CAD rendering); (c) lifting a target object in a real-environment trial.
Fig. 2: C 3 -JEPA architecture overview. Stage I maps synchronized multi-view RGB to bound task-object slots and free context slots; Stage II injects the control latent through AdaLN-Zero and predicts the state autoregressively. Right: the predictive interface.
Fig. 3: Binding-guided patch-to-slot grounding in C 3 -JEPA. Left to right: synchronized multi-view RGB frames from one real recording; grouping slot maps for the target (slot 0) and gripper (slot 1); the binding overlay, where weak anchors select patch support; the VideoSAUR-decoded slot masks that feed temporal state formation; and a PCA view of the patch features.
Visual encoder (frozen DINOv3-L backbone)
Slots
6 total: 3 binding (target UUV, gripper, and a binding-only slot for the fixed ROV frame), 3 free context
TABLE I: Training configuration of the visual encoder and the control-conditioned JEPA predictor.
(a) Deployed checkpoints
Model
Feat. MSE
Target group
Target dec.
Grip group
Grip dec.
Binding + SIGReg
0.0186
0.897
0.915
0.881
0.933
Held-out fusion
0.0183
0.895
0.903
0.877
0.921
Unguided self-supervised (no binding)
0.0157
0.085
0.018
0.116
0.118
(b) Equal-budget fusion ladder
Binding, weighted-sum fusion (no cross-view)
0.0231
0.799
0.713
0.805
0.814
TABLE II: Representation checkpoints. Higher Dice and lower feature MSE are better. (a) Deployed full-data checkpoints; the unguided row is the equal-budget baseline without binding anchors. (b) Equal-budget fusion ladder.
Fig. 4: Conditional rollout diagnostic on the 229-recording interaction set. (a,b) Terminal relative-motion error versus free-rollout horizon for persistence, C 3 -JEPA, the unroll-regularized variant, and the long-context variants ( 16/32 frames); (c) latent MSE against persistence; (d) 5 s control-conditioning ablation, recorded versus zeroed control, with bootstrap 95% CIs.
Fig. 5: UUV (slot 0) mask motion under the world model at two ranges: far ( ≈3.4 m, 3 s rollout, top) and close ( ≈0.8 m, recursive 5 s, bottom). Each panel pairs recorded (blue) and predicted (orange) mask-centroid tracks with the bottom-camera frame.
Representation
Joint MAE ( ∘ )
Rel. pos. MAE (m)
Yaw MAE ( ∘ )
C 3 -JEPA (binding + slots + SIGReg)
4.90
0.60
6.50
LeWM (latent, SIGReg only)
9.81
0.79
13.82
TABLE III: Downstream probe transfer on the grasping scene (recording-disjoint split; identical probe architecture and training). Lower is better. LeWM is trained from scratch on the same recordings and control signals as C 3 -JEPA.
Fig. 6: Field validation. (a,b) Mask quality on the withheld camera for the mean-token, recovery and mask-term arms; dashed: that camera’s true tokens. (c) Token error of that recovery against the mean-token floor, over the seven slots (bars) and the target (rules). (d) Per-segment token error of a 36-step free rollout against persistence. (e) Error reduction over each reference. (f) Mask quality through the rollout.
Fig. 9: Field experimental conditions. (a) The heavy-duty underwater salvage ROV used in this study. (b) The test pool: 60 m by 20 m, 10 m deep. (c) The aluminum pipe on the pool floor. (d) The steel frame on the pool floor.
Deployment A
Deployment B
(77 segments)
(36 segments)
Hidden-view token error
0.615→0.147
0.962→0.332
Hidden-view mask IoU
0.21 / 0.48 / 0.52
0.20 / 0.50 / 0.58
Recovered share of range
88%
81%
5 s prediction vs. persistence
−23%
−31%
TABLE IV: Field-validation summary. Deployment A: navigation above a submerged steel frame; Deployment B: contact with a submerged aluminum pipe.
Fig. 7: Hidden-camera reconstruction on the two deployments (one row each): the field frame from the withheld camera, the mask decoded from its own tokens, and the mask recovered from the other camera.
Fig. 8: Cross-view decoding on a field recording (CAM A top row, CAM B bottom row): the original frame, the mask from that camera’s own tokens, and the merge decoder mask from cross-view fusion.
Underwater robots rely on complementary sensors whose reliability changes abruptly with water visibility and vehicle motion. We introduce AquaJEPA, a sensor-configurable family of action-conditioned joint-embedding predictive models spanning full multimodal, camera-only, sonar-only, and sensor-dropout configurations. Its members share a latent objective and receding-horizon control interface that predict future representations and physical dynamics from camera, forward-looking sonar, proprioception, and thruster commands. Trained from scratch on one hour of action-labelled data, the family is evaluated in Stonefish on 120 fresh paired scenarios spanning unseen layouts, visibility changes, dynamics shifts, and scheduled DVL loss. AquaJEPA-base achieves the strongest aggregate closed-loop performance, improving success over state-only by 12.5 percentage points and reducing final error by 0.189 m; both paired 95% intervals exclude zero. In a separate three-seed evaluation, it reduces paired final error relative to AquaJEPA-S by 0.118 m, with the same direction for every seed. AquaJEPA-robust more than halves prediction error during camera and camera-DVL blackouts. These results show that full multimodal prediction improves over state-only control and the sonar-only family member in this benchmark, while sensor-dropout training provides robustness under sensor loss.
Controlling an agent with vision requires being able to separate useful information from irrelevant background information. JEPA-style latent world models seem like a natural approach for this, as they do not perform pixel-level reconstruction; however, they are still sensitive to these distractor signals and experience latent collapse. In this work, we introduce Controllability Factorized JEPA (CF-JEPA), a JEPA-style world model which splits the latent space into controllable and uncontrollable subspaces. This factorization allows us to capture all the distractor information into the uncontrollable region, while we use the control-relevant latent information for our task. With this, we show comparable performance across 2D and 3D control tasks under nominal conditions and improved performance under distracted conditions, where CF-JEPA is the only model that does not experience latent collapse. We also validate our model under distracted conditions for a simulated robot task, highlighting the practical application of such a scheme.
Morgan Byrd, Robert Wright, Sehoon Ha
Georgia Institute of Technology, Atlanta, GA, 30308, USA · Georgia Tech Research Institute, Atlanta, GA, 30308, USA
Predictive world models enable agents to model scene dynamics and reason about the consequences of their actions. Inspired by human perception, object-centric world models capture scene dynamics using object-level representations, which can be used for downstream applications such as action planning. However, most object-centric world models and reinforcement learning (RL) approaches learn reactive policies that are fixed at inference time, limiting generalization to novel situations. We propose Slot-MPC, an object-centric world modeling framework that enables planning through Model Predictive Control (MPC). Slot-MPC leverages vision encoders to learn slot-based representations, which encode individual objects in the scene, and uses these structured representations to learn an action-conditioned object-centric dynamics model. At inference time, the learned dynamics model enables action planning via MPC, allowing agents to adapt to previously unseen situations. Since the learned world model is differentiable, we can use gradient-based MPC to directly optimize actions, which is computationally more efficient than relying on gradient-free, sampling-based MPC methods. Experiments on simulated robotic manipulation tasks show that Slot-MPC improves both task performance and planning efficiency compared to non-object-centric world model baselines. In the considered offline setting with limited state-action coverage, we find that gradient-based MPC performs better than gradient-free, sampling-based MPC. Our results demonstrate that explicitly structured, object-centric representations provide a strong inductive bias for controllable and generalizable decision-making. Code and additional results are available at https://slot-mpc.github.io.
Jonathan Spieler, Angel Villar-Corrales, Sven Behnke
Autonomous Intelligent Systems, Computer Science Institute VI - Intelligent Systems and Robotics, Center for Robotics and the Lamarr Institute for Machine Learning and Artificial Intelligence, University of Bonn, Germany