We present Underwater C3-JEPA (cross-view, control-conditioned, context-extended), an object-centric multi-view predictive world model for near-field heavy-load underwater ROV salvage. Without contact sensors, it predicts in latent space how the task-object state evolves through contact interaction and under the hydrodynamic lag of the vehicle, from synchronized multi-view RGB observations and vehicle control signals. C3-JEPA encodes multi-camera observations into task-object and context tokens, fuses cross-camera evidence through held-out-view attention, and directly predicts future states conditioned on control. Weak binding anchors the target and gripper at low annotation cost, while SIGReg sharpens the geometric representation. Experiments show that the learned representation transfers substantially more task-relevant information to downstream probes than a reconstruction-free latent baseline, while keeping the predictor lightweight. The resulting predictive interface supports model-predictive-control (MPC) candidate evaluation and imagined-rollout behavior-agent training. Validation on real underwater video shows the same architecture recovering a withheld camera's object state and staying ahead of persistence, so the recipe transfers beyond simulation.
Figures & tables
Fig. 1: The heavy-duty underwater salvage ROV used in this study: (a) isometric view; (b) grasping a target object (CAD rendering); (c) lifting a target object in a real-environment trial.
Fig. 2: C 3 -JEPA architecture overview. Stage I maps synchronized multi-view RGB to bound task-object slots and free context slots; Stage II injects the control latent through AdaLN-Zero and predicts the state autoregressively. Right: the predictive interface.
Fig. 3: Binding-guided patch-to-slot grounding in C 3 -JEPA. Left to right: synchronized multi-view RGB frames from one real recording; grouping slot maps for the target (slot 0) and gripper (slot 1); the binding overlay, where weak anchors select patch support; the VideoSAUR-decoded slot masks that feed temporal state formation; and a PCA view of the patch features.
Visual encoder (frozen DINOv3-L backbone)
Slots
6 total: 3 binding (target UUV, gripper, and a binding-only slot for the fixed ROV frame), 3 free context
TABLE I: Training configuration of the visual encoder and the control-conditioned JEPA predictor.
(a) Deployed checkpoints
Model
Feat. MSE
Target group
Target dec.
Grip group
Grip dec.
Binding + SIGReg
0.0186
0.897
0.915
0.881
0.933
Held-out fusion
0.0183
0.895
0.903
0.877
0.921
Unguided self-supervised (no binding)
0.0157
0.085
0.018
0.116
0.118
(b) Equal-budget fusion ladder
Binding, weighted-sum fusion (no cross-view)
0.0231
0.799
0.713
0.805
0.814
TABLE II: Representation checkpoints. Higher Dice and lower feature MSE are better. (a) Deployed full-data checkpoints; the unguided row is the equal-budget baseline without binding anchors. (b) Equal-budget fusion ladder.
Fig. 4: Conditional rollout diagnostic on the 229-recording interaction set. (a,b) Terminal relative-motion error versus free-rollout horizon for persistence, C 3 -JEPA, the unroll-regularized variant, and the long-context variants ( 16/32 frames); (c) latent MSE against persistence; (d) 5 s control-conditioning ablation, recorded versus zeroed control, with bootstrap 95% CIs.
Fig. 5: UUV (slot 0) mask motion under the world model at two ranges: far ( ≈3.4 m, 3 s rollout, top) and close ( ≈0.8 m, recursive 5 s, bottom). Each panel pairs recorded (blue) and predicted (orange) mask-centroid tracks with the bottom-camera frame.
Representation
Joint MAE ( ∘ )
Rel. pos. MAE (m)
Yaw MAE ( ∘ )
C 3 -JEPA (binding + slots + SIGReg)
4.90
0.60
6.50
LeWM (latent, SIGReg only)
9.81
0.79
13.82
TABLE III: Downstream probe transfer on the grasping scene (recording-disjoint split; identical probe architecture and training). Lower is better. LeWM is trained from scratch on the same recordings and control signals as C 3 -JEPA.
Fig. 6: Field validation. (a,b) Mask quality on the withheld camera for the mean-token, recovery and mask-term arms; dashed: that camera’s true tokens. (c) Token error of that recovery against the mean-token floor, over the seven slots (bars) and the target (rules). (d) Per-segment token error of a 36-step free rollout against persistence. (e) Error reduction over each reference. (f) Mask quality through the rollout.
Fig. 9: Field experimental conditions. (a) The heavy-duty underwater salvage ROV used in this study. (b) The test pool: 60 m by 20 m, 10 m deep. (c) The aluminum pipe on the pool floor. (d) The steel frame on the pool floor.
Deployment A
Deployment B
(77 segments)
(36 segments)
Hidden-view token error
0.615→0.147
0.962→0.332
Hidden-view mask IoU
0.21 / 0.48 / 0.52
0.20 / 0.50 / 0.58
Recovered share of range
88%
81%
5 s prediction vs. persistence
−23%
−31%
TABLE IV: Field-validation summary. Deployment A: navigation above a submerged steel frame; Deployment B: contact with a submerged aluminum pipe.
Fig. 7: Hidden-camera reconstruction on the two deployments (one row each): the field frame from the withheld camera, the mask decoded from its own tokens, and the mask recovered from the other camera.
Fig. 8: Cross-view decoding on a field recording (CAM A top row, CAM B bottom row): the original frame, the mask from that camera’s own tokens, and the merge decoder mask from cross-view fusion.
Autonomous Intelligent Systems, Computer Science Institute VI - Intelligent Systems and Robotics, Center for Robotics and the Lamarr Institute for Machine Learning and Artificial Intelligence, University of Bonn, Germany