Reconstructing and animating water scenery from nature produces compelling and immersive visual experiences. Previous work examined this task from the perspective of 2D video textures, with the goal of creating a looping video. In our work, we tackle the problem from a 3D perspective, creating a looping 4D dynamic reconstruction which can be interactively rendered from novel viewpoints from a single non-looping 2D source video. We represent motion as a 3D static \textit{Eulerian} motion field that advects canonical Gaussian splats that are cyclically reborn at fixed time periods, supervised using rendering losses. To model non-periodic and stochastic dynamics present in real-world scenes, we add a non-periodic, time-varying residual term to capture deviations from the static Eulerian motion field. We show quantitatively and qualitatively that our framework enables photorealistic animation of water scenes better than prior art.
Figures & tables
Figure 2 : Method Overview. (Stage 1) Given a monocular input video, we first create a static reconstruction in the form of Canonical Gaussian Splats (Sec. 3.4 ). (Stage 2) We use depth computed from the static reconstruction and camera poses to lift predicted 2D optical flow to 3D, and use (3D position, 3D scene flow) pairs at every frame to pre-train a static Eulerian motion field (Sec. 3.4 ). (Stage 3) To synthesize input images at each time step, we advect Canonical Gaussian Splats using our motion field through forward Euler integration with our looping algorithm (Sec. 3.2 , Alg. 1 ) and additionally apply a non-periodic residual field D to account for the non-periodic nature of water (Sec. 3.3 ). Advected splats are rendered into synthesized images; they are compared to the ground truth via rendering losses that are back-propagated to optimize Canonical Gaussian Splat parameters and the motion field (Sec. 3.5 ). Given the trained loopable model, we can generate infinite looping animations by repeating every L frames using Alg. 2 and setting pn=0 .
Figure 3 : Why per-Gaussian random start times? Five Gaussians at different canonical positions (dashed circles) are advected by a downward Eulerian motion field V (gray arrows) with cycle length L=15 . Each is tn=(i−tn0)modL steps below its canonical (rebirth) position and snaps back on rebirth ( tn=0 ). (a) With a shared start time, every Gaussian has the same tn=imodL and resets simultaneously at i=0,15,… : the whole scene moves as one rigid clump (red band) and teleports back to the top in a single frame (jump artifact). (b) With random start times ( tn0=0,3,6,9,12 for the red , blue , green , violet , orange Gaussians) the resets are staggered across the cycle, so at any frame only a few Gaussians reset while the rest flow smoothly.
Figure 4 : Qualitative comparison of reconstruction compared to baselines. Our method produces much sharper result than the baselines in the water regions. While MovieS appears sharp in row2, its overall view synthesis quality is poor ( refer to videos in the supp. material ).
Full image
Water region
Method
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
KID ↓
FVD ↓
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
KID ↓
FVD ↓
MoVieS [ 15 ]
15.19
0.331
0.524
100.87
0.027
1564.90
15.35
0.291
0.526
87.21
0.014
1157.00
MoSca [ 11 ]
19.87
0.509
0.565
117.54
0.037
828.87
23.31
0.723
0.577
104.51
0.024
534.47
4DGS [ 39 ]
23.01
0.693
0.397
84.00
0.026
634.19
25.71
0.803
0.436
87.58
0.022
417.40
AmbGS [ 34 ]
22.52
0.706
0.344
61.32
0.014
536.55
25.49
0.806
0.392
69.15
0.014
314.24
Ours
23.05
0.716
0.315
39.63
0.007
210.01
25.37
0.798
0.375
46.50
0.008
156.69
Table 1: Quantitative comparison of methods. ↓ ( ↑ ) indicates lower (higher) is better. We report metrics over the full image as well as restricted to the water region, which isolates the dynamic content our method targets. FID, KID and FVD follow the distribution-based evaluation protocol of Lift4D [ 18 ] . Our method achieves the best overall reconstruction accuracy in comparison to the baselines and leads the baselines in spatial and temporal distribution alignment by a large margin.
Figure 5 : User study. Bar height is the percentage of times users rated our method’s visual quality higher than the competing method (“Winrate”), for three settings: Appearance (camera moves, scene frozen), Motion (camera fixed, scene advances), and Reference (both advance, shown alongside the real footage). Aggregated over 25 participants and 7 scenes ( 2100 judgments). Our method was preferred over all baseline methods in all settings.
Figure 6 : Scene flow comparison. We visualize tracks formed by canonical Gaussian Splats by iteratively querying our Eulerian motion field with residual applied, on two scenes (rows), against 4DGS [ 39 ] , AmbGS [ 34 ] , MoSca [ 11 ] and MoVieS [ 15 ] . For the baselines, which do not expose an explicit Eulerian motion field, we extract the corresponding tracks by querying each method’s own motion representation at the same set of points and time steps. Our Eulerian motion field yields long, coherent trajectories that follow natural water flow, whereas the baselines produce short, incoherent, or near-static tracks.
Figure 7 : Visualizing temporal progression of water. We compare our method to the baselines at different time steps. Our method reproduces natural flow of water, while 4DGS [ 39 ] AmbGS [ 34 ] and MoSca [ 11 ] only demonstrate slight shift in color over time.
Cycle length
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
KID ↓
FVD ↓
L=35
22.78
0.702
0.329
42.18
0.008
247.76
L=25
22.87
0.708
0.321
39.35
0.007
210.96
L=15 (ours)
23.05
0.716
0.315
39.63
0.007
210.01
Table 2 : Cycle-length ablation. Effect of the cycle length L on reconstruction quality. Metrics follow the same protocol as Tab. 3 .
Method
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
KID ↓
FVD ↓
w/o Eulerian motion field
23.42
0.733
0.319
50.86
0.011
422.66
w/o flow initialization
22.50
0.704
0.325
55.44
0.012
473.29
w/o non-periodic residual
22.97
0.713
0.321
47.22
0.010
309.37
Static reconstruction
23.11
0.723
0.319
60.88
0.015
522.97
w/o Alg. 1
22.86
0.702
0.329
47.01
0.010
334.30
Full model
23.05
0.716
0.315
39.63
0.007
210.01
Table 3 : Ablations of our method. Each row removes a single component of our full model or replaces it with a static reconstruction. Best per column is highlighted in red (bold) and second best in orange.
Figure 8 : Ablation track visualization. Tracks formed by iteratively querying each ablation variant’s motion representation on two scenes, shown at the final frame of a 15-frame propagation. The full model produces coherent water-following trajectories; model with the Eulerian motion field or flow initialization removed barely reconstructs and water flow motion, and removing the non-periodic residual produces less natural and stochastic tracks.
Figure 9 : Random start time prevents ”looping artifacts” For each scene we fix the camera, advance time for 60 frames (4 periods with 15 frames in each period), and stack a single horizontal scanline (red, in the sampled-row frames (b)) into an x – t image (c); time runs downward. GT (a) is the reference. Without random start times (bottom row per scene, “w/o Alg. 1”) the Gaussians pulse as one body, resulting in the slice being crossed by horizontal bands that span the full width at a single instant at loop boundaries. With random start times (top row per scene, “Ours”) water particles spread stochastically across loops. Per-frame metrics barely register this ( −0.19 dB PSNR), while FVD rises significantly.
Figure 10 : Effect of changing cycle length L=(15,25,35) . Columns vary the cycle length L ; rows show two time steps in the loop. It can be observed that L=25 produces better dynamics (more variation in the waterfall states) compared to L=15, while L=35 results in blurrier overall reconstruction due to worse convergence.
Figure 11 : Extensions. With the help of SFM tools such as MegaSam [ 12 ] , our method can be extended to reconstruct motion of water surfaces, e.g . , lakes. We also demonstrate that our Eulerian motion field can represent other media such as fire by extending it to a time-varying model.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Alg. 1
Alg. 2
Speedup
15 frames ( L=15 , 1 cycle)
Deformation only
0.673 s (44.88 ms/frame)
0.224 s (14.93 ms/frame)
3.01×
End-to-end (incl. raster)
0.719 s (47.95 ms/frame)
0.270 s (17.98 ms/frame)
2.67×
450 frames ( L=15 , 30 cycles)
Deformation only
23.483 s (52.19 ms/frame)
5.605 s (12.46 ms/frame)
4.19×
End-to-end (incl. raster)
28.232 s (62.74 ms/frame)
7.072 s (15.72 ms/frame)
3.99×
Appendix
Table B1 : Animation run time. Wall-clock time to render 15 frames ( L=15 , 1 cycle) and 450 frames ( L=15 , 30 cycles), for both the pure deformation pass and end-to-end (deformation + rasterization), using the training-time propagation (Alg. 1) versus our inference-time propagation (Alg. 2). Reusing per-cycle propagations across all frames of a cycle yields a ∼3 – 4× speedup, with a larger gain as the number of cycles grows.
Figure B1 : The seven captured water scenes , each shown at the middle frame of its sequence. Captions give the number of frames (train/test), the maximum camera baseline, and the maximum rotation between viewing directions. Every scene is a hand-held fisheye capture rectified to a pinhole camera; test frames are three held-out temporal segments per scene.
Scene
Frames
Train
Test
Baseline
Path
Rotation
Water
Greenacre Park
750
660
90
7.0 m
15 m
111 ∘
34%
Garden
750
660
90
6.8 m
11 m
101 ∘
39%
Parley Falls
750
660
90
8.4 m
17 m
83 ∘
39%
Creek
750
660
90
3.4 m
11 m
93 ∘
38%
Sculpture
750
660
90
10.4 m
14 m
105 ∘
44%
Waterstep
800
710
90
12.1 m
16 m
101 ∘
63%
Appendix
Table B2 : Per-scene capture statistics. Baseline is the maximum distance between any two camera centres, Path the total trajectory length, Rotation the maximum angle between any two optical axes, and Water the mean fraction of pixels inside the water mask.
Figure C2 : Component ablations. Our full model reconstructs the best dynamism of water with all the proposed components. Without initializing the motion field or learning a position-based deformation field from scratch converges to very small motion (water particles become blurry strands), while no opacity and color deformation prevents our looping representation from fitting to the non-looping input video, producing blurry results.
Figure D3 : User-study interface. (a) instructions shown at the start of the study and (b) forced-choice questions asked on every comparison page.