Reconstructing and animating water scenery from nature produces compelling and immersive visual experiences. Previous work examined this task from the perspective of 2D video textures, with the goal of creating a looping video. In our work, we tackle the problem from a 3D perspective, creating a looping 4D dynamic reconstruction which can be interactively rendered from novel viewpoints from a single non-looping 2D source video. We represent motion as a 3D static \textit{Eulerian} motion field that advects canonical Gaussian splats that are cyclically reborn at fixed time periods, supervised using rendering losses. To model non-periodic and stochastic dynamics present in real-world scenes, we add a non-periodic, time-varying residual term to capture deviations from the static Eulerian motion field. We show quantitatively and qualitatively that our framework enables photorealistic animation of water scenes better than prior art.
Figures & tables
Figure 2 : Method Overview. (Stage 1) Given a monocular input video, we first create a static reconstruction in the form of Canonical Gaussian Splats (Sec. 3.4 ). (Stage 2) We use depth computed from the static reconstruction and camera poses to lift predicted 2D optical flow to 3D, and use (3D position, 3D scene flow) pairs at every frame to pre-train a static Eulerian motion field (Sec. 3.4 ). (Stage 3) To synthesize input images at each time step, we advect Canonical Gaussian Splats using our motion field through forward Euler integration with our looping algorithm (Sec. 3.2 , Alg. 1 ) and additionally apply a non-periodic residual field D to account for the non-periodic nature of water (Sec. 3.3 ). Advected splats are rendered into synthesized images; they are compared to the ground truth via rendering losses that are back-propagated to optimize Canonical Gaussian Splat parameters and the motion field (Sec. 3.5 ). Given the trained loopable model, we can generate infinite looping animations by repeating every L frames using Alg. 2 and setting pn=0 .
Figure 3 : Why per-Gaussian random start times? Five Gaussians at different canonical positions (dashed circles) are advected by a downward Eulerian motion field V (gray arrows) with cycle length L=15 . Each is tn=(i−tn0)modL steps below its canonical (rebirth) position and snaps back on rebirth ( tn=0 ). (a) With a shared start time, every Gaussian has the same tn=imodL and resets simultaneously at i=0,15,… : the whole scene moves as one rigid clump (red band) and teleports back to the top in a single frame (jump artifact). (b) With random start times ( tn0=0,3,6,9,12 for the red , blue , green , violet , orange Gaussians) the resets are staggered across the cycle, so at any frame only a few Gaussians reset while the rest flow smoothly.
Figure 4 : Qualitative comparison of reconstruction compared to baselines. Our method produces much sharper result than the baselines in the water regions. While MovieS appears sharp in row2, its overall view synthesis quality is poor ( refer to videos in the supp. material ).
Full image
Water region
Method
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
KID ↓
FVD ↓
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
KID ↓
FVD ↓
MoVieS [ 15 ]
15.19
0.331
0.524
100.87
0.027
1564.90
15.35
0.291
0.526
87.21
0.014
1157.00
MoSca [ 11 ]
19.87
0.509
0.565
117.54
0.037
828.87
23.31
0.723
0.577
104.51
0.024
534.47
4DGS [ 39 ]
23.01
0.693
0.397
84.00
0.026
634.19
25.71
0.803
0.436
87.58
0.022
417.40
AmbGS [ 34 ]
22.52
0.706
0.344
61.32
0.014
536.55
25.49
0.806
0.392
69.15
0.014
314.24
Ours
23.05
0.716
0.315
39.63
0.007
210.01
25.37
0.798
0.375
46.50
0.008
156.69
Table 1: Quantitative comparison of methods. ↓ ( ↑ ) indicates lower (higher) is better. We report metrics over the full image as well as restricted to the water region, which isolates the dynamic content our method targets. FID, KID and FVD follow the distribution-based evaluation protocol of Lift4D [ 18 ] . Our method achieves the best overall reconstruction accuracy in comparison to the baselines and leads the baselines in spatial and temporal distribution alignment by a large margin.
Figure 5 : User study. Bar height is the percentage of times users rated our method’s visual quality higher than the competing method (“Winrate”), for three settings: Appearance (camera moves, scene frozen), Motion (camera fixed, scene advances), and Reference (both advance, shown alongside the real footage). Aggregated over 25 participants and 7 scenes ( 2100 judgments). Our method was preferred over all baseline methods in all settings.
Figure 6 : Scene flow comparison. We visualize tracks formed by canonical Gaussian Splats by iteratively querying our Eulerian motion field with residual applied, on two scenes (rows), against 4DGS [ 39 ] , AmbGS [ 34 ] , MoSca [ 11 ] and MoVieS [ 15 ] . For the baselines, which do not expose an explicit Eulerian motion field, we extract the corresponding tracks by querying each method’s own motion representation at the same set of points and time steps. Our Eulerian motion field yields long, coherent trajectories that follow natural water flow, whereas the baselines produce short, incoherent, or near-static tracks.
Figure 7 : Visualizing temporal progression of water. We compare our method to the baselines at different time steps. Our method reproduces natural flow of water, while 4DGS [ 39 ] AmbGS [ 34 ] and MoSca [ 11 ] only demonstrate slight shift in color over time.
Cycle length
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
KID ↓
FVD ↓
L=35
22.78
0.702
0.329
42.18
0.008
247.76
L=25
22.87
0.708
0.321
39.35
0.007
210.96
L=15 (ours)
23.05
0.716
0.315
39.63
0.007
210.01
Table 2 : Cycle-length ablation. Effect of the cycle length L on reconstruction quality. Metrics follow the same protocol as Tab. 3 .
Method
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
KID ↓
FVD ↓
w/o Eulerian motion field
23.42
0.733
0.319
50.86
0.011
422.66
w/o flow initialization
22.50
0.704
0.325
55.44
0.012
473.29
w/o non-periodic residual
22.97
0.713
0.321
47.22
0.010
309.37
Static reconstruction
23.11
0.723
0.319
60.88
0.015
522.97
w/o Alg. 1
22.86
0.702
0.329
47.01
0.010
334.30
Full model
23.05
0.716
0.315
39.63
0.007
210.01
Table 3 : Ablations of our method. Each row removes a single component of our full model or replaces it with a static reconstruction. Best per column is highlighted in red (bold) and second best in orange.
Figure 8 : Ablation track visualization. Tracks formed by iteratively querying each ablation variant’s motion representation on two scenes, shown at the final frame of a 15-frame propagation. The full model produces coherent water-following trajectories; model with the Eulerian motion field or flow initialization removed barely reconstructs and water flow motion, and removing the non-periodic residual produces less natural and stochastic tracks.
Figure 9 : Random start time prevents ”looping artifacts” For each scene we fix the camera, advance time for 60 frames (4 periods with 15 frames in each period), and stack a single horizontal scanline (red, in the sampled-row frames (b)) into an x – t image (c); time runs downward. GT (a) is the reference. Without random start times (bottom row per scene, “w/o Alg. 1”) the Gaussians pulse as one body, resulting in the slice being crossed by horizontal bands that span the full width at a single instant at loop boundaries. With random start times (top row per scene, “Ours”) water particles spread stochastically across loops. Per-frame metrics barely register this ( −0.19 dB PSNR), while FVD rises significantly.
Figure 10 : Effect of changing cycle length L=(15,25,35) . Columns vary the cycle length L ; rows show two time steps in the loop. It can be observed that L=25 produces better dynamics (more variation in the waterfall states) compared to L=15, while L=35 results in blurrier overall reconstruction due to worse convergence.
Figure 11 : Extensions. With the help of SFM tools such as MegaSam [ 12 ] , our method can be extended to reconstruct motion of water surfaces, e.g . , lakes. We also demonstrate that our Eulerian motion field can represent other media such as fire by extending it to a time-varying model.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Alg. 1
Alg. 2
Speedup
15 frames ( L=15 , 1 cycle)
Deformation only
0.673 s (44.88 ms/frame)
0.224 s (14.93 ms/frame)
3.01×
End-to-end (incl. raster)
0.719 s (47.95 ms/frame)
0.270 s (17.98 ms/frame)
2.67×
450 frames ( L=15 , 30 cycles)
Deformation only
23.483 s (52.19 ms/frame)
5.605 s (12.46 ms/frame)
4.19×
End-to-end (incl. raster)
28.232 s (62.74 ms/frame)
7.072 s (15.72 ms/frame)
3.99×
Appendix
Table B1 : Animation run time. Wall-clock time to render 15 frames ( L=15 , 1 cycle) and 450 frames ( L=15 , 30 cycles), for both the pure deformation pass and end-to-end (deformation + rasterization), using the training-time propagation (Alg. 1) versus our inference-time propagation (Alg. 2). Reusing per-cycle propagations across all frames of a cycle yields a ∼3 – 4× speedup, with a larger gain as the number of cycles grows.
Figure B1 : The seven captured water scenes , each shown at the middle frame of its sequence. Captions give the number of frames (train/test), the maximum camera baseline, and the maximum rotation between viewing directions. Every scene is a hand-held fisheye capture rectified to a pinhole camera; test frames are three held-out temporal segments per scene.
Scene
Frames
Train
Test
Baseline
Path
Rotation
Water
Greenacre Park
750
660
90
7.0 m
15 m
111 ∘
34%
Garden
750
660
90
6.8 m
11 m
101 ∘
39%
Parley Falls
750
660
90
8.4 m
17 m
83 ∘
39%
Creek
750
660
90
3.4 m
11 m
93 ∘
38%
Sculpture
750
660
90
10.4 m
14 m
105 ∘
44%
Waterstep
800
710
90
12.1 m
16 m
101 ∘
63%
Appendix
Table B2 : Per-scene capture statistics. Baseline is the maximum distance between any two camera centres, Path the total trajectory length, Rotation the maximum angle between any two optical axes, and Water the mean fraction of pixels inside the water mask.
Figure C2 : Component ablations. Our full model reconstructs the best dynamism of water with all the proposed components. Without initializing the motion field or learning a position-based deformation field from scratch converges to very small motion (water particles become blurry strands), while no opacity and color deformation prevents our looping representation from fitting to the non-looping input video, producing blurry results.
Figure D3 : User-study interface. (a) instructions shown at the start of the study and (b) forced-choice questions asked on every comparison page.
Recent advancements in image animation have utilized diffusion models to breathe life into static images. However, existing controllable frameworks typically rely on Lagrangian motion guidance, where optical flow is estimated relative to the initial frame. This paper revisits the same optical-flow primitive through a more local supervision design: we use adjacent-frame Eulerian motion fields to guide generation, where the motion signal always describes a short temporal hop. This shift enables parallelized training and provides bounded-error supervision throughout the generation process. To mitigate the drift artifacts common in adjacent frame generation, we introduce a Bidirectional Geometric Consistency mechanism, which computes a forward-backward cycle check to mathematically identify and mask occluded regions, preventing the model from learning incorrect warping objectives. Extensive experiments demonstrate that our approach accelerates training, preserves temporal coherence, and reduces dynamic artifacts compared to reference-based baselines. The code, model, and data have been made available at https://nguyentthong.github.io/eulerian/
Thong Nguyen, Khoi M. Le, Cong-Duy Nguyen +3
National University of Singapore, Singapore · Centre for AI Research, VinUniversity, Vietnam · Nanyang Technological University, Singapore
Novel view rendering of large and complex reconstructed scenes is becoming increasingly photorealistic. However, most reconstructions remain static and lack the ambient motion that makes environments immersive. We present AniGS, a method for scene-level animation of 3D Gaussian Splatting (3DGS) reconstructions that adds subtle, distributed dynamics, e.g., vegetation motion, while preserving rigid structures. Unlike existing 3D animation techniques which are limited to object-centric subjects or small regions, AniGS is designed for large, cluttered, navigable scenes. AniGS represents the scene with a canonical 3DGS and models motion using a time-conditioned deformation field. To animate the entire scene, we leverage a pretrained video diffusion model and introduce an iterative dataset--model update strategy that progressively expands viewpoint coverage and repeatedly updates camera-fixed training videos using a render-and-refine scheme. To prevent artifacts from unintended motion in static areas, we further introduce a composed video-to-video refinement scheme that restricts motion to desired regions. Experiments on five real-world, large-scale outdoor scenes demonstrate that AniGS produces natural ambient dynamics and high-quality novel view videos, enabling more immersive viewing experiences of reconstructed environments.
Yen-Chi Cheng, Chen Gao, Chuhan Chen +8
University of Illinois Urbana-Champaign, Champaign-Urbana, Illinois, USA · Meta · Waymo, Bellevue, Washington, USA +2
We introduce LivingWorld, an interactive framework for generating 4D worlds with environmental dynamics from a single image. While recent advances in 3D scene generation enable large-scale environment creation, most approaches focus primarily on reconstructing static geometry, leaving scene-scale environmental dynamics such as clouds, water, or smoke largely unexplored. Modeling such dynamics is challenging because motion must remain coherent across an expanding scene while supporting low-latency user feedback. LivingWorld addresses this challenge by progressively constructing a globally coherent motion field as the scene expands. To maintain global consistency during expansion, we introduce a geometry-aware alignment module that resolves directional and scale ambiguities across views. We further represent motion using a compact hash-based motion field, enabling efficient querying and stable propagation of dynamics throughout the scene. This representation also supports bidirectional motion propagation during rendering, producing long and temporally coherent 4D sequences without relying on expensive video-based refinement. On a single RTX 5090 GPU, generating each new scene expansion step requires 9 seconds, followed by 3 seconds for motion alignment and motion field updates, enabling interactive 4D world generation with globally coherent environmental dynamics. Video demonstrations are available at paper.pnu-cvsp.com/LivingWorld.
Hyeongju Mun, In-Hwan Jin, Sohyeong Kim +1
Department of Electrical and Electronic Engineering, Pusan National University