Recent depth foundation models like Depth Anything 3 (DA3) achieve remarkable multi-view depth estimation but assume static 3D scenes, limiting their applicability to real-world dynamic environments. Existing training-free 4D methods like Easi3R and VGGT4D rely on correspondence-trained backbones whose attention encodes cross-frame matching, a property absent in depth-only models like DA3. We present Dyna3, a training-free framework that extends DA3 for 4D dynamic scene reconstruction without any fine-tuning. Our key insight is that DA3's cross-view features, though trained only for depth consistency, implicitly encode motion-discriminative signals when combined with best-match feature search across frames. Its static surfaces find consistent matches globally, while dynamic objects cannot. We further adopt vision-language models (VLM) to automatically generate scene-specific semantic prompts for SAM 3, enabling precise instance-level segmentation that distinguishes which objects move from what objects exist. For reconstruction, we decouple the scene into a cross-frame aligned static background and per-frame dynamic point clouds. Experiments on four datasets demonstrate that Dyna3 surpasses correspondence-trained methods with +5.5pp J-Mean over state-of-the-art VGGT4D on dynamic object segmentation, while achieving up to 13x faster pose estimation and 3x faster 4D reconstruction with 4 to 8x lower memory. Dyna3 could therefore enable much denser temporal sampling that prior methods cannot support.
Figures & tables
Figure 1 : Dyna3 extends a frozen depth foundation model (DA3) to training-free 4D reconstruction. Left: DA3’s attention misses moving objects. Dyna3 uses VLM-guided progressive mask refinement to reconstruct static scene, per-frame dynamic points, and camera poses. Right: Dyna3 is more accurate than Easi3R and VGGT4D, 3.2× faster and 4.5× less memory than VGGT4D.
Figure 2 : Overview of the Dyna3.
DAVIS-16
DAVIS-17
Method
J-Mean ↑
F-Mean ↑
J-Mean ↑
F-Mean ↑
Easi3R [ 4 ]
48.92
45.29
49.65
44.09
VGGT4D [ 8 ]
59.50
55.47
56.45
51.09
Dyna3 (Ours)
65.04
66.93
58.73
62.46
Table 1 : Dynamic object segmentation on DAVIS-2016 and DAVIS-2017. J-Mean and F-Mean denote region similarity and boundary accuracy (%), respectively. Best results are bold . Dyna3 achieves state-of-the-art performance on both datasets, with particularly large gains on boundary accuracy.
Method
ATE ↓
RTE ↓
RRE ↓
Time(s/vid) ↓
Mem(GB) ↓
MonST3R [ 38 ]
0.156
0.103
12.041
82.93
42.4
Easi3R [ 4 ]
0.063
0.046
2.523
138.27
45.7
VGGT4D [ 8 ]
0.016
0.020
0.612
45.53
130.1
Dyna3 (Ours)
0.121
0.024
2.506
15.64
24.7
Table 2 : Camera pose estimation on TUM-dynamics. ATE and RTE are in meters, RRE in degrees. Best and second-best results are bold and underlined . Dyna3 achieves competitive pose accuracy while being significantly more efficient, running 2.9 × faster with 5.3 × less memory than VGGT4D.
Method
ATE ↓
RTE ↓
RRE ↓
Time(s/vid) ↓
Mem(GB) ↓
Easi3R [ 4 ]
0.151
0.051
0.277
220.05
27.7
VGGT4D [ 8 ]
0.076
0.046
0.273
28.03
45.2
Dyna3 (Ours)
0.197
0.040
0.319
3.26
10.2
Table 3 : Camera pose estimation on Sintel. Camera pose estimation on Sintel. ATE and RTE are in meters, RRE in degrees. Best and second-best results are bold and underlined . Dyna3 achieves the best relative translation error while running 8.6 × faster and using 4.4 × less memory than VGGT4D.
Stride
ATE ↓
RTE ↓
RRE ↓
Time (vs VGGT4D) ↓
Mem (vs VGGT4D) ↓
30
0.023
0.140
13.674
3.63 (12.4 × faster)
16.4 (7.9 × smaller)
10
0.021
0.063
9.232
7.57 (13.4 × faster)
21.5 (6.7 × smaller)
3
0.021
0.024
2.506
15.64 (VGGT4D OOM)
24.7 (VGGT4D OOM)
Table 4 : Effect of temporal stride on TUM-dynamics. Denser sampling ( i.e. smaller stride) significantly improves relative pose accuracy. At stride 3, VGGT4D runs out of memory while Dyna3 completes successfully, demonstrating the practical advantage of efficient inference.
Stride
ATE ↓
RTE ↓
RRE ↓
Time (vs VGGT4D) ↓
Mem (vs VGGT4D) ↓
3
0.208
0.106
2.383
1.15 (9.4 × faster)
6.4 (5.0 × smaller)
2
0.208
0.092
1.698
1.61 (10.0 × faster)
8.6 (5.1 × smaller)
1
0.197
0.040
0.319
3.26 (8.6 × faster)
10.2 (4.4 × smaller)
Table 5 : Effect of temporal stride on Sintel. Denser sampling ( i.e. smaller stride) dramatically reduces rotation error (RRE improves 7.5 × from stride 3 to stride 1). Dyna3 remains 8-10 × faster than VGGT4D across all stride settings.
Method
Accuracy ↓
Completeness ↓
Distance ↓
Time ↓
Mem. ↓
Mean
Med.
Mean
Med.
Mean
Med.
DAS3R [ 35 ]
0.192
0.142
0.250
0.108
0.428
0.336
-
-
CUT3R [ 23 ]
0.073
0.054
0.133
0.049
0.328
0.224
-
-
MonST3R [ 38 ]
0.090
0.033
0.113
0.064
0.279
0.234
-
-
Easi3R [ 4 ]
0.070
0.044
0.060
0.033
0.194
0.132
7.14s (10.2 × slower)
11.85GB (3.2 × larger)
VGGT4D [ 8 ]
0.022
0.004
0.051
0.012
0.123
0.050
2.26s (3.2 × slower)
16.70GB (4.5 × larger)
Table 6 : 4D reconstruction on DyCheck. Accuracy measures distance from predicted to ground truth points; Completeness measures the reverse direction; Distance combines both. Best and second-best results are bold and underlined . Dyna3 achieves the second-best reconstruction quality while being 3.2 × faster and using 4.5 × less memory than VGGT4D, and outperforms all fine-tuned methods (MonST3R, CUT3R, DAS3R) despite being training-free.
Variant
J-Mean ↑
F-Mean ↑
DA3 baseline (no motion pipeline)
8.16
8.43
+ Motion Prior Extraction
29.65
27.34
+ Fine Instance Selection
33.51
32.18
+ VLM Prompting (w/o SAM3)
53.35
54.29
+ SAM3 Segmentation ( Full Dyna3 )
65.04
66.93
Table 7 : Ablation study on Dyna3 components on DAVIS-2016. We progressively add components to validate their contribution. Each component provides consistent improvements in dynamic object segmentation.
Figure 3 : Comparison of DA3 cross-attention maps versus our best-match feature search on DAVIS-2016. DA3’s attention (Column 3) produces uniformly distributed responses across the scene, failing to isolate dynamic objects (Column 2). In contrast, our method (Column 4) successfully concentrates on moving regions. This pattern holds consistently across all 20 validation videos.
Figure 5 : Progressive dynamic mask refinement on DAVIS-2016. Column 1-2: input frame and ground truth. Column 3: DA3 cross-attention map computed following Easi3R’s formulation, which fails to isolate dynamic objects due to DA3’s depth-focused training. Columns 4-6: Dyna3’s progressive refinement from motion map Si (best-match feature search), to refined motion map S~i (masked by SAM3 instance regions), to final merged dynamic mask Mi (after motion-aware instance selection). Per-video J-Mean scores demonstrate strong performance on clear subjects while revealing challenges on small objects.
Loss Component
DUSt3R
VGGT
DA3
Ldepth (depth supervision)
✓
✓
✓
Lcamera (pose supervision)
✓
✓
✓
Lpmap (direct pointmap)
✓
✓
–
Lpcd (point cloud)
–
–
✓
Ltrack (correspondence)
✓
✓
–
Table 9 : Training loss comparison across architectures.
MAIS, Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Tencent Hunyuan +3