Interp3R: Continuous-time 3D Geometry Estimation with Frames and Events
Authors: Shuang Guo, Filbert Febryanto, Lei Sun, Luc Van Gool, Guillermo Gallego
Organizations: TU Berlin, Berlin, Germany · Robotics Institute Germany, Berlin, Germany · INSAIT, Sofia University “St. Kliment Ohridski”, Sofia, Bulgaria · Einstein Center Digital Future and SCIoI Excellence Cluster, Berlin, Germany
In recent years, 3D visual foundation models, pioneered by pointmap-based approaches such as DUSt3R, have attracted a lot of interest, achieving impressive accuracy and strong generalization across diverse scenes. However, these methods are inherently limited to recovering scene geometry only at the discrete time instants when images are captured, leaving the scene evolution during the blind time between consecutive frames largely unexplored. We introduce Interp3R, to the best of our knowledge, the first method that enhances pointmap-based models to estimate depth and camera poses at arbitrary time instants. It leverages asynchronous event data to interpolate pointmaps produced by frame-based models, enabling temporally continuous geometric representations. Depth and camera poses are then jointly recovered by aligning the interpolated pointmaps together with those predicted by the underlying frame-based models into a consistent spatial framework. We train Interp3R exclusively on a synthetic dataset, yet demonstrate strong generalization across six datasets, both synthetic and real. Compared with the best two-stage baseline, Interp3R reduces absolute relative depth error by 15%-32% on DSEC and absolute trajectory error by up to 51% on EDS.
Figures & tables
Figure 1: The proposed method takes two images I0,I1 and the events between them E as input (Left), and recovers scene geometry as pointmaps at an arbitrary time instant τ between the images (Right). All available pointmaps can then be used to estimate depth maps and camera poses.
Figure 2: Pipeline for pointmap prediction and interpolation. The pointmap model first takes as input two frames and outputs the source pointmaps. Interp3R then interpolates the pointmap at the arbitrary target time τ∈(0,1) from two directions, with the first part ( 0→τ , “forward”) and the second part ( 1→τ , “backward”) event data.
Figure 3: Coarse-to-fine global alignment. Green nodes Θt collect the camera intrinsics, pose and depth variables at time t . Each Xmn is a per-pixel 3D pointmap of Im expressed in In ’s camera coordinate. In the coarse alignment (blue), the pointmap pairs connect and constrain only the two source nodes at t=0 and t=1 . In the fine alignment (red), the interpolated node at t=τ is connected to both source nodes through the red pointmap pairs and all node variables are refined.
Dataset
Method
skip = 1
skip = 3
skip = 7
skip = 15
Abs Rel ↓
δ<1.25↑
Abs Rel ↓
δ<1.25↑
Abs Rel ↓
δ<1.25↑
Abs Rel ↓
δ<1.25↑
PointOdyssey
RIFE
0.097
0.925
0.106
0.911
0.146
0.891
0.196
0.846
TimeLens
0.087
0.925
0.099
0.912
0.142
0.858
0.215
0.777
VDM-EVFI
0.108
0.889
0.113
0.875
0.118
0.861
0.137
0.839
Interp3R (Ours)
0.0607
0.9541
0.0609
0.9572
0.0719
0.9434
0.1051
0.8863
Sintel
RIFE
0.381
0.538
0.478
0.532
0.515
0.522
0.580
0.507
Table 1: Depth estimation results. Best and second-best are highlighted.
Dataset
Method
skip = 1
skip = 3
skip = 7
skip = 15
ATE ↓
RTE ↓
RRE ↓
ATE ↓
RTE ↓
RRE ↓
ATE ↓
RTE ↓
RRE ↓
ATE ↓
RTE ↓
RRE ↓
Sintel
RIFE
0.301
0.104
1.259
0.292
0.115
1.290
0.328
0.146
1.675
0.434
0.157
2.547
TimeLens
0.321
0.114
1.265
0.385
0.174
1.412
0.566
0.175
2.410
0.599
0.141
3.719
VDM-EVFI
0.319
0.127
1.314
0.339
0.128
1.193
0.405
0.148
1.351
0.379
0.118
1.496
Interp3R (Ours)
0.2143
0.0885
0.8152
0.2919
0.1401
0.9076
0.2188
0.1014
1.2142
0.2790
0.1538
1.8689
Bonn
RIFE
0.011
0.007
0.749
0.012
0.007
0.746
0.013
0.007
0.827
0.015
0.008
0.782
Table 2: Pose evaluation results. Best and second-best are highlighted.
Figure 4: Qualitative comparison of depth estimation on the Bonn dataset (skip = 3).
Figure 5: Qualitative estimated depth results on the EDS dataset (skip = 3).
Method
Depth estimation (DSEC)
Camera pose estimation (EDS)
skip = 1
skip = 3
skip = 7
skip = 1
skip = 3
skip = 7
Abs Rel ↓
δ<1.25↑
Abs Rel ↓
δ<1.25↑
Abs Rel ↓
δ<1.25↑
ATE ↓
RTE ↓
RRE ↓
ATE ↓
RTE ↓
RRE ↓
ATE ↓
RTE ↓
RRE ↓
w/o events
0.1034
0.8902
0.1077
0.8927
0.1226
0.8555
0.0284
0.0096
0.4543
0.0332
0.0148
0.4411
0.0897
0.0562
0.4661
MonST3R + Interp3R
0.1217
0.8552
0.1112
0.8803
0.1180
0.8869
0.0237
0.0053
0.3370
0.0314
0.0078
0.3726
0.0605
0.0331
0.4011
Align3R + Interp3R
0.0993
0.8967
0.0953
0.9104
0.1048
0.8830
0.0242
0.0064
0.3456
0.0226
0.0065
0.3386
0.0403
0.0167
0.3464
Table 3: Ablation studies. Depth estimation on DSEC and camera pose estimation on EDS.
Figure 6: 3D reconstruction of a driving sequence fro m the DSEC dataset (skip = 3).
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Continuous-time depth results with Interp3R. The depth maps at t=0 and t=1 are shown together with predictions at t=1/π , t=0.5 and t=10/4 , illustrating inference beyond a fixed, uniformly spaced temporal grid.
Method
Stage
Number of interpolations M
1
3
5
7
Interp3R
Inference
0.285
0.426
0.635
0.725
Alignment
7.747
10.342
12.331
14.973
Baseline
Inference
0.360
1.030
1.999
3.287
Alignment
12.878
14.126
16.553
20.511
Appendix
Table 5: Computational effort. Inference and alignment times in seconds for M interpolations.
Figure 8: Qualitative results of depth estimation on the DAVIS bear sequence (skip = 3).
Figure 9: Qualitative depth comparison on the Bonn crowd2 sequence (skip = 3).
Figure 10: Qualitative depth comparison on the EDS 01_peanuts_light sequence.
Figure 11: Qualitative depth comparison on the EDS 02_rocket_earth_light sequence.
Figure 12: Camera trajectories on the ECD boxes_6dof sequence. Dashed gray lines indicate ground truth; solid blue lines indicate estimates.
Figure 12: Camera trajectories on ECD boxes_6dof (continued).
Figure 13: Camera trajectories on the ECD dynamic_6dof sequence. Dashed gray lines indicate ground truth; solid blue lines indicate estimates.
Figure 14: Camera trajectories on the EDS 01_peanuts_light sequence. Dashed gray lines indicate ground truth; solid blue lines indicate estimates.
Figure 15: 3D reconstruction on the ECD dynamic_6dof sequence (skip = 3).
Figure 16: 3D reconstruction on the dsec_zurich_city_09d sequence (skip = 3).