Interp3R: Continuous-time 3D Geometry Estimation with Frames and Events
Authors: Shuang Guo, Filbert Febryanto, Lei Sun, Luc Van Gool, Guillermo Gallego
Organizations: TU Berlin, Berlin, Germany · Robotics Institute Germany, Berlin, Germany · INSAIT, Sofia University “St. Kliment Ohridski”, Sofia, Bulgaria · Einstein Center Digital Future and SCIoI Excellence Cluster, Berlin, Germany
In recent years, 3D visual foundation models, pioneered by pointmap-based approaches such as DUSt3R, have attracted a lot of interest, achieving impressive accuracy and strong generalization across diverse scenes. However, these methods are inherently limited to recovering scene geometry only at the discrete time instants when images are captured, leaving the scene evolution during the blind time between consecutive frames largely unexplored. We introduce Interp3R, to the best of our knowledge, the first method that enhances pointmap-based models to estimate depth and camera poses at arbitrary time instants. It leverages asynchronous event data to interpolate pointmaps produced by frame-based models, enabling temporally continuous geometric representations. Depth and camera poses are then jointly recovered by aligning the interpolated pointmaps together with those predicted by the underlying frame-based models into a consistent spatial framework. We train Interp3R exclusively on a synthetic dataset, yet demonstrate strong generalization across six datasets, both synthetic and real. Compared with the best two-stage baseline, Interp3R reduces absolute relative depth error by 15%-32% on DSEC and absolute trajectory error by up to 51% on EDS.
Figures & tables
Figure 1: The proposed method takes two images I0,I1 and the events between them E as input (Left), and recovers scene geometry as pointmaps at an arbitrary time instant τ between the images (Right). All available pointmaps can then be used to estimate depth maps and camera poses.
Figure 2: Pipeline for pointmap prediction and interpolation. The pointmap model first takes as input two frames and outputs the source pointmaps. Interp3R then interpolates the pointmap at the arbitrary target time τ∈(0,1) from two directions, with the first part ( 0→τ , “forward”) and the second part ( 1→τ , “backward”) event data.
Figure 3: Coarse-to-fine global alignment. Green nodes Θt collect the camera intrinsics, pose and depth variables at time t . Each Xmn is a per-pixel 3D pointmap of Im expressed in In ’s camera coordinate. In the coarse alignment (blue), the pointmap pairs connect and constrain only the two source nodes at t=0 and t=1 . In the fine alignment (red), the interpolated node at t=τ is connected to both source nodes through the red pointmap pairs and all node variables are refined.
Dataset
Method
skip = 1
skip = 3
skip = 7
skip = 15
Abs Rel ↓
δ<1.25↑
Abs Rel ↓
δ<1.25↑
Abs Rel ↓
δ<1.25↑
Abs Rel ↓
δ<1.25↑
PointOdyssey
RIFE
0.097
0.925
0.106
0.911
0.146
0.891
0.196
0.846
TimeLens
0.087
0.925
0.099
0.912
0.142
0.858
0.215
0.777
VDM-EVFI
0.108
0.889
0.113
0.875
0.118
0.861
0.137
0.839
Interp3R (Ours)
0.0607
0.9541
0.0609
0.9572
0.0719
0.9434
0.1051
0.8863
Sintel
RIFE
0.381
0.538
0.478
0.532
0.515
0.522
0.580
0.507
Table 1: Depth estimation results. Best and second-best are highlighted.
Dataset
Method
skip = 1
skip = 3
skip = 7
skip = 15
ATE ↓
RTE ↓
RRE ↓
ATE ↓
RTE ↓
RRE ↓
ATE ↓
RTE ↓
RRE ↓
ATE ↓
RTE ↓
RRE ↓
Sintel
RIFE
0.301
0.104
1.259
0.292
0.115
1.290
0.328
0.146
1.675
0.434
0.157
2.547
TimeLens
0.321
0.114
1.265
0.385
0.174
1.412
0.566
0.175
2.410
0.599
0.141
3.719
VDM-EVFI
0.319
0.127
1.314
0.339
0.128
1.193
0.405
0.148
1.351
0.379
0.118
1.496
Interp3R (Ours)
0.2143
0.0885
0.8152
0.2919
0.1401
0.9076
0.2188
0.1014
1.2142
0.2790
0.1538
1.8689
Bonn
RIFE
0.011
0.007
0.749
0.012
0.007
0.746
0.013
0.007
0.827
0.015
0.008
0.782
Table 2: Pose evaluation results. Best and second-best are highlighted.
Figure 4: Qualitative comparison of depth estimation on the Bonn dataset (skip = 3).
Figure 5: Qualitative estimated depth results on the EDS dataset (skip = 3).
Method
Depth estimation (DSEC)
Camera pose estimation (EDS)
skip = 1
skip = 3
skip = 7
skip = 1
skip = 3
skip = 7
Abs Rel ↓
δ<1.25↑
Abs Rel ↓
δ<1.25↑
Abs Rel ↓
δ<1.25↑
ATE ↓
RTE ↓
RRE ↓
ATE ↓
RTE ↓
RRE ↓
ATE ↓
RTE ↓
RRE ↓
w/o events
0.1034
0.8902
0.1077
0.8927
0.1226
0.8555
0.0284
0.0096
0.4543
0.0332
0.0148
0.4411
0.0897
0.0562
0.4661
MonST3R + Interp3R
0.1217
0.8552
0.1112
0.8803
0.1180
0.8869
0.0237
0.0053
0.3370
0.0314
0.0078
0.3726
0.0605
0.0331
0.4011
Align3R + Interp3R
0.0993
0.8967
0.0953
0.9104
0.1048
0.8830
0.0242
0.0064
0.3456
0.0226
0.0065
0.3386
0.0403
0.0167
0.3464
Table 3: Ablation studies. Depth estimation on DSEC and camera pose estimation on EDS.
Figure 6: 3D reconstruction of a driving sequence fro m the DSEC dataset (skip = 3).
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Continuous-time depth results with Interp3R. The depth maps at t=0 and t=1 are shown together with predictions at t=1/π , t=0.5 and t=10/4 , illustrating inference beyond a fixed, uniformly spaced temporal grid.
Method
Stage
Number of interpolations M
1
3
5
7
Interp3R
Inference
0.285
0.426
0.635
0.725
Alignment
7.747
10.342
12.331
14.973
Baseline
Inference
0.360
1.030
1.999
3.287
Alignment
12.878
14.126
16.553
20.511
Appendix
Table 5: Computational effort. Inference and alignment times in seconds for M interpolations.
Figure 8: Qualitative results of depth estimation on the DAVIS bear sequence (skip = 3).
Figure 9: Qualitative depth comparison on the Bonn crowd2 sequence (skip = 3).
Figure 10: Qualitative depth comparison on the EDS 01_peanuts_light sequence.
Figure 11: Qualitative depth comparison on the EDS 02_rocket_earth_light sequence.
Figure 12: Camera trajectories on the ECD boxes_6dof sequence. Dashed gray lines indicate ground truth; solid blue lines indicate estimates.
Figure 12: Camera trajectories on ECD boxes_6dof (continued).
Figure 13: Camera trajectories on the ECD dynamic_6dof sequence. Dashed gray lines indicate ground truth; solid blue lines indicate estimates.
Figure 14: Camera trajectories on the EDS 01_peanuts_light sequence. Dashed gray lines indicate ground truth; solid blue lines indicate estimates.
Figure 15: 3D reconstruction on the ECD dynamic_6dof sequence (skip = 3).
Figure 16: 3D reconstruction on the dsec_zurich_city_09d sequence (skip = 3).
We present SCOPE (Scale-Consistent One-Pass Estimation of 3D Geometry), a novel approach for estimating 3D geometry from extended monocular video sequences, where existing methods struggle to maintain both geometric accuracy and temporal consistency across hundreds of frames. Our approach generates affine-invariant 3D point maps with shared parameters across entire sequences, enabling consistent scale-invariant representations. We introduce three key innovations: viewpoint-invariant geometry aligning multi-perspective points in a unified reference frame; appearance-invariant learning enforcing consistency across exponential timescales; and frequency-modulated positioning enabling extrapolation to sequences vastly exceeding training length. Experiments across diverse datasets demonstrate significant improvements, reducing relative point map error by 24.2% and temporal alignment error by 34.9% on ScanNet compared to state-of-the-art methods. Our approach handles challenging scenarios with complex camera trajectories and lighting variations while efficiently processing extended sequences in a single pass. Project page: https://scope3d.github.io/.
Zheng Zhang, Lihe Yang, Tianyu Yang +6
The University of Hong Kong, Hong Kong · Alibaba Group, Hong Kong · Alibaba Group, China +2
Recent feed-forward geometry foundation models have demonstrated impressive generalization by recovering depth and poses in a single forward pass. However, these models are typically constrained by a global coordinate frame assumption. This dependency becomes a significant bottleneck for long-context and streaming reconstruction, as it forces the network to maintain an arbitrary temporal origin and handle translation magnitudes that grow unbounded over time. Our solution, which we call R3, employs relative regression. We employ a lightweight MLP to predict confidence-weighted relative constraints. These confidences serve as a unified anchor: weighting losses during training and guiding pose aggregation during inference. R3 supports both full-context offline reconstruction and causal, bounded-memory streaming. Our evaluation in both offline and streaming settings validates the effectiveness of our relative mechanism. Project page: https://kevinxu02.github.io/r3-site
Congrong Xu, Huachen Gao, Xingyu Chen +3
University of Michigan · Westlake University · NVIDIA
Video depth estimation extends monocular prediction into the temporal domain to ensure coherence. However, existing methods often suffer from spatial blurring in fine-detail regions and temporal inconsistencies. We argue that current approaches, which primarily rely on temporal smoothing via Transformers, struggle to maintain strict 3D geometric consistency-particularly under rotations or drastic view changes. To address this, we propose GemDepth, a framework built on the insight that an explicit awareness of camera motion and global 3D structure is a prerequisite for 3D consistency. Distinctively, GemDepth introduces a Geometry-Embedding Module (GEM) that predicts inter-frame camera poses to generate implicit geometric embeddings. This injection of motion priors equips the network with intrinsic 3D perception and alignment capabilities. Guided by these geometric cues, our Alternating Spatio-Temporal Transformer (ASTT) captures latent point-level correspondences to simultaneously enhance spatial precision for sharp details and enforce rigorous temporal consistency. Furthermore, GemDepth employs a data-efficient training strategy, effectively bridging the gap between high efficiency and robust geometric consistency. As shown in Fig.2, comprehensive evaluations demonstrate that GemDepth achieves state-of-the-art performance across multiple datasets, particularly in complex dynamic scenarios. The code is publicly available at: https://github.com/Yuecheng919/GemDepth.
Yuecheng Liu, Junda Cheng, Longliang Liu +4
1Huazhong University of Science & Technology · 2Carizon · 3Optics Valley Laboratory