Recent event-based depth estimation methods successfully transfer geometric priors from vision foundation models via cross-modal distillation. However, their reliance on synchronized RGB-event pairs or depth annotations during training severely restricts practical deployment. To overcome this bottleneck, we propose SFE-VGGT, a novel source-free framework that distills the geometric priors of VGGT to the event domain without any paired RGB observations. Our core idea is to reconstruct surrogate frames directly from the target event stream to act as a frozen geometric teacher, entirely eliminating the need for genuine source RGB data. Crucially, as these surrogate frames inherently yield imperfect and spatially varying supervision, directly distilling from them propagates artifacts. To resolve this, we introduce a novel reliability-aware distillation strategy. This includes Density-Aware Feature Distillation to emphasize informative event regions, and Confidence-Weighted Depth Distillation to dynamically regulate supervision based on relative teacher-student prediction confidence. Meanwhile, we propose a Cross-Frame Relational Consistency loss that enforces temporal geometric stability using reliable inter-frame correspondences, bypassing the need for temporally consistent teacher's depth. Extensive experiments demonstrate that, despite source-free, our SFE-VGGT closely matches the accuracy of RGB-dependent baselines under standard conditions and significantly surpasses them in challenging nighttime scenarios. Across MVSEC nighttime sequences, SFE-VGGT reduces the average 10 m depth error by 15.3% compared with EventVGGT. Moreover, our method exhibits robust zero-shot generalization across real-world datasets, proving that highly effective geometric priors can be transferred to event cameras using strictly source-free supervision.
Figures & tables
Fig. 2: Overview of the proposed SFE-VGGT framework. During training, the event data is processed by a pretrained E2VID model to construct surrogate frames for the frozen VGGT teacher, while the student directly processes event representations. Knowledge is transferred through (a) Density-Aware Feature Distillation (DAFD), which weights patch-level feature alignment by event density, (b) Confidence-Weighted Depth Distillation (CWDD), which weights depth supervision using relative prediction confidence, and (c) Cross-Frame Relational Consistency (CFRC), which enforces consistent depth relations across matched points between frames. During inference, only the event-based student is retained.
Method
Training
Inference
10m ↓
20m ↓
30m ↓
Source-free
Depth GT
Input
E2Depth [ 5 ]
–
✓
E
1.79
5.35
8.31
RAMNet [ 6 ]
–
✓
E+I
0.81
2.26
3.58
HMNet [ 25 ]
–
✓
E+I
0.55
1.80
3.27
ER-F2D [ 7 ]
–
✓
E+I
0.67
1.69
2.81
SRFNet [ 2 ]
–
✓
E+I
1.27
1.68
2.76
TABLE I: Mean absolute depth error on EventScape in meters. E denotes event-only input and E+I denotes event and RGB image input. “Depth GT” indicates whether metric ground-truth depth values are used as training supervision. Best and second-best results are highlighted.
Method
Training
Inference
Night1
Night2
Night3
Day1
Source-free
Depth GT
Input
10m ↓
20m ↓
30m ↓
10m ↓
20m ↓
30m ↓
10m ↓
20m ↓
30m ↓
10m ↓
20m ↓
30m ↓
E2Depth [ 5 ]
–
✓
E
3.38
3.82
4.46
1.67
2.63
3.58
1.42
2.33
3.18
1.67
2.64
3.13
RAMNet [ 6 ]
–
✓
E+I
2.50
3.19
3.82
1.21
2.31
3.28
1.01
2.34
3.43
1.39
2.17
2.76
EvT+ [ 22 ]
–
✓
E+I
1.45
2.10
2.88
1.48
2.13
2.90
1.38
2.03
2.77
1.24
1.91
2.36
HMNet [ 25 ]
–
✓
E+I
1.50
2.48
3.19
1.36
2.25
2.96
1.27
2.17
2.86
1.22
2.21
2.68
ER-F2D [ 7 ]
–
✓
E+I
1.58
2.24
2.78
1.54
2.23
2.95
1.24
1.96
2.81
1.34
2.25
2.62
TABLE II: Mean absolute depth error on MVSEC in meters. E denotes event-only input and E+I denotes event and RGB image input. “Depth GT” indicates whether metric ground-truth depth values are used as training supervision. Among distillation-based methods, the Best and second-best results are highlighted. Overall best results across all methods are shown in bold .
Fig. 3: Qualitative results on MVSEC dataset.
Method
Input
10m ↓
20m ↓
30m ↓
RAMNet [ 6 ]
E+I
2.62
11.26
19.11
SRFNet [ 2 ]
E+I
1.50
3.57
6.12
EventDAM [ 4 ]
E
1.20
2.60
5.18
EventVGGT [ 8 ]
E
0.54
0.89
1.33
SFE-VGGT
E
1.04
1.44
2.35
TABLE III: Zero-shot depth estimation on DENSE. Mean absolute depth error is reported in meters.
Method
Input
10m ↓
20m ↓
30m ↓
δ1↑
δ2↑
δ3↑
VGGT [ 9 ]
E
2.42
2.71
3.44
0.42
0.69
0.84
VGGT [ 9 ]
I
2.31
2.68
3.33
0.41
0.71
0.88
EventVGGT [ 8 ]
E
1.67
2.02
2.61
0.57
0.82
0.93
SFE-VGGT
E
1.41
2.04
2.92
0.48
0.71
0.83
TABLE IV: Direct VGGT transfer analysis on MVSEC Night1. Mean absolute depth error is reported in meters.
LDAFD
LCWDD
LCFRC
10m ↓
20m ↓
30m ↓
δ1↑
δ2↑
δ3↑
✓
0.620
0.951
1.337
0.740
0.880
0.948
✓
✓
0.622
0.947
1.327
0.739
0.875
0.942
✓
✓
0.604
0.948
1.336
0.752
0.887
0.947
✓
✓
✓
0.617
0.941
1.322
0.743
0.880
0.948
TABLE V: Ablation of the proposed training objectives.
Settings
10m ↓
20m ↓
30m ↓
w/o event-density weighting ( LDAFD )
0.633
0.965
1.347
w/o relative-confidence weighting ( LCWDD )
0.631
0.954
1.334
w/o RANSAC ( LCFRC )
0.630
0.944
1.325
Full model (all activated)
0.617
0.941
1.322
TABLE VI: Ablation of key components within the proposed objectives.