Embodied systems need geometric perception that exploits available observations beyond images alone. Recent feed-forward 3D models incorporate geometric priors, including camera poses, intrinsics, and depth. However, handling noisy poses, preserving accurate priors, and recovering physical scale require more than simply accepting these inputs. We introduce \emph{Vision-Prior Geometry Grounded Transformer} (VPGGT), a VGGT-based framework that extends OmniVGGT for prior-aware embodied perception. We formulate sensor-motivated pose corruptions from ground-truth trajectories for training and introduce a parameter-free \emph{prior residual connection} (PRC) to mitigate \emph{prior dilution}, where predictions are less accurate than their supplied pose priors. Our noise formulation targets camera poses; supplied intrinsics and depth receive no additional corruption. We further introduce \emph{Metric Global Attention}, which conditions a global scale token on available pose and depth scales and predicts a shared metric scaling factor for the geometric outputs. Experiments across four datasets show that \emph{PRC} improves translation-direction accuracy and joint pose AUC over a matched training baseline when camera priors are provided for all views, under both exact and corrupted poses. These results support explicit prior access during refinement as a useful addition to feature-level conditioning.
Figures & tables
Figure 1: VPGGT is an embodied perception foundation model that combines images with optional camera parameters and depth observations to predict camera parameters, depth maps, and pointmaps with a shared metric scale. By leveraging geometric priors supplied by upstream sensing or estimation, the model circumvents the direct processing of raw sensor streams.
Figure 2: Prior dilution during camera refinement. A0 is released OmniVGGT; A1 uses the same checkpoint with PRC enabled, without additional training. All views receive camera priors; no depth is supplied. Colors distinguish methods and markers distinguish datasets. Dashed lines show input-prior scores. Reset is a state intervention, not a learned iteration.
Figure 3: Overview of VPGGT. Images and optional camera or depth priors are encoded and processed by a shared alternating-attention backbone. PRC places available camera priors in the camera head’s refinement state after its initial prediction. A Scale Encoder conditions a global token on available metric-scale statistics, using learned placeholders when they are absent. The camera, depth, and pointmap heads predict geometry, and a scale head supplies a shared factor that converts translations, depths, and points to metric units.
Figure 4: Prior Residual Connection. PRC retains feature-level conditioning and supplies aligned camera priors as the state from which subsequent residual updates are predicted.
Training data
Translation
Rotation
TartanAir
Time-dependent
Time-dependent
TUM RGB-D
Time-dependent
Time-dependent
Virtual KITTI 2
Step-dependent
Time-dependent
C3VD
Independent
Independent
Table 2: Pose-corruption profiles. Drift is applied to relative-motion increments.
TUM RGB-D
C3VD
Method
R↑
T↑
A↑
AR ↓
δ1↑
ATE ↓
R↑
T↑
A↑
AR ↓
δ1↑
ATE ↓
RGB only
OmniVGGT
99.67
38.00
61.94
0.0782
92.57
0.00554
87.51
14.67
40.88
0.5114
29.89
0.1962
VPGGT
96.22
50.44
69.38
0.0607
94.21
0.00444
100.00
27.41
62.58
0.0563
97.86
0.0759
w/ Ce
OmniVGGT
100.00
70.78
78.95
0.0779
92.18
0.00243
98.62
45.23
71.44
0.4848
30.96
0.0769
Table 3: Pose and depth on TUM RGB-D and C3VD. R , T , and A denote RRA@ 5∘ , RTA@ 5∘ , and AUC@ 30∘ in percent. AR is median-aligned depth AbsRel; δ1 is depth accuracy at threshold 1.25, in percent. ATE uses Sim(3) alignment, in meters for TUM and millimeters for C3VD. Bold marks the unique best network within each input setting; ties at displayed precision are unmarked.
Sewerage
Ocean
Method
R↑
T↑
A↑
AR ↓
δ1↑
ATE ↓
R↑
T↑
A↑
AR ↓
δ1↑
ATE ↓
RGB only
VGGT
94.67
67.44
84.51
0.0867
89.71
0.0892
91.22
42.78
75.51
0.0868
92.95
0.1786
OmniVGGT
88.56
69.11
84.63
0.0940
87.92
0.0804
90.67
45.00
78.14
0.0777
94.89
0.1591
VPGGT
83.33
53.56
78.84
0.1009
86.96
0.1148
90.33
42.78
75.68
0.0804
95.30
0.2079
w/ Ce
Table 4: Pose and depth on TartanAir-V2 subsets. Metrics and boldface follow Table 3 ; ATE is in meters. Sewerage uses P001/P002 and Ocean uses P000/P001 in both Easy and Hard, with five windows per sequence. These results cover the selected subsets, not the full benchmark.
Input
TUM RGB-D
C3VD
RGB
0.1003/5.07/0.02226
0.0974/7.67/0.204
Ce
0.0977/4.88/0.02058
0.0900/6.80/0.173
Cn
0.0976/4.78/0.01811
0.0899/6.72/0.176
Table 5: Metric reconstruction of full VPGGT. Entries are metric AbsRel/ EscaleX (%)/ATE-SE(3). No post-hoc scale fitting to evaluation ground truth is used. ATE is in millimeters for C3VD and meters for the other datasets. All three metrics are lower-is-better; different input settings are not ranked against one another.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Unit
TartanAir
TUM
VKITTI2
C3VD
Heading-rate bias std.
rad/s
4.37×10−3
π/10800
π/36000
–
Heading white std.
deg
0.1100
0.08
0.03
–
Other-axis white std.
deg
0.02
0.03
0.02
–
Velocity-bias std.
m/s
0.076
0.0045
–
–
Translation white std.
m
0.001
0.0009
–
–
Translation scale-bias std.
–
–
–
0.01
–
Appendix
Table 6: Base pose-corruption parameters before severity scaling. Time-scaled white increments use a reference interval of 0.1 s.
Figure 5: Pose-prior corruption in the training data. Two examples are shown for each dataset. Within each panel, the left plot displays the complete trajectory, while the right plot displays a selected segment of the same ground-truth and corrupted trajectories. Colors identify the ground truth and corruption realizations, as indicated by the plot legends. Ground-truth and corrupted poses share a coordinate system, without independently fitted trajectory or scale alignment. Spatial units and segment information are provided in the plots.
Dataset
Nominal FPS
Maximum depth (m)
PNG divisor
TartanAir-V2
10
500
–
Virtual KITTI 2
10
655
100
TUM RGB-D
30
10
–
C3VD
30
0.15
–
Appendix
Table 7: Dataset-specific evaluation settings. Depth limits apply to ground-truth scoring support. The PNG divisor is used only for encoded PNG depth.
D
C
Pose
Model
R@2 ↑
R@5 ↑
R@15 ↑
T@2 ↑
T@5 ↑
T@15 ↑
A@15 ↑
A@30 ↑
0
0
–
C1
58.52
88.69
96.46
10.49
46.48
88.31
58.07
75.01
0
0
–
C2
58.97
88.37
96.12
10.69
46.35
87.78
57.77
74.63
30
0
–
C1
60.92
88.91
96.28
9.30
42.18
86.72
55.78
73.39
30
0
–
C2
60.56
88.48
96.24
9.32
41.98
86.39
55.50
73.12
50
0
–
C1
62.41
89.00
96.03
9.04
41.86
86.26
55.49
73.01
50
0
–
C2
62.32
88.78
95.96
8.99
41.66
85.77
55.12
72.77
Appendix
Table 8: Complete matched-training results on TartanAir-V2. C1: OmniVGGT-trained; C2: VPGGT (PRC only). Bold marks the better displayed score within each matched pair. Angular thresholds are in degrees; angular scores and δ1 are percentages.
D
C
Pose
Model
R@2 ↑
R@5 ↑
R@15 ↑
T@2 ↑
T@5 ↑
T@15 ↑
A@15 ↑
A@30 ↑
0
0
–
C1
100.00
100.00
100.00
97.87
99.65
99.81
98.54
99.21
0
0
–
C2
100.00
100.00
100.00
96.98
99.68
99.84
98.43
99.15
30
0
–
C1
100.00
100.00
100.00
96.92
99.65
99.84
98.40
99.14
30
0
–
C2
100.00
100.00
100.00
95.46
99.68
99.84
98.17
99.01
50
0
–
C1
99.43
100.00
100.00
95.75
99.40
99.75
97.97
98.89
50
0
–
C2
98.89
100.00
100.00
93.97
99.14
99.65
97.46
98.57
Appendix
Table 9: Complete matched-training results on Virtual KITTI 2. C1: OmniVGGT-trained; C2: VPGGT (PRC only). Bold marks the better displayed score within each matched pair. Angular thresholds are in degrees; angular scores and δ1 are percentages.
D
C
Pose
Model
R@2 ↑
R@5 ↑
R@15 ↑
T@2 ↑
T@5 ↑
T@15 ↑
A@15 ↑
A@30 ↑
0
0
–
C1
94.78
96.33
97.22
22.22
51.22
75.78
56.78
69.17
0
0
–
C2
95.44
97.22
97.22
20.22
50.33
76.56
56.53
69.33
30
0
–
C1
94.11
97.00
97.22
17.44
47.33
73.00
52.87
66.46
30
0
–
C2
96.56
97.22
97.22
17.11
48.56
73.44
52.90
66.66
50
0
–
C1
94.78
95.56
97.22
17.11
45.78
73.56
52.44
66.21
50
0
–
C2
96.89
97.22
97.22
16.56
45.89
73.33
52.10
66.46
Appendix
Table 10: Complete matched-training results on TUM RGB-D. C1: OmniVGGT-trained; C2: VPGGT (PRC only). Bold marks the better displayed score within each matched pair. Angular thresholds are in degrees; angular scores and δ1 are percentages.
D
C
Pose
Model
R@2 ↑
R@5 ↑
R@15 ↑
T@2 ↑
T@5 ↑
T@15 ↑
A@15 ↑
A@30 ↑
0
0
–
C1
100.00
100.00
100.00
8.64
29.19
71.85
43.97
61.55
0
0
–
C2
100.00
100.00
100.00
7.65
27.75
72.30
42.73
61.49
30
0
–
C1
99.95
100.00
100.00
6.32
25.33
60.30
36.48
54.68
30
0
–
C2
100.00
100.00
100.00
6.17
26.12
61.38
36.58
54.92
50
0
–
C1
100.00
100.00
100.00
7.75
30.52
63.41
39.78
57.00
50
0
–
C2
100.00
100.00
100.00
8.40
30.72
64.69
40.18
57.29
Appendix
Table 11: Complete matched-training results on C3VD. C1: OmniVGGT-trained; C2: VPGGT (PRC only). Bold marks the better displayed score within each matched pair. Angular thresholds are in degrees; angular scores and δ1 are percentages.