Embodied systems need geometric perception that exploits available observations beyond images alone. Recent feed-forward 3D models incorporate geometric priors, including camera poses, intrinsics, and depth. However, handling noisy poses, preserving accurate priors, and recovering physical scale require more than simply accepting these inputs. We introduce \emph{Vision-Prior Geometry Grounded Transformer} (VPGGT), a VGGT-based framework that extends OmniVGGT for prior-aware embodied perception. We formulate sensor-motivated pose corruptions from ground-truth trajectories for training and introduce a parameter-free \emph{prior residual connection} (PRC) to mitigate \emph{prior dilution}, where predictions are less accurate than their supplied pose priors. Our noise formulation targets camera poses; supplied intrinsics and depth receive no additional corruption. We further introduce \emph{Metric Global Attention}, which conditions a global scale token on available pose and depth scales and predicts a shared metric scaling factor for the geometric outputs. Experiments across four datasets show that \emph{PRC} improves translation-direction accuracy and joint pose AUC over a matched training baseline when camera priors are provided for all views, under both exact and corrupted poses. These results support explicit prior access during refinement as a useful addition to feature-level conditioning.
Figures & tables
Figure 1: VPGGT is an embodied perception foundation model that combines images with optional camera parameters and depth observations to predict camera parameters, depth maps, and pointmaps with a shared metric scale. By leveraging geometric priors supplied by upstream sensing or estimation, the model circumvents the direct processing of raw sensor streams.
Figure 2: Prior dilution during camera refinement. A0 is released OmniVGGT; A1 uses the same checkpoint with PRC enabled, without additional training. All views receive camera priors; no depth is supplied. Colors distinguish methods and markers distinguish datasets. Dashed lines show input-prior scores. Reset is a state intervention, not a learned iteration.
Figure 3: Overview of VPGGT. Images and optional camera or depth priors are encoded and processed by a shared alternating-attention backbone. PRC places available camera priors in the camera head’s refinement state after its initial prediction. A Scale Encoder conditions a global token on available metric-scale statistics, using learned placeholders when they are absent. The camera, depth, and pointmap heads predict geometry, and a scale head supplies a shared factor that converts translations, depths, and points to metric units.
Figure 4: Prior Residual Connection. PRC retains feature-level conditioning and supplies aligned camera priors as the state from which subsequent residual updates are predicted.
Training data
Translation
Rotation
TartanAir
Time-dependent
Time-dependent
TUM RGB-D
Time-dependent
Time-dependent
Virtual KITTI 2
Step-dependent
Time-dependent
C3VD
Independent
Independent
Table 2: Pose-corruption profiles. Drift is applied to relative-motion increments.
TUM RGB-D
C3VD
Method
R↑
T↑
A↑
AR ↓
δ1↑
ATE ↓
R↑
T↑
A↑
AR ↓
δ1↑
ATE ↓
RGB only
OmniVGGT
99.67
38.00
61.94
0.0782
92.57
0.00554
87.51
14.67
40.88
0.5114
29.89
0.1962
VPGGT
96.22
50.44
69.38
0.0607
94.21
0.00444
100.00
27.41
62.58
0.0563
97.86
0.0759
w/ Ce
OmniVGGT
100.00
70.78
78.95
0.0779
92.18
0.00243
98.62
45.23
71.44
0.4848
30.96
0.0769
Table 3: Pose and depth on TUM RGB-D and C3VD. R , T , and A denote RRA@ 5∘ , RTA@ 5∘ , and AUC@ 30∘ in percent. AR is median-aligned depth AbsRel; δ1 is depth accuracy at threshold 1.25, in percent. ATE uses Sim(3) alignment, in meters for TUM and millimeters for C3VD. Bold marks the unique best network within each input setting; ties at displayed precision are unmarked.
Sewerage
Ocean
Method
R↑
T↑
A↑
AR ↓
δ1↑
ATE ↓
R↑
T↑
A↑
AR ↓
δ1↑
ATE ↓
RGB only
VGGT
94.67
67.44
84.51
0.0867
89.71
0.0892
91.22
42.78
75.51
0.0868
92.95
0.1786
OmniVGGT
88.56
69.11
84.63
0.0940
87.92
0.0804
90.67
45.00
78.14
0.0777
94.89
0.1591
VPGGT
83.33
53.56
78.84
0.1009
86.96
0.1148
90.33
42.78
75.68
0.0804
95.30
0.2079
w/ Ce
Table 4: Pose and depth on TartanAir-V2 subsets. Metrics and boldface follow Table 3 ; ATE is in meters. Sewerage uses P001/P002 and Ocean uses P000/P001 in both Easy and Hard, with five windows per sequence. These results cover the selected subsets, not the full benchmark.
Input
TUM RGB-D
C3VD
RGB
0.1003/5.07/0.02226
0.0974/7.67/0.204
Ce
0.0977/4.88/0.02058
0.0900/6.80/0.173
Cn
0.0976/4.78/0.01811
0.0899/6.72/0.176
Table 5: Metric reconstruction of full VPGGT. Entries are metric AbsRel/ EscaleX (%)/ATE-SE(3). No post-hoc scale fitting to evaluation ground truth is used. ATE is in millimeters for C3VD and meters for the other datasets. All three metrics are lower-is-better; different input settings are not ranked against one another.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Unit
TartanAir
TUM
VKITTI2
C3VD
Heading-rate bias std.
rad/s
4.37×10−3
π/10800
π/36000
–
Heading white std.
deg
0.1100
0.08
0.03
–
Other-axis white std.
deg
0.02
0.03
0.02
–
Velocity-bias std.
m/s
0.076
0.0045
–
–
Translation white std.
m
0.001
0.0009
–
–
Translation scale-bias std.
–
–
–
0.01
–
Appendix
Table 6: Base pose-corruption parameters before severity scaling. Time-scaled white increments use a reference interval of 0.1 s.
Figure 5: Pose-prior corruption in the training data. Two examples are shown for each dataset. Within each panel, the left plot displays the complete trajectory, while the right plot displays a selected segment of the same ground-truth and corrupted trajectories. Colors identify the ground truth and corruption realizations, as indicated by the plot legends. Ground-truth and corrupted poses share a coordinate system, without independently fitted trajectory or scale alignment. Spatial units and segment information are provided in the plots.
Dataset
Nominal FPS
Maximum depth (m)
PNG divisor
TartanAir-V2
10
500
–
Virtual KITTI 2
10
655
100
TUM RGB-D
30
10
–
C3VD
30
0.15
–
Appendix
Table 7: Dataset-specific evaluation settings. Depth limits apply to ground-truth scoring support. The PNG divisor is used only for encoded PNG depth.
D
C
Pose
Model
R@2 ↑
R@5 ↑
R@15 ↑
T@2 ↑
T@5 ↑
T@15 ↑
A@15 ↑
A@30 ↑
0
0
–
C1
58.52
88.69
96.46
10.49
46.48
88.31
58.07
75.01
0
0
–
C2
58.97
88.37
96.12
10.69
46.35
87.78
57.77
74.63
30
0
–
C1
60.92
88.91
96.28
9.30
42.18
86.72
55.78
73.39
30
0
–
C2
60.56
88.48
96.24
9.32
41.98
86.39
55.50
73.12
50
0
–
C1
62.41
89.00
96.03
9.04
41.86
86.26
55.49
73.01
50
0
–
C2
62.32
88.78
95.96
8.99
41.66
85.77
55.12
72.77
Appendix
Table 8: Complete matched-training results on TartanAir-V2. C1: OmniVGGT-trained; C2: VPGGT (PRC only). Bold marks the better displayed score within each matched pair. Angular thresholds are in degrees; angular scores and δ1 are percentages.
D
C
Pose
Model
R@2 ↑
R@5 ↑
R@15 ↑
T@2 ↑
T@5 ↑
T@15 ↑
A@15 ↑
A@30 ↑
0
0
–
C1
100.00
100.00
100.00
97.87
99.65
99.81
98.54
99.21
0
0
–
C2
100.00
100.00
100.00
96.98
99.68
99.84
98.43
99.15
30
0
–
C1
100.00
100.00
100.00
96.92
99.65
99.84
98.40
99.14
30
0
–
C2
100.00
100.00
100.00
95.46
99.68
99.84
98.17
99.01
50
0
–
C1
99.43
100.00
100.00
95.75
99.40
99.75
97.97
98.89
50
0
–
C2
98.89
100.00
100.00
93.97
99.14
99.65
97.46
98.57
Appendix
Table 9: Complete matched-training results on Virtual KITTI 2. C1: OmniVGGT-trained; C2: VPGGT (PRC only). Bold marks the better displayed score within each matched pair. Angular thresholds are in degrees; angular scores and δ1 are percentages.
D
C
Pose
Model
R@2 ↑
R@5 ↑
R@15 ↑
T@2 ↑
T@5 ↑
T@15 ↑
A@15 ↑
A@30 ↑
0
0
–
C1
94.78
96.33
97.22
22.22
51.22
75.78
56.78
69.17
0
0
–
C2
95.44
97.22
97.22
20.22
50.33
76.56
56.53
69.33
30
0
–
C1
94.11
97.00
97.22
17.44
47.33
73.00
52.87
66.46
30
0
–
C2
96.56
97.22
97.22
17.11
48.56
73.44
52.90
66.66
50
0
–
C1
94.78
95.56
97.22
17.11
45.78
73.56
52.44
66.21
50
0
–
C2
96.89
97.22
97.22
16.56
45.89
73.33
52.10
66.46
Appendix
Table 10: Complete matched-training results on TUM RGB-D. C1: OmniVGGT-trained; C2: VPGGT (PRC only). Bold marks the better displayed score within each matched pair. Angular thresholds are in degrees; angular scores and δ1 are percentages.
D
C
Pose
Model
R@2 ↑
R@5 ↑
R@15 ↑
T@2 ↑
T@5 ↑
T@15 ↑
A@15 ↑
A@30 ↑
0
0
–
C1
100.00
100.00
100.00
8.64
29.19
71.85
43.97
61.55
0
0
–
C2
100.00
100.00
100.00
7.65
27.75
72.30
42.73
61.49
30
0
–
C1
99.95
100.00
100.00
6.32
25.33
60.30
36.48
54.68
30
0
–
C2
100.00
100.00
100.00
6.17
26.12
61.38
36.58
54.92
50
0
–
C1
100.00
100.00
100.00
7.75
30.52
63.41
39.78
57.00
50
0
–
C2
100.00
100.00
100.00
8.40
30.72
64.69
40.18
57.29
Appendix
Table 11: Complete matched-training results on C3VD. C1: OmniVGGT-trained; C2: VPGGT (PRC only). Bold marks the better displayed score within each matched pair. Angular thresholds are in degrees; angular scores and δ1 are percentages.
Vision-Language-Action (VLA) models drive next-generation autonomous systems, but training them requires scalable, high-quality annotations from complex environments. Current cloud pipelines rely on generic vision-language models (VLMs) that lack geometric reasoning and domain semantics due to their 2D image-text pretraining. To address this mismatch, we propose XEmbodied, a cloud-side foundation model that endows VLMs with intrinsic 3D geometric awareness and interaction with physical cues (e.g., occupancy grids, 3D boxes). Instead of treating geometry as auxiliary input, XEmbodied integrates geometric representations via a structured 3D Adapter and distills physical signals into context tokens using an Efficient Image-Embodied Adapter. Through progressive domain curriculum and reinforcement learning post-training, XEmbodied preserves general capabilities while demonstrating robust performance across 18 public benchmarks. It significantly improves spatial reasoning, traffic semantics, embodied affordance, and out-of-distribution generalization for large-scale scenario mining and embodied VQA.
Kangan Qian, ChuChu Xie, Yang Zhong +13
Tsinghua University · Automotive and Robotics, Xiaomi Corporation · National University of Singapore +2
Visual imitation learning enables robots to acquire visuomotor skills directly from images, yet RGB observations lack explicit geometric cues, making learned policies brittle to camera perturbations. To address this, we propose \textbf{Ray-conditioned Vision Transformer Encoder (RayViT)}, a lightweight architecture that injects camera geometry into pretrained ViT backbones. RayViT represents camera geometry as a Plücker ray map, patchifies it into ray features, and uses gated cross-attention to produce a ray-conditioned class token. These ray features are added as dense positional embeddings, while the ray class token replaces the original ViT class token to provide a geometry-aware summary representation. We combine this approach with an auxiliary cosine similarity loss to consistently improve the performance and robustness for geometry-aware tokens. Experiments on sim- and real-robot tasks demonstrate that RayViT improves robustness by approximately 13 percentage points under camera perturbations in multi-task RoboCasa benchmark and by 1.78 average completed stages in real-world multi-task success rate compared to baselines.
Qian Wang, Longrui Chen, Peiran Sun +8
Karlsruhe Institute of Technology, Germany · University of Leeds, United Kingdom
Estimating 3D attributes directly from images has advanced rapidly with the Visual Geometry Grounded Transformer (VGGT), which predicts camera parameters, depth maps, and point clouds in a single forward pass. However, its 1.2B-parameter scale severely limits deployment on resource-constrained platforms such as UAVs and mobile AR devices. To address this limitation, we introduce QVGGT, a tailored quantization framework designed to compress VGGT. Our approach starts from the observation that transformer blocks within VGGT exhibit heterogeneous sensitivity to quantization. We thus analyze per-block quantization sensitivity and propose a selective mixed-precision strategy that allocates higher precision to the most fragile transformer blocks. To address the amplification of quantization error caused by high-variance camera and register tokens, we further introduce token filtering with camera information compensation, which removes these outliers from activation calibration and restores their geometric cues using a PCA-derived global compensation token. Finally, we develop a task-aware scale search mechanism that evaluates candidate quantization scales not only through layer reconstruction but also through multi-head supervision and cross-head geometric consistency among camera poses, depth maps, and point maps. Extensive experiments on multiple geometry perception benchmarks demonstrate that QVGGT achieves near-lossless W4A16 quantization, preserving the accuracy of all 3D prediction heads while delivering 3∼4.9× memory reduction and up to 2.8× real hardware speedup over FP32. Our approach makes high-fidelity 3D perception feasible on edge devices, enabling practical deployment of feed-forward 3D reconstruction models in real-world constrained environments.
Zhizhen Pan, Hesong Wang, Huan Wang
Westlake University · Beijing University of Posts and Telecommunications