eVGGT: An Efficient Geometry-Aware Vision Encoder for Visuomotor Policies
Authors: An Dinh Vuong, Minh Nhat Vu, Ian Reid
Organizations: Department of Computer Vision, Mohammed bin Zayed University of Artificial Intelligence, UAE · AIT Austrian Institute of Technology GmbH, Austria
Geometry-grounded vision models, such as VGGT, have emerged as robust visual encoders, providing essential geometric priors for robotic manipulation. However, the high computational cost of these models often leads to slow inference, limiting their practical applications in real-world robotics. This paper introduces eVGGT, a lightweight geometry-aware vision encoder distilled from the high-performing VGGT. Our findings demonstrate two primary advantages: i) integrating eVGGT into imitation learning frameworks (including ACT and Diffusion Policy) yields up to a 6.3% improvement in success rate over standard 2D encoders across bimanual and single-arm tasks in both simulation and real-world settings with variable viewpoints; ii) eVGGT achieves a nearly 5 times speedup and a 63% reduction in memory usage compared to state-of-the-art geometry-aware encoders while maintaining comparable task performance. These results suggest that eVGGT substantially alleviates the performance-latency bottleneck that has limited geometry-aware visuomotor policies in real-world deployment.
Figures & tables
Fig. 1: Performance-efficiency trade-off of vision encoders. We compare DP [ 3 ] , the standard visuomotor policy, paired with different vision backbones. Success rate ( y -axis) vs. inference throughput in Hz ( x -axis); bubble area denotes VRAM (GB). The dashed line marks the 20 Hz real-time threshold. DP+VGGT is robust but slow and memory-heavy (8.14 GB, 4.5 Hz); DP+eVGGT achieves the best trade-off with the highest success rate, enabling real-time deployment at lower memory (3.01 GB, 20.9 Hz).
Fig. 2: VGGT architecture. VGGT is a transformer-based vision encoder that produces geometry-aware features from multi-view input.
Method
Modality
Beat Block Hammer
Adjust Bottle
Lift Pot
Place Burger Fries
Press Stapler
Handover Mic
Average
Seen
Unseen
Avg
Seen
Unseen
Avg
Seen
Unseen
Avg
Seen
Unseen
Avg
Seen
Unseen
Avg
Seen
Unseen
Avg
Seen
Unseen
Avg
DP3 [ 51 ]
Point cloud
0.082
0.018
0.050
0.622
0.296
0.459
0.102
0.038
0.070
0.184
0.036
0.110
0.152
0.238
0.195
0.724
0.316
0.520
0.311
0.157
0.234
RDT [ 21 ]
RGB
0.090
0.092
0.091
0.496
0.354
0.425
0.118
0.036
0.077
0.196
0.134
0.165
0.198
0.122
0.160
0.606
0.204
0.405
0.284
0.157
0.221
ACT [ 52 ]
RGB
0.074
0.000
0.037
0.162
0.118
0.140
0.034
0.006
0.020
0.042
0.000
0.021
0.024
0.016
0.020
0.022
0.000
0.011
0.060
0.023
0.042
ACT+eVGGT (ours)
RGB
0.066
0.012
0.039
0.198
0.104
0.151
0.016
0.022
0.019
0.098
0.000
0.049
0.136
0.032
0.084
0.038
0.014
0.026
0.092
0.031
0.061
DP [ 3 ]
RGB
0.142
0.014
0.078
0.462
0.238
0.350
0.062
0.020
0.041
0.164
0.038
0.101
0.182
0.094
0.138
0.502
0.118
0.310
0.252
0.087
0.170
TABLE I: Success rates on RoboTwin simulator
Method
Modality
Pick Cube
Push Cube
Stack Cube
Average
DP [ 3 ]
RGBD
0.952
0.922
0.524
0.799
RDT [ 21 ]
RGB
0.774
1.000
0.702
0.825
ACT [ 52 ]
RGB
0.712
0.904
0.474
0.697
ACT+eVGGT (ours)
RGB
0.756
0.888
0.496
0.713
DP [ 3 ]
RGB
0.848
0.876
0.618
0.781
DP+eVGGT (ours)
RGB
0.964
0.912
0.654
0.843
TABLE II: Success rates on ManiSkill Simulator
Modality & Efficiency
Success Rate
Method
Modality
Inference Speed (Hz)
Memory (GB)
Beat Block Hammer
Adjust Bottle
Average
3D
DP3 [ 51 ]
Point cloud (simulator)
-
2.54
0.082
0.622
0.352
DP3 [ 51 ]
Point cloud (VGGT [ 42 ] )
3.5
10.47
0.042
0.252
0.147
2D
DP [ 3 ] +ResNet
RGB
25.4
2.00
0.142
0.462
0.302
DP [ 3 ] +ViT
RGB
24.8
2.11
0.124
0.374
0.249
geo-aware
DP+Metric3Dv2 [ 14 ]
RGB
17.1
4.39
0.138
0.392
0.265
TABLE III: Comparison with Different Vision Methods
Fig. 5: Success rates when scaling views. We compare DP+eVGGT against the DP+ResNet baseline across 1 to 4 RGB views on RoboTwin.
Fig. 6: Ablation study: Reconstruction accuracy vs. Success rates. We compare AbsRel across three model configurations and their corresponding success rates when paired with DP.
Fig. 7: Qualitative 3D reconstruction results. Qualitative comparison between eVGGT and its teacher VGGT on 3D reconstruction. For point clouds, we omit the ground truth since it is provided as semantically labeled points rather than a comparable geometric reference. For depth, results are shown on RoboTwin (top) and ManiSkill (bottom).
Fig. 8: Robot Setup. We conduct robot experiments with three camera views to push the cube to the target location.