eVGGT: An Efficient Geometry-Aware Vision Encoder for Visuomotor Policies
Authors: An Dinh Vuong, Minh Nhat Vu, Ian Reid
Organizations: Department of Computer Vision, Mohammed bin Zayed University of Artificial Intelligence, UAE · AIT Austrian Institute of Technology GmbH, Austria
Geometry-grounded vision models, such as VGGT, have emerged as robust visual encoders, providing essential geometric priors for robotic manipulation. However, the high computational cost of these models often leads to slow inference, limiting their practical applications in real-world robotics. This paper introduces eVGGT, a lightweight geometry-aware vision encoder distilled from the high-performing VGGT. Our findings demonstrate two primary advantages: i) integrating eVGGT into imitation learning frameworks (including ACT and Diffusion Policy) yields up to a 6.3% improvement in success rate over standard 2D encoders across bimanual and single-arm tasks in both simulation and real-world settings with variable viewpoints; ii) eVGGT achieves a nearly 5 times speedup and a 63% reduction in memory usage compared to state-of-the-art geometry-aware encoders while maintaining comparable task performance. These results suggest that eVGGT substantially alleviates the performance-latency bottleneck that has limited geometry-aware visuomotor policies in real-world deployment.
Figures & tables
Fig. 1: Performance-efficiency trade-off of vision encoders. We compare DP [ 3 ] , the standard visuomotor policy, paired with different vision backbones. Success rate ( y -axis) vs. inference throughput in Hz ( x -axis); bubble area denotes VRAM (GB). The dashed line marks the 20 Hz real-time threshold. DP+VGGT is robust but slow and memory-heavy (8.14 GB, 4.5 Hz); DP+eVGGT achieves the best trade-off with the highest success rate, enabling real-time deployment at lower memory (3.01 GB, 20.9 Hz).
Fig. 2: VGGT architecture. VGGT is a transformer-based vision encoder that produces geometry-aware features from multi-view input.
Method
Modality
Beat Block Hammer
Adjust Bottle
Lift Pot
Place Burger Fries
Press Stapler
Handover Mic
Average
Seen
Unseen
Avg
Seen
Unseen
Avg
Seen
Unseen
Avg
Seen
Unseen
Avg
Seen
Unseen
Avg
Seen
Unseen
Avg
Seen
Unseen
Avg
DP3 [ 51 ]
Point cloud
0.082
0.018
0.050
0.622
0.296
0.459
0.102
0.038
0.070
0.184
0.036
0.110
0.152
0.238
0.195
0.724
0.316
0.520
0.311
0.157
0.234
RDT [ 21 ]
RGB
0.090
0.092
0.091
0.496
0.354
0.425
0.118
0.036
0.077
0.196
0.134
0.165
0.198
0.122
0.160
0.606
0.204
0.405
0.284
0.157
0.221
ACT [ 52 ]
RGB
0.074
0.000
0.037
0.162
0.118
0.140
0.034
0.006
0.020
0.042
0.000
0.021
0.024
0.016
0.020
0.022
0.000
0.011
0.060
0.023
0.042
ACT+eVGGT (ours)
RGB
0.066
0.012
0.039
0.198
0.104
0.151
0.016
0.022
0.019
0.098
0.000
0.049
0.136
0.032
0.084
0.038
0.014
0.026
0.092
0.031
0.061
DP [ 3 ]
RGB
0.142
0.014
0.078
0.462
0.238
0.350
0.062
0.020
0.041
0.164
0.038
0.101
0.182
0.094
0.138
0.502
0.118
0.310
0.252
0.087
0.170
TABLE I: Success rates on RoboTwin simulator
Method
Modality
Pick Cube
Push Cube
Stack Cube
Average
DP [ 3 ]
RGBD
0.952
0.922
0.524
0.799
RDT [ 21 ]
RGB
0.774
1.000
0.702
0.825
ACT [ 52 ]
RGB
0.712
0.904
0.474
0.697
ACT+eVGGT (ours)
RGB
0.756
0.888
0.496
0.713
DP [ 3 ]
RGB
0.848
0.876
0.618
0.781
DP+eVGGT (ours)
RGB
0.964
0.912
0.654
0.843
TABLE II: Success rates on ManiSkill Simulator
Modality & Efficiency
Success Rate
Method
Modality
Inference Speed (Hz)
Memory (GB)
Beat Block Hammer
Adjust Bottle
Average
3D
DP3 [ 51 ]
Point cloud (simulator)
-
2.54
0.082
0.622
0.352
DP3 [ 51 ]
Point cloud (VGGT [ 42 ] )
3.5
10.47
0.042
0.252
0.147
2D
DP [ 3 ] +ResNet
RGB
25.4
2.00
0.142
0.462
0.302
DP [ 3 ] +ViT
RGB
24.8
2.11
0.124
0.374
0.249
geo-aware
DP+Metric3Dv2 [ 14 ]
RGB
17.1
4.39
0.138
0.392
0.265
TABLE III: Comparison with Different Vision Methods
Fig. 5: Success rates when scaling views. We compare DP+eVGGT against the DP+ResNet baseline across 1 to 4 RGB views on RoboTwin.
Fig. 6: Ablation study: Reconstruction accuracy vs. Success rates. We compare AbsRel across three model configurations and their corresponding success rates when paired with DP.
Fig. 7: Qualitative 3D reconstruction results. Qualitative comparison between eVGGT and its teacher VGGT on 3D reconstruction. For point clouds, we omit the ground truth since it is provided as semantically labeled points rather than a comparable geometric reference. For depth, results are shown on RoboTwin (top) and ManiSkill (bottom).
Fig. 8: Robot Setup. We conduct robot experiments with three camera views to push the cube to the target location.
Visual imitation learning enables robots to acquire visuomotor skills directly from images, yet RGB observations lack explicit geometric cues, making learned policies brittle to camera perturbations. To address this, we propose \textbf{Ray-conditioned Vision Transformer Encoder (RayViT)}, a lightweight architecture that injects camera geometry into pretrained ViT backbones. RayViT represents camera geometry as a Plücker ray map, patchifies it into ray features, and uses gated cross-attention to produce a ray-conditioned class token. These ray features are added as dense positional embeddings, while the ray class token replaces the original ViT class token to provide a geometry-aware summary representation. We combine this approach with an auxiliary cosine similarity loss to consistently improve the performance and robustness for geometry-aware tokens. Experiments on sim- and real-robot tasks demonstrate that RayViT improves robustness by approximately 13 percentage points under camera perturbations in multi-task RoboCasa benchmark and by 1.78 average completed stages in real-world multi-task success rate compared to baselines.
Qian Wang, Longrui Chen, Peiran Sun +8
Karlsruhe Institute of Technology, Germany · University of Leeds, United Kingdom
Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-view images. Despite their promising performance, these models scale quadratically with the number of input views due to their global attention mechanism, resulting in substantial latency for long sequence inputs. There have been some recent efforts to accelerate VGGT, but they primarily focus on reducing \emph{token redundancy} through token merging or key/value sparsification. Our work resolves this bottleneck from a different perspective by investigating \emph{architectural redundancy} in visual geometry transformers. We show that the multi-head attention modules in VGGT's global-attention layers contain substantial architectural redundancy, with only a subset of heads carrying critical geometric information. In light of this observation, we propose VGGT-Prime, a compute-adaptive mixture-of-heads model that resolves this redundancy to accelerate visual geometry transformers while maintaining competitive reconstruction quality. The key idea of VGGT-Prime is to estimate the appropriate computation level for each global-attention head using a lightweight router and then dynamically assign each head to different computation modes. Extensive experiments on multiple datasets demonstrate that VGGT-Prime can achieve an {8×} inference speedup over VGGT while maintaining competitive performance on camera pose, depth, and point-cloud predictions. We further show that VGGT-Prime is complementary to existing acceleration methods, such as token merging, further improving inference speed by up to 14× over VGGT. An overview of our work is available on our \href{https://vggt-prime.github.io}{project page}.
Abteen Arab, Guile Wu, Chengjie Huang +1
Huawei Noah’s Ark Lab · University of British Columbia
Vision-Language-Action (VLA) models exhibit strong generalization for robotic manipulation, yet their high inference latency limits real time deployment. We identify two primary sources of temporal redundancy in existing VLA pipelines: repeated visual encoding of highly similar consecutive frames and multi step iterative sampling in diffusion based policies. To address this, we propose a system level acceleration strategy that reduces computation in both perception and action generation. On the perception side, we incrementally update only tokens corresponding to dynamic scene regions instead of re-encoding entire frames. On the policy side, we compress diffusion sampling into a compact 2-step schedule through efficiency oriented training while preserving action precision. Experiments on Libero, RobotWin, and Real Robot Platforms demonstrate over 2 times speedup while maintaining high performance, achieving up to 98% success rate on general manipulation benchmarks. Our codes will be released on Github.
Yuzhou Wu, Yuxin Zheng, Muchun Niu +6
1Tianji KernalMind co ltd · 2Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · 3Shanghai Jiao Tong University, Shanghai, China +2