cs.ROSep 19, 2025

eVGGT: An Efficient Geometry-Aware Vision Encoder for Visuomotor Policies

Authors: An Dinh Vuong, Minh Nhat Vu, Ian Reid

Organizations: Department of Computer Vision, Mohammed bin Zayed University of Artificial Intelligence, UAE · AIT Austrian Institute of Technology GmbH, Austria

Abstract

Geometry-grounded vision models, such as VGGT, have emerged as robust visual encoders, providing essential geometric priors for robotic manipulation. However, the high computational cost of these models often leads to slow inference, limiting their practical applications in real-world robotics. This paper introduces eVGGT, a lightweight geometry-aware vision encoder distilled from the high-performing VGGT. Our findings demonstrate two primary advantages: i) integrating eVGGT into imitation learning frameworks (including ACT and Diffusion Policy) yields up to a 6.3% improvement in success rate over standard 2D encoders across bimanual and single-arm tasks in both simulation and real-world settings with variable viewpoints; ii) eVGGT achieves a nearly 5 times speedup and a 63% reduction in memory usage compared to state-of-the-art geometry-aware encoders while maintaining comparable task performance. These results suggest that eVGGT substantially alleviates the performance-latency bottleneck that has limited geometry-aware visuomotor policies in real-world deployment.

Figures & tables

Explore similar work

CardsList
  1. RayViT: Ray-Conditioned Visual Representations for Viewpoint-Robust Imitation Learning

    Jul 31, 2026Qian Wang, Longrui Chen, Peiran Sun +8Visual Geometry Grounded TransformerRobotic Perception

  2. VGGT-Prime: Compute-Adaptive Mixture-of-Heads for Efficient Visual Geometry Transformers

    Sep 20, 2026Abteen Arab, Guile Wu, Chengjie Huang +1Visual Geometry Grounded TransformerSelf-Supervised Vision Transformers

  3. Reducing Temporal Redundancy for Efficient Vision-Language-Action Inference

    Jul 14, 2026Yuzhou Wu, Yuxin Zheng, Muchun Niu +6Diffusion-Based Vision-Language-ActionsVision-Language-Action Framework