cs.CVOct 7, 2026

Hardware-aware Calibrated Clustered Attention for Efficient Visual Geometric Transformers

Authors: Weitian Wang, Shubham Rai, Cecilia De La Parra, Akash Kumar

Organizations: Robert Bosch GmbH, Germany · Ruhr University Bochum, Germany

Abstract

The Visual Geometry Grounded Transformer (VGGT) marks a significant leap forward in 3D scene reconstruction, as it is the first model that directly infers all key 3D attributes (camera poses, depths, and dense geometry) jointly in one pass. However, this joint inference mechanism requires global attention layers with extremely long sequences that causes a significant latency bottleneck. In this paper, we propose blockwise clustered attention (BC attention) to accelerate the global attention layers in VGGT. By limiting the clustering within HW-friendly neighborhood blocks, BC attention reduces the computation overhead of query clustering as well as the costly data movement between on- and off-chip memory. This enables BC attention to scale to long sequences and deliver practical latency improvements on GPUs. Moreover, we introduce a hashing hyperplane calibration method and a threshold-based error compensation method to reduce clustering errors efficiently, which is a bottleneck in the current clustered attention mechanism. Overall, our experiments on GPU demonstrate that calibrated BC attention accelerates the global attention layers by 2.10-2.63×\times and the whole backbone by 1.77-2.35×\times with negligible loss (1%) for large scenes. With a small performance loss (< 5%), calibrated BC attention further achieves a 2.26-2.87×\times latency improvement on the global attention layers and a 1.90-2.55×\times improvement on the backbone.

Figures & tables

Explore similar work

CardsList
  1. VGGT-Prime: Compute-Adaptive Mixture-of-Heads for Efficient Visual Geometry Transformers

    Sep 20, 2026Abteen Arab, Guile Wu, Chengjie Huang +13D ViTsMulti-View 3D Reconstruction

  2. PrePARE: Pre-AA Token Pruning for Frozen Multi-View Geometry Transformers

    May 8, 2026Haotang Li, Zhenyu Qi, Shaohan Henry Wang +5Multi-View GeometryEfficient ViTs

  3. RegimeVGGT: Layer-Wise Spatially Preserving Redundancy Removal for Visual Geometry Grounded Transformer

    Jun 16, 2026Jinhao You, Shuo Lyu, Zhuohang Lyu +53D ViTsMulti-View 3D Reconstruction