The Visual Geometry Grounded Transformer (VGGT) marks a significant leap forward in 3D scene reconstruction, as it is the first model that directly infers all key 3D attributes (camera poses, depths, and dense geometry) jointly in one pass. However, this joint inference mechanism requires global attention layers with extremely long sequences that causes a significant latency bottleneck. In this paper, we propose blockwise clustered attention (BC attention) to accelerate the global attention layers in VGGT. By limiting the clustering within HW-friendly neighborhood blocks, BC attention reduces the computation overhead of query clustering as well as the costly data movement between on- and off-chip memory. This enables BC attention to scale to long sequences and deliver practical latency improvements on GPUs. Moreover, we introduce a hashing hyperplane calibration method and a threshold-based error compensation method to reduce clustering errors efficiently, which is a bottleneck in the current clustered attention mechanism. Overall, our experiments on GPU demonstrate that calibrated BC attention accelerates the global attention layers by 2.10-2.63× and the whole backbone by 1.77-2.35× with negligible loss (1%) for large scenes. With a small performance loss (< 5%), calibrated BC attention further achieves a 2.26-2.87× latency improvement on the global attention layers and a 1.90-2.55× improvement on the backbone.
Figures & tables
Fig. 1 : Architecture overview of VGGT [ 1 ] . The global attention layers that perform attention computations across tokens from all views cause a significant latency bottleneck.
Fig. 2 : Attention score distribution comparison between VGGT and Llama 3.1 8B [ 16 ] . VGGT has nearly uniform attention distribution in early and late layers. For middle layers, VGGT still has a more flat attention distribution than Llama.
Fig. 3 : Cosine similarity matrix between queries from the first three frames after adding the positional embedding (RoPE).
Fig. 4 : Pipeline of hyperplanes calibration through differentiable surrogate
Fig. 5 : Top 10% outlier thresholds of 4 random scenes across layers at different depths
Table 6
Dataset
ETH3D
Metrics
Acc. ↓
Comp. ↓
Overall ↓
Standard Attention
Baseline
0.114
0.125
0.120
Calibrated BC Attention
γ=4
0.115
0.127
0.121
TABLE III : Task Performance on MapAnything [ 24 ]
Fig. 6 : Comparison of reconstruction results using standard attention, blockwise clustered attention with calibrated and random hyperplanes. The reconstruction result using random hyperplanes exhibits noticeable noise compared to the other two results.
Block Size
64
128
256
ETH3D Acc. ↓
0.906
0.891
0.129
TABLE IV : Task Performance using Different Block Size