cs.CVOct 7, 2026

Hardware-aware Calibrated Clustered Attention for Efficient Visual Geometric Transformers

Authors: Weitian Wang, Shubham Rai, Cecilia De La Parra, Akash Kumar

Organizations: Robert Bosch GmbH, Germany · Ruhr University Bochum, Germany

Abstract

The Visual Geometry Grounded Transformer (VGGT) marks a significant leap forward in 3D scene reconstruction, as it is the first model that directly infers all key 3D attributes (camera poses, depths, and dense geometry) jointly in one pass. However, this joint inference mechanism requires global attention layers with extremely long sequences that causes a significant latency bottleneck. In this paper, we propose blockwise clustered attention (BC attention) to accelerate the global attention layers in VGGT. By limiting the clustering within HW-friendly neighborhood blocks, BC attention reduces the computation overhead of query clustering as well as the costly data movement between on- and off-chip memory. This enables BC attention to scale to long sequences and deliver practical latency improvements on GPUs. Moreover, we introduce a hashing hyperplane calibration method and a threshold-based error compensation method to reduce clustering errors efficiently, which is a bottleneck in the current clustered attention mechanism. Overall, our experiments on GPU demonstrate that calibrated BC attention accelerates the global attention layers by 2.10-2.63×\times and the whole backbone by 1.77-2.35×\times with negligible loss (1%) for large scenes. With a small performance loss (< 5%), calibrated BC attention further achieves a 2.26-2.87×\times latency improvement on the global attention layers and a 1.90-2.55×\times improvement on the backbone.

Figures & tables

Explore similar work

Sep 20, 2026cs.CV

VGGT-Prime: Compute-Adaptive Mixture-of-Heads for Efficient Visual Geometry Transformers

Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-view images. Despite their promising performance, these models scale quadratically with the number of input views due to their global attention mechanism, resulting in substantial latency for long sequence inputs. There have been some recent efforts to accelerate VGGT, but they primarily focus on reducing \emph{token redundancy} through token merging or key/value sparsification. Our work resolves this bottleneck from a different perspective by investigating \emph{architectural redundancy} in visual geometry transformers. We show that the multi-head attention modules in VGGT's global-attention layers contain substantial architectural redundancy, with only a subset of heads carrying critical geometric information. In light of this observation, we propose VGGT-Prime, a compute-adaptive mixture-of-heads model that resolves this redundancy to accelerate visual geometry transformers while maintaining competitive reconstruction quality. The key idea of VGGT-Prime is to estimate the appropriate computation level for each global-attention head using a lightweight router and then dynamically assign each head to different computation modes. Extensive experiments on multiple datasets demonstrate that VGGT-Prime can achieve an {8×8\times} inference speedup over VGGT while maintaining competitive performance on camera pose, depth, and point-cloud predictions. We further show that VGGT-Prime is complementary to existing acceleration methods, such as token merging, further improving inference speed by up to 14×14{\times} over VGGT. An overview of our work is available on our project page.
May 8, 2026cs.CV

PrePARE: Pre-AA Token Pruning for Frozen Multi-View Geometry Transformers

Multi-view geometry transformers are feed-forward 3D foundation models that jointly predict depth maps, point maps, and camera poses for N images in a single forward pass. Most of them build on an alternating-attention (AA) stack whose cross-view attention runs over the patch tokens of all N views in every block, so its cost grows quadratically in N: on 300-frame ScanNetv2 clips, VGGT peaks at 72.3 GiB, and MapAnything does not fit on a 93 GiB card. Existing methods reduce tokens or attention inside the AA stack, so the full patch grid still enters it, and the pre-AA interface, where the encoder hands its tokens to the first AA block, has been overlooked. We propose PrePARE, which prunes patch tokens once at the pre-AA interface and restores the dense grid once after the last AA block. Every block of the frozen AA stack then runs on the reduced tokens, while the prediction heads receive the full grid. A Token Scorer, supervised by how the unpruned AA stack uses each token, and a Feature-guided Restoration module are the only trained parts. The Token Routing module is rule-based. On ScanNetv2 at N = 300, PrePARE reduces the peak memory of VGGT by 80%, from 72.3 to 14.4 GiB, below every in-AA method we ran, and runs 5.4 times faster, while its per-scene Chamfer distance is lower than the unpruned model's on 34 of 50 scenes (sign test p = 0.015). On pi^3 and Depth Anything 3 it saves 15% and 20% of peak memory, where the in-AA methods save at most 9%, and at 1000 frames it is the only method we ran that fits Depth Anything 3 and MapAnything on one GPU.
Jun 16, 2026cs.CV

RegimeVGGT: Layer-Wise Spatially Preserving Redundancy Removal for Visual Geometry Grounded Transformer

Visual Geometry Grounded Transformer (VGGT) recovers dense 3D scene structure from multi-view images in one forward pass, but quadratic cross-frame attention limits its scalability. Existing training-free accelerators reduce computation uniformly along one axis, missing layer heterogeneity. Our spectral, probing, and causal analyses reveal three regimes: shallow layers lack cross-view structure, middle layers drive cross-view alignment, and deep layers are redundant for dense geometry yet their cross-frame attention remains essential for pose. RegimeVGGT applies layer-wise U-shaped compression along two axes: Saliency-Guided Banded Merging protects geometry- and edge-salient tokens, while Selectively Protected K/V Downsampling preserves cross-frame spatial coverage and the pose-critical path through a phase-shifted spatial grid, a reference-frame anchor, and uncompressed camera/register tokens. Training-free, RegimeVGGT achieves a 6.7x speedup over VGGT* at matched reconstruction quality.