cs.CVSep 29, 2026

Does the VGGT Family Need All Its Layers?

Authors: Fengyi Zhang, Holger Caesar, Xiangyu Sun, Zheng Zhang, Zi Huang, Yadan Luo

Organizations: The University of Queensland · Delft University of Technology · Harbin Institute of Technology

Abstract

Which layers of a feed-forward geometry model are needed to preserve both camera poses and dense 3D structure? We study layer redundancy in VGGT, π3π^3, and VGGT-ΩΩ: 3,018 pruned configurations, scored on seven camera-pose and dense-geometry metrics across four indoor and outdoor datasets. Four findings follow: (i) Removable layers cluster in two redundancy regions: a dominant early region and a narrower late one, while deletions spanning the intervening layers are consistently more disruptive. This recurring pattern holds across models, datasets, and metrics, and contrasts with the middle-to-late redundancy commonly reported in the literature. (ii) Within these regions, we observe that the joint degradation from deleting two intervals is approximately the sum of their individual degradations, reducing the number of model evaluations for pruning search from O(L4)O(L^4) to O(L2)O(L^2), where LL is the aggregator depth. (iii) We find that CKA provides a cheaper representation-based proxy for interval degradation, offering a practical trade-off between pruning quality and calibration cost. (iv) Closed-form linear calibration recovers accuracy after pruning without end-to-end retraining. A least-squares analysis shows that using a shared map for special and patch tokens generally incurs excess reconstruction loss, motivating token-aware recovery. Recovery maps fitted on just 100 calibration scenes generalize to held-out scenes and unseen datasets. The resulting models reduce aggregator parameters by up to 44% while maintaining accuracy comparable to their intact counterparts. Code and experimental results will be available at our project page: https://xian-bei.github.io/vggt-family-layer-redundancy/

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jun 16, 2026cs.CV

RegimeVGGT: Layer-Wise Spatially Preserving Redundancy Removal for Visual Geometry Grounded Transformer

Visual Geometry Grounded Transformer (VGGT) recovers dense 3D scene structure from multi-view images in one forward pass, but quadratic cross-frame attention limits its scalability. Existing training-free accelerators reduce computation uniformly along one axis, missing layer heterogeneity. Our spectral, probing, and causal analyses reveal three regimes: shallow layers lack cross-view structure, middle layers drive cross-view alignment, and deep layers are redundant for dense geometry yet their cross-frame attention remains essential for pose. RegimeVGGT applies layer-wise U-shaped compression along two axes: Saliency-Guided Banded Merging protects geometry- and edge-salient tokens, while Selectively Protected K/V Downsampling preserves cross-frame spatial coverage and the pose-critical path through a phase-shifted spatial grid, a reference-frame anchor, and uncompressed camera/register tokens. Training-free, RegimeVGGT achieves a 6.7x speedup over VGGT* at matched reconstruction quality.
May 8, 2026cs.CV

PrePARE: Pre-AA Token Pruning for Frozen Multi-View Geometry Transformers

Multi-view geometry transformers are feed-forward 3D foundation models that jointly predict depth maps, point maps, and camera poses for N images in a single forward pass. Most of them build on an alternating-attention (AA) stack whose cross-view attention runs over the patch tokens of all N views in every block, so its cost grows quadratically in N: on 300-frame ScanNetv2 clips, VGGT peaks at 72.3 GiB, and MapAnything does not fit on a 93 GiB card. Existing methods reduce tokens or attention inside the AA stack, so the full patch grid still enters it, and the pre-AA interface, where the encoder hands its tokens to the first AA block, has been overlooked. We propose PrePARE, which prunes patch tokens once at the pre-AA interface and restores the dense grid once after the last AA block. Every block of the frozen AA stack then runs on the reduced tokens, while the prediction heads receive the full grid. A Token Scorer, supervised by how the unpruned AA stack uses each token, and a Feature-guided Restoration module are the only trained parts. The Token Routing module is rule-based. On ScanNetv2 at N = 300, PrePARE reduces the peak memory of VGGT by 80%, from 72.3 to 14.4 GiB, below every in-AA method we ran, and runs 5.4 times faster, while its per-scene Chamfer distance is lower than the unpruned model's on 34 of 50 scenes (sign test p = 0.015). On pi^3 and Depth Anything 3 it saves 15% and 20% of peak memory, where the in-AA methods save at most 9%, and at 1000 frames it is the only method we ran that fits Depth Anything 3 and MapAnything on one GPU.
May 14, 2026cs.CV

VGGT-ΩΩ

Recent feed-forward reconstruction models, such as VGGT, have proven competitive with traditional optimization-based reconstructors while also providing geometry-aware features useful for other tasks. Here, we show that the quality of these models scales predictably with model and data size. We do so by introducing VGGT-ΩΩ, which substantially improves reconstruction accuracy, efficiency, and capabilities for both static and dynamic scenes. To enable training this model at an unprecedented scale, we introduce architectural changes that improve training efficiency, a high-quality data annotation pipeline that supports dynamic scenes, and a self-supervised learning protocol. We simplify VGGT's architecture by using a single dense prediction head with multi-task supervision and removing the expensive high-resolution convolutional layers. We also use registers to aggregate scene information into a compact representation and introduce register attention, which restricts inter-frame information exchange to these registers, in part replacing global attention. In this way, during training, VGGT-ΩΩ uses only about 30% of the GPU memory of its predecessor, allowing us to train with 15x more supervised data than prior work and to leverage vast amounts of unlabeled video data. VGGT-ΩΩ achieves strong results for reconstruction of static and dynamic scenes across multiple benchmarks, for example, improving over the previous best camera estimation accuracy on Sintel by 77%. We also show that the learned registers can improve vision-language-action models and support alignment with language, suggesting that reconstruction can be a powerful and scalable proxy task for spatial understanding. Project Page: http://vggt-omega.github.io/