cs.CVSep 29, 2026

Does the VGGT Family Need All Its Layers?

Authors: Fengyi Zhang, Holger Caesar, Xiangyu Sun, Zheng Zhang, Zi Huang, Yadan Luo

Organizations: The University of Queensland · Delft University of Technology · Harbin Institute of Technology

Abstract

Which layers of a feed-forward geometry model are needed to preserve both camera poses and dense 3D structure? We study layer redundancy in VGGT, π3π^3, and VGGT-ΩΩ: 3,018 pruned configurations, scored on seven camera-pose and dense-geometry metrics across four indoor and outdoor datasets. Four findings follow: (i) Removable layers cluster in two redundancy regions: a dominant early region and a narrower late one, while deletions spanning the intervening layers are consistently more disruptive. This recurring pattern holds across models, datasets, and metrics, and contrasts with the middle-to-late redundancy commonly reported in the literature. (ii) Within these regions, we observe that the joint degradation from deleting two intervals is approximately the sum of their individual degradations, reducing the number of model evaluations for pruning search from O(L4)O(L^4) to O(L2)O(L^2), where LL is the aggregator depth. (iii) We find that CKA provides a cheaper representation-based proxy for interval degradation, offering a practical trade-off between pruning quality and calibration cost. (iv) Closed-form linear calibration recovers accuracy after pruning without end-to-end retraining. A least-squares analysis shows that using a shared map for special and patch tokens generally incurs excess reconstruction loss, motivating token-aware recovery. Recovery maps fitted on just 100 calibration scenes generalize to held-out scenes and unseen datasets. The resulting models reduce aggregator parameters by up to 44% while maintaining accuracy comparable to their intact counterparts. Code and experimental results will be available at our project page: https://xian-bei.github.io/vggt-family-layer-redundancy/

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RegimeVGGT: Layer-Wise Spatially Preserving Redundancy Removal for Visual Geometry Grounded Transformer

    Jun 16, 2026Jinhao You, Shuo Lyu, Zhuohang Lyu +5Visual Geometry Grounded TransformerSelf-Supervised Vision Transformers

  2. PrePARE: Pre-AA Token Pruning for Frozen Multi-View Geometry Transformers

    May 8, 2026Haotang Li, Zhenyu Qi, Shaohan Henry Wang +5Visual Geometry Grounded TransformerVisual Token Pruning

  3. VGGT-ΩΩ

    May 14, 2026Jianyuan Wang, Minghao Chen, Shangzhan Zhang +7Visual Geometry Grounded TransformerScene Reconstruction