cs.CVOct 1, 2026

VASC: Value-Aware Sparse Attention with Cross-Layer Memory for Efficient 3D Reconstruction

Authors: Junyi Wu, Fanqing Kong, Leyang Chen, Shaoqiu Zhang, Yulun Zhang

Organizations: Shanghai Jiao Tong University

Abstract

Feed-forward 3D vision models such as VGGT have achieved remarkable progress, unifying camera estimation and dense scene reconstruction in a single pass. However, their quadratic global attention makes long image sequences expensive, while existing sparse methods may favor highly attended yet value-redundant regions. To address these limitations, we introduce VASC, a training-free sparse attention method combining value-aware block selection and execution-aware cross-layer memory. Our value-aware block selection integrates pooled query--key relevance with neighboring value contrast, reducing redundancy while preserving query-relevant and distinctive content. Cross-layer memory tracks unserved demand across layers and updates this state according to actual execution, enabling previously underserved blocks to compete under a fixed computation budget. Experiments on 7Scenes and NeuralRGB-D with VGGT and π3π^3 demonstrate improved pose estimation and reconstruction quality compared with FasterVGGT, together with up to 2.29×2.29\times faster inference than dense VGGT. Code is available at https://github.com/kosakayamahoo-design/VASC.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ReSS: Residual-Restoring Sparse Attention for 3D Vision Transformers

    Sep 28, 2026Yongsung Kim, Jaehoon Lee, Minjun Park +3Visual Geometry Grounded TransformerDynamic Sparse Attention

  2. Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval

    May 10, 2026Zichen Zou, Xiaosong Jia, Zuxuan Wu +1Visual Geometry Grounded Transformer3D Reconstruction

  3. RegimeVGGT: Layer-Wise Spatially Preserving Redundancy Removal for Visual Geometry Grounded Transformer

    Jun 16, 2026Jinhao You, Shuo Lyu, Zhuohang Lyu +5Visual Geometry Grounded TransformerSelf-Supervised Vision Transformers