cs.CVOct 5, 2026

VGGT-Bridge: Beyond Sequential Pose Graphs via Coarse-Stride Skip Edges

Authors: Sungjae Choi, Hanna Bae, Sunghyun Baek, Junmo Kim

Organizations: Korea Advanced Institute of Science and Technology, South Korea

Abstract

Feed-forward visual geometry transformers such as VGGT reconstruct dense 3D structure from images in a single forward pass, simplifying multi-view 3D reconstruction. However, their quadratic attention complexity makes them difficult to scale to long sequences with thousands of frames. Chunk-and-align frameworks address this by splitting a long sequence into overlapping chunks and stitching their local reconstructions into a pose graph. Yet existing methods connect only sequentially adjacent chunks, so small per-frame errors accumulate along the chain into large-scale drift. To move beyond sequential edges, we propose VGGT-Bridge, which adds long-range skip edges that directly constrain non-adjacent chunks without retraining. By running VGGT on sparsely sampled coarse chunks, each coarse chunk bridges distant fine chunks into a single direct constraint. We further turn VGGT's first-frame scale bias into a drift correction by feeding selected coarse chunks in reverse, and a loop-aware policy keeps this reversal compatible with existing loop closures. VGGT-Bridge reduces ATE by 28.3% on KITTI Odometry, 18.8% on Virtual KITTI, and 10.0% on Waymo Open over the SwiftVGGT baseline, achieving the best performance among all chunk-and-align methods.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Mamba-VGGT: Persistent Long-Sequence Video Geometry Grounded Transformer via External Sliding Window Mamba Memory

    May 17, 2026Tianchen Deng, Zhenxiang Xiong, Nailin Wang +4Visual Geometry Grounded Transformer3D Reconstruction

  2. Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval

    May 10, 2026Zichen Zou, Xiaosong Jia, Zuxuan Wu +1Visual Geometry Grounded Transformer3D Reconstruction

  3. RegimeVGGT: Layer-Wise Spatially Preserving Redundancy Removal for Visual Geometry Grounded Transformer

    Jun 16, 2026Jinhao You, Shuo Lyu, Zhuohang Lyu +5Visual Geometry Grounded TransformerSelf-Supervised Vision Transformers