Feed-forward visual geometry transformers such as VGGT reconstruct dense 3D structure from images in a single forward pass, simplifying multi-view 3D reconstruction. However, their quadratic attention complexity makes them difficult to scale to long sequences with thousands of frames. Chunk-and-align frameworks address this by splitting a long sequence into overlapping chunks and stitching their local reconstructions into a pose graph. Yet existing methods connect only sequentially adjacent chunks, so small per-frame errors accumulate along the chain into large-scale drift. To move beyond sequential edges, we propose VGGT-Bridge, which adds long-range skip edges that directly constrain non-adjacent chunks without retraining. By running VGGT on sparsely sampled coarse chunks, each coarse chunk bridges distant fine chunks into a single direct constraint. We further turn VGGT's first-frame scale bias into a drift correction by feeding selected coarse chunks in reverse, and a loop-aware policy keeps this reversal compatible with existing loop closures. VGGT-Bridge reduces ATE by 28.3% on KITTI Odometry, 18.8% on Virtual KITTI, and 10.0% on Waymo Open over the SwiftVGGT baseline, achieving the best performance among all chunk-and-align methods.
Figures & tables
Figure 1 : Prior methods connect only sequentially adjacent chunks ( S1→S2→⋯ ), so per-frame errors accumulate along the chain into large-scale drift. VGGT-Bridge inserts skip edges (pink arcs, e.g . , S1→S3 ) that directly constrain non-adjacent chunks under the chunk-and-align pipeline. On KITTI Odometry Seq. 08, these long-range constraints lower ATE RMSE to 47.1 m.
Figure 2 : Method overview. Gray, blue, and yellow boxes denote input frames, sparsely sampled coarse frames, and alignment overlaps, respectively. Standard chunk-and-align pipelines construct overlapping fine chunks and form sequential edges Eseq (blue) between adjacent chunks, with optional loop edges Eloop (purple) between revisited chunks. VGGT-Bridge additionally runs VGGT on sparse coarse chunks and composes two cross-scale Sim(3) alignments through each coarse chunk into a long-range skip edge Eskip (rose) between non-adjacent fine chunks. Forward/reverse anchor markers show the selected processing direction of each coarse chunk, chosen by a loop-aware policy to complement the available loop closures. All edge sets are jointly optimized in a lightweight Sim(3) pose graph to recover globally consistent chunk poses.
Method
Calib.
00
01
02
03
04
05
06
07
08
09
10
Avg. ( ↓ )
# frames
4541
1101
4661
801
271
2761
1101
1101
4071
1591
1201
Length (m)
2453
5067
3724
561
394
2206
1233
650
3223
1705
920
Speed (m/fr)
0.82
2.23
1.09
0.70
1.45
0.80
1.12
0.59
0.79
1.07
0.77
Loop
✓
×
✓
×
×
✓
✓
✓
×
✓
×
ORB-SLAM2 [ 28 ]
✓
6.03
508.34
14.76
1.02
1.57
4.04
11.16
2.19
38.85
8.39
6.63
54.82
DPVO [ 40 ]
✓
113.21
12.69
123.40
2.09
0.68
58.96
54.78
19.26
115.90
75.10
13.63
53.61
Table 1 : Performance comparison on KITTI Odometry dataset. Bold / underline mark the best/second-best among the calibration-free methods. “Calib.” = uses camera intrinsics, “TL” = tracking lost, “OOM” = out of memory. Results with * are reproduced by us.
Method (segment ID prefix)
Calib.
1634
1838
3156
3461
3711
4058
4604
5200
6104
Avg. ( ↓ )
# frames
198
199
199
199
196
199
198
199
198
Length (m)
160
42
165
351
273
86
266
135
63
Speed (m/fr)
0.81
0.21
0.83
1.77
1.39
0.43
1.34
0.68
0.32
Traffic
Low
High
Low
Low
Med.
Low
Med.
Low
High
DROID-SLAM [ 39 ]
✓
3.705
0.301
0.447
8.653
9.320
7.621
4.170
TL
0.264
4.396
MASt3R-SLAM [ 29 ]
×
4.500
0.556
1.833
12.544
8.601
1.412
5.428
7.910
1.195
5.560
Table 2 : Performance comparison on Waymo Open. Bold / underline mark the best/second-best among calibration-free methods. “Calib.” = uses camera intrinsics. Results with * are reproduced by us.
Method
Calib.
Scene01
Scene02
Scene06
Scene18
Scene20
Avg. ( ↓ )
DROID-SLAM [ 39 ]
✓
1.137
0.064
0.038
2.205
4.157
1.520
MASt3R-SLAM [ 29 ]
×
TL
–
CUT3R [ 45 ]
×
48.362
20.119
0.772
17.149
103.692
38.019
Fast3R [ 52 ]
×
OOM
–
VGGT-Long [ 8 ]
×
1.049
0.703
0.438
1.397
6.683
2.054
SwiftVGGT * [ 20 ]
×
5.178
0.332
0.186
1.817
4.960
2.495
Table 3 : Performance comparison on Virtual KITTI. Bold / underline mark the best/second-best among calibration-free methods. “Calib.” = uses camera intrinsics, “TL” = tracking lost, “OOM” = out of memory (on at least one weather variant within the scene). Results with * are reproduced by us.
Figure 3 : Camera trajectory and dense 3D reconstruction comparison on KITTI Odometry.
Figure 4 : Scale bias on KITTI Odometry 00–10 ( n=11 sequences). Left: mean ± std of the pairwise Umeyama logs measured during chunk alignment, for sequential and coarse chunks. The arrow marks the shift Δ from forward to reverse anchoring, with the percentage of edges where Δ>0 . Right: residual local scale error ∣logs∣ (bars) and the ATE RMSE (lines) of the final trajectory, for SwiftVGGT and our method. The scale error is measured on sliding 75 -frame windows against the ground truth.
Translation ATE (m)
Mean
Configuration
00
01
02
03
04
05
06
07
08
09
10
Trans
Rot
Scale
# frames
4541
1101
4661
801
271
2761
1101
1101
4071
1591
1201
Loop
✓
–
✓
–
–
✓
✓
✓
–
✓
–
SwiftVGGT
9.92
114.96
55.45
10.38
5.48
12.50
8.38
3.50
69.74
30.92
25.55
31.53
13.38
0.157
+ coarse-stride chunk
13.00
69.68
53.43
9.26
4.44
12.30
7.28
4.68
72.43
24.06
25.45
26.91
12.15
0.139
+ reverse coarse chunk
15.01
56.35
49.80
9.50
3.97
10.88
8.87
3.16
45.53
27.59
20.74
22.85
12.41
0.115
Table 4 : Component ablation on KITTI Odometry. Each row cumulatively adds one component on top of the previous. Trans and Rot are the translation (m) and rotation ( ∘ ) ATE, and Scale is the mean ∣logs∣ of Fig. 4 (right).
Table 9
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
KITTI Odometry
Virtual KITTI
Waymo Open
Fine chunk size B
75
75
60
Fine overlap O
30
30
24
Coarse ( r=3 )
B3
35
35
35
O3
14
14
14
Coarse ( r=5 )
B5
25
25
25
O5
10
10
10
Appendix
Table 7 : Hyperparameter settings.
VGGT-Bridge
Dataset
# seq.
Mean logs
SwiftVGGT
wˉ=0.3
wˉ=0.7
wˉ=1.0
KITTI Odometry
11
−0.0370±0.0225
31.53
22.61
22.82
23.06
Virtual KITTI
30
−0.0761±0.0896
2.495
2.027
1.853
1.790
Waymo Open
9
−0.0470±0.0341
1.557
1.401
1.406
1.408
EuRoC MAV
11
−0.0387±0.0221
1.722
1.698
1.632
1.638
7-Scenes
7
−0.0167±0.0169
0.087
0.083
0.082
0.082
Appendix
Table 8 : Anchoring bias and skip-edge weight across five benchmarks. Mean logs is the mean ± std of the fine sequential edges over the sequences, as in Fig. 4 (left) of the main paper. The remaining columns report the mean translation ATE RMSE. Bold marks the best result in each row.
Method
00
01
02
03
04
05
06
07
08
09
10
Mean
SwiftVGGT [ 20 ]
9.92
114.96
55.45
10.38
5.48
12.50
8.38
3.50
69.74
30.92
25.55
31.53
VGGT-Bridge
13.74
56.34
48.70
9.33
4.00
11.02
6.93
3.15
47.15
26.24
22.05
22.61
Δ (%)
+38.5
−51.0
−12.2
−10.1
−27.0
−11.8
−17.3
−10.0
−32.4
−15.1
−13.7
−28.3
VGGT-Long [ 8 ]
10.81
66.46
56.00
5.56
3.93
10.11
6.90
2.83
57.45
31.17
22.11
24.85
+ VGGT-Bridge
12.08
41.47
49.21
6.57
3.20
10.15
7.65
4.30
47.61
23.92
20.50
20.61
Δ (%)
+11.7
−37.6
−12.1
+18.2
−18.6
+0.4
+10.9
+51.9
−17.1
−23.3
−7.3
−17.1
Appendix
Table 9 : Cross-framework transfer of VGGT-Bridge on VGGT-Long. (per-sequence translation ATE RMSE (m) on KITTI Odometry). The same skip-edge construction is added to two backbones, and Δ is the change relative to each backbone.
VGGT-Long [ 8 ]
SwiftVGGT [ 20 ]
VGGT-Bridge (ours)
Seq
Acc
Comp
CD
Acc
Comp
CD
Acc
Comp
CD
00
1.62
7.31
4.47
1.95
7.24
4.60
2.11
8.78
5.45
01
42.73
16.83
29.78
34.65
20.09
27.37
20.91
7.86
14.38
02
18.72
30.28
24.50
17.82
32.57
25.19
13.72
29.19
21.45
04
0.80
6.80
3.80
1.59
6.53
4.06
1.59
6.84
4.21
05
1.24
7.17
4.21
4.38
9.29
6.84
1.73
7.55
4.64
Appendix
Table 10 : Reconstruction quality on KITTI Odometry. Acc / Comp / CD (m), lower is better. Averaged over the ten sequences with available raw LiDAR scan, omitting Seq. 03. Best per metric per sequence in bold .
VGGT-Long [ 8 ]
SwiftVGGT [ 20 ]
VGGT-Bridge (ours)
Segment
Acc
Comp
CD
Acc
Comp
CD
Acc
Comp
CD
1634
0.87
2.11
1.49
1.02
1.81
1.42
0.94
1.90
1.42
1838
0.45
5.87
3.16
1.38
5.85
3.62
1.35
5.44
3.39
3156
0.66
2.58
1.62
1.00
2.26
1.63
0.99
2.11
1.55
3461
0.93
4.47
2.70
1.38
5.53
3.45
1.36
5.50
3.43
3711
1.96
4.93
3.44
2.60
4.26
3.43
2.44
4.37
3.41
Appendix
Table 11 : Reconstruction quality on Waymo Open. Acc / Comp / CD (m), lower is better. Best per metric per segment in bold .
Scene 01
Scene 02
Method
Clone
Fog
Morning
Overcast
Rain
Sunset
Clone
Fog
Morning
Overcast
Rain
Sunset
DROID-SLAM † [ 39 ]
1.03
1.87
0.99
1.01
0.78
1.15
0.10
0.04
0.05
0.05
0.04
0.11
MASt3R-SLAM [ 29 ]
TL
CUT3R [ 45 ]
43.30
62.19
50.61
38.73
51.55
43.79
23.77
9.95
28.41
24.64
7.96
25.97
Fast3R [ 52 ]
OOM
VGGT-Long [ 8 ]
0.76
0.87
0.93
0.67
1.80
1.26
0.72
0.71
0.72
0.68
0.69
0.69
Appendix
Table 12 : Per-scene, per-condition ATE RMSE (m) on Virtual KITTI. Best among the calibration-free methods in bold . “ † ” uses camera intrinsics, all other methods are calibration-free. “TL” = tracking lost, “OOM” = out of memory. Results with * are reproduced by us.
KITTI Seq. 07 (1,101 frames)
KITTI Seq. 08 (4,071 frames)
Stage
SwiftVGGT
Ours
Δ
SwiftVGGT
Ours
Δ
Loop detection
3.1
3.0
−0.1
10.2
10.1
−0.1
VGGT chunk inference
58.1
84.4
+26.3
197.6
307.9
+110.3
Loop-centric chunk inference
1.1
1.3
+0.2
0.0
0.0
+0.0
Loop-closure alignment
0.9
1.1
+0.2
0.0
0.0
+0.0
Sequential-edge composition
29.9
30.0
+0.1
76.0
76.0
+0.0
Appendix
Table 13 : Per-stage runtime breakdown on KITTI Odometry Seq. 07 and Seq. 08. Seq. 07 has one detected loop pair and activates the loop pipeline, while Seq. 08 has none detected. The loop pipeline is identical in both methods, so only the stages VGGT-Bridge modifies differ.
GPU
Host
Method
Mean
Max
Mean
Max
SwiftVGGT [ 20 ]
10.30
14.49
15.74
29.90
VGGT-Bridge
10.30
14.49
19.81
38.90
Appendix
Table 14 : Peak memory on KITTI Odometry 00–10 (GiB), as the mean and the maximum over the eleven sequences. GPU is the peak memory reserved by PyTorch, and host is the peak resident memory.
SwiftVGGT, O
SwiftVGGT
VGGT-Bridge
30
45
60
+B′=90
O=30
Frames to VGGT
1.00×
1.49×
2.94×
2.00×
1.53×
Mean ATE (m)
31.53
32.61
34.70
32.91
22.61
Appendix
Table 15 : Compute-matched comparison on KITTI Odometry 00–10. Frames to VGGT is the total number of frames fed to VGGT, relative to SwiftVGGT.
Method
00
01
02
03
04
05
06
07
08
09
10
Avg. ( ↓ )
VGGT-Long [ 8 ]
9.87
111.06
37.56
4.89
3.75
9.09
7.47
4.02
62.86
47.48
25.49
29.41
SwiftVGGT [ 20 ]
8.17
102.53
36.49
8.12
4.88
11.94
8.88
5.01
64.68
44.13
26.18
29.18
VGGT-Bridge (ours)
11.61
61.21
38.18
7.50
3.91
9.92
7.71
4.63
49.29
35.19
25.84
23.18
Appendix
Table 16 : Full-transformation ATE on KITTI Odometry. Bold marks the best result.
Method (segment ID prefix)
1634
1838
3156
3461
3711
4058
4604
5200
6104
Avg. ( ↓ )
VGGT-Long [ 8 ]
3.09
2.54
2.34
4.01
4.04
3.14
2.87
3.41
2.33
3.08
SwiftVGGT [ 20 ]
3.11
2.45
2.35
2.72
4.08
3.11
2.84
2.83
2.21
2.85
VGGT-Bridge (ours)
3.20
2.01
2.36
2.28
4.53
2.94
3.03
2.66
2.22
2.80
Appendix
Table 17 : Full-transformation ATE on Waymo Open. Bold marks the best result.
Configuration
Mean ATE
Δ vs default
Default ( R={3,5} , W=60 , τ=0.7 )
22.61
—
Coarse stride set R
R={3} (drop stride 5)
24.52
+1.91
R={5} (drop stride 3)
27.34
+4.73
R={3,5,7} (add stride 7)
22.90
+0.29
Loop-density window W
Appendix
Table 18 : Hyperparameter sweeps on KITTI Odometry 00–10. Δ is the change relative to the default VGGT-Bridge configuration ( R={3,5} , W=60 , τ=0.7 ). Translation-only ATE RMSE (m) reported, as in the main tables.
Figure 5 : Additional camera trajectory and dense 3D reconstruction comparison on KITTI Odometry.
Figure 6 : Visualization of 3D reconstruction on Waymo Open. VGGT-Bridge reconstructions shown in RGB color (top) and colored by height, with red for low and blue for high (bottom).
Figure 7 : Visualization of 3D reconstruction on Virtual KITTI. VGGT-Bridge reconstructions shown in RGB color (top) and colored by height, with red for low and blue for high (bottom).
Visual Geometry Grounded Transformers (VGGT) have set new benchmarks in high-fidelity 3D scene reconstruction. However, as the sequence length increases, these models suffer from catastrophic geometric forgetting and accumulation drift, primarily due to the quadratic complexity of global attention which necessitates truncated temporal windows. To overcome the resulting geometric drift, we present Mamba-VGGT, an enhanced VGGT framework capable of persistent long-range reasoning. Our key contribution is a Sliding Window Mamba (SWM) memory module that maintains an explicit external memory token across temporal windows. This module leverages selective state-space modeling to distill and propagate global geometric priors, effectively bypassing the memory constraints of traditional transformers. To integrate these long-term temporal cues without disrupting the highly optimized spatial features of the pre-trained VGGT, we propose a Zero-Init Spatial Memory Injector. Utilizing zero-convolutional layers, this injector adaptively fuses persistent memory into the patch token stream, ensuring structural stability and seamless feature alignment. Extensive experiments demonstrate that our approach significantly outperforms existing VGGT-based methods in maintaining spatial consistency and reducing trajectory accumulation errors. Our work provides a scalable, linear-complexity solution for geometry-grounded world modeling in extensive 3D environments.
Tianchen Deng, Zhenxiang Xiong, Nailin Wang +4
Shanghai Jiao Tong University · Nanyang Technological University · ETH Zurich +1
Visual Geometry Grounded Transformer (VGGT) advances 3D reconstruction via scalable Transformer architecture, but the quadratic complexity of global attention prevents long context application. StreamVGGT enables streaming with causal attention, yet its KV cache grows linearly with frames, causing memory overflow and quality degradation. We present RetrieveVGGT, a training-free framework, which formulates context construction for VGGT as a retrieval problem. By retrieving a fixed number of relevant frames at each step, VGGT maintains a controllable memory budget, which is close to its training context length. Interestingly, we find that the similarity between current frame queries and cached history frame keys at the first global attention layer of VGGT is already a strong indicator of relevance, eliminating the need for additional learned scoring. To enhance information diversity similar to a recommender system, we propose Segment Sampling so that the retrieval spans distinct relevant segments rather than a single high-similarity region. We design a pose-aware spatial memory mechanism that organizes history frames according to their already estimated camera poses, enabling location-aware retrieval. Extensive experiments demonstrate that RetrieveVGGT achieves state-of-the-art performance, outperforming StreamVGGT, TTT3R, and InfiniteVGGT while maintaining constant memory usage regardless of sequence length. Code is available at https://github.com/zzctmd/RetrieveVGGT.
Zichen Zou, Xiaosong Jia, Zuxuan Wu +1
Institute of Trustworthy Embodied AI (TEAI), Fudan University · 2Shanghai Key Laboratory of Multimodal Embodied AI
Visual Geometry Grounded Transformer (VGGT) recovers dense 3D scene structure from multi-view images in one forward pass, but quadratic cross-frame attention limits its scalability. Existing training-free accelerators reduce computation uniformly along one axis, missing layer heterogeneity. Our spectral, probing, and causal analyses reveal three regimes: shallow layers lack cross-view structure, middle layers drive cross-view alignment, and deep layers are redundant for dense geometry yet their cross-frame attention remains essential for pose. RegimeVGGT applies layer-wise U-shaped compression along two axes: Saliency-Guided Banded Merging protects geometry- and edge-salient tokens, while Selectively Protected K/V Downsampling preserves cross-frame spatial coverage and the pose-critical path through a phase-shifted spatial grid, a reference-frame anchor, and uncompressed camera/register tokens. Training-free, RegimeVGGT achieves a 6.7x speedup over VGGT* at matched reconstruction quality.
Jinhao You, Shuo Lyu, Zhuohang Lyu +5
University of Pennsylvania · University of California, Irvine · Nanyang Technological University +1