Feed-forward visual geometry transformers such as VGGT reconstruct dense 3D structure from images in a single forward pass, simplifying multi-view 3D reconstruction. However, their quadratic attention complexity makes them difficult to scale to long sequences with thousands of frames. Chunk-and-align frameworks address this by splitting a long sequence into overlapping chunks and stitching their local reconstructions into a pose graph. Yet existing methods connect only sequentially adjacent chunks, so small per-frame errors accumulate along the chain into large-scale drift. To move beyond sequential edges, we propose VGGT-Bridge, which adds long-range skip edges that directly constrain non-adjacent chunks without retraining. By running VGGT on sparsely sampled coarse chunks, each coarse chunk bridges distant fine chunks into a single direct constraint. We further turn VGGT's first-frame scale bias into a drift correction by feeding selected coarse chunks in reverse, and a loop-aware policy keeps this reversal compatible with existing loop closures. VGGT-Bridge reduces ATE by 28.3% on KITTI Odometry, 18.8% on Virtual KITTI, and 10.0% on Waymo Open over the SwiftVGGT baseline, achieving the best performance among all chunk-and-align methods.
Figures & tables
Figure 1 : Prior methods connect only sequentially adjacent chunks ( S1→S2→⋯ ), so per-frame errors accumulate along the chain into large-scale drift. VGGT-Bridge inserts skip edges (pink arcs, e.g . , S1→S3 ) that directly constrain non-adjacent chunks under the chunk-and-align pipeline. On KITTI Odometry Seq. 08, these long-range constraints lower ATE RMSE to 47.1 m.
Figure 2 : Method overview. Gray, blue, and yellow boxes denote input frames, sparsely sampled coarse frames, and alignment overlaps, respectively. Standard chunk-and-align pipelines construct overlapping fine chunks and form sequential edges Eseq (blue) between adjacent chunks, with optional loop edges Eloop (purple) between revisited chunks. VGGT-Bridge additionally runs VGGT on sparse coarse chunks and composes two cross-scale Sim(3) alignments through each coarse chunk into a long-range skip edge Eskip (rose) between non-adjacent fine chunks. Forward/reverse anchor markers show the selected processing direction of each coarse chunk, chosen by a loop-aware policy to complement the available loop closures. All edge sets are jointly optimized in a lightweight Sim(3) pose graph to recover globally consistent chunk poses.
Method
Calib.
00
01
02
03
04
05
06
07
08
09
10
Avg. ( ↓ )
# frames
4541
1101
4661
801
271
2761
1101
1101
4071
1591
1201
Length (m)
2453
5067
3724
561
394
2206
1233
650
3223
1705
920
Speed (m/fr)
0.82
2.23
1.09
0.70
1.45
0.80
1.12
0.59
0.79
1.07
0.77
Loop
✓
×
✓
×
×
✓
✓
✓
×
✓
×
ORB-SLAM2 [ 28 ]
✓
6.03
508.34
14.76
1.02
1.57
4.04
11.16
2.19
38.85
8.39
6.63
54.82
DPVO [ 40 ]
✓
113.21
12.69
123.40
2.09
0.68
58.96
54.78
19.26
115.90
75.10
13.63
53.61
Table 1 : Performance comparison on KITTI Odometry dataset. Bold / underline mark the best/second-best among the calibration-free methods. “Calib.” = uses camera intrinsics, “TL” = tracking lost, “OOM” = out of memory. Results with * are reproduced by us.
Method (segment ID prefix)
Calib.
1634
1838
3156
3461
3711
4058
4604
5200
6104
Avg. ( ↓ )
# frames
198
199
199
199
196
199
198
199
198
Length (m)
160
42
165
351
273
86
266
135
63
Speed (m/fr)
0.81
0.21
0.83
1.77
1.39
0.43
1.34
0.68
0.32
Traffic
Low
High
Low
Low
Med.
Low
Med.
Low
High
DROID-SLAM [ 39 ]
✓
3.705
0.301
0.447
8.653
9.320
7.621
4.170
TL
0.264
4.396
MASt3R-SLAM [ 29 ]
×
4.500
0.556
1.833
12.544
8.601
1.412
5.428
7.910
1.195
5.560
Table 2 : Performance comparison on Waymo Open. Bold / underline mark the best/second-best among calibration-free methods. “Calib.” = uses camera intrinsics. Results with * are reproduced by us.
Method
Calib.
Scene01
Scene02
Scene06
Scene18
Scene20
Avg. ( ↓ )
DROID-SLAM [ 39 ]
✓
1.137
0.064
0.038
2.205
4.157
1.520
MASt3R-SLAM [ 29 ]
×
TL
–
CUT3R [ 45 ]
×
48.362
20.119
0.772
17.149
103.692
38.019
Fast3R [ 52 ]
×
OOM
–
VGGT-Long [ 8 ]
×
1.049
0.703
0.438
1.397
6.683
2.054
SwiftVGGT * [ 20 ]
×
5.178
0.332
0.186
1.817
4.960
2.495
Table 3 : Performance comparison on Virtual KITTI. Bold / underline mark the best/second-best among calibration-free methods. “Calib.” = uses camera intrinsics, “TL” = tracking lost, “OOM” = out of memory (on at least one weather variant within the scene). Results with * are reproduced by us.
Figure 3 : Camera trajectory and dense 3D reconstruction comparison on KITTI Odometry.
Figure 4 : Scale bias on KITTI Odometry 00–10 ( n=11 sequences). Left: mean ± std of the pairwise Umeyama logs measured during chunk alignment, for sequential and coarse chunks. The arrow marks the shift Δ from forward to reverse anchoring, with the percentage of edges where Δ>0 . Right: residual local scale error ∣logs∣ (bars) and the ATE RMSE (lines) of the final trajectory, for SwiftVGGT and our method. The scale error is measured on sliding 75 -frame windows against the ground truth.
Translation ATE (m)
Mean
Configuration
00
01
02
03
04
05
06
07
08
09
10
Trans
Rot
Scale
# frames
4541
1101
4661
801
271
2761
1101
1101
4071
1591
1201
Loop
✓
–
✓
–
–
✓
✓
✓
–
✓
–
SwiftVGGT
9.92
114.96
55.45
10.38
5.48
12.50
8.38
3.50
69.74
30.92
25.55
31.53
13.38
0.157
+ coarse-stride chunk
13.00
69.68
53.43
9.26
4.44
12.30
7.28
4.68
72.43
24.06
25.45
26.91
12.15
0.139
+ reverse coarse chunk
15.01
56.35
49.80
9.50
3.97
10.88
8.87
3.16
45.53
27.59
20.74
22.85
12.41
0.115
Table 4 : Component ablation on KITTI Odometry. Each row cumulatively adds one component on top of the previous. Trans and Rot are the translation (m) and rotation ( ∘ ) ATE, and Scale is the mean ∣logs∣ of Fig. 4 (right).
Table 9
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
KITTI Odometry
Virtual KITTI
Waymo Open
Fine chunk size B
75
75
60
Fine overlap O
30
30
24
Coarse ( r=3 )
B3
35
35
35
O3
14
14
14
Coarse ( r=5 )
B5
25
25
25
O5
10
10
10
Appendix
Table 7 : Hyperparameter settings.
VGGT-Bridge
Dataset
# seq.
Mean logs
SwiftVGGT
wˉ=0.3
wˉ=0.7
wˉ=1.0
KITTI Odometry
11
−0.0370±0.0225
31.53
22.61
22.82
23.06
Virtual KITTI
30
−0.0761±0.0896
2.495
2.027
1.853
1.790
Waymo Open
9
−0.0470±0.0341
1.557
1.401
1.406
1.408
EuRoC MAV
11
−0.0387±0.0221
1.722
1.698
1.632
1.638
7-Scenes
7
−0.0167±0.0169
0.087
0.083
0.082
0.082
Appendix
Table 8 : Anchoring bias and skip-edge weight across five benchmarks. Mean logs is the mean ± std of the fine sequential edges over the sequences, as in Fig. 4 (left) of the main paper. The remaining columns report the mean translation ATE RMSE. Bold marks the best result in each row.
Method
00
01
02
03
04
05
06
07
08
09
10
Mean
SwiftVGGT [ 20 ]
9.92
114.96
55.45
10.38
5.48
12.50
8.38
3.50
69.74
30.92
25.55
31.53
VGGT-Bridge
13.74
56.34
48.70
9.33
4.00
11.02
6.93
3.15
47.15
26.24
22.05
22.61
Δ (%)
+38.5
−51.0
−12.2
−10.1
−27.0
−11.8
−17.3
−10.0
−32.4
−15.1
−13.7
−28.3
VGGT-Long [ 8 ]
10.81
66.46
56.00
5.56
3.93
10.11
6.90
2.83
57.45
31.17
22.11
24.85
+ VGGT-Bridge
12.08
41.47
49.21
6.57
3.20
10.15
7.65
4.30
47.61
23.92
20.50
20.61
Δ (%)
+11.7
−37.6
−12.1
+18.2
−18.6
+0.4
+10.9
+51.9
−17.1
−23.3
−7.3
−17.1
Appendix
Table 9 : Cross-framework transfer of VGGT-Bridge on VGGT-Long. (per-sequence translation ATE RMSE (m) on KITTI Odometry). The same skip-edge construction is added to two backbones, and Δ is the change relative to each backbone.
VGGT-Long [ 8 ]
SwiftVGGT [ 20 ]
VGGT-Bridge (ours)
Seq
Acc
Comp
CD
Acc
Comp
CD
Acc
Comp
CD
00
1.62
7.31
4.47
1.95
7.24
4.60
2.11
8.78
5.45
01
42.73
16.83
29.78
34.65
20.09
27.37
20.91
7.86
14.38
02
18.72
30.28
24.50
17.82
32.57
25.19
13.72
29.19
21.45
04
0.80
6.80
3.80
1.59
6.53
4.06
1.59
6.84
4.21
05
1.24
7.17
4.21
4.38
9.29
6.84
1.73
7.55
4.64
Appendix
Table 10 : Reconstruction quality on KITTI Odometry. Acc / Comp / CD (m), lower is better. Averaged over the ten sequences with available raw LiDAR scan, omitting Seq. 03. Best per metric per sequence in bold .
VGGT-Long [ 8 ]
SwiftVGGT [ 20 ]
VGGT-Bridge (ours)
Segment
Acc
Comp
CD
Acc
Comp
CD
Acc
Comp
CD
1634
0.87
2.11
1.49
1.02
1.81
1.42
0.94
1.90
1.42
1838
0.45
5.87
3.16
1.38
5.85
3.62
1.35
5.44
3.39
3156
0.66
2.58
1.62
1.00
2.26
1.63
0.99
2.11
1.55
3461
0.93
4.47
2.70
1.38
5.53
3.45
1.36
5.50
3.43
3711
1.96
4.93
3.44
2.60
4.26
3.43
2.44
4.37
3.41
Appendix
Table 11 : Reconstruction quality on Waymo Open. Acc / Comp / CD (m), lower is better. Best per metric per segment in bold .
Scene 01
Scene 02
Method
Clone
Fog
Morning
Overcast
Rain
Sunset
Clone
Fog
Morning
Overcast
Rain
Sunset
DROID-SLAM † [ 39 ]
1.03
1.87
0.99
1.01
0.78
1.15
0.10
0.04
0.05
0.05
0.04
0.11
MASt3R-SLAM [ 29 ]
TL
CUT3R [ 45 ]
43.30
62.19
50.61
38.73
51.55
43.79
23.77
9.95
28.41
24.64
7.96
25.97
Fast3R [ 52 ]
OOM
VGGT-Long [ 8 ]
0.76
0.87
0.93
0.67
1.80
1.26
0.72
0.71
0.72
0.68
0.69
0.69
Appendix
Table 12 : Per-scene, per-condition ATE RMSE (m) on Virtual KITTI. Best among the calibration-free methods in bold . “ † ” uses camera intrinsics, all other methods are calibration-free. “TL” = tracking lost, “OOM” = out of memory. Results with * are reproduced by us.
KITTI Seq. 07 (1,101 frames)
KITTI Seq. 08 (4,071 frames)
Stage
SwiftVGGT
Ours
Δ
SwiftVGGT
Ours
Δ
Loop detection
3.1
3.0
−0.1
10.2
10.1
−0.1
VGGT chunk inference
58.1
84.4
+26.3
197.6
307.9
+110.3
Loop-centric chunk inference
1.1
1.3
+0.2
0.0
0.0
+0.0
Loop-closure alignment
0.9
1.1
+0.2
0.0
0.0
+0.0
Sequential-edge composition
29.9
30.0
+0.1
76.0
76.0
+0.0
Appendix
Table 13 : Per-stage runtime breakdown on KITTI Odometry Seq. 07 and Seq. 08. Seq. 07 has one detected loop pair and activates the loop pipeline, while Seq. 08 has none detected. The loop pipeline is identical in both methods, so only the stages VGGT-Bridge modifies differ.
GPU
Host
Method
Mean
Max
Mean
Max
SwiftVGGT [ 20 ]
10.30
14.49
15.74
29.90
VGGT-Bridge
10.30
14.49
19.81
38.90
Appendix
Table 14 : Peak memory on KITTI Odometry 00–10 (GiB), as the mean and the maximum over the eleven sequences. GPU is the peak memory reserved by PyTorch, and host is the peak resident memory.
SwiftVGGT, O
SwiftVGGT
VGGT-Bridge
30
45
60
+B′=90
O=30
Frames to VGGT
1.00×
1.49×
2.94×
2.00×
1.53×
Mean ATE (m)
31.53
32.61
34.70
32.91
22.61
Appendix
Table 15 : Compute-matched comparison on KITTI Odometry 00–10. Frames to VGGT is the total number of frames fed to VGGT, relative to SwiftVGGT.
Method
00
01
02
03
04
05
06
07
08
09
10
Avg. ( ↓ )
VGGT-Long [ 8 ]
9.87
111.06
37.56
4.89
3.75
9.09
7.47
4.02
62.86
47.48
25.49
29.41
SwiftVGGT [ 20 ]
8.17
102.53
36.49
8.12
4.88
11.94
8.88
5.01
64.68
44.13
26.18
29.18
VGGT-Bridge (ours)
11.61
61.21
38.18
7.50
3.91
9.92
7.71
4.63
49.29
35.19
25.84
23.18
Appendix
Table 16 : Full-transformation ATE on KITTI Odometry. Bold marks the best result.
Method (segment ID prefix)
1634
1838
3156
3461
3711
4058
4604
5200
6104
Avg. ( ↓ )
VGGT-Long [ 8 ]
3.09
2.54
2.34
4.01
4.04
3.14
2.87
3.41
2.33
3.08
SwiftVGGT [ 20 ]
3.11
2.45
2.35
2.72
4.08
3.11
2.84
2.83
2.21
2.85
VGGT-Bridge (ours)
3.20
2.01
2.36
2.28
4.53
2.94
3.03
2.66
2.22
2.80
Appendix
Table 17 : Full-transformation ATE on Waymo Open. Bold marks the best result.
Configuration
Mean ATE
Δ vs default
Default ( R={3,5} , W=60 , τ=0.7 )
22.61
—
Coarse stride set R
R={3} (drop stride 5)
24.52
+1.91
R={5} (drop stride 3)
27.34
+4.73
R={3,5,7} (add stride 7)
22.90
+0.29
Loop-density window W
Appendix
Table 18 : Hyperparameter sweeps on KITTI Odometry 00–10. Δ is the change relative to the default VGGT-Bridge configuration ( R={3,5} , W=60 , τ=0.7 ). Translation-only ATE RMSE (m) reported, as in the main tables.
Figure 5 : Additional camera trajectory and dense 3D reconstruction comparison on KITTI Odometry.
Figure 6 : Visualization of 3D reconstruction on Waymo Open. VGGT-Bridge reconstructions shown in RGB color (top) and colored by height, with red for low and blue for high (bottom).
Figure 7 : Visualization of 3D reconstruction on Virtual KITTI. VGGT-Bridge reconstructions shown in RGB color (top) and colored by height, with red for low and blue for high (bottom).