3D vision transformers such as VGGT predict camera poses and scene geometry from multi-view images in a single forward pass, but their global attention over all concatenated view tokens dominates computation as the number of views grows. To reduce this cost, SparseVGGT and HeSS sparsify attention at the block level, and both retain blocks with high attention probability. However, we observe that attention probability poorly predicts how much the model's behavior actually changes when a block is removed, and we show that this mismatch is why performance collapses as sparsity increases. In this paper, we propose ReSS (ReSidual-ReStoring Sparse Attention), which recasts block selection from a problem of maximizing the retained attention mass to one of minimizing the drift that sparsification leaves in the residual stream. We introduce a drift score that quantifies how much each block shifts the residual, and, since the drift of a drop set depends on the directions of the contribution vectors rather than on their magnitudes alone, an iterative residual restoration procedure that refines the drop set as a whole. Across three backbones and five datasets, ReSS preserves dense performance better than prior methods at matched sparsity. Two further results support drift as the quantity that governs the cost of sparsification: maximizing drift degrades performance faster than random selection, and plotted against realized drift instead of sparsity, all methods fall approximately onto a single curve. Code is available at https://github.com/libary753/ReSS.
Figures & tables
Figure 1 : Sparse selection by residual restoration. Attention writes its output into the residual stream, and sparsification shifts what it writes; that shift is the residual drift. ReSS restores it by selecting the block set that minimizes this drift.
Figure 2 : Rank agreement with masking-induced residual changes. (Left) Rank of each selection score (vertical) against the rank of the actual residual change caused by masking each block (horizontal); the diagonal indicates perfect agreement. (Right) Rank correlation across global attention layers.
Figure 3 : Overview of ReSS. (a) Masking a key block shifts what attention writes into the residual stream. (b) The drift score measures this shift per block; the ch blocks with the highest scores form the initial selected set. (c) Iterative restoration then reduces the drift of the drop set as a whole, swapping blocks while preserving the budget.
Figure 4 : Qualitative results. Two DTU scans reconstructed at five sparsity levels, on VGGT [ 24 ] and π3 [ 26 ] . Columns are the measured sparsity, matched across methods. Points are drawn in green where their distance to the ground truth exceeds 5 mm, and each panel reports the fraction of such points.
Figure 5 : Quantitative results. Reconstruction (recon) and camera pose estimation (pose) across sparsity levels, on the VGGT [ 24 ] (top) and π3 [ 26 ] (bottom) backbones. Dashed lines denote dense attention; arrows mark whether higher ( ↑ ) or lower ( ↓ ) is better. As sparsity grows, ReSS preserves dense performance better than SparseVGGT [ 23 ] and HeSS [ 11 ] .
Figure 6 : Quality against latency. Each point is one operating point on ScanNet++ with the VGGT backbone; the dashed line marks dense attention. Latency covers the full forward pass including all selection overhead. LLM criteria reach dense quality only at dense latency or beyond; video methods lose either quality (SVG) or speed (SVG2, SVG-EAR).
Figure 7 : Number of restoration rounds. Quality saturates after three to four rounds, while the realized drift keeps decreasing beyond that point. The grey line marks the setting we adopt. Dashed line denotes dense attention.
Figure 8 : Drift minimization vs maximization. With the budget and sparsity fixed, we flip the sign of the objective, either only in the restoration stage (swap only) or from the initial selection (full). Maximizing drift collapses performance faster than random selection, supporting drift as a valid selection criterion.
Latency
Scoring mem.
Peak mem.
(s)
(MiB)
(GiB)
SparseVGGT
7.58
—
—
ReSS (full)
7.82 ( +0.24 )
205
16.4
w/o fused kernel
8.84 ( +1.26 )
205
16.4
w/o expansion
8.48 ( +0.90 )
7760
22.4
w/o cached Mh
7.87 ( +0.29 )
278
16.4
Table 1: Cost analysis. Selection cost at 100 views on ScanNet++, sparsity 0.745 . Parenthesized: overhead over SparseVGGT [ 23 ] . RTX 4090, bfloat16.
Figure 9 : Drift determines quality. Performance plotted against the realized drift J of each selection instead of sparsity, on the VGGT [ 24 ] backbone over DTU and ETH3D. Different methods, including the drift-maximizing variants, collapse onto a single curve.
Reconstruction
Pose (AUC@ 30∘ )
DTU ↓
ETH3D ↑
DTU ↑
ETH3D ↑
ReSS (full)
0.791
0.743
0.996
0.795
− head budget
0.983
0.696
0.980
0.762
− restoration
1.151
0.582
0.980
0.702
− drift score
1.481
0.581
0.958
0.643
+ CDF threshold
1.338
0.538
0.963
0.623
Table 2: Ablation study. Components are removed ( − ) or added ( + ) cumulatively, one per row, from ReSS down to SparseVGGT [ 23 ] ; the last row coincides with SparseVGGT. VGGT [ 24 ] backbone, all rows matched at the highest operating point.
dense
s=0.07
s=0.35
s=0.73
SparseVGGT
HeSS
Ours
SparseVGGT
HeSS
Ours
SparseVGGT
HeSS
Ours
VGGT
DTU ↓
0.588
0.597
0.590
0.593
0.727
0.718
0.623
1.338
0.982
0.791
VGGT
ETH3D ↑
0.756
0.754
0.756
0.751
0.739
0.729
0.768
0.538
0.600
0.743
VGGT
HiRoom ↑
0.790
0.781
0.773
0.793
0.738
0.766
0.802
0.472
0.575
0.755
VGGT
ScanNet++ ↑
0.676
0.674
0.681
0.678
0.623
0.629
0.676
0.508
0.527
0.626
VGGT
7-Scenes ↑
0.560
0.560
0.563
0.560
0.553
0.561
0.561
0.525
0.535
0.555
Table S1: Absolute numbers behind Fig. 5 of the main paper. Three of the five operating points, showing the reconstruction metric (Chamfer on DTU, F1 elsewhere). The best value within each compression group is in bold. Sparsity is measured, not requested.
dense
s=0.10
s=0.48
s=0.80
SparseVGGT
HeSS
Ours
SparseVGGT
HeSS
Ours
SparseVGGT
HeSS
Ours
DA3
DTU ↓
0.850
0.852
0.875
0.854
0.985
1.464
0.901
1.625
2.205
1.195
DA3
ETH3D ↑
0.853
0.857
0.848
0.851
0.847
0.851
0.854
0.786
0.835
0.832
DA3
HiRoom ↑
0.894
0.892
0.893
0.891
0.877
0.846
0.855
0.760
0.718
0.796
DA3
ScanNet++ ↑
0.772
0.772
0.772
0.772
0.768
0.766
0.771
0.743
0.755
0.771
DA3
7-Scenes ↑
0.569
0.570
0.569
0.570
0.572
0.572
0.570
0.564
0.571
0.572
Table S2: Absolute numbers for DepthAnything3. The same runs as Fig. S1 . Read as in Tab. S1 .
Figure S1 : DepthAnything3 on the five DA3-Bench datasets. Read as in Fig. 5 : the vertical axis is the metric in its own units with the dashed line at dense, and the horizontal axis is measured sparsity. The trend matches the two backbones in the main paper, but the margins between methods are smaller. On 7-Scenes the methods differ only slightly, so that axis is magnified accordingly.
rounds
track-best
sparsity
CD ↓
AUC@ 30∘↑
5
–
0.7274
0.8403
0.9939
5
✓
0.7274
0.7910
0.9958
10
–
0.7274
0.7952
0.9950
10
✓
0.7274
0.7934
0.9964
Table S3: Track-best on and off. DTU, VGGT, 22 scans, highest operating point. All four arms have the same measured sparsity, so the rows are directly comparable.
DTU
ScanNet++
λ
CD ↓
AUC@ 30∘↑
AUC@ 5∘↑
F1 ↑
0.0
2.3221
0.9425
0.7137
0.7204
0.5
2.1541
0.9729
0.8390
0.7416
1.0
2.2054
0.9774
0.8650
0.7551
Table S4: λ sweep for DepthAnything3. 22 DTU scenes and 20 ScanNet++ scenes at the highest compression operating point; measured sparsity 0.7963 on DTU and 0.7946 on ScanNet++, identical across the three settings. The best value is in bold.
Rule in the original paper
Why it fails on 3D ViTs
Our treatment
Causal attention
Global attention is bidirectional, with no query–key ordering constraint.
Run without a causal mask; order-dependent rules are reinterpreted to keep their intent ( neutral : preserves the rule’s intent).
SpargeAttn [ 31 ] similarity gate τsim=0.6
Nearly every VGGT block is judged self-dissimilar and reverted to dense: measured sparsity 0.006 at the default cumulative threshold of the official implementation ( 0.98 ), and only 0.062 even at 0.1 .
Disable the gate and sweep the cumulative threshold alone ( favors the baseline : no compression curve exists otherwise). Without the gate, the selection effectively reduces to the SparseVGGT rule of a cumulative threshold on pooled query/key inner products.
FlexPrefill [ 12 ] representative queries = last 128
In a causal LM those queries are the only ones that have seen the full context; under bidirectional attention they are a biased sample covering part of the last frame.
Sample uniformly across the sequence ( favors the baseline : the literal rule gives a worse estimate).
In a causal LM only non-negative offsets exist, so the slash family is the lower triangle; under bidirectional attention both signs carry mass, and searching the causal range alone would discard half of the candidate lines.
Search the full signed range −(Tk−1)…(Tq−1) ( neutral : preserves the rule’s intent).
FlexPrefill [ 12 ] always-keep = first/last key block of each query block
In a causal LM the last key block is the diagonal block; under bidirectional attention it merely points at the end of the sequence, and the rule loses its intent.
Keep the first block and the diagonal block, i.e., the blocks the rule pointed at in a causal LM ( neutral : preserves the rule’s intent).
FlexPrefill [ 12 ] vertical/slash line granularity
The paper’s Algorithm 3 defines the lines at token resolution, whereas the released implementation sum-pools the vertical and slash scores into blocks of 128 before selecting; neither granularity is the 64 -wide key block of the shared kernel.
Select at token resolution, following the paper, and keep a key block if a selected line passes through it ( direction not established : rasterizing adds blocks relative to a token-exact mask, whereas the released implementation keeps coarser, wider blocks).
Table S5: Porting decisions. Each row is a point where the original rule fails on 3D ViTs, together with our treatment; the parenthetical states whom the treatment favors, where this can be established. The favorable decisions are cases where, without the treatment, the baseline either could not be evaluated at all or would be penalized as an artifact of the port.
Chamfer ↓
Pose AUC@ 30∘↑
Sparsity
0.35
0.73
0.35
0.73
VGGT (dense)
0.588
0.588
0.999
0.999
SparseVGGT [ 23 ]
0.727
1.338
0.991
0.963
HeSS [ 11 ]
0.718
0.982
0.995
0.982
XAttention [ 28 ]
1.663
1.961
0.959
0.915
SpargeAttn [ 31 ]
0.781
1.964
0.989
0.903
Table S6: Comparison with sparse attention for LLMs. DTU, VGGT [ 24 ] backbone, 22 scans, at two matched sparsity levels. The last block lists three methods proposed for LLM long-context prefill, ported onto the same block grid and the same sparse kernel as ReSS, so that only the block selection criterion differs.
Figure S2 : Comparison with sparse attention for LLMs. DTU, VGGT [ 24 ] backbone. Dashed line denotes dense attention. ReSS preserves dense performance better than the three methods ported from LLM sparse attention.
t (s)
at 50 iter
dense
10.47 (constant)
10.47
SparseVGGT
7.58 (constant)
7.58
SVG2
12.40+0.199⋅iter
22.35
Table S7: Latency against the k-means iteration count. VGGT, 100 views, RTX 4090. The SVG2 row is a linear fit over the sweep; dense and SparseVGGT do not depend on the iteration count, which confirms that the sweep touches nothing but k-means.
Figure S3 : Qualitative results on DTU and ETH3D (VGGT). Each block is one scene; rows are the methods and columns the measured sparsity, matched across methods. A point is drawn in green when its distance to the ground truth exceeds a per-benchmark threshold (5 mm on DTU, and the F1 threshold of 0.25 m on ETH3D), and each panel reports the fraction of points above it.
Figure S4 : Qualitative results on HiRoom and ScanNet++ (VGGT). Read as in Fig. S3 . The accuracy threshold is 0.05 m on both benchmarks.