A correctly reconstructed distant point appears uncertain even when a frozen 3D model processes equivalent inputs because its output frame rotates fractionally. This exposes a weakness of trainingfree perturbation uncertainty: when outputs contain an unobserved symmetry, run-to-run variation potentially reflects symmetry rather than error. Existing alternatives have trade-offs: built-in confidence is outperformed in most evaluated conditions, while trained evidential heads require modelspecific supervision. For point maps, this research derives a closed-form, error-independent variance term that grows with scene extent and potentially overwhelms the desired signal. Simulation reproduces the effect; all 30 real VGGT view-sets tested exhibit its predicted ∥xp∥2 signature. ReVar3R robustly registers predictions to a common similarity frame before computing per-point variance, without retraining or modifying the frozen model. Optional calibration and fusion use a held-out split. Across VGGT, π3, and MASt3R on six datasets, the same estimator on every backbone lowers AUSE below built-in confidence in 15 of 18 conditions. The staged evaluation yields 11 of 18 wins for the label-free core, 12/18 for label-free equal-weight fusion, 14/18 with held-out weights, and 15/18 when the built-in signal is included. Against a trained evidential head, the result is a trade-off: the head calibrates magnitude better and leads in its training domain, whereas ReVar3R transfers across backbones without adaptation. Its ranking improves point filtering, but it does not detect stable systematic bias, aid novel-view synthesis, or transfer calibration across domains.
Figure 2: Gauge contamination in simulation. (a) A growing injected Sim(3) gauge degrades unaligned perturbation variance (AUSE ↓ ); gauge-aligned variance stays near oracle quality. (b) At a fixed gauge, the degradation grows with the scene’s distance from the origin, the lever arm in the ∥x∥2 term of Eq. ( 3 ).
Figure 3: Main comparison over six datasets. AUSE improvement (built-in − ReVar3R) per backbone and dataset; blue is a gain and red a loss, darker is larger. Positive cells ( 15 of 18 ) are conditions where ReVar3R lowers AUSE.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Model and geometry
f
frozen feed-forward 3D model (weights never modified)
V
number of input views
H,W
height and width of each per-view point map
X∈RV×H×W×3
per-view point maps emitted by f
p,v
pixel (correspondence) index; view index
Appendix
Table 1: Notation. Definitions used throughout the paper, grouped by model, perturbation ensemble, gauge, estimator, and metric.
Sweep
setting
ρunaligned
ρaligned
AUSE un
AUSE al
gauge mag.
0.0
0.991
0.990
0.007
0.008
0.01
0.936
0.989
0.033
0.008
0.03
0.774
0.989
0.117
0.008
0.10
0.395
0.990
0.469
0.007
distance
D=0
0.962
0.989
0.020
0.008
D=20
0.852
0.989
0.090
0.008
Appendix
Table 2: Synthetic gauge contamination. Spearman against true uncertainty and AUSE show that aligned estimates remain invariant while unaligned estimates degrade with gauge magnitude and scene distance. Component rows use magnitude 0.05 ; the N -sweep reports unaligned Spearman and shows persistent contamination.
Figure 4: The gauge term grows with distance within real scenes. After per-view-set normalization and pooling over 30 view-sets from six datasets (eight scenes in all), Uunaligned−Ualigned rises monotonically with point distance ∣xp∣ . Its rank correlation with ∣xp∣2 is positive in 30/30 view-sets (median +0.59 ), matching the ∥xp∥2 term in Eq. ( 3 ).
Centre AUSE
Centre Spearman
Rot. AUSE
ρ(∥c∥2)
Backbone
Dataset
naive
align
naive
align
naive
align
naive
π3
7-Scenes
0.217
0.175
0.276
0.351
0.299
0.314
0.38∗
π3
DTU
0.210
0.119
−0.006
0.521
0.156
0.088
0.15
π3
ETH3D
0.534
0.519
0.157
0.007
0.401
0.445
0.07
VGGT
7-Scenes
0.556
0.461
−0.297
−0.301
0.322
0.291
0.70∗
VGGT
DTU
0.083
0.048
0.121
0.386
0.355
0.305
0.66∗
Appendix
Table 3: The gauge confound transfers to SE(3) camera poses, but ranking gains are mixed. Naive and gauge-aligned pose uncertainty rank pose error over all cameras pooled per condition (AUSE lower is better; Spearman higher is better). Bold marks the better AUSE. Centre alignment lowers AUSE in all six conditions, whereas outlier-robust centre Spearman improves in three (both DTU and π3 /7-Scenes); VGGT/7-Scenes and VGGT/ETH3D remain weak negative rankers. The final column measures Eq. ( 3 )’s ∥c∥2 signature through the rank correlation between naive centre uncertainty and squared camera-centre magnitude, which is strong where the frame moves.
Level
What it adds
Labels
Core
aligned single-family variance (a ranking)
none
+ fusion
rank-space fusion of families with the built-in confidence
held-out split
+ calibration
isotonic map from score to error magnitude (ECE)
held-out split
Appendix
Table 4: The three levels of ReVar3R. The training-free core produces the AUSE/Spearman ranking. Fusion and calibration use a small held-out labelled split; neither changes a single family’s ranking, and calibration alone changes only magnitude. This separation distinguishes ranking from calibration in the Trust3R comparison (Section 6.5 ).
Figure 5: Component decomposition. Mean AUSE reduction (positive = improvement) and, in parentheses, the share of conditions each component helps, split into in-distribution and shifted conditions. Alignment supplies the largest in-distribution gain, held-out learned fusion is the most consistent component, and equal fusion exhibits occasional degradation under shift. The fifth group isolates admitting built-in confidence as a fusion input: it is of the same order as the fusion step itself, which is why the label-free core ( 11/18 ) is reported alongside the fused 15/18 .
Model
Dataset
built-in
unaligned
core (train-free)
full (fused)
VGGT
7-Scenes
0.712
0.655
0.536
0.491
VGGT
DTU
0.248
0.313
0.236
0.171
VGGT
ETH3D
0.740
0.520
0.578
0.635
VGGT
BlendedMVS
0.786
0.597
0.276
0.407
π3
7-Scenes
0.517
0.541
0.439
0.363
π3
DTU
0.273
0.660
0.276
0.220
Appendix
Table 5: Training-free core vs. built-in. Decomposition of the gain (AUSE ↓ ) across the four original datasets (twelve conditions): built-in confidence; unaligned single-family perturbation variance; the training-free core ReVar3R estimator (robust-aligned single-family variance, no fusion, no built-in confidence); and the full held-out fusion. The training-free core beats built-in in 8/12 here and 11/18 with the competitor benchmarks (Table 10 ), versus 5/18 for unaligned perturbation. The gain comes from alignment, not fusion with the baseline.
Aggregator
ID
OOD
overall
Trace variance (used)
0.359
0.464
0.394
Median absolute deviation
0.440
0.507
0.462
Largest eigenvalue
0.365
0.468
0.399
Appendix
Table 6: Aggregation statistic (full grid). Mean AUSE over 3 models ×6 datasets for three aligned-photometric aggregators; lower is better. Trace variance is strongest, median spread is clearly weaker, and the largest eigenvalue is marginally weaker. The trace column reproduces the rung-7 core values of Table 5 on the four original datasets; TUM and KITTI are recomputed here over ten view-sets rather than the eight used in Table 10 .
AUSE
axis
variant
all
ID
OOD
Spearman
helps
group
none
0.543
0.526
0.578
+0.20
3/18
translation
0.500
0.459
0.581
+0.27
8/18
rigid SE(3)
0.460
0.420
0.539
+0.31
7/18
similarity Sim(3)
0.392
0.347
0.483
+0.36
11/18
reference
run 0
0.392
0.347
0.483
+0.36
11/18
Appendix
Table 7: Alignment ablations: group, reference, and scope. Photometric-core means over 18 conditions (AUSE ↓ , Spearman ↑ ). Similarity alignment performs best, reference choices are indistinguishable, and per-view scope beats scene-level scope. “helps” counts conditions beating built-in AUSE; similarity/run-0/per-view is the default core and appears identically in all three blocks. This ablation uses the first 8 view-sets in every condition, whereas Table 5 uses 10 for the four original datasets (TUM and KITTI have eight in both). With the larger sample, unaligned variance beats built-in confidence in 5/18 conditions rather than 3/18 : π3 on ETH3D and VGGT on BlendedMVS change sign.
Figure 6: Alignment removes the dominant nuisance. For each backbone (group) and dataset (colour), AUSE ↓ of photometric perturbation variance without alignment (hollow circle) and of the same runs after gauge alignment (filled circle, the ReVar3R core); each arrow runs from unaligned to aligned, so an arrow sloping down is a reduction. The pair of one condition is offset horizontally so the arrow stays readable where the effect is small. The labels under the groups give the mean reduction over the six datasets (VGGT 0.165 , MASt3R 0.131 , π30.105 ). Alignment lowers AUSE in 15 of the 18 conditions; the exceptions are ETH3D for VGGT and π3 , and BlendedMVS for π3 , the last by 0.009 . Table 5 lists the values for the four original datasets.
input-noise
photometric
Model
Dataset
built-in
unaligned
aligned
aligned
MASt3R
DTU
0.463
0.491
0.358
0.343
MASt3R
ETH3D
0.570
0.568
0.540
0.576
MASt3R
7-Scenes
0.587
0.494
0.374
0.389
π3
DTU
0.273
0.613
0.272
0.269
π3
ETH3D
0.528
0.490
0.453
0.481
Appendix
Table 8: Aligned sampling baseline at matched budget. AUSE ↓ for the input-noise ensemble with and without gauge alignment, beside aligned photometric jitter, all scored identically. Bold marks the best of the three uncertainty estimators per row. Same pass budget, registration and scoring throughout, so the columns differ only in whether the ensemble is aligned and in which perturbation generates it.
AUSE ↓
Spearman ↑
Model
Dataset
built-in
ReVar3R
built-in
ReVar3R
VGGT
7-Scenes
0.712
0.491
0.06
0.32
VGGT
DTU
0.248
0.171
0.31
0.44
VGGT
ETH3D
0.740
0.635
−0.11
0.13
VGGT
BlendedMVS
0.786
0.407
−0.16
0.43
π3
7-Scenes
0.517
0.363
0.36
0.48
Appendix
Table 9: Per-condition results at native resolution. AUSE ( ↓ ) and Spearman ( ↑ ) compare built-in confidence with full ReVar3R. Improvements concentrate where built-in ranking is weak; regressions occur where it is already comparatively strong.
Dataset
Model
built-in
unaligned
core (train-free)
full (fused)
TUM
VGGT
0.296 / +0.48
0.819 / − 0.01
0.446 / +0.39
0.199 / +0.51
TUM
π3
0.273 / +0.25
0.286 / +0.65
0.212 / +0.51
0.198 / +0.46
TUM
MASt3R
0.226 / +0.63
0.433 / +0.24
0.270 / +0.44
0.182 / +0.56
KITTI
VGGT
0.401 / +0.30
0.592 / +0.32
0.432 / +0.58
0.237 / +0.62
KITTI
π3
0.385 / +0.45
0.451 / +0.53
0.337 / +0.64
0.273 / +0.74
KITTI
MASt3R
0.346 / +0.53
0.477 / +0.45
0.313 / +0.67
0.215 / +0.77
Appendix
Table 10: Competitor benchmarks (TUM, KITTI). AUSE ↓ / Spearman ↑ for built-in confidence, unaligned single-family perturbation, the training-free core (aligned single-family variance), and the full fused estimator. Full ReVar3R beats built-in AUSE in 6/6 conditions; the core wins in 3/6 ( π3 on both benchmarks and MASt3R/KITTI).
7-Scenes
TUM
ETH3D
DTU
KITTI
sets
23
12
24
30
12
scenes
7
2
9
30
1
pooled AUSE ↓
built-in
0.535
0.232
0.585
0.557
0.359
ReVar3R
0.293
0.188
0.667
0.403
0.216
Trust3R
0.268
0.119
0.477
0.637
0.319
per-set median AUSE ↓
ReVar3R
0.272
0.148
0.471
0.228
0.209
Appendix
Table 11: ReVar3R vs. a trained evidential head (Trust3R). Frozen MASt3R at 224 px on identical point sets and all available view-sets. AUSE ↓ , Spearman ↑ ; Trust3R uses its paper-default epistemic readout and ReVar3R entries are means over five seeds. “scenes” counts the independent scenes available for inference.
AURC ↓
AUSE ↓ (unnormalized)
built-in
ReVar3R
input
Trust3R
built-in
ReVar3R
input
Trust3R
Dataset
noise
noise
7-Scenes
0.249
0.189
0.224
0.179
0.127
0.068
0.102
0.067
TUM
0.079
0.082
0.101
0.067
0.027
0.030
0.050
0.019
ETH3D
3.812
3.535
3.408
3.513
2.114
1.837
1.710
1.834
DTU
34.95
25.93
33.20
33.59
16.83
7.81
15.08
17.87
Appendix
Table 12: Trust3R’s metric definition, recomputed on the present point sets. Risk-coverage AURC ↓ and unnormalized AUSE ↓ , both in scene units, so methods are comparable within a row but values are not comparable across datasets. Ten view-sets per dataset (eight for TUM and KITTI), seed 0 , Trust3R total readout, and an exact sort-based curve in place of their log-binned estimator; it is a cross-check on the metric definition, not a reproduction of their protocol.
Signal
7-Scenes
DTU
ETH3D
Built-in conf.
0.702 / 0.09
0.402 / 0.24
0.754 / −0.09
Learned probe (frozen features)
0.791 / 0.09
0.245 / 0.56
0.644 / 0.14
ReVar3R (aligned perturbation)
0.546 / 0.29
0.170 / 0.46
0.510 / 0.15
Appendix
Table 13: Frozen-feature probe on VGGT. AUSE ↓ / Spearman ↑ for built-in confidence, a learned probe, and ReVar3R. ReVar3R has lower AUSE throughout, while the probe has higher Spearman on DTU.
Figure 7: Better ranking yields better geometry. Mean error of kept points ↓ vs. removal fraction. Dropping the points ReVar3R ranks most-uncertain retains lower-error geometry than built-in confidence, a naive input-noise ensemble, or random removal, closing 30 – 57% of the built-in-to-oracle gap at 50% removal. All panels use frozen MASt3R: (a) and (b) the 224 px protocol with ten view-sets, (c) the native-resolution KITTI run with eight, where the 224 px Trust3R checkpoint does not apply. Trust3R filters its own refined point map and starts from a different zero-removal error (DTU 31.7 vs. 35.0 ), so read its shape, not its offset.
Figure 8: Calibration improves within domain. Raw versus isotonic-calibrated ECE ( ↓ ) from a within-domain calibration split. Raw perturbation scores are poorly scaled, with DTU and ETH3D saturated; a small held-out isotonic map corrects ECE for ReVar3R and the trained head. On 7-Scenes, Trust3R’s native standard deviation is already calibrated.
Figure 9: Calibration does not transfer across domains. ECE ( ↓ ) for an isotonic map fit on the row domain and evaluated on the column domain, for ReVar3R (left) and Trust3R (right). Both methods calibrate well on the diagonal and poorly off it; neither transfers magnitude calibration across domains.
Figure 10: Sample-count sweep. AUSE ( ↓ ) versus perturbed passes N for each backbone and dataset; shading gives standard deviation across view-sets, and the dashed line marks the N=5 default. In-distribution AUSE flattens by N=5 , while shifted ETH3D is nearly flat throughout.
Method
AUSE ↓
Spearman ↑
ECE ↓
oracle (bound)
0.000
1.00
−
built-in conf.
0.585
0.31
−
best ranking
random removal
0.743
0.00
−
reference
input-noise ensemble
0.766
−0.31
−
worse than random
Trust3R (trained)
0.475
0.25
0.10
best AUSE and calibration
ReVar3R
0.667
0.13
0.38
neither
Appendix
Table 14: Distribution-shift breakdown on ETH3D with frozen MASt3R at 224 px, over 24 view-sets spanning nine scenes. Bold marks the best non-oracle value per column.
Method
Training
Params
Passes
Peak mem.
Latency
Built-in conf.
none
0
0 (free)
≤2.91 GB
∼0.92 s/pair
Input-noise ensemble
none
0
∼5
∼ base
subset of 9 s/set
ReVar3R
none
0
∼10 – 11
2.91 GB
9.09 s/set ( ∼9.9× )
Trust3R
10 ep. / 150 k pairs
40.4 M
1
3.37 GB
0.919 s/pair
Appendix
Table 15: Cost and efficiency on the shared backbone. ReVar3R adds no parameters or training and uses comparable peak memory, but requires multiple passes and roughly an order of magnitude more latency than the single-pass trained head. Latency is wall-clock time on one 16 GB GPU.
Figure 11: Qualitative uncertainty maps from VGGT (seed 0), one typical view-set per dataset (median within-set AUSE gap between built-in confidence and the full estimator). Columns: input, ground-truth error, built-in confidence as uncertainty, unaligned photometric variance, and ReVar3R: the core (the same runs, gauge-aligned) and the full estimator (held-out rank fusion with view-order variance and built-in confidence). Each map shows every pixel’s percentile rank within the view among pixels with ground truth, from low (dark) to high (light) on the scale at the right; grey pixels have no ground truth. The number under each map is that view’s own AUSE.
Model
Dataset
built-in
ReVar3R
wins
MASt3R
BlendedMVS
0.282
0.349±0.011
0/5
MASt3R
DTU
0.463
0.330±0.012
5/5
MASt3R
ETH3D
0.570
0.556±0.005
5/5
MASt3R
KITTI
0.358
0.240±0.007
5/5
MASt3R
7-Scenes
0.587
0.353±0.005
5/5
MASt3R
TUM
0.231
0.172±0.004
5/5
Appendix
Table 16: Full estimator over five seeds. Per-condition AUSE ↓ of the fused estimator, mean ± standard deviation over five seeds, against built-in confidence; “wins” counts the seeds in which ReVar3R is lower.
trim drop
ρ / AUSE
Cauchy ×
ρ / AUSE
depth ext.
ρ / AUSE
0%
0.98/0.011
0.25
0.99/0.008
2.0
0.99/0.008
5%
0.99/0.008
0.5
0.99/0.007
1.0
0.99/0.008
10%
0.99/0.008
1.0
0.99/0.008
0.3
0.99/0.008
20%
0.99/0.007
2.0
0.99/0.008
0.1
0.99/0.008
30%
0.99/0.007
4.0
0.99/0.008
0.03
0.99/0.008
Appendix
Table 17: Alignment robustness in simulation. Aligned-estimator Spearman ↑ / AUSE ↓ against known σtrue remains stable across trimming-fraction, Cauchy-scale, and near-planar sweeps. The default drops 10% and sets Cauchy scale to the median residual.