A correctly reconstructed distant point appears uncertain even when a frozen 3D model processes equivalent inputs because its output frame rotates fractionally. This exposes a weakness of trainingfree perturbation uncertainty: when outputs contain an unobserved symmetry, run-to-run variation potentially reflects symmetry rather than error. Existing alternatives have trade-offs: built-in confidence is outperformed in most evaluated conditions, while trained evidential heads require modelspecific supervision. For point maps, this research derives a closed-form, error-independent variance term that grows with scene extent and potentially overwhelms the desired signal. Simulation reproduces the effect; all 30 real VGGT view-sets tested exhibit its predicted ∥xp∥2 signature. ReVar3R robustly registers predictions to a common similarity frame before computing per-point variance, without retraining or modifying the frozen model. Optional calibration and fusion use a held-out split. Across VGGT, π3, and MASt3R on six datasets, the same estimator on every backbone lowers AUSE below built-in confidence in 15 of 18 conditions. The staged evaluation yields 11 of 18 wins for the label-free core, 12/18 for label-free equal-weight fusion, 14/18 with held-out weights, and 15/18 when the built-in signal is included. Against a trained evidential head, the result is a trade-off: the head calibrates magnitude better and leads in its training domain, whereas ReVar3R transfers across backbones without adaptation. Its ranking improves point filtering, but it does not detect stable systematic bias, aid novel-view synthesis, or transfer calibration across domains.
Figure 2: Gauge contamination in simulation. (a) A growing injected Sim(3) gauge degrades unaligned perturbation variance (AUSE ↓ ); gauge-aligned variance stays near oracle quality. (b) At a fixed gauge, the degradation grows with the scene’s distance from the origin, the lever arm in the ∥x∥2 term of Eq. ( 3 ).
Figure 3: Main comparison over six datasets. AUSE improvement (built-in − ReVar3R) per backbone and dataset; blue is a gain and red a loss, darker is larger. Positive cells ( 15 of 18 ) are conditions where ReVar3R lowers AUSE.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Model and geometry
f
frozen feed-forward 3D model (weights never modified)
V
number of input views
H,W
height and width of each per-view point map
X∈RV×H×W×3
per-view point maps emitted by f
p,v
pixel (correspondence) index; view index
Appendix
Table 1: Notation. Definitions used throughout the paper, grouped by model, perturbation ensemble, gauge, estimator, and metric.
Sweep
setting
ρunaligned
ρaligned
AUSE un
AUSE al
gauge mag.
0.0
0.991
0.990
0.007
0.008
0.01
0.936
0.989
0.033
0.008
0.03
0.774
0.989
0.117
0.008
0.10
0.395
0.990
0.469
0.007
distance
D=0
0.962
0.989
0.020
0.008
D=20
0.852
0.989
0.090
0.008
Appendix
Table 2: Synthetic gauge contamination. Spearman against true uncertainty and AUSE show that aligned estimates remain invariant while unaligned estimates degrade with gauge magnitude and scene distance. Component rows use magnitude 0.05 ; the N -sweep reports unaligned Spearman and shows persistent contamination.
Figure 4: The gauge term grows with distance within real scenes. After per-view-set normalization and pooling over 30 view-sets from six datasets (eight scenes in all), Uunaligned−Ualigned rises monotonically with point distance ∣xp∣ . Its rank correlation with ∣xp∣2 is positive in 30/30 view-sets (median +0.59 ), matching the ∥xp∥2 term in Eq. ( 3 ).
Centre AUSE
Centre Spearman
Rot. AUSE
ρ(∥c∥2)
Backbone
Dataset
naive
align
naive
align
naive
align
naive
π3
7-Scenes
0.217
0.175
0.276
0.351
0.299
0.314
0.38∗
π3
DTU
0.210
0.119
−0.006
0.521
0.156
0.088
0.15
π3
ETH3D
0.534
0.519
0.157
0.007
0.401
0.445
0.07
VGGT
7-Scenes
0.556
0.461
−0.297
−0.301
0.322
0.291
0.70∗
VGGT
DTU
0.083
0.048
0.121
0.386
0.355
0.305
0.66∗
Appendix
Table 3: The gauge confound transfers to SE(3) camera poses, but ranking gains are mixed. Naive and gauge-aligned pose uncertainty rank pose error over all cameras pooled per condition (AUSE lower is better; Spearman higher is better). Bold marks the better AUSE. Centre alignment lowers AUSE in all six conditions, whereas outlier-robust centre Spearman improves in three (both DTU and π3 /7-Scenes); VGGT/7-Scenes and VGGT/ETH3D remain weak negative rankers. The final column measures Eq. ( 3 )’s ∥c∥2 signature through the rank correlation between naive centre uncertainty and squared camera-centre magnitude, which is strong where the frame moves.
Level
What it adds
Labels
Core
aligned single-family variance (a ranking)
none
+ fusion
rank-space fusion of families with the built-in confidence
held-out split
+ calibration
isotonic map from score to error magnitude (ECE)
held-out split
Appendix
Table 4: The three levels of ReVar3R. The training-free core produces the AUSE/Spearman ranking. Fusion and calibration use a small held-out labelled split; neither changes a single family’s ranking, and calibration alone changes only magnitude. This separation distinguishes ranking from calibration in the Trust3R comparison (Section 6.5 ).
Figure 5: Component decomposition. Mean AUSE reduction (positive = improvement) and, in parentheses, the share of conditions each component helps, split into in-distribution and shifted conditions. Alignment supplies the largest in-distribution gain, held-out learned fusion is the most consistent component, and equal fusion exhibits occasional degradation under shift. The fifth group isolates admitting built-in confidence as a fusion input: it is of the same order as the fusion step itself, which is why the label-free core ( 11/18 ) is reported alongside the fused 15/18 .
Model
Dataset
built-in
unaligned
core (train-free)
full (fused)
VGGT
7-Scenes
0.712
0.655
0.536
0.491
VGGT
DTU
0.248
0.313
0.236
0.171
VGGT
ETH3D
0.740
0.520
0.578
0.635
VGGT
BlendedMVS
0.786
0.597
0.276
0.407
π3
7-Scenes
0.517
0.541
0.439
0.363
π3
DTU
0.273
0.660
0.276
0.220
Appendix
Table 5: Training-free core vs. built-in. Decomposition of the gain (AUSE ↓ ) across the four original datasets (twelve conditions): built-in confidence; unaligned single-family perturbation variance; the training-free core ReVar3R estimator (robust-aligned single-family variance, no fusion, no built-in confidence); and the full held-out fusion. The training-free core beats built-in in 8/12 here and 11/18 with the competitor benchmarks (Table 10 ), versus 5/18 for unaligned perturbation. The gain comes from alignment, not fusion with the baseline.
Aggregator
ID
OOD
overall
Trace variance (used)
0.359
0.464
0.394
Median absolute deviation
0.440
0.507
0.462
Largest eigenvalue
0.365
0.468
0.399
Appendix
Table 6: Aggregation statistic (full grid). Mean AUSE over 3 models ×6 datasets for three aligned-photometric aggregators; lower is better. Trace variance is strongest, median spread is clearly weaker, and the largest eigenvalue is marginally weaker. The trace column reproduces the rung-7 core values of Table 5 on the four original datasets; TUM and KITTI are recomputed here over ten view-sets rather than the eight used in Table 10 .
AUSE
axis
variant
all
ID
OOD
Spearman
helps
group
none
0.543
0.526
0.578
+0.20
3/18
translation
0.500
0.459
0.581
+0.27
8/18
rigid SE(3)
0.460
0.420
0.539
+0.31
7/18
similarity Sim(3)
0.392
0.347
0.483
+0.36
11/18
reference
run 0
0.392
0.347
0.483
+0.36
11/18
Appendix
Table 7: Alignment ablations: group, reference, and scope. Photometric-core means over 18 conditions (AUSE ↓ , Spearman ↑ ). Similarity alignment performs best, reference choices are indistinguishable, and per-view scope beats scene-level scope. “helps” counts conditions beating built-in AUSE; similarity/run-0/per-view is the default core and appears identically in all three blocks. This ablation uses the first 8 view-sets in every condition, whereas Table 5 uses 10 for the four original datasets (TUM and KITTI have eight in both). With the larger sample, unaligned variance beats built-in confidence in 5/18 conditions rather than 3/18 : π3 on ETH3D and VGGT on BlendedMVS change sign.
Figure 6: Alignment removes the dominant nuisance. For each backbone (group) and dataset (colour), AUSE ↓ of photometric perturbation variance without alignment (hollow circle) and of the same runs after gauge alignment (filled circle, the ReVar3R core); each arrow runs from unaligned to aligned, so an arrow sloping down is a reduction. The pair of one condition is offset horizontally so the arrow stays readable where the effect is small. The labels under the groups give the mean reduction over the six datasets (VGGT 0.165 , MASt3R 0.131 , π30.105 ). Alignment lowers AUSE in 15 of the 18 conditions; the exceptions are ETH3D for VGGT and π3 , and BlendedMVS for π3 , the last by 0.009 . Table 5 lists the values for the four original datasets.
input-noise
photometric
Model
Dataset
built-in
unaligned
aligned
aligned
MASt3R
DTU
0.463
0.491
0.358
0.343
MASt3R
ETH3D
0.570
0.568
0.540
0.576
MASt3R
7-Scenes
0.587
0.494
0.374
0.389
π3
DTU
0.273
0.613
0.272
0.269
π3
ETH3D
0.528
0.490
0.453
0.481
Appendix
Table 8: Aligned sampling baseline at matched budget. AUSE ↓ for the input-noise ensemble with and without gauge alignment, beside aligned photometric jitter, all scored identically. Bold marks the best of the three uncertainty estimators per row. Same pass budget, registration and scoring throughout, so the columns differ only in whether the ensemble is aligned and in which perturbation generates it.
AUSE ↓
Spearman ↑
Model
Dataset
built-in
ReVar3R
built-in
ReVar3R
VGGT
7-Scenes
0.712
0.491
0.06
0.32
VGGT
DTU
0.248
0.171
0.31
0.44
VGGT
ETH3D
0.740
0.635
−0.11
0.13
VGGT
BlendedMVS
0.786
0.407
−0.16
0.43
π3
7-Scenes
0.517
0.363
0.36
0.48
Appendix
Table 9: Per-condition results at native resolution. AUSE ( ↓ ) and Spearman ( ↑ ) compare built-in confidence with full ReVar3R. Improvements concentrate where built-in ranking is weak; regressions occur where it is already comparatively strong.
Dataset
Model
built-in
unaligned
core (train-free)
full (fused)
TUM
VGGT
0.296 / +0.48
0.819 / − 0.01
0.446 / +0.39
0.199 / +0.51
TUM
π3
0.273 / +0.25
0.286 / +0.65
0.212 / +0.51
0.198 / +0.46
TUM
MASt3R
0.226 / +0.63
0.433 / +0.24
0.270 / +0.44
0.182 / +0.56
KITTI
VGGT
0.401 / +0.30
0.592 / +0.32
0.432 / +0.58
0.237 / +0.62
KITTI
π3
0.385 / +0.45
0.451 / +0.53
0.337 / +0.64
0.273 / +0.74
KITTI
MASt3R
0.346 / +0.53
0.477 / +0.45
0.313 / +0.67
0.215 / +0.77
Appendix
Table 10: Competitor benchmarks (TUM, KITTI). AUSE ↓ / Spearman ↑ for built-in confidence, unaligned single-family perturbation, the training-free core (aligned single-family variance), and the full fused estimator. Full ReVar3R beats built-in AUSE in 6/6 conditions; the core wins in 3/6 ( π3 on both benchmarks and MASt3R/KITTI).
7-Scenes
TUM
ETH3D
DTU
KITTI
sets
23
12
24
30
12
scenes
7
2
9
30
1
pooled AUSE ↓
built-in
0.535
0.232
0.585
0.557
0.359
ReVar3R
0.293
0.188
0.667
0.403
0.216
Trust3R
0.268
0.119
0.477
0.637
0.319
per-set median AUSE ↓
ReVar3R
0.272
0.148
0.471
0.228
0.209
Appendix
Table 11: ReVar3R vs. a trained evidential head (Trust3R). Frozen MASt3R at 224 px on identical point sets and all available view-sets. AUSE ↓ , Spearman ↑ ; Trust3R uses its paper-default epistemic readout and ReVar3R entries are means over five seeds. “scenes” counts the independent scenes available for inference.
AURC ↓
AUSE ↓ (unnormalized)
built-in
ReVar3R
input
Trust3R
built-in
ReVar3R
input
Trust3R
Dataset
noise
noise
7-Scenes
0.249
0.189
0.224
0.179
0.127
0.068
0.102
0.067
TUM
0.079
0.082
0.101
0.067
0.027
0.030
0.050
0.019
ETH3D
3.812
3.535
3.408
3.513
2.114
1.837
1.710
1.834
DTU
34.95
25.93
33.20
33.59
16.83
7.81
15.08
17.87
Appendix
Table 12: Trust3R’s metric definition, recomputed on the present point sets. Risk-coverage AURC ↓ and unnormalized AUSE ↓ , both in scene units, so methods are comparable within a row but values are not comparable across datasets. Ten view-sets per dataset (eight for TUM and KITTI), seed 0 , Trust3R total readout, and an exact sort-based curve in place of their log-binned estimator; it is a cross-check on the metric definition, not a reproduction of their protocol.
Signal
7-Scenes
DTU
ETH3D
Built-in conf.
0.702 / 0.09
0.402 / 0.24
0.754 / −0.09
Learned probe (frozen features)
0.791 / 0.09
0.245 / 0.56
0.644 / 0.14
ReVar3R (aligned perturbation)
0.546 / 0.29
0.170 / 0.46
0.510 / 0.15
Appendix
Table 13: Frozen-feature probe on VGGT. AUSE ↓ / Spearman ↑ for built-in confidence, a learned probe, and ReVar3R. ReVar3R has lower AUSE throughout, while the probe has higher Spearman on DTU.
Figure 7: Better ranking yields better geometry. Mean error of kept points ↓ vs. removal fraction. Dropping the points ReVar3R ranks most-uncertain retains lower-error geometry than built-in confidence, a naive input-noise ensemble, or random removal, closing 30 – 57% of the built-in-to-oracle gap at 50% removal. All panels use frozen MASt3R: (a) and (b) the 224 px protocol with ten view-sets, (c) the native-resolution KITTI run with eight, where the 224 px Trust3R checkpoint does not apply. Trust3R filters its own refined point map and starts from a different zero-removal error (DTU 31.7 vs. 35.0 ), so read its shape, not its offset.
Figure 8: Calibration improves within domain. Raw versus isotonic-calibrated ECE ( ↓ ) from a within-domain calibration split. Raw perturbation scores are poorly scaled, with DTU and ETH3D saturated; a small held-out isotonic map corrects ECE for ReVar3R and the trained head. On 7-Scenes, Trust3R’s native standard deviation is already calibrated.
Figure 9: Calibration does not transfer across domains. ECE ( ↓ ) for an isotonic map fit on the row domain and evaluated on the column domain, for ReVar3R (left) and Trust3R (right). Both methods calibrate well on the diagonal and poorly off it; neither transfers magnitude calibration across domains.
Figure 10: Sample-count sweep. AUSE ( ↓ ) versus perturbed passes N for each backbone and dataset; shading gives standard deviation across view-sets, and the dashed line marks the N=5 default. In-distribution AUSE flattens by N=5 , while shifted ETH3D is nearly flat throughout.
Method
AUSE ↓
Spearman ↑
ECE ↓
oracle (bound)
0.000
1.00
−
built-in conf.
0.585
0.31
−
best ranking
random removal
0.743
0.00
−
reference
input-noise ensemble
0.766
−0.31
−
worse than random
Trust3R (trained)
0.475
0.25
0.10
best AUSE and calibration
ReVar3R
0.667
0.13
0.38
neither
Appendix
Table 14: Distribution-shift breakdown on ETH3D with frozen MASt3R at 224 px, over 24 view-sets spanning nine scenes. Bold marks the best non-oracle value per column.
Method
Training
Params
Passes
Peak mem.
Latency
Built-in conf.
none
0
0 (free)
≤2.91 GB
∼0.92 s/pair
Input-noise ensemble
none
0
∼5
∼ base
subset of 9 s/set
ReVar3R
none
0
∼10 – 11
2.91 GB
9.09 s/set ( ∼9.9× )
Trust3R
10 ep. / 150 k pairs
40.4 M
1
3.37 GB
0.919 s/pair
Appendix
Table 15: Cost and efficiency on the shared backbone. ReVar3R adds no parameters or training and uses comparable peak memory, but requires multiple passes and roughly an order of magnitude more latency than the single-pass trained head. Latency is wall-clock time on one 16 GB GPU.
Figure 11: Qualitative uncertainty maps from VGGT (seed 0), one typical view-set per dataset (median within-set AUSE gap between built-in confidence and the full estimator). Columns: input, ground-truth error, built-in confidence as uncertainty, unaligned photometric variance, and ReVar3R: the core (the same runs, gauge-aligned) and the full estimator (held-out rank fusion with view-order variance and built-in confidence). Each map shows every pixel’s percentile rank within the view among pixels with ground truth, from low (dark) to high (light) on the scale at the right; grey pixels have no ground truth. The number under each map is that view’s own AUSE.
Model
Dataset
built-in
ReVar3R
wins
MASt3R
BlendedMVS
0.282
0.349±0.011
0/5
MASt3R
DTU
0.463
0.330±0.012
5/5
MASt3R
ETH3D
0.570
0.556±0.005
5/5
MASt3R
KITTI
0.358
0.240±0.007
5/5
MASt3R
7-Scenes
0.587
0.353±0.005
5/5
MASt3R
TUM
0.231
0.172±0.004
5/5
Appendix
Table 16: Full estimator over five seeds. Per-condition AUSE ↓ of the fused estimator, mean ± standard deviation over five seeds, against built-in confidence; “wins” counts the seeds in which ReVar3R is lower.
trim drop
ρ / AUSE
Cauchy ×
ρ / AUSE
depth ext.
ρ / AUSE
0%
0.98/0.011
0.25
0.99/0.008
2.0
0.99/0.008
5%
0.99/0.008
0.5
0.99/0.007
1.0
0.99/0.008
10%
0.99/0.008
1.0
0.99/0.008
0.3
0.99/0.008
20%
0.99/0.007
2.0
0.99/0.008
0.1
0.99/0.008
30%
0.99/0.007
4.0
0.99/0.008
0.03
0.99/0.008
Appendix
Table 17: Alignment robustness in simulation. Aligned-estimator Spearman ↑ / AUSE ↓ against known σtrue remains stable across trimming-fraction, Cauchy-scale, and near-planar sweeps. The default drops 10% and sets Cauchy scale to the median residual.
Geometric foundation models hold promise for unconstrained dense geometry prediction from uncalibrated images. However, in current feed-forward designs, their predicted confidence scores are heuristic, lack probabilistic interpretation, and often fail to indicate where and how much the predicted geometry can be trusted. To address this gap, we present Trust3R, a lightweight evidential uncertainty framework for feed-forward 3D reconstruction. Trust3R combines gated residual mean refinement with a Normal-Inverse-Wishart evidential head, yielding a closed-form multivariate Student-t distribution for per-point geometric uncertainty. This design provides probabilistically grounded pointmap uncertainty estimates while adding moderate inference overhead. We evaluate on diverse indoor and outdoor benchmarks and compare against MASt3R's built-in confidence map as well as common uncertainty-aware baselines spanning single-pass heteroscedastic regression and sampling-based methods such as MC dropout and deep ensembles. Experimental results show that Trust3R consistently improves risk-coverage and sparsification, and generally improves geometric accuracy. These gains are reflected in stronger uncertainty ranking across benchmarks, with 25% lower AURC and 41% lower AUSE on ScanNet++, providing a practical reliability signal for uncertainty-aware weighting in downstream geometry pipelines. The project page and code are available at https://trust3r-z.github.io/.
Zihao Zhu, Wenyuan Zhao, Nuo Chen +2
Department of Electrical and Computer Engineering, Texas A&M University, College Station, TX, USA
Feed-forward 3D reconstruction models output a per-pixel confidence that is used by downstream systems as an uncertainty signal. The confidence is trained to serve as a weight in the training loss of models. Whether the confidence can be used as an uncertainty magnitude has not been measured. We audit seven backbones on 13 datasets and score the confidence on four properties, i.e., ranking of error, ratio of error to uncertainty on average, slope of this ratio across the confidence range, and coverage of the implied error distribution. Although the confidence ranks error quite well, the uncertainty decoded from the confidence is too small compared to the actual error. The uncertainty has the right size only under the exact training conditions. The median case is off by at least 2.4x across all seven models, while the uncertainty is further off the more confident the model is. Our work shows that the overconfidence appears on unseen scenes even when the model reaches its loss's optimum. As a post-hoc repair we fit a power law on the confidence with two constants per backbone--dataset pair. The repair brings all four audited properties to target at the dataset level, while leaving ranking untouched. Fitted with the target dataset held out, the constants bring the median case from 2.4x off to 1.35x. The repair does not hold below the dataset level, where two-thirds of held-out scenes are still more than five points off in coverage. We attribute what the repair cannot reach to the model, which carries neither the scale of the error nor the shape of its distribution across predictions. We release the audit protocol, its results, and the fitted constants per backbone-dataset pair.
Nanxing Nick Deng, Qing Cheng, Niclas Zeller +1
Technical University of Munich · Munich Center for Machine Learning (MCML) · Karlsruhe University of Applied Sciences
Modern feed-forward 3D reconstruction methods like VGGT predict pixel-aligned pointmaps in camera-centric coordinate frames. However, this choice of coordinate frame is not always optimal. We propose instead to predict pointmaps in upright, gravity-aligned frames that exploit strong structural cues present in many real-world scenes. Unlike camera-centric frames, gravity-aligned frames share a common vertical axis across viewpoints, reducing the rotational degrees of freedom needed to relate pointmaps to one another. To this end, we introduce the Gravity Grounded Geometry Transformer (G3T), fine-tuned from existing models on gravity-aligned 3D data. G3T produces highly accurate gravity-aware predictions, including upright pointmaps and camera-to-gravity poses. We further introduce G3T-Long, a submap-based incremental 3D reconstruction pipeline that leverages the reduced rotational degrees of freedom afforded by upright frames to achieve significantly improved reconstruction accuracy.