How can a 3D reconstruction system acquire and retain useful information to understand the geometry of a scene from partial views under a limited computation budget? Existing active view acquisition methods typically estimate uncertainty over observed or instantiated geometry, limiting their ability to reason about unseen structure, while long-horizon reconstruction methods often retain redundant observations. We introduce Matisse, a training-free framework that unifies active reconstruction and keyframe selection by leveraging evidence provided by a pretrained generative 3D model. Matisse estimates Evidential Uncertainty from cross-attention evidence associated with 3D latent tokens and derives an Evidential Information Gain to guide both view acquisition and keyframe selection based on the expected reduction in posterior entropy. Matisse supports multi-object scenes through occlusion-aware, object-balanced aggregation and propagates uncertainty through intermediate latents to avoid full reconstruction during planning. Matisse reduces Chamfer distance by 12.7%, 3.8%, and 9.2% on GSO30, YCB-V, and Replica, respectively, relative to the best baseline on each dataset, and achieves a 1.50× end-to-end speedup over the best active reconstruction baseline on GSO30 with the same reconstruction backend. In the GSO30 keyframe selection experiment for long-horizon reconstruction, Matisse achieves comparable Chamfer distance using 14% of the input views compared with Stream3D.
Figures & tables
Figure 1: Overview of Matisse. (a) Active reconstruction workflow. Given the current observation, Matisse estimates uncertainty over observed and unobserved geometry, selects the next best view, and as more frames are acquired, it decides which keyframes to retain to achieve a more complete reconstruction. (b) Real robot experiment. The aggregated depth point cloud with selected camera views and object bounding boxes (left), alongside the real captured scene image (upper right) and reconstructed object meshes (lower right).
Figure 2: Qualitative comparisons on GSO30. Matisse reconstructs geometry and appearance more faithfully than the other active reconstruction baselines.
Figure 3: GSO30 keyframe selection with exponential plateau fits.
Geometry
Appearance
Efficiency
Backend
Method
CD (mm) ↓
IoU ↑
P-FID ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Runtime (s) ↓
Feed-forward
TRELLIS.2 ( Xiang et al., 2026a )
16.768
0.665
66.609
18.658
0.915
0.087
99.969 ± 31.493
TRELLIS.2+M.D. ( Xiang et al., 2026a )
18.357
0.634
76.751
18.409
0.914
0.092
145.346 ± 52.187
SAM3D ( Chen et al., 2026 )
7.043
0.869
22.561
20.421
0.922
0.065
14.513 ± 3.845
3DGS
FisherRF ( Jiang et al., 2024 )
18.494
0.332
137.698
16.145
0.915
0.104
109.827 ± 12.062
GauSS-MI ( Xie et al., 2025 )
12.482
0.421
116.364
10.145
0.802
0.196
95.851 ± 9.837
Table 2: YCB-V results. Best/second-best values are green/yellow.
Geometry
Appearance
Method
CD (mm) ↓
IoU ↑
P-FID ↓
PSNR ↑
SSIM ↑
LPIPS ↓
w/o Information Gain
7.403
0.866
23.633
20.489
0.921
0.066
w/o Explorative Factor
6.460
0.884
20.938
20.831
0.923
0.063
Matisse-G (Ours)
6.451
0.885
21.450
20.849
0.923
0.062
Matisse-RH (Ours)
6.474
0.885
21.192
20.912
0.923
0.063
Table 3: YCB-V component ablations. Best/second-best values are green/yellow.
Table 6: Sensitivity analysis to prior pseudo-evidence α on GSO30. Green and yellow mark the best and second-best settings; the selected α=1 row is gray.
Figure 4: View-budget comparison on YCB-V scene 48. Blue points show MAGICIAN’s measured Chamfer distance from 5 to 50 acquired frames; the blue curve is a descriptive exponential decay fit with a plateau. The orange star marks Matisse’s 5-frame result. The broken vertical axis separates the two CD ranges. Lower CD is better.
YCB-V scene
Method
48
49
50
51
52
53
54
55
56
57
58
59
Mean
Stream3D (24 views)
8.775
4.872
12.294
6.670
6.833
4.348
7.891
5.648
7.436
6.530
6.785
5.015
6.994
Matisse-G (4 views)
6.609
4.588
13.716
6.562
6.722
4.023
7.699
5.685
6.797
5.031
6.482
5.983
6.760
Appendix
Table 7: Additional keyframe-selection results on YCB-V. Entries are scene-wise CD (mm); the final column is the mean over all 55 instances. Lower is better, and bold marks the better method.
Figure 5: Qualitative comparisons on YCB-V. Rows show scene 50 from the fifth camera and scene 59 from the third camera. Columns show ground truth, Random, FisherRF, GauSS-MI, GAVIS, MAGICIAN, and Matisse-G; all reconstruction methods use the Stream3D backend. Predictions are vertex-colored meshes with registration and selected object-orientation corrections for visualization. Ground-truth geometry appears only in the GT column.
Figure 6: Qualitative mesh comparisons on Replica. Rows show office_3 (top) and office_4 (bottom). Columns show ground truth, Random, FisherRF, GauSS-MI, GAVIS, MAGICIAN, and Matisse-G. All reconstruction methods use the Stream3D backend. Registered meshes of the selected objects are shown from a common viewpoint within each row.
Figure 7: Analysis of Matisse’s predicted uncertainty. (a) The pagoda region outside the input image’s field of view exhibits higher uncertainty than visible regions. (b) Uncertainty remains similar at rotationally corresponding locations on the bottle. The concentric rings in the polar plot indicate approximate invariance across yaw angles, consistent with the object’s rotational symmetry. (c) Higher mean uncertainty over occupied tokens is consistent with lower model confidence as Gaussian blur removes visual detail from the bottle image.
Active 3D reconstruction relies on active view selection to maximize reconstruction fidelity under limited capture budgets. However, most existing methods rely on surrogate signals such as parameter uncertainty or geometric heuristics, but these signals are often misaligned with the ultimate goal: the fidelity of rendered predictions. We propose GO-PRE, a goal-oriented next-best-view selection framework that explicitly targets information gain in the prediction space. Specifically, we formulate the objective as maximizing the reduction of the average marginal predictive entropy over a user-specified target view manifold. GO-PRE supports interactive goal specification and yields an efficient acquisition rule that enables real-time computation of information gain. Extensive experiments across benchmarks demonstrate that GO-PRE consistently improves active reconstruction performance and provides more reliable uncertainty quantification compared to state-of-the-art methods.
Yan Song, Zhihao Li, Chenglong Li +3
College of Intelligent Robotics and Advanced Manufacturing, Fudan University, Shanghai, China · School of Software, Shandong University, Jinan, China · School of Data Science and Engineering, East China Normal University, Shanghai, China
Existing active reconstruction systems with Gaussian-splatting maps select observations greedily, optimizing a single next-best-view (NBV) at each step and connecting the chosen views by short-horizon path planning. This greedy decoupling disregards the global structure of scene information, producing inefficient trajectories that waste sensing capacity in transit between selected views. In this work, we study active reconstruction as an ergodic coverage problem: the time-averaged spatial statistics of the sensor trajectory should match a target information distribution induced by the current map. Our approach derives this target distribution online from uncertainty and visibility, and calculates ergodic trajectories via a kernel-ergodic horizon planner with gradient flow and footprint depletion, closing the loop between mapping and trajectory optimization. We thoroughly evaluate TRACE on the Replica dataset against the Next-Best-View (NBV) baselines, improving PSNR by 1.5 dB. Code: https://github.com/spikelab-jhu/trace-active-reconstruction.
Ziyue Zheng, Linli Shi, Bingkun He +2
1Johns Hopkins University, USA · University of Pennsylvania, USA
Active 3D reconstruction of moving objects requires selecting informative viewpoints while accounting for object motion uncertainty during the decision-to-execution delay. Existing methods address only parts of this problem: next-best-view (NBV) planners for object reconstruction typically optimize surface coverage but assume static objects, while motion-aware active perception for moving targets accounts for target motion but prioritizes tracking or visibility over reconstruction coverage. This work presents a motion-uncertainty-aware NBV framework for reconstructing an unknown rigid object undergoing planar motion, using noisy planar position measurements of the object and depth observations from a mobile robot. The key idea is to evaluate each candidate viewpoint by its expected observation quality over plausible future object states induced by motion and measurement uncertainty, rather than at a single predicted object pose. To obtain this predictive belief, a fixed-lag Gaussian Process smoother estimates and predicts the object state from noisy position measurements. The resulting belief is used to generate candidate viewpoints around the predicted object location, filter them by reachability, and estimate their expected coverage-driven scores. Simulation and real-world experiments demonstrate improved reconstruction completeness over non-predictive NBV and prediction-only tracking methods, bridging coverage-driven active reconstruction and prediction-driven tracking.
Karen Li, Mattia Mantovani, Robert J. Wood +2
Harvard University, Cambridge, MA 02134, USA. · University of Modena and Reggio Emilia, 42122 Reggio Emilia, Italy.