How can a 3D reconstruction system acquire and retain useful information to understand the geometry of a scene from partial views under a limited computation budget? Existing active view acquisition methods typically estimate uncertainty over observed or instantiated geometry, limiting their ability to reason about unseen structure, while long-horizon reconstruction methods often retain redundant observations. We introduce Matisse, a training-free framework that unifies active reconstruction and keyframe selection by leveraging evidence provided by a pretrained generative 3D model. Matisse estimates Evidential Uncertainty from cross-attention evidence associated with 3D latent tokens and derives an Evidential Information Gain to guide both view acquisition and keyframe selection based on the expected reduction in posterior entropy. Matisse supports multi-object scenes through occlusion-aware, object-balanced aggregation and propagates uncertainty through intermediate latents to avoid full reconstruction during planning. Matisse reduces Chamfer distance by 12.7%, 3.8%, and 9.2% on GSO30, YCB-V, and Replica, respectively, relative to the best baseline on each dataset, and achieves a 1.50× end-to-end speedup over the best active reconstruction baseline on GSO30 with the same reconstruction backend. In the GSO30 keyframe selection experiment for long-horizon reconstruction, Matisse achieves comparable Chamfer distance using 14% of the input views compared with Stream3D.
Figures & tables
Figure 1: Overview of Matisse. (a) Active reconstruction workflow. Given the current observation, Matisse estimates uncertainty over observed and unobserved geometry, selects the next best view, and as more frames are acquired, it decides which keyframes to retain to achieve a more complete reconstruction. (b) Real robot experiment. The aggregated depth point cloud with selected camera views and object bounding boxes (left), alongside the real captured scene image (upper right) and reconstructed object meshes (lower right).
Figure 2: Qualitative comparisons on GSO30. Matisse reconstructs geometry and appearance more faithfully than the other active reconstruction baselines.
Figure 3: GSO30 keyframe selection with exponential plateau fits.
Geometry
Appearance
Efficiency
Backend
Method
CD (mm) ↓
IoU ↑
P-FID ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Runtime (s) ↓
Feed-forward
TRELLIS.2 ( Xiang et al., 2026a )
16.768
0.665
66.609
18.658
0.915
0.087
99.969 ± 31.493
TRELLIS.2+M.D. ( Xiang et al., 2026a )
18.357
0.634
76.751
18.409
0.914
0.092
145.346 ± 52.187
SAM3D ( Chen et al., 2026 )
7.043
0.869
22.561
20.421
0.922
0.065
14.513 ± 3.845
3DGS
FisherRF ( Jiang et al., 2024 )
18.494
0.332
137.698
16.145
0.915
0.104
109.827 ± 12.062
GauSS-MI ( Xie et al., 2025 )
12.482
0.421
116.364
10.145
0.802
0.196
95.851 ± 9.837
Table 2: YCB-V results. Best/second-best values are green/yellow.
Geometry
Appearance
Method
CD (mm) ↓
IoU ↑
P-FID ↓
PSNR ↑
SSIM ↑
LPIPS ↓
w/o Information Gain
7.403
0.866
23.633
20.489
0.921
0.066
w/o Explorative Factor
6.460
0.884
20.938
20.831
0.923
0.063
Matisse-G (Ours)
6.451
0.885
21.450
20.849
0.923
0.062
Matisse-RH (Ours)
6.474
0.885
21.192
20.912
0.923
0.063
Table 3: YCB-V component ablations. Best/second-best values are green/yellow.
Table 6: Sensitivity analysis to prior pseudo-evidence α on GSO30. Green and yellow mark the best and second-best settings; the selected α=1 row is gray.
Figure 4: View-budget comparison on YCB-V scene 48. Blue points show MAGICIAN’s measured Chamfer distance from 5 to 50 acquired frames; the blue curve is a descriptive exponential decay fit with a plateau. The orange star marks Matisse’s 5-frame result. The broken vertical axis separates the two CD ranges. Lower CD is better.
YCB-V scene
Method
48
49
50
51
52
53
54
55
56
57
58
59
Mean
Stream3D (24 views)
8.775
4.872
12.294
6.670
6.833
4.348
7.891
5.648
7.436
6.530
6.785
5.015
6.994
Matisse-G (4 views)
6.609
4.588
13.716
6.562
6.722
4.023
7.699
5.685
6.797
5.031
6.482
5.983
6.760
Appendix
Table 7: Additional keyframe-selection results on YCB-V. Entries are scene-wise CD (mm); the final column is the mean over all 55 instances. Lower is better, and bold marks the better method.
Figure 5: Qualitative comparisons on YCB-V. Rows show scene 50 from the fifth camera and scene 59 from the third camera. Columns show ground truth, Random, FisherRF, GauSS-MI, GAVIS, MAGICIAN, and Matisse-G; all reconstruction methods use the Stream3D backend. Predictions are vertex-colored meshes with registration and selected object-orientation corrections for visualization. Ground-truth geometry appears only in the GT column.
Figure 6: Qualitative mesh comparisons on Replica. Rows show office_3 (top) and office_4 (bottom). Columns show ground truth, Random, FisherRF, GauSS-MI, GAVIS, MAGICIAN, and Matisse-G. All reconstruction methods use the Stream3D backend. Registered meshes of the selected objects are shown from a common viewpoint within each row.
Figure 7: Analysis of Matisse’s predicted uncertainty. (a) The pagoda region outside the input image’s field of view exhibits higher uncertainty than visible regions. (b) Uncertainty remains similar at rotationally corresponding locations on the bottle. The concentric rings in the polar plot indicate approximate invariance across yaw angles, consistent with the object’s rotational symmetry. (c) Higher mean uncertainty over occupied tokens is consistent with lower model confidence as Gaussian blur removes visual detail from the bottle image.
College of Intelligent Robotics and Advanced Manufacturing, Fudan University, Shanghai, China · School of Software, Shandong University, Jinan, China · School of Data Science and Engineering, East China Normal University, Shanghai, China