3D scene inpainting aims to recover missing or occluded regions in edited 3D scenes, while ensuring geometric and textural consistency. Existing approaches, however, typically require accurately calibrated camera poses, which restricts their applicability in casual, in-the-wild scenarios and introduces additional preprocessing overhead. To overcome this limitation, we present FreeInpaint, a novel feed-forward framework that generates complete and 3D-consistent scenes directly from unposed multi-view images with masked regions. At its core, FreeInpaint extends a 3D foundation model to propagate masked regions from a reference view to other unposed views, bridging 3D reconstruction and scene inpainting while preserving the model's native ability to recover camera poses and scene geometry. Our method addresses two key challenges in adapting feed-forward 3D foundation models to masked inputs. First, masked regions can corrupt cross-view correspondence reasoning, degrading pose estimation and geometry recovery. To address this, we introduce a Learnable Mask Attention mechanism that preserves the spatial anchoring of reliable observations while allowing masked regions to progressively absorb useful context in deeper layers. Second, under severe occlusions, a single forward pass often lacks sufficient appearance evidence for high-fidelity completion. Therefore, we propose a Support Token Refinement strategy, which injects diffusion-generated support evidence as confidence-weighted auxiliary tokens to refine under-observed regions while preserving the original spatial anchor. Extensive experiments across diverse datasets demonstrate that FreeInpaint achieves superior inpainting quality, eliminating the reliance on pre-computed camera poses while keeping a fast inference speed. The project page is https://rorisis.github.io/FreeInpaint/.
Figures & tables
Figure 1: FreeInpaint takes a single reference image and multiple unposed masked source images as inputs. It jointly estimates the camera parameters and reconstructs a high-fidelity, multi-view consistent 3D Gaussian field.
Figure 2: The pipeline of our FreeInpaint. Our method adapts DA3 for pose-free 3D inpainting by introducing learnable mask attention, which prevents masked regions from corrupting both pose and geometry estimation (Sec. 3.2 ). Additionally, we propose an inference-time support token refinement strategy to progressively inpaint content in occluded blind spots lacking visual cues (Sec. 3.3 ), while fine-tuning the DPT heads to adapt the model for inpainting tasks (Sec. 3.4 ).
Figure 3: Support Token Refinement. We dynamically detect unobserved blind spots and designate them as anchor views. A diffusion prior is then employed to rapidly inpaint the missing content, while spatial and texture priors ensure multi-view consistency during the generation process.
Methods
SPIn-NeRF [ 27 ]
360-USID [ 41 ]
LLFF [ 25 ]
Time ↓
PSNR ↑
LPIPS ↓
PSNR ↑
LPIPS ↓
C-KID ↓
C-FID ↓
SPIn-NeRF [ 27 ]
16.32
0.4122
17.59
0.4126
0.5978
388.29
∼ 5h
NeRFiller [ 40 ]
16.86
0.4183
15.03
0.5036
0.5835
360.02
∼ 30m
InFusion [ 21 ]
13.99
0.5216
15.62
0.3716
0.6652
428.18
∼ 30m
GScream [ 39 ]
16.85
0.3644
16.29
0.4415
0.6300
409.16
∼ 2h
AuraFusion360 [ 41 ]
14.51
0.7077
18.34
0.2982
0.6378
402.18
∼ 1h
Table 1: Quantitative comparisons with per-scene optimization methods on SPIn-NeRF [ 27 ] , 360-USID [ 41 ] , and LLFF [ 25 ] datasets. Inference time is measured on a single A6000 GPU. We report the cumulative time for multi-stage pipelines, while our method requires only ∼ 0.4s for a single forward pass. * For scenes with severe occlusions, a dynamic Support Token Refinement is triggered, slightly increasing the time to ∼ 1.8s (see Sec. 4.2), which remains orders of magnitude faster than optimization baselines.
Figure 4: Qualitative results with per-scene optimization methods on SPIn-NeRF [ 27 ] , 360-USID [ 41 ] , and LLFF [ 25 ] datasets. FreeInpaint successfully synthesizes high-fidelity and multi-view consistent novel views without requiring camera poses.
Methods
SPIn-NeRF [ 27 ]
360-USID [ 41 ]
LLFF [ 25 ]
PSNR ↑
LPIPS ↓
FID ↓
PSNR ↑
LPIPS ↓
FID ↓
C-KID ↓
C-FID ↓
LAMA [ 35 ] +DA3
16.67
0.4034
196.01
17.76
0.3211
186.83
0.5629
361.75
MVInpainter [ 5 ] +DA3
16.79
0.3733
164.22
17.24
0.3891
275.06
0.5724
372.76
Ours
17.79
0.2819
148.05
18.58
0.2526
199.61
0.5613
343.97
Table 2: Quantitative comparisons with feed-forward unposed methods on SPIn-NeRF [ 27 ] , 360-USID [ 41 ] , and LLFF [ 25 ] datasets.
Figure 5: Qualitative results on SPIn-NeRF [ 27 ] and 360-USID [ 41 ] datasets.
Table 8
Figure 6: Qualitative results of the object replacement task. We compare FreeInpaint with DiGA3D [ 29 ] and InFusion [ 21 ] methods.
Figure 7: Qualitative ablation study on the 360-USID dataset. Fine-tuning (+ Fine-Tuning) alone still yields blurriness. Integrating our learnable mask attention (+ Mask Attn.) significantly sharpens structural details by suppressing invalid features. The support token refinement (+ Token Refine.) further recovers high-fidelity textures in occluded regions.
Mask Attention
PSNR ↑
LPIPS ↓
ATE ↓
RPE t ↓
RPE r ↓
No mask
16.96
0.2854
0.0462
0.0803
0.7766
Hard mask
17.97
0.2581
0.0377
0.0764
0.7430
Learnable mask
18.58
0.2526
0.0340
0.0644
0.7069
Table 5: Effect of Learnable Mask Attention. We compare different mask-aware attention strategies under masked inputs.
Table 12
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Ablation on the Support View Generation. Our full method effectively reconstructs the occluded regions with high fidelity, preventing the geometric or textural inconsistencies seen in the ablated baselines.
Max Iters.
PSNR ↑
LPIPS ↓
Time ↓
1 (Initial pass)
18.02
0.2738
0.41
2
18.29
0.2703
1.12
3
18.46
0.2623
1.83
4
18.48
0.2620
2.54
Appendix
Table 8: Ablation studies on parameters of support token refinement. We evaluate the impact of different iteration numbers and visibility score thresholds ( τ ) on the 360-USID [ 41 ] and IMFine [ 34 ] datasets. Note that the reported metrics are averaged across both datasets. “Iters.” refers to the total number of feed-forward passes through our model (e.g., 3 Iters. equals 1 initial pass plus up to 2 refinement loops). The default settings used in our main experiments are highlighted in gray .
Figure 9: Visualization of Fine-Tuning Strategy on the SPIn-NeRF [ 27 ] dataset.
Figure 10: Visualization of Mask Generation Strategy.
Table 17
Mask Attention
PSNR ↑
LPIPS ↓
ATE ↓
RPE t ↓
RPE r ↓
MAT-style mask [ 16 ]
18.13
0.2620
0.0361
0.0743
0.7459
MLLAM-style mask [ 1 ]
18.06
0.2651
0.0374
0.0775
0.7468
Our Learnable mask
18.58
0.2526
0.0340
0.0644
0.7069
Appendix
Table 11: Effect of different mask-aware attention strategies under the same DA3 backbone and training protocol.
Total Input Views
Method
SPIn-NeRF
360-USID
PSNR ↑
LPIPS ↓
FID ↓
PSNR ↑
LPIPS ↓
FID ↓
4
InstaInpaint
16.89
0.2844
149.11
14.66
0.4934
259.39
FreeInpaint (Ours)
17.79
0.2819
148.05
17.19
0.3148
226.30
8
InstaInpaint
16.72
0.3013
144.31
15.60
0.4707
228.08
FreeInpaint (Ours)
18.32
0.2702
143.88
18.58
0.2526
199.61
16
InstaInpaint
17.04
0.3630
168.35
15.51
0.5012
265.28
Appendix
Table 12: View-count scalability comparison with InstaInpaint [ 46 ] on SPIn-NeRF [ 27 ] and 360-USID [ 41 ] . “Total Input Views” includes one reference view and the remaining masked source views.
Method
Pose Input
GS25 [ 13 ]
Co3D [ 32 ]
C-KID ↓
C-FID ↓
C-KID ↓
C-FID ↓
LaMa [ 35 ] +DA3
✗
0.3030
305.23
0.3870
360.72
MVInpainter [ 5 ] +DA3
✗
0.3624
356.07
0.3975
361.37
InstaInpaint [ 46 ]
✓
0.3406
330.84
0.3568
330.41
FreeInpaint (Ours)
✗
0.2889
292.28
0.3295
325.44
Appendix
Table 13: Quantitative evaluation on unseen in-the-wild scenes from GS25 [ 13 ] and Co3D [ 32 ] .
Figure 11: Additional Visualization on the NeRFiller [ 40 ] dataset. As illustrated, in extreme scenarios where highly sparse input views are combined with large-scale masks, the model struggles to retrieve sufficient contextual information. This occasionally degrades the inpainting quality and introduces some visual artifacts.
Figure 12: Additional Qualitative Results on the GS25 dataset [ 13 ] under unposed 8-view in-the-wild settings.
Figure 13: Additional qualitative results on view-count scalability.
Figure 14: Additional Qualitative Results on the SPIn-NeRF [ 27 ] , LLFF [ 25 ] , and IMFine [ 34 ] datasets.
Figure 15: Additional Qualitative Results on GS25 [ 13 ] and Co3D [ 32 ] datasets.