Inpainting 3D Gaussian Splatting scenes, a key challenge in 3D editing, requires generating plausible content within a masked region of 3D space. Prior approaches rely on 2D diffusion models to produce one or several inpainted reference views, making them susceptible to challenges associated with multi-view inconsistency and lengthy optimization times. Departing from these approaches, we introduce a reference-free Gaussian splatting inpainting method operating natively in 3D. Our method leverages the embedding space of a large, pre-trained 3D prior, combined with a structure completion network to feed a generative prior which reconstructs the missing region's geometry and appearance. Performing content generation entirely in 3D, it avoids the need to reconcile inconsistencies of multiple inpainted reference images, and is faster than related 2D-based methods. We demonstrate the effectiveness of our method qualitatively and quantitatively, through extensive experiments and a user study. To the best of our knowledge, this work is the first Gaussian splatting inpainting method to operate in the learned representation space of a 3D-native generative prior without relying on inpainted reference views.
Figures & tables
Figure 1: Overview of Wings . Given the input 3DGS encoded as SLAT and a mask M , our structure completion network predicts masked-region occupancy. The SLAT s are inpainted directly in 3D, conditioned on the context latents and decoded to Gaussian splats.
Figure 2: Attention masking. The attention mask acts on the query-key dot product QK⊤ , effectively zeroing out target-to-target self-attention interactions. We note that, since the mask is not applied to the value tensor V , the output of self-attention is not null for target latents.
Method
Ours
Inpaint360GS
Gaussian Grouping
InFusion
Expected rank ↓
Ours
–
75.11
84.19
93.01
1.48 ± 0.07
Inpaint360GS
24.89
–
77.04
85.04
2.13 ± 0.07
Gaussian Grouping
15.81
22.96
–
67.17
2.94 ± 0.08
InFusion
6.99
14.96
32.83
–
3.45 ± 0.08
Table 1: User study. Each cell contains the tie-adjusted win-rate (higher is better) in %, row-vs-column. See App. A.2 for details.
Method
LPIPS ↓
FID ↓
Local-FID ↓
KID (×103)↓
Human Pref. (%)↑
Runtime (s) ↓
Gauss. Grouping
0.255
66.98
66.53
18.02
34.8
780.0
InFusion
0.250
88.71
87.74
40.31
18.6
212.9
Inpaint360GS
0.236
51.17
52.43
6.16
63.6
202.1
Ours
0.235
49.96
49.21
6.08
82.9
176.8
Table 2: Averaged metrics across datasets. Bold and underline indicate best- and second-best.
Figure 3: Qualitative comparison on the COR-NeRF, Inpaint360 & 360-USID datasets (top to bottom, two rows each, best viewed zoomed in). See the supplemental video to assess multiview-consistency.
Method
LPIPS ↓
FID ↓
Local-FID ↓
KID ( ×103 ) ↓
Ours
0.235
49.961
49.214
6.080
Ours (w/o attention masking)
0.236
52.806
51.718
9.497
Ours (w/o axis alignment)
0.241
55.336
54.661
10.030
Ours (w/o structure prediction model)
0.238
57.260
59.181
11.893
Ours (w/o decoder adaptation)
0.250
68.776
63.755
24.867
Table 3: Perceptual metrics for our method and its ablations.
Figure 4: Sparse structure reconstruction metrics against masked-to-context voxel ratio. Our proposed structure completion network strongly outperforms TRELLIS GS + RePaint for all context ratios.
Figure 5: Ablation visualization. We show the context and mask regions as overlays on the input image. Each column shows the completed structure (inset, top row), the PCA of the inpainted features after the integration process (principal three components in RGB, middle row), and the final decoded result in the bottom row.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Inpainting results on the “Red bull” scene from the Inpaint360 dataset. Even though Inpaint360GS achieves better (lower) FID and KID scores, our method produces a more coherent continuation of the striped surface, and is preferred by 91.7% of human evaluators.
Figure 7: Inpainting results on the “Car” scene from the Inpaint360 dataset. Similar to Figure 6 , Inpaint360GS performs better on the pixel-aligned metrics PSNR and SSIM, even though our method produces a more plausible and notably sharper inpainting result, making it the preferred choice for 100% of human evaluators.
Figure 8: Inpainting results on the “Plant” scene from the 360-USID dataset. Here, Ours performs better on the distributional FID and KID metrics, despite being preferred by only 37.5% of human evaluators compared to Inpaint360GS.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
Gaussian Grouping
20.58
0.617
0.255
InFusion
20.68
0.628
0.250
Inpaint360GS
21.94
0.644
0.236
Ours
21.95
0.640
0.235
Appendix
Table 4: Aggregate image-fidelity metrics, averaged across all datasets. Bold and underline indicate the best and second-best result in each column, respectively.
Stage
Ours
Ours (XS)
Context encoding
20.40
20.13
Decoder adaptation
103.02
23.58
Structure prediction
0.56
0.56
Conditional latent sampling
51.20
29.59
Decoding and merging
1.66
1.67
Total runtime (s) ↓
176.84
75.39
Appendix
Table 5: Per-stage runtime of our method.
Figure 9: Train and validation viewpoint PSNR during scene-conditioned decoder adaptation vs. number of training cameras.
Figure 10: Illustration of the generative prior’s bias towards generating axis-aligned patterns. (a) While self-attention masking neither degrades nor improves perceptual loss during inpainting for these examples, inpainting perceptually worsens between 0 and 90° rotations. (b) The flow velocity mean squared error also increases as the patterned SLAT grids are rotated, compared to an axis-aligned velocity at 0°.
Figure 11: Qualitative visualization of the effect of our axis-alignment heuristic on inpainting results. Without it, the generation of context-consistent intricate patterns, being biased by the flow velocity towards alignment, is degraded.
Figure 12: Quadric extrapolation against our learned structure prediction network for held-out test examples. Our model is able to successfully predict occupancy from complex context structure.
Figure 13: We ablate our proposed structure completion network with the TRELLIS backbone (see main text) and additionally show a comparison against a quadric-extrapolation baseline of the active ground truth voxels, which performs worse in all cases.
Single-reference
Distributional
Dataset
Scene
Human (%)
PSNR
SSIM
LPIPS
FID
Local-FID
KID
Local-KID
COR-NeRF
aircap
75.0
backpack
100.0
campingbox
88.9
dolls
66.7
gunnysack
100.0
Appendix
Table 6: Scene-level comparison of Ours and Inpaint360GS according to automated metrics and human preferences. Metric cells display within-metric normalized differences, where positive values take a green colour and favour Ours. The Human column reports the tie-adjusted preference for Ours. Across the 29 scenes, humans prefer Ours on 21, Inpaint360GS on 7 and rate the two as equal on one.
Reference-based
Distributional
Human
Scene
Method
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
Local-FID ↓
KID ↓
Human Pref. ↑
Aircap
Gaussian Grouping
20.35
0.599
0.259
44.04
41.31
4.26
48.3±10.7
InFusion
19.99
0.598
0.260
42.70
40.44
5.17
23.0±10.7
Inpaint360GS
21.41
0.612
0.251
37.65
34.47
2.28
63.4±11.2
Ours
21.13
0.617
0.247
36.55
32.36
2.12
67.8±13.9
Backpack
Gaussian Grouping
21.34
0.856
0.211
60.66
57.02
13.40
16.2±5.4
Appendix
Table 7: Per-scene metrics on COR-NeRF. KID is reported ×103 . Human Pref. is each method’s expected score against an average-skill opponent under the Davidson model, reported with a participant-bootstrap 95% confidence interval. Bold denotes the best method for each scene and metric.
Reference-based
Distributional
Human
Scene
Method
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
Local-FID ↓
KID ↓
Human Pref. ↑
Bag
Gaussian Grouping
24.84
0.835
0.148
35.46
35.38
−0.55
25.9±6.9
InFusion
25.16
0.842
0.157
60.34
70.92
15.04
10.7±4.3
Inpaint360GS
27.47
0.856
0.133
29.08
25.85
−3.55
75.0±7.8
Ours
27.20
0.848
0.136
30.28
27.20
−2.44
88.8±5.1
Car
Gaussian Grouping
18.91
0.759
0.184
127.19
122.27
66.05
29.9±5.7
Appendix
Table 8: Per-scene metrics on Inpaint360. KID is reported ×103 . Human Pref. is each method’s expected score against an average-skill opponent under the Davidson model, reported with a participant-bootstrap 95% confidence interval. Bold denotes the best method for each scene and metric.
Reference-based
Distributional
Human
Scene
Method
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
Local-FID ↓
KID ↓
Human Pref. ↑
Carton
Gaussian Grouping
20.46
0.459
0.324
44.84
68.15
15.06
36.1±10.5
InFusion
21.01
0.498
0.309
87.81
128.00
59.90
12.1±5.2
Inpaint360GS
21.38
0.499
0.302
37.63
63.23
11.52
72.0±10.3
Ours
21.44
0.495
0.298
34.51
53.13
8.04
83.7±8.2
Cone
Gaussian Grouping
16.49
0.183
0.319
55.10
67.60
2.71
77.4±10.0
Appendix
Table 9: Per-scene metrics on 360-USID. KID is reported ×103 . Human Pref. is each method’s expected score against an average-skill opponent under the Davidson model, reported with a participant-bootstrap 95% confidence interval. Bold denotes the best method for each scene and metric.
Figure 14: Qualitative comparison of our method and the implemented baselines on the COR-NeRF dataset. Our Wings framework provides sharp, consistent inpainting across the majority of the scenes in this dataset.
Figure 15: Qualitative comparison of our method and the implemented baselines on the Inpaint360 dataset.
Figure 16: Qualitative comparison of our method and the implemented baselines on the 360-USID dataset.
The tasks of object removal and inpainting 3D Gaussian Splatting (3DGS) scenes face challenges such as 3D consistency across camera views. In comparing 2D inpainters and their suitability for the 3D domain, we find that reconstruction-based inpainters outperform generative diffusion models in 3D consistency. Integrating these 2D inpainters into different single-step methods for creating and finetuning 3DGS scenes, our results indicate that initializing the scene from scratch produces higher quality results than finetuning the existing scene. Using a state-of-the-art generative 2D inpainter, we create a straightforward baseline to underline the importance of object removal before inpainting in the 3D setting. Since 360° datasets rarely include real-world ground truths, and challenging occlusion scenarios are equally sparse, we introduce a novel multi-object scene with recorded ground truth data and many views with object occlusions.
Finn Dröge, Cecilia Curreli, Abhishek Saroha +1
1Technical University of Munich · 2Munich Center for Machine Learning
Recent advances in 3D scene editing have leveraged iterative diffusion models to update input views. However, this process is computationally expensive and struggles to produce sharp details. Meanwhile, ``hallucination drift'' frequently introduces multi-view inconsistencies, leading to structural artifacts when rendering novel viewpoints. To address this problem, we present 3D-GIMP (3D Gaussian Inpainting Meets Patch Matching), a novel hybrid paradigm designed for high-fidelity object removal in 3D Gaussian Splatting. Instead of diffusing every view, 3D-GIMP performs a single generative inpainting on a key reference view, which serves as an appearance prior. We then introduce a 3D-aware PatchMatch algorithm to propagate these reference textures across all remaining views via correspondence matching, effectively bypassing the stochastic nature of frame-by-frame diffusion. By prioritizing reconstructive consistency over iterative generation, 3D-GIMP maintains high-frequency details across arbitrary resolutions while ensuring a mathematically consistent 3D reconstruction. Our experiments demonstrate that 3D-GIMP not only achieves competitive inpainting quality as previous methods using diffusion in multiple views, but also outperforms these methods in rendering speed and view consistency.
3D scene inpainting is essential for reconstructing areas corrupted by occlusions or limited viewpoints. While recent methods leverage Gaussian Splatting (GS) for efficient 3D editing, they often depend on precise multi-view segmentation masks and are inherently constrained to object removal tasks. We propose CoIn, a novel framework that bridges 2D inpainting models and 3DGS through a multi-stage consistency pipeline. Our approach first generates initial inpainted images using a diffusion model, enabling the use of arbitrary-shaped masks and diverse tasks like object insertion. We then introduce Reference Adaptive GS with Feature Attention to reconstruct a coarse 3D scene by adaptively weighing towards a reference view (2D -> 3D). This 3D representation provides geometric guidance to the diffusion process via GS-based Reference Feature Warping, ensuring multi-view consistency (3D -> 2D). Finally, a Texture-Enhancing Discriminator refines the 3D scene to achieve high photometric realism (2D -> 3D). Experiments show that CoIn, effectively leveraging bidirectional information flow, achieves state-of-the-art performance and effectively handles both object removal and object insertion with flexible mask input.
Hana Kim, Minje Kim, Tae-Kyun Kim
LG Electronics, Seoul, South Korea · KAIST, Daejeon, South Korea