3D object removal aims to remove target objects from reconstructed scenes and complete the geometry and appearance of occluded regions. Existing NeRF- and 3DGS-based methods typically inpaint 2D images to guide 3D completion. However, complex multi-object layouts limit the surrounding context visible in each view, making 2D inpainting prone to artifacts. Inconsistent completions across views also introduce conflicting supervision and blurry reconstructions. We propose 3D Gaussian Multi-Object Removal via Texture-Space Inpainting (3DTexMOR). Our key idea is to perform inpainting in a unified texture space shared by all views. By combining complementary observations, this space provides richer context for recovering missing regions and promotes cross-view appearance consistency. We aggregate multi-view observations into texture maps, inpaint the missing regions, and reproject the completed maps into camera views to supervise Gaussian scene completion. To avoid the influence of view-dependent highlights and reflections, we decompose appearance and aggregate view-independent intrinsic attributes instead of RGB colors. We further introduce geometrically regularized Gaussian completion to constrain the geometry of the completed regions. Extensive experiments demonstrate visually plausible completions and state-of-the-art multi-object removal performance, improving PSNR by 5.8 dB and reducing LPIPS by at least 22% compared with existing methods.
Figures & tables
Figure 1: We propose 3DTexMOR, a novel framework for high-quality 3D multi-object removal. By performing inpainting in a unified texture space, 3DTexMOR recovers the regions occluded by the removed objects with consistent appearance and geometry. Compared with existing methods, 3DTexMOR produces more consistent completion scenes with fewer artifacts and less blurriness.
Figure 2: Overview of the 3DTexMOR framework. Our Gaussian scene representation incorporates object identity features and intrinsic appearance attributes. The texture-space inpainting module aggregates complementary multi-view intrinsic observations into a shared texture space, completes the masked regions, and reprojects the completed maps to provide shared supervision across camera views. The geometrically regularized Gaussian completion module lifts the completed textures into a 3D Gaussian representation on the estimated support plane and regularizes its geometry during optimization. The final scene combines the original Gaussians representing the unremoved scene content with the optimized completion Gaussians.
Figure 3: Weighted aggregation of multi-view observations into a unified texture space.
Syn-Diffuse dataset
Syn-Reflective dataset
Method
PSNR / mPSNR ↑
SSIM / mSSIM ↑
LPIPS / mLPIPS ↓
FID / mFID ↓
PSNR / mPSNR ↑
SSIM / mSSIM ↑
LPIPS / mLPIPS ↓
FID / mFID ↓
Ours
30.85 / 28.02
0.881 / 0.924
0.181 / 0.127
48.3 / 48.6
29.58 / 29.08
0.976 / 0.985
0.121 / 0.117
51.8 / 56.3
GOR-IS
25.04 / 20.94
0.878 / 0.910
0.232 / 0.179
199.0 / 142.9
28.90 / 26.81
0.973 / 0.979
0.118 / 0.109
76.8 / 87.9
GPGS
21.18 / 16.67
0.821 / 0.873
0.268 / 0.202
228.4 / 171.6
22.17 / 17.54
0.858 / 0.868
0.257 / 0.192
217.6 / 164.2
3DGIC
22.49 / 17.50
0.805 / 0.867
0.289 / 0.198
211.7 / 162.8
23.61 / 18.38
0.845 / 0.864
0.275 / 0.188
201.1 / 154.7
AuraFusion360
20.44 / 16.92
0.799 / 0.866
0.290 / 0.210
256.5 / 180.3
21.46 / 17.77
0.839 / 0.845
0.276 / 0.200
243.7 / 171.3
Table 1: Quantitative comparison on the Syn-Diffuse and the Syn-Reflective datasets. We report full-image metrics and mask-domain metrics. Higher PSNR/SSIM and lower LPIPS/FID are better. Red highlights the best result and yellow highlights the second-best result in each group.
Figure 4: Visual comparisons with prior works on real-world and synthetic scenes, showing 3DTexMOR produces more consistent completion, with fewer artifacts and less blurriness.
Figure 5: Ablation of texture-space inpainting (inpaint.). We replace texture-space inpainting with image-space inpainting and compare the inpainted results and the renderings. Image-space inpainting has difficulty recovering occluded regions from limited observations, leading to more artifacts.
Variant
PSNR/mPSNR ↑
SSIM/mSSIM ↑
LPIPS/mLPIPS ↓
FID/mFID ↓
image-space
23.49/19.17
0.823/0.882
0.266/0.205
185.1/124.8
texture-space
30.22/28.56
0.929/0.955
0.151/0.122
50.1/52.5
Table 2: Quantitative comparison between image-space and texture-space inpainting on the synthetic datasets. All metrics are averaged over both Syn-Diffuse and Syn-Reflective. The visual comparison is shown in Fig. 5 .
Figure 6: Ablation of intrinsic decomposition and geometric regularization. Starting from a base framework with only texture-space inpainting enabled, we gradually add intrinsic decomposition, the anchor loss, and the surface loss. Intrinsic decomposition provides view-independent attributes for reducing artifacts caused by glossy surfaces. The anchor and surface losses discourage Gaussian drift from the completed regions, encouraging plausible geometry.
Variant
PSNR/mPSNR ↑
SSIM/mSSIM ↑
LPIPS/mLPIPS ↓
FID/mFID ↓
Base framework
27.52/23.12
0.874/0.919
0.209/0.168
94.3/71.6
+ Intrinsic decomposition
28.15/27.68
0.902/0.925
0.172/0.142
64.7/60.0
+ Anchor loss
28.81/27.84
0.911/0.928
0.165/0.138
61.2/58.4
+ Surface loss
30.22/28.56
0.929/0.955
0.151/0.122
50.1/52.5
Table 3: Component ablation with texture-space inpainting enabled, corresponding to Fig. 6 . Components are added gradually to evaluate the effects of intrinsic decomposition and geometric regularization. Intrinsic decomposition reduces reflection artifacts on glossy surfaces, while the anchor and surface losses progressively improve the geometry of the completed regions.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure B.1: Visual comparison of multi-object removal.
Figure B.2: Texture-space and image-space inpainting for single-object and multi-object removal.
Recent advances in Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have made it standard practice to reconstruct 3D scenes from multi-view images. Removing objects from such 3D representations is a fundamental editing task that requires complete and seamless inpainting of occluded regions, ensuring consistency in geometry and appearance. Although existing methods have made notable progress in improving inpainting consistency, they often neglect global lighting effects, leading to physically implausible results. Moreover, these methods struggle with view-dependent non-Lambertian surfaces, where appearance varies across viewpoints, leading to unreliable inpainting. In this paper, we present 3D Gaussian Object Removal in the Intrinsic Space (GOR-IS), a novel framework for physically consistent and visually coherent 3D object removal. Our approach decomposes the scene into intrinsic components and explicitly models light transport to maintain global lighting effects consistency. Furthermore, we introduce an intrinsic-space inpainting module that operates directly in the material and lighting domains, effectively addressing the challenges posed by non-Lambertian surfaces. Extensive experiments on both synthetic and real-world datasets demonstrate that our framework substantially improves the physical consistency and visual coherence of object removal, outperforming existing methods by 13% in perceptual similarity (LPIPS) and 2dB in peak signal-to-noise ratio (PSNR). Code is publicly available at https://applezyh.github.io/GOR-IS-project-page/
Yonghao Zhao, Yupeng Gao, Jian Yang +2
VCIP, College of Computer Science, Nankai University · School of Intelligence Science and Technology, Nanjing University
Removing unwanted objects from reconstructed 3D scenes is an important task in computer vision, supporting applications in AR/VR, robotics, and digital content creation. Existing methods typically complete the entire masked region in a single step and without effectively utilizing semantic information from other views, leading to difficulties in handling complex geometric details and textures. In this work, we propose a novel framework that integrates Semantic-guided Block Matching (SBM) and Region-Wise Progressive Refinement (RPR) for high-quality 3D object removal. First, we leverage DINOv2 to encode semantic guidance from multi-view observations, and the best match tokens are decoded to complete missing regions in the target view while maintaining cross-view consistency. Second, we introduce a RPR strategy that segments the target mask into multiple subregions and selectively refines those with poor visual quality. Our method is built upon Gaussian Splatting, ensuring high-fidelity scene reconstruction with efficient computation. Experimental results demonstrate that our approach outperforms existing Gaussian-based methods in terms of perceptual quality and coherence in 3D object removal.
Xianliang Huang, Chen Xiao, Yuanxiang Ni +5
PICO, ByteDance Inc. · Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences · Fudan University
The tasks of object removal and inpainting 3D Gaussian Splatting (3DGS) scenes face challenges such as 3D consistency across camera views. In comparing 2D inpainters and their suitability for the 3D domain, we find that reconstruction-based inpainters outperform generative diffusion models in 3D consistency. Integrating these 2D inpainters into different single-step methods for creating and finetuning 3DGS scenes, our results indicate that initializing the scene from scratch produces higher quality results than finetuning the existing scene. Using a state-of-the-art generative 2D inpainter, we create a straightforward baseline to underline the importance of object removal before inpainting in the 3D setting. Since 360° datasets rarely include real-world ground truths, and challenging occlusion scenarios are equally sparse, we introduce a novel multi-object scene with recorded ground truth data and many views with object occlusions.
Finn Dröge, Cecilia Curreli, Abhishek Saroha +1
1Technical University of Munich · 2Munich Center for Machine Learning