Text-driven 3D Gaussian editing commonly does not distinguish the editing reliability of rendered views, although different viewpoints provide supervision of substantially different quality. Views that clearly show the scene and match the edit instruction provide reliable guidance, while less informative views may weaken the edit when all views are treated equally. We present View Matters, a view-importance-aware framework that conducts editing around reliable keyframes. Keyframe Importance Estimation (KIE) identifies reliable views using geometric visibility, semantic distinctiveness, and edit relevance. Keyframe-Guided Editing (KGE) then propagates their editing signals asymmetrically to non-keyframes without noisy reverse influence, while Importance-Aware Optimization (IAO) preserves this reliability preference during 3DGS optimization. Across 23 scene-prompt pairs, View Matters achieves the highest average CLIP text-image similarity of 0.2822 and directional similarity of 0.2564 among the evaluated methods, with a four-minute editing time. Additional adjacent-view analysis indicates that the fidelity-oriented editing process maintains cross-view coherence.
Figures & tables
Figure 1: Overview of View Matters. Given a 3DGS scene and an editing instruction, KIE estimates view reliability and selects representative keyframes. KGE lets these keyframes establish a shared context and asymmetrically guide less-reliable non-keyframes without reverse interference. IAO then preserves the reliability hierarchy by assigning stronger supervision to keyframes during 3DGS optimization. The resulting scene follows the requested edit faithfully while maintaining cross-view coherence.
Figure 2: Illustration of the proposed keyframe-guided attention mechanism. Our method modifies the attention computation inside a frozen IP2P denoising network without introducing additional trainable parameters. Keyframes first perform global interaction to establish a shared editing context. Each non-keyframe then attends to its associated keyframe together with the immediately preceding frame, yielding an asymmetrical propagation scheme that uses keyframes as the dominant guidance source while maintaining local continuity across views.
Figure 3: The two examples span complementary editing regimes: a localized edit of a human subject (top) and a global transformation of an outdoor scene (bottom). Our method more faithfully follows both instructions while preserving editing-irrelevant content and the underlying scene structure. Additional comparisons across scenes and edit instructions are provided in Appendix C .
Method
T–I ↑
Direction ↑
Time ↓
IN2N
0.2599
0.1767
28 min
VICA
0.2439
0.1565
25 min
GaussianEditor
0.2769
0.2227
8 min
GaussCtrl
0.2589
0.1572
10 min
EditSplat
0.2719
0.2152
7 min
DGE
0.2753
0.2179
5 min
Table 1: Quantitative comparison with representative text-driven 3D editing methods. Best results are in bold, and second-best results are underlined.
Figure 4: With the same candidate views and keyframe budget, random and importance-based selection yield different keyframes (top) and editing results (bottom). Our selection provides more reliable guidance and achieves a more faithful target transformation, highlighting the importance of view selection in multi-view editing.
Figure 5: Sensitivity of view-importance weights. The three cues are complementary, with Geo-dom. (0.6,0.2,0.2) adopted as the default for its best overall fidelity.
KIE
KGE
IAO
T-I ↑
Direction ↑
0.2702
0.2229
✓
0.2754
0.2464
✓
✓
0.2761
0.2486
✓
✓
✓
0.2822
0.2564
Table 2: Ablation study on View Matters. ✓ denotes the inclusion of each module. The full model consistently achieves the best results.
Figure 6: Representative adjacent-view consistency analysis on two ordered rendering sequences.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Number of keyframes
T-I ↑
Dir ↑
Training Time (s)
Max Alloc. (GB)
Max Res. (GB)
2
0.2763
0.2781
263.46
17.78
25.38
4
0.2793
0.2840
266.04
17.78
25.37
5
0.2786
0.2821
268.34
17.78
25.39
10
0.2780
0.2819
324.58
24.54
37.25
Appendix
Table 3: Effectiveness and resource consumption under different keyframe settings.
Figure 7: Sensitivity to the keyframe-loss weight λk . Relative changes in CLIP T–I and CLIP Directional Similarity are reported against λk=1 for the coarse and fine-grained sweeps.
Figure 8: Qualitative ablation results. The progressive visual improvements illustrate the complementary effects of KGE, KIE, and IAO. Red and blue boxes highlight different local regions for comparison.
Figure 9: Qualitative comparison with C3Editor and 3D-consistent
Figure 10: Qualitative results across subjects and edit scopes. The examples cover human subjects, complete scenes, and non-human objects, including both localized edits and global transformations. For each instruction, two rendered viewpoints show that our method follows the target semantics while preserving editing-irrelevant content and the underlying scene structure.
Figure 11: Additional qualitative results across scenes and editing instructions. Each group presents the source scene and edited renderings from multiple viewpoints. Our method handles diverse identity, material, appearance, style, and color transformations while maintaining stable edited content across views.
Scene
Edit Instruction
Target Prompt
Source Prompt
Face
“Turn him into the Tolkien Elf.”
“A Tolkien Elf man with curly hair.”
“A man with curly hair in a grey jacket.”
Face
“Turn him into an Einstein.”
“Einstein with curly hair in a grey jacket.”
“A man with curly hair in a grey jacket.”
Face
“Turn his face into a skull.”
“A skull with curly hair in a grey jacket.”
“A man with curly hair in a grey jacket.”
Face
“Turn him into spiderman with a mask.”
“A spider man with a mask and curly hair.”
“A man with curly hair in a grey jacket.”
Person
“Turn the man into a clown.”
“A clown standing next to a wall wearing a blue T-shirt.”
“A man standing next to a wall wearing a blue T-shirt.”
Person
“Turn him into a Super Mario.”
“A photo of a Super Mario.”
“A photo of a person.”
Appendix
Table 4: Detailed evaluation dataset: 23 scene–prompt pairs. The source and target prompts are used for computing CLIP and CLIP directional scores.
Text-driven 3D scene editing with 3D Gaussian Splatting (3DGS) typically applies a 2D diffusion editor to views rendered from fixed training cameras, limiting both the spatial coverage of edits and the user's freedom to target specific objects in complex scenes. We present LB-Edit, a framework that addresses two coupled problems: where to place editing cameras for localized edits, and how to make per-view edits agree with one another so that the 3D scene remains consistent after fine-tuning. First, Attention-Guided Editing Camera Placement (ACP) probes the diffusion model's self- and cross-attention at multiple candidate camera distances to find where attention is well-contained in the region of interest, then places a compact, geometrically diverse editing camera set at that attention-optimal distance. Second, Multi-view Attention Alignment (MAA) steers the editor toward the same edit across views along two axes: it aligns appearance by sharing self-attention features via token-level correspondence, and aligns spatial location by lifting cross-attention maps onto the 3D Gaussians as a shared 3D attention field, suppressing both appearance and spatial drift. Experiments on multi-object and single-object scenes show that our method achieves the highest user preference in instruction fidelity, multi-view consistency, and editing locality, using as few as 5 editing views and reducing latency by up to 7x over existing methods.
Recent advancements in diffusion and flow models have greatly improved text-based image editing, yet methods that edit images independently often produce geometrically and photometrically inconsistent results across different views of the same scene. Such inconsistencies are particularly problematic for editing of 3D representations such as NeRFs or Gaussian splat models. We propose a training-free guidance framework that enforces multi-view consistency during the image editing process. The key idea is that corresponding points should look similar after editing. To achieve this, we introduce a consistency loss that guides the denoising process toward coherent edits. The framework is flexible and can be combined with widely varying image editing methods, supporting both dense and sparse multi-view editing setups. Experimental results show that our approach significantly improves 3D consistency compared to existing multi-view editing methods. We also show that this increased consistency enables high-quality Gaussian splat editing with sharp details and strong fidelity to user-specified text prompts. Please refer to our project page for video results: https://3d-consistent-editing.github.io/
Josef Bengtson, David Nilsson, Dong In Lee +2
Chalmers University of Technology · Korea University
Recent advances in text-guided image editing and 3D Gaussian Splatting (3DGS) have enabled high-quality 3D scene manipulation. However, existing pipelines rely on iterative edit-and-fit optimization at test time, alternating between 2D diffusion editing and 3D reconstruction. This process is computationally expensive, scene-specific, and prone to cross-view inconsistencies. We propose a feed-forward framework for cross-view consistent 3D scene editing from sparse views. Instead of enforcing consistency through iterative 3D refinement, we introduce a cross-view regularization scheme in the image domain during training. By jointly supervising multi-view edits with geometric alignment constraints, our model produces view-consistent results without per-scene optimization at inference. The edited views are then lifted into 3D via a feedforward 3DGS model, yielding a coherent 3DGS representation in a single forward pass. Experiments demonstrate competitive editing fidelity and substantially improved cross-view consistency compared to optimization-based methods, while reducing inference time by orders of magnitude.