Text-driven 3D Gaussian editing commonly does not distinguish the editing reliability of rendered views, although different viewpoints provide supervision of substantially different quality. Views that clearly show the scene and match the edit instruction provide reliable guidance, while less informative views may weaken the edit when all views are treated equally. We present View Matters, a view-importance-aware framework that conducts editing around reliable keyframes. Keyframe Importance Estimation (KIE) identifies reliable views using geometric visibility, semantic distinctiveness, and edit relevance. Keyframe-Guided Editing (KGE) then propagates their editing signals asymmetrically to non-keyframes without noisy reverse influence, while Importance-Aware Optimization (IAO) preserves this reliability preference during 3DGS optimization. Across 23 scene-prompt pairs, View Matters achieves the highest average CLIP text-image similarity of 0.2822 and directional similarity of 0.2564 among the evaluated methods, with a four-minute editing time. Additional adjacent-view analysis indicates that the fidelity-oriented editing process maintains cross-view coherence.
Figures & tables
Figure 1: Overview of View Matters. Given a 3DGS scene and an editing instruction, KIE estimates view reliability and selects representative keyframes. KGE lets these keyframes establish a shared context and asymmetrically guide less-reliable non-keyframes without reverse interference. IAO then preserves the reliability hierarchy by assigning stronger supervision to keyframes during 3DGS optimization. The resulting scene follows the requested edit faithfully while maintaining cross-view coherence.
Figure 2: Illustration of the proposed keyframe-guided attention mechanism. Our method modifies the attention computation inside a frozen IP2P denoising network without introducing additional trainable parameters. Keyframes first perform global interaction to establish a shared editing context. Each non-keyframe then attends to its associated keyframe together with the immediately preceding frame, yielding an asymmetrical propagation scheme that uses keyframes as the dominant guidance source while maintaining local continuity across views.
Figure 3: The two examples span complementary editing regimes: a localized edit of a human subject (top) and a global transformation of an outdoor scene (bottom). Our method more faithfully follows both instructions while preserving editing-irrelevant content and the underlying scene structure. Additional comparisons across scenes and edit instructions are provided in Appendix C .
Method
T–I ↑
Direction ↑
Time ↓
IN2N
0.2599
0.1767
28 min
VICA
0.2439
0.1565
25 min
GaussianEditor
0.2769
0.2227
8 min
GaussCtrl
0.2589
0.1572
10 min
EditSplat
0.2719
0.2152
7 min
DGE
0.2753
0.2179
5 min
Table 1: Quantitative comparison with representative text-driven 3D editing methods. Best results are in bold, and second-best results are underlined.
Figure 4: With the same candidate views and keyframe budget, random and importance-based selection yield different keyframes (top) and editing results (bottom). Our selection provides more reliable guidance and achieves a more faithful target transformation, highlighting the importance of view selection in multi-view editing.
Figure 5: Sensitivity of view-importance weights. The three cues are complementary, with Geo-dom. (0.6,0.2,0.2) adopted as the default for its best overall fidelity.
KIE
KGE
IAO
T-I ↑
Direction ↑
0.2702
0.2229
✓
0.2754
0.2464
✓
✓
0.2761
0.2486
✓
✓
✓
0.2822
0.2564
Table 2: Ablation study on View Matters. ✓ denotes the inclusion of each module. The full model consistently achieves the best results.
Figure 6: Representative adjacent-view consistency analysis on two ordered rendering sequences.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Number of keyframes
T-I ↑
Dir ↑
Training Time (s)
Max Alloc. (GB)
Max Res. (GB)
2
0.2763
0.2781
263.46
17.78
25.38
4
0.2793
0.2840
266.04
17.78
25.37
5
0.2786
0.2821
268.34
17.78
25.39
10
0.2780
0.2819
324.58
24.54
37.25
Appendix
Table 3: Effectiveness and resource consumption under different keyframe settings.
Figure 7: Sensitivity to the keyframe-loss weight λk . Relative changes in CLIP T–I and CLIP Directional Similarity are reported against λk=1 for the coarse and fine-grained sweeps.
Figure 8: Qualitative ablation results. The progressive visual improvements illustrate the complementary effects of KGE, KIE, and IAO. Red and blue boxes highlight different local regions for comparison.
Figure 9: Qualitative comparison with C3Editor and 3D-consistent
Figure 10: Qualitative results across subjects and edit scopes. The examples cover human subjects, complete scenes, and non-human objects, including both localized edits and global transformations. For each instruction, two rendered viewpoints show that our method follows the target semantics while preserving editing-irrelevant content and the underlying scene structure.
Figure 11: Additional qualitative results across scenes and editing instructions. Each group presents the source scene and edited renderings from multiple viewpoints. Our method handles diverse identity, material, appearance, style, and color transformations while maintaining stable edited content across views.
Scene
Edit Instruction
Target Prompt
Source Prompt
Face
“Turn him into the Tolkien Elf.”
“A Tolkien Elf man with curly hair.”
“A man with curly hair in a grey jacket.”
Face
“Turn him into an Einstein.”
“Einstein with curly hair in a grey jacket.”
“A man with curly hair in a grey jacket.”
Face
“Turn his face into a skull.”
“A skull with curly hair in a grey jacket.”
“A man with curly hair in a grey jacket.”
Face
“Turn him into spiderman with a mask.”
“A spider man with a mask and curly hair.”
“A man with curly hair in a grey jacket.”
Person
“Turn the man into a clown.”
“A clown standing next to a wall wearing a blue T-shirt.”
“A man standing next to a wall wearing a blue T-shirt.”
Person
“Turn him into a Super Mario.”
“A photo of a Super Mario.”
“A photo of a person.”
Appendix
Table 4: Detailed evaluation dataset: 23 scene–prompt pairs. The source and target prompts are used for computing CLIP and CLIP directional scores.