3D stylization is central to game development, virtual reality, and digital arts, where the demand for diverse assets calls for scalable methods that support fast, high-fidelity manipulation. Existing text-to-3D stylization methods typically distill from 2D image editors, requiring time-intensive per-asset optimization and exhibiting multi-view inconsistency due to the limitations of current text-to-image models, which makes them impractical for large-scale production. In this paper, we introduce GaussianBlender, a pioneering feed-forward framework for text-driven 3D stylization that performs edits instantly at inference. Our method learns structured, disentangled latent spaces with controlled information sharing for geometry and appearance from spatially-grouped 3D Gaussians. A latent diffusion model then applies text-conditioned edits on these learned representations. Comprehensive evaluations show that GaussianBlender not only delivers instant, high-fidelity, geometry-preserving, multi-view consistent stylization, but also surpasses methods that require per-instance test-time optimization - unlocking practical, democratized 3D stylization at scale.
Figures & tables
Figure 1 : Given a 3D Gaussian splat asset and an edit prompt, InStyle - a diffusion-based feed-forward style editor - delivers geometry-preserving, multi-view-consistent appearance stylization with a single-pass inference (0.26 s per asset), without per-asset test-time optimization. InStyle enables (a) library-scale stylization across diverse assets and (b) interactive, user-driven customization, (c) while generalizing to real-world captures - unlocking practical, democratized 3D stylization at scale.
Figure 2 : Overview of our method. (1) Latent space learning: Given input Gaussians, our method first groups them based on spatial proximity and encodes into group-structured disentangled latent spaces, with controlled cross-branch feature-sharing. (2) Latent diffusion pre-training: A denoiser dϱ then learns to denoise the noisy appearance latent zcst conditioned on text embedding C . (3) Latent editing: Once 3D priors are captured, dϱ is further trained to learn an editing function g(⋅) that maps latent zcs to a modified latent zce based on text embedding Ce , guided by the geometry latent zgs . At inference , InStyle generates modified high-quality, 3D-consistent assets from text prompts in a single feed-forward pass instantly, fully eliminating test-time optimization. Trainable models at each stage are denoted.
Figure 3 : Qualitative comparison with state-of-the-art methods. Baselines often produce over-saturated, dramatic stylizations that distort 3D structure ( e.g . , blurred boundaries; rows 2,3,5, col. 4), minimal edits that are barely perceptible (row 2, col. 2,3), or severe geometric artifacts. In contrast, InStyle delivers high-fidelity, text-aligned 3D stylizations with strong geometry preservation in a single feed-forward pass.
Figure 4 : Cross-dataset generalization on OmniObject3D [ 74 ] . Our framework demonstrates strong style editing performance on out-of-distribution 3D assets.
Figure 5 : Generalization to complex real-world scenes from SceneSplat [ 36 ] . Using a simple extract-edit-reinsert pipeline, InStyle produces coherent stylizations despite capture artifacts, clutter, and occlusions.
Method
Text Alignment ↑
Structural Consistency ↑
Visual Quality ↑
IN2N
16.30%
20.24%
21.68%
IGS2GS
14.67%
17.48%
20.21%
GaussCtrl
20.10%
15.84%
17.48%
GaussianEditor
15.21%
4.37%
6.01%
Ours
33.69%
42.05%
34.60%
Table 2 : User study results. InStyle is consistently preferred over the baselines across the three criteria.
LightGaussian
Input 3DGS
w/o disent.
w/ disent. (Ours)
PSNR ↑
35.2117
32.7136
33.6635
34.3299
Table 3 : Effect of disentanglement on 3DGS parameter reconstruction. LightGaussian: original LightGaussian [ 17 ] reconstruction; Input 3DGS : downsampled LightGaussian 3DGS used as our model input; w/o disent. : single-branch variant; w/ disent. : dual-branch Gaussian VAE.
Figure 6 : Visual assessment of Gaussian VAE reconstructions. The proposed dual-branch Gaussian VAE yields sharper reconstructions with better geometric fidelity and color accuracy.
CLIPsim↑
CLIPdir↑
(1) w/o LLD
0.227
0.183
(2) w/o feat. sharing
0.221
0.171
(3) Ours (full)
0.260
0.223
Table 4 : Assessment of framework components. We report CLIPsim↑ and CLIPdir↑ to compare different variants.
Figure 7 : Effect of style catalog size. InStyle maintains stylization performance as the style catalog size grows from 1 to 30.
Figure 8 : Application. (a) Interactive scene editing: InStyle can be adopted for interactive object-level scene editing. (b) Style transfer via appearance-latent exchange: Transferring the appearance code from a source to a target asset modifies appearance while preserving the geometric structure.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Evaluation Benchmark
“Make its colours look like rainbow”, “Make it Starry Night Van Gogh painting style”, “Make it wooden”, “Make it in cyberpunk style”, “Make it pop art neon duotone”, “Make it marble”, “Make it Barbie style”, “Make it golden”, “Make it botanical themed”, “Make it Halloween themed”
Training (single-prompt)
Each prompt in the evaluation benchmark is used to train a separate model.
Training (multi-prompt)
“Make its colours look like rainbow”, “Make it Starry Night Van Gogh painting style”, “Make it wooden”, “Make it in cyberpunk style”, “Make it pop art neon duotone”, “Make it marble”, “Make it Barbie style”, “Make it golden”, “Make it botanical themed”, “Make it Halloween themed”, “Make it Hello Kitty themed”, “Make it look like made of steel”, “Make it The Scream Edvard Munch painting”, “Make it in sunset palette”, “Make it Pirates of the Caribbean style”, “Make it look like a statue”, “Make it Garfield themed”, “Make it look like made of bronze”, “Make it Dalmatian style”, "Make it gothic style", "Make it ocean palette", "Make it graffiti street art style", “Make it watercolor painting style”, "Turn it into red", “Turn it into orange”, "Turn it into pink", "Turn it into blue", "Turn it into yellow", "Turn it into green", "Turn it into black"
Appendix
Table 5 : Edit prompts used in the evaluation benchmark and for training the single-prompt and multi-prompt models.
Criterion
Text Alignment
Which stylized result follows the prompt most closely? Consider whether the style change is too subtle (almost no visible change), or too drastic (over-saturated).
Structural Consistency
Which edited result best preserves the original asset’s 3D structure while applying only the requested appearance stylization? The object’s shape should not change after stylization (e.g., no changes to the object’s physical structure; for characters, the face, identity and features, must remain unchanged).
Visual Quality
Which edited result looks best overall?
Appendix
Table 6 : The criteria used in the user study. The definitions listed here are also provided to participants.
Figure 9 : Visual results of expanding the style catalog. InStyle maintains high-quality 3D stylization and geometry preservation as the number of training prompts increases.
Figure 10 : Additional visual results of our method. InStyle delivers instant, geometry-preserving, and multi-view-consistent 3D stylizations in a single feed-forward pass.
Figure 11 : Additional visual results of our method. InStyle delivers instant, geometry-preserving, and multi-view-consistent 3D stylizations in a single feed-forward pass.
CLIPsim↑
CLIPdir↑
Structure Dist. ↓
MV-L1
0.251
0.210
0.0064
DMD + MV-L1 (Ours)
0.260
0.223
0.0050
Appendix
Table 7 : Effect of DMD over MV-L1 Supervision. DMD improves metrics over MV-L1 supervision.
Figure 12 : Effect of DMD over MV-L1 Supervision. Compared with the MV-L1 baseline, adding the DMD objective yields sharper, more vivid stylization results.
CLIPsim↑
CLIPdir↑
Structure Dist. ↓
(1) sT∈[5.5,9.5]
0.260
0.223
0.0050
(2) sT∈[9.5,12.5]
0.263
0.226
0.0089
(3) sT∈[12.5,15.5]
0.268
0.232
0.0172
Appendix
Table 8 : Effect of guidance scale. Higher text guidance scales yield improved CLIP metrics with minimal loss in structural consistency.
Figure 13 : Latent perturbation and appearance swapping. Perturbations affect the corresponding appearance or geometry factor, while appearance swaps transfer appearance while preserving geometry.
Figure 14 : Failure cases. Our model struggles to represent complex scenes with fine-grained details using 16384 Gaussians, leading to lower-quality results. This can be addressed by increasing the number of Gaussians to better capture such details.
The growing demand for rapid and scalable 3D asset creation has driven interest in feed-forward 3D reconstruction methods, with 3D Gaussian Splatting (3DGS) emerging as an effective scene representation. While recent approaches have demonstrated pose-free reconstruction from unposed image collections, integrating stylization or appearance control into such pipelines remains underexplored. Existing attempts largely rely on image-based conditioning, which limits both controllability and flexibility. In this work, we introduce AnyStyle, a feed-forward 3D reconstruction and stylization framework that enables pose-free, zero-shot stylization through multimodal conditioning. Our method supports both textual and visual style inputs, allowing users to control the scene appearance using natural language descriptions or reference images. We propose a modular stylization architecture that requires only minimal architectural modifications and can be integrated into existing feed-forward 3D reconstruction backbones. Experiments demonstrate that AnyStyle improves style controllability over prior feed-forward stylization methods while preserving high-quality geometric reconstruction. A user study further confirms that AnyStyle achieves superior stylization quality compared to an existing state-of-the-art approach. Repository: https://github.com/joaxkal/AnyStyle.
Joanna Kaleta, Bartosz Świrta, Kacper Kania +3
Warsaw University of Technology · Sano Centre for Computational Medicine · IDEAS NCBR +2
With the growth of gaming, animation, and virtual reality industries, the demand for efficient generation of stylized 3D assets is rapidly increasing. However, existing approaches still struggle to jointly preserve style fidelity, geometric consistency, and generation efficiency, as most of them still rely on indirect 2D-to-3D stylization pipelines. This motivates a native 3D stylization framework that can explicitly disentangle style from geometry while remaining efficient. To this end, we propose DreamStyle3D, an efficient framework for stylized 3D asset generation built on a Decoupled Dual Cross-Attention mechanism. Our method explicitly separates geometric and stylistic features to enable efficient style injection while preserving structural consistency, and further adopts a lightweight training strategy to enhance style consistency and model generalization. In addition, we build an automated data pipeline and construct a dataset of about 15K content-style-stylized triplets for training and evaluation. Extensive experiments demonstrate that our DreamStyle3D can generate high-fidelity, geometrically consistent stylized 3D assets within 10 seconds, substantially improving efficiency while maintaining superior style quality and offering a new solution for 3D content creation. The project is available at https://github.com/NK-JittorCV/nk-3D/tree/main/models/DreamStyle3D.
Kai Wang, Ziheng Ouyang, Xuying Zhang +2
VCIP, Nankai University Tianjin, China · Zhongguancun Academy Beijing, China · Joy Future Academy, JD Beijing, China +1
Text-guided style editing of 3D assets is essential for adapting existing objects to diverse visual aesthetics in digital content creation. Despite rapid progress in 3D shape modeling, faithfully stylizing an existing asset remains challenging when the desired stylization involves fine-grained structural ornamentation, which requires the model to preserve the source geometry and object identity, while coherently integrating new style-specific details. We propose \textbf{OrnaStyler}, a zero-shot framework for text-guided ornament-aware 3D stylization. Built upon rectified flow-based generative modeling, OrnaStyler introduces an inversion-guided editing strategy that recovers content-aware latent representations at both geometry and appearance levels in a staged manner to facilitate faithful editing. Our core idea is to explicitly model the spatial configuration of stylistic elements, thereby mitigating the fundamental tension between content preservation and style expression in the voxel space. Specifically, at the geometry level, we manipulate voxel representations through flow inversion to synthesize ornament-enhanced structures while preserving the spatial identity of the source asset. Then, at the appearance level, we introduce an adjacency-aware feature inpainting mechanism to harmonize newly generated ornaments with the original content, yielding coherent geometry-appearance integration. Our approach operates solely in the inference phase and enables selective editing over geometric augmentation or appearance stylization. Extensive experiments on both generated and real-world 3D assets against prior methods demonstrate that OrnaStyler achieves state-of-the-art editing performance in terms of content preservation, style fidelity, and overall visual realism. Code is available at: https://github.com/tomohiro0427/OrnaStyler
Tomohiro Aizawa, Shigeru Kuriyama, Chunzhi Gu
University of Fukui, Japan · Toyohashi University of Technology, Japan · AI Lab, CyberAgent, Inc., Japan