We introduce StyleFields, a DeepSDF-based architecture for high-fidelity 3D reconstruction that enables controllable geometric style mixing: the coarse structure of one object can be combined with the fine-scale details of another. The core idea is depth-aware modulation: instead of a single global code, we inject latents via multi-level Adaptive Instance Normalization at several decoder depths, and supervise matching auxiliary heads with a coarse-to-fine schedule while gradually growing network depth. This aligns early layers with global shape and later layers with high-frequency detail, achieving content-style decoupling without part labels or adversarial training. StyleFields delivers faithful reconstructions, convincing cross-instance hybrids, and consistent gains in ablations over injection depth and supervision granularity. We further demonstrate a practical application in automotive aerodynamics: a learned surrogate drag predictor serves as a differentiable objective to optimize reconstructed cars, allowing targeted edits of global form or surface details by freezing the complementary latent stream. StyleFields offers a simple, effective recipe for controllable implicit reconstruction and downstream performance-driven design.
Figures & tables
Figure 1 : StyleFields architecture with depth-aware AdaIN and progressive decoder growth. Fading colors indicate the curriculum on depth. Earlier ResMLP blocks model coarse structure, while later (faded) ones are gradually activated for fine detail.
Figure 2 : Qualitative comparison of 3D reconstruction on Cars , Chairs , and Planes . Rows show two shape instances per category; columns show methods.
Figure 3 : Global latent interpolation between two shapes across different 3D reconstruction baselines and StyleFields .
Figure 4 : StyleFields cross-instance style mixing between two car shapes (source on the left, target on the right). Each column shows a hybrid produced by selectively substituting style codes of Shape 1 with codes from Shape 2. (A) Coarse styles (w0+,w1+) from Shape 1 and fine styles (w2+,w3+) from Shape 2. (B) Mid-level styles (w1+,w2+) from Shape 1 and coarsest and finest styles (w0+,w3+) from Shape 2. (C) Fine styles (w2+,w3+) from Shape 1 and coarse styles (w0+,w1+) from Shape 2.
Figure 5 : Drag Optimization. Optimizing the global latent z induces global shape drift and often converges to similar solutions across different initial shapes. In contrast, optimizing per-scale style codes w+ enables targeted edits at specific geometric scales (coarse vs. fine). Surface color represents the predicted air pressure (blue to red).
Cars
Chairs
Planes
Method
CD ( ↓ )
IoU (%, ↑ )
CD ( ↓ )
IoU (%, ↑ )
CD ( ↓ )
IoU (%, ↑ )
DeepSDF [ 22 ]
1.6×10−4
97.27
8.8×10−4
89.74
8.0×10−4
89.28
Curriculum DeepSDF [ 10 ]
1.6×10−4
97.19
1.2×10−3
89.33
9.2×10−4
90.57
Curriculum + Style DeepSDF
1.3×10−4
97.87
5.7×10−4
90.12
7.8×10−4
91.26
DECOLLAGE [ 2 ]
1.6×10−3
78.97
3.0×10−3
49.96
3.3×10−3
39.12
DAE-NET [ 5 ]
2.8×10−3
81.91
1.0×10−2
56.37
−
−
Table 1 : Quantitative reconstruction on Cars , Chairs , and Planes . Lower CD is better; higher IoU is better. All methods share splits, sampling, and extraction.
Table 2 : Drag values for two cars under early and late W+ style optimization.
Figure 6 : Drag-aware edits for four shapes. In each visualization, red regions highlight surface areas modified by the optimization to reduce drag.
Figure 7 : t-SNE comparisons for anchor shapes using three latent representations: early w+ , global latent z , and late w+ . In each row, the anchor shape is highlighted by the red box.
Figure 8 : Additional anchor-wise comparisons showing the same qualitative behavior across latent spaces.
Figure 9 : Further examples illustrating the consistency of the latent-space organization across different anchor shapes.
Figure 10 : Additional cross-instance style-mixing examples for analyzing the role of late W+ . Across examples, substituting deeper style codes while keeping the early style codes fixed tends to preserve the anchor’s coarse silhouette and global proportions more strongly than edits involving the single global latent, while still altering finer geometric characteristics.
Figure 11 : Five more cross-instance style-mixing examples continuing the same trend.
Figure 12 : Additional examples showing that deeper-code substitution yields more localized, detail-oriented changes, whereas changing the early style codes produces larger variation in coarse structure.
Figure 13 : Final style-mixing example supporting the interpretation that late W+ predominantly influences finer geometric detail rather than serving as a second entangled global code.
Figure 14 : Dense global-latent interpolation for cars (example 1). Each method is shown with interpolation coefficients t∈{0,0.1,…,1.0} . The denser sampling makes clear that interpolation through a single global latent code changes coarse structure and fine detail simultaneously, producing averaged intermediate geometries rather than disentangled coarse-to-fine transitions.
Figure 15 : Dense global-latent interpolation for cars (example 2). As in the main paper, none of the methods disentangle low-frequency and high-frequency geometry when interpolation is performed only in the global latent z ; both levels of structure evolve jointly along the path.
Figure 16 : Dense global-latent interpolation for chairs. While interpolation in the single global latent z remains entangled for all methods, StyleFields produces substantially cleaner and more structurally plausible intermediate chairs. In particular, seat, backrest, and leg geometry remain more coherent throughout the trajectory, whereas the baselines exhibit stronger degradation and instability in the intermediate shapes.
Figure 17 : Dense global-latent interpolation for planes. The trajectories are visually smoother than in the chair case, but the same entanglement remains: coarse airplane structure and finer geometric details change together along the interpolation path, rather than being controlled independently by scale.
Cars
Chairs
Planes
Method
# Params
CD ( ↓ )
IoU ( ↑ )
CD ( ↓ )
IoU ( ↑ )
CD ( ↓ )
IoU ( ↑ )
DeepSDF
1.7M
1.6×10−4
97.27
8.8×10−4
89.74
8.0×10−4
89.28
Curriculum DeepSDF
2.1M
1.6×10−4
97.19
1.2×10−3
89.33
9.2×10−4
90.57
Curriculum + Style DeepSDF Reconstruct with Z and ∣J∣=4
4.2M
1.3×10−4
97.87
5.7×10−4
90.12
7.8×10−4
91.26
StyleFields : Reconstruct with Z with ∣J∣=1
3.2M
1.6×10−4
97.25
8.6×10−4
90.05
8.1×10−4
89.15
StyleFields : Reconstruct with Z with ∣J∣=2
3.6M
1.5×10−4
97.27
6.4×10−4
91.17
7.9×10−4
91.32
Table 3 : Ablation on the number of depth-wise style groups in J across Cars, Chairs, and Planes. Here, ∣J∣ denotes the number of modulation groups used to partition the decoder into coarse-to-fine style-controlled ranges. We also report the corresponding number of learnable style parameters introduced by each setting. Lower CD is better and higher IoU is better.
Figure 18 : Additional StyleFields cross-instance style mixing between two chair shapes (source on the left, target on the right). Each column shows a hybrid produced by selectively substituting style codes of Shape 1 with codes from Shape 2. (A) Coarse styles (w0+,w1+) from Shape 1 and fine styles (w2+,w3+) from Shape 2. (B) Mid-level styles (w1+,w2+) from Shape 1 and coarsest and finest styles (w0+,w3+) from Shape 2. (C) Fine styles (w2+,w3+) from Shape 1 and coarse styles (w0+,w1+) from Shape 2.
Figure 19 : More cross-instance style mixing examples on Chairs. The same substitution patterns as in Fig. 18 are used. Replacing shallow style groups leads to larger changes in global chair structure, while preserving them retains more of the source shape’s broad configuration and yields more localized modifications.
Figure 20 : Additional StyleFields cross-instance style mixing between two airplane shapes (source on the left, target on the right). Each column shows a hybrid produced by selectively substituting style codes of Shape 1 with codes from Shape 2. (A) Coarse styles (w0+,w1+) from Shape 1 and fine styles (w2+,w3+) from Shape 2. (B) Mid-level styles (w1+,w2+) from Shape 1 and coarsest and finest styles (w0+,w3+) from Shape 2. (C) Fine styles (w2+,w3+) from Shape 1 and coarse styles (w0+,w1+) from Shape 2.
The growing demand for rapid and scalable 3D asset creation has driven interest in feed-forward 3D reconstruction methods, with 3D Gaussian Splatting (3DGS) emerging as an effective scene representation. While recent approaches have demonstrated pose-free reconstruction from unposed image collections, integrating stylization or appearance control into such pipelines remains underexplored. Existing attempts largely rely on image-based conditioning, which limits both controllability and flexibility. In this work, we introduce AnyStyle, a feed-forward 3D reconstruction and stylization framework that enables pose-free, zero-shot stylization through multimodal conditioning. Our method supports both textual and visual style inputs, allowing users to control the scene appearance using natural language descriptions or reference images. We propose a modular stylization architecture that requires only minimal architectural modifications and can be integrated into existing feed-forward 3D reconstruction backbones. Experiments demonstrate that AnyStyle improves style controllability over prior feed-forward stylization methods while preserving high-quality geometric reconstruction. A user study further confirms that AnyStyle achieves superior stylization quality compared to an existing state-of-the-art approach. Repository: https://github.com/joaxkal/AnyStyle.
Joanna Kaleta, Bartosz Świrta, Kacper Kania +3
Warsaw University of Technology · Sano Centre for Computational Medicine · IDEAS NCBR +2
This study addresses the partial-to-complete geometry reconstruction of deformable objects (DOs) from point-cloud observations toward precise DO manipulation. Recent DO reconstruction approaches often adopt implicit neural representations (INRs) to model continuous surfaces as well as capture structural variability. However, these methods typically rely on object-specific shape priors that improve training stability and limit generalization. To figure it out, we introduce ParCo-SDF, a two-stage partial-to-complete signed distance field (SDF) reconstruction framework consisting of temporal geometry encoding followed by FiLM-conditioned SDF prediction. The temporal encoder captures structural similarity across DO sequence, enabling prior-free stable training. FiLM-based conditioning preserves reconstruction expressivity while reducing network complexity. We evaluate the proposed method against a state-of-the-art DO surface reconstruction baseline on a rubber band manipulation dataset, demonstrating robust and high-fidelity reconstruction under severe occlusions.
Deokmin Hwang, Minseok Song, Daehyung Park
School of Computing, Korea Advanced Institute of Science and Technology, Korea
Implicit Neural Representations (INRs) have become the standard for continuous 2D shape modeling, but they suffer from black-box uneditability, vulnerability to noise, and high parameter counts that severely hinder deployment on edge devices. We introduce Fluid-SDF, a highly compressed, differentiable Constructive Solid Geometry (CSG) framework that models shapes using explicit geometric primitives blended via a smooth minimum function. By replacing traditional multi-layer perceptrons (MLPs) with a parameterized primitive engine, Fluid-SDF reconstructs complex, non-convex topologies using strictly under 100 parameters, achieving comparable or superior intersection-over-union (mIoU) to standard neural baselines. Furthermore, we demonstrate that Fluid-SDF acts as a powerful geometric prior, inherently resisting high-frequency dataset noise where capacity-matched neural networks catastrophically overfit. Finally, unlike standard INRs, Fluid-SDF's explicit parameter space allows for direct, zero-shot user editing of local and global shape features without retraining. By bypassing expensive on-device gradient updates entirely, Fluid-SDF is uniquely suited for mobile AI, augmented reality, and resource-constrained embedded environments