Authors: Neil De La Fuente, Joan Lafuente, Mukhammadali Sayfiddinov, Felicia Scharitzer, Marc Pollefeys, Ata Celen, Sayan Deb Sarkar, Elisabetta Fedele
Organizations: ETH Zürich · Microsoft · Stanford University
Current 3D generation methods lack explicit local control: geometric adherence is often defined by a global control strength, and appearance cannot be specified locally. We present SpaceFlow, a training-free pipeline for locally controllable 3D generation from text descriptions and a collection of geometric primitives. Each primitive serves as a proxy for an object part and is assigned a local control level, enabling users to specify whether regions should strictly follow the input shape or allow generative completion. During structure generation, we enforce these spatial constraints within the generative flow process. For appearance synthesis, the generated structure is segmented and matched to the primitives. Each generated part is conditioned only on its assigned text or image cue, thereby limiting cross-part leakage. Regional geometry metrics demonstrate that SpaceFlow preserves the specified geometry in high-control regions and enables plausible shape variation in low-control areas. A user study further indicates that the resulting balance between geometric fidelity and generative freedom remains competitive in overall quality. When evaluating appearance on fixed geometry, text-conditioned routing achieves state-of-the-art prompt faithfulness and color/material accuracy. Qualitative results additionally show localized routing of image cues. The project page is available at SpaceFlow3D.github.io.
Figures & tables
Figure 2 : SpaceFlow pipeline. Given editable geometric controls with local geometric conditioning strengths τi , SpaceFlow encodes the high-control regions and injects them into the sparse-structure flow process over a controlled time interval. The resulting structure is segmented with PartField [ 55 ] , then per-part text or image conditions are routed to each part, and the latent is finally passed to an appearance flow model and refined with self-similarity guidance, producing a 3D asset with localized geometric and appearance control.
Figure 3 : Local spatial control. Starting from the same "hourglass" superquadric primitive set and prompt, high control (τ=10) preserves the corresponding part’s input geometry, while low control (τ=3) allows generative variation. The four settings showcase independent control of the frame and sand.
Figure 4 : Spatial and appearance control in isolation. (a) Spatial control. Uniform control with SpaceControl either loses specified structure ( τ=3 ) or constrains the full object ( τ=10 ), whereas SpaceFlow preserves the high-control regions and allows the generative prior to complete the low-control regions. (b) Appearance control on fixed geometry. Global conditioning can miss part-level cues or spread them beyond their targets; SpaceFlow routes each color or material cue to its target part while retaining coherent textures.
Prompt and geometric fidelity
Realism
Overall
Baseline
Win rate
95% CI
Win rate
95% CI
Win rate
95% CI
SpaceControl ( τ=3 )
78.3%
[68.3, 85.8]
43.4%
[33.2, 54.1]
75.9%
[65.7, 83.8]
SpaceControl ( τ=10 )
66.3%
[55.6, 75.5]
85.5%
[76.4, 91.5]
68.7%
[58.1, 77.6]
Table 1 : Localized geometric control. Win rates (%) of SpaceFlow ( τlow=3 , τhigh=10 ) against SpaceControl [ 28 ] using a uniform global control strength. Brackets report 95% Wilson CIs over 83 assets.
Method
CDhigh↓
ΔCD↑
ΔDs→10↑
CLIP ↑
SpaceControl ( τ=3 )
15.99
-0.14
0.01
0.232
SpaceControl ( τ=10 )
0.94
0.11
0.00 †
0.226
SpaceFlow ( Ours )
3.65
6.05
0.15
0.233
Table 2 : Quantitative evaluation of spatial control selectivity and semantic alignment. We report the regional Chamfer Distance delta ( ΔCD=CDlow−CDhigh , ×103 ), the feature deviation delta to the reference ( ΔDs→10=Dlows→10−Dhighs→10 ), and CLIP text-image similarity. Positive Δ values indicate effective localized control selectivity between high- and low-control regions. †ΔD is 0.00 by definition for τ=10 as D10→10=0 .
Figure 5 : Qualitative comparison of spatial and appearance control across diverse assets. Control (top): Input superquadric proxies with designated local control strengths (orange: high control τi=10 ; gray: low control τi=3 ) and localized part-level appearance prompts. SpaceControl [ 28 ] : Baselines under uniform control strengths ( τ=3 vs. τ=10 ), illustrating the trade-off between generative reinterpretation and strict adherence to the proxy geometry. Texture Baselines: Appearance baselines (TRELLIS [ 3 ] and GuideFlow3D [ 30 ] ) built directly on our generated structure to isolate texture synthesis. Local Control (Ours): SpaceFlow provides localized geometric control (generating text-aligned geometric details in low-control regions while preserving high-control shapes) alongside accurate part-level texture routing.
Figure 6 : User study showing overall preference for SpaceFlow over SpaceControl [ 28 ] with uniform τ=3 and τ=10 guidance across 337 evaluation trials.
Prompt faithfulness
Texture detail
Overall
Baseline
Win rate
95% CI
Win rate
95% CI
Win rate
95% CI
TRELLIS
63.9%
[53.1, 73.4]
31.3%
[22.4, 41.9]
59.0%
[48.3, 69.0]
GuideFlow3D
58.5%
[47.7, 68.6]
37.5%
[27.7, 48.5]
57.3%
[46.5, 67.5]
Table 3 : Pairwise appearance evaluation. Win rates (%) of SpaceFlow against fixed-structure appearance baselines. Brackets report 95% Wilson CIs over the set of test assets.
Method
PF ↑
CM ↑
TD ↑
Overall ↑
TRELLIS
6.13
5.79
6.87
6.42
GuideFlow3D
6.40
5.96
6.76
6.46
SpaceFlow (ours)
7.07
6.43
6.49
6.78
Table 4 : Quantitative evaluation of appearance control. Mean scores ( 1 – 10 ) averaged over 83 cases: prompt faithfulness (PF), color/material accuracy (CM), texture detail (TD), and overall. All variants share the same structure; metrics are appearance-only.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure S1 : Illustration of a single Conditioned Flow iteration. The process receives the low-control region mask Mlow (derived from the bounding boxes), the high-control target latent z1high , and the current latent state zt . The flow model predicts the next state zt+1 . Inside the high-control region (masked by 1−Mlow ), a weighted average of the target latent z1high and the predicted latent zt+1 is computed. This is combined with the unconstrained prediction zt+1 in the low-control region (masked by Mlow ) to form zt+1′ . Finally, noise is added to zt+1′ up to timestep t (repeated K times) to update the latent zt for the subsequent step.
Figure S2 : Routed cross-attention. Each text or image cue is encoded by a frozen CLIP (text) or DINOv2 (image) encoder into its own key/value pair Ki,Vi . Within every cross-attention layer of the appearance transformer, the latent queries Q are partitioned by the routing volume so that the voxels of a part attend only to their assigned cue—e.g. “red wings” routes to the wings while the global “plane” prompt covers the remaining voxels—and the per-region outputs Hi are recombined into Hout . This confines each appearance cue to its target part and reduces cross-part texture leakage.
Figure S3 : Part-wise appearance pipeline. From the generated sparse structure and the part-wise text or image cues, we decode the structure into a mesh and extract a PartField triplane feature field. Active structure voxels are clustered into semantic localities; each cluster centroid is assigned to its nearest input superquadric, and the resulting map is rasterized into a condition-routing volume that stays part-consistent even where low-control geometry drifts from the scaffold. The appearance rectified-flow transformer then synthesizes the SLAT appearance latents under this routing.
Setting
Value
Structure backbone
TRELLIS text sparse-structure flow model
Structure latent tensor shape
8×16×16×16
Local control strengths
τlow=3 , τhigh=10
High-control blending weight ( α )
0.18 for local- τ
Number of structure flow steps
12
Flow solver
Euler flow-matching sampler with classifier-free guidance interval
Appendix
Table S1 : Structure-generation settings used in our spatial-control experiments.
Setting
Value
Appearance backbone
TRELLIS SLAT flow model and decoders
Conditioning used in quantitative evaluation
Global and part-specific text prompts
Text-condition encoder
CLIP ViT-L/14
Image-condition encoder
DINOv2 ViT-L/14, used for image-conditioned standalone evaluation and qualitative examples
Part representation
PartField triplanes
PartField descriptor dimension
448
Appendix
Table S2 : Appearance-generation and self-similarity guidance settings. Text-conditioned appearance is evaluated pairwise and with standalone VLM scores; image-conditioned appearance is evaluated standalone using the criteria in Fig. S18 and is additionally illustrated qualitatively.
Category
Count
Share
Representative Examples
Vehicles & Transport
15
18.1%
Airplane, Space Shuttle, Satellite, Sailboat, F1 Car, Truck
Tin Robot, Wooden Dolls, Bonsai, Cactus, Elephant, Trophy
Furniture & Seating
13
15.7%
Office Chair, Bar Stool, Armchairs, Park Bench, Pool Table
Tools & Utility
10
12.0%
Fire Hydrant, Hammer, Sewing Machine, Toolbox, Iron, Bucket
Lighting & Lanterns
8
9.6%
Desk Lamps, Chandelier, Stone Lantern, Candle Lantern
Appendix
Table S3 : Object category distribution in the evaluation benchmark ( N=83 ). The benchmark spans 7 broad semantic categories designed to evaluate multi-part geometric control and local appearance conditioning.
Figure S4 : Dataset assets gallery. Ten representative assets from our 83-asset benchmark. Orange primitives designate high-control regions requiring strict geometric preservation, while white/gray primitives designate low-control regions intended for generative completion. Subtitles denote the target object concept.
Setting
Value / Specification
Judge Model
GPT-5.6 Sol
Sampling Temperature
Default ( T=1.0 )
Number of Passes
3 independent passes per condition
Pairwise Aggregation
2-of-3 majority vote (re-salted slot order)
Standalone Aggregation
Mean across 3 passes
Inter-Pass Agreement
Pairwise agreement: 0.898 ( Spearman )
Appendix
Table S4 : VLM-judge configuration and evaluation settings. Summary of model parameters, multi-pass aggregation rules, and rendering specifications.
Figure S5 : User study interface screenshot. Screenshot of the evaluation interface to study participants during a trial. The interface displays the text prompt, the input geometric control shape with color-coded region guidelines, and two randomized, anonymous outputs for comparison.
Figure S6 : Geometric deviation across local control strengths. We fix the high-control strength at τhigh=10 and vary the low-control strength over τlow∈{1,3,5,7,9} . We report the regional symmetric mean-squared Chamfer distance to the input primitive surfaces ( ×103 ; lower is better), averaged asset-wise over the 83-asset dataset. Lines show the mean and shaded bands indicate 95% bootstrap confidence intervals. Low-control regions (gray) deviate most under weak control, and the gap to high-control regions (orange) closes as the two control strengths converge.
Figure S7 : Qualitative local control strength continuum. We fix the protected primitives at τhigh=10 and vary the editable-region strength over τlow∈{1,3,5,7,9} . Orange input primitives denote high control and gray input primitives denote low control. Lower values permit greater prompt-driven geometric freedom, whereas increasing τlow progressively strengthens adherence to the input geometry. Outputs are shown without texture to isolate geometric behavior.
Figure S8 : Visualization of spatial feature distance. Left: input superquadrics (top; orange denotes high control and gray denotes low control) and the asset generated by SpaceFlow (bottom). Right: voxel-wise cosine distance between the generated asset’s PartField features and those of the nearest spatial match in the uniformly controlled ( τ=10 ) reference. Feature deviations concentrate primarily in the low-control region.
Figure S9 : Primitive-to-part appearance routing. In each pair (a–h), we show the input structure (left) and the resulting 323 routing volume (right). The mapping is defined color-wise: regions sharing the same color between the input primitives and the generated structure represent a 1-to-1 mapping to that condition (with gray assigned to the global prompt). Subcaptions provide the flattened prompt containing all localized appearance cues.
Figure S10 : Semantic-geometric conflict. Synthesizing a “person” over an input “bench” structure (left). Lacking a joint prior for conflicting modalities, the model generates unnatural distortions by embedding the bench’s high-control geometry inside the generated “person” (right).
Figure S11 : Failure cases of local appearance routing. For each asset we show the derived routing volume (left; colors denote distinct appearance conditions) and the generated result (right). (a–b) The handle cluster wraps around the rim, so the wooden handle appearance bleeds onto the metal pan body. (c–d) The seat cluster covers only part of the cushion, leaving the rest textured by the competing condition.
Figure S12 : Additional qualitative results for image-conditioned appearance control. Each panel shows an input geometric structure (left), the object-level text prompt (bottom), and two reference images specifying the desired material, texture or style for the object or for particular parts. Dotted arrows indicate the part-level associations between the reference images and the corresponding target regions. We recommend zooming in to inspect the fine-grained structural and appearance details; for example, in the metal-toolbox example, the generated asset includes the small front latches visible in the reference image and renders them with a consistent metallic appearance. These examples illustrate how SpaceFlow transfers reference-driven appearance cues to the corresponding regions while retaining the overall structure of the input object.
Figure S13 : Extended qualitative gallery (Part 1). Text-conditioned SpaceFlow results across diverse object categories. Each cell pairs the input superquadric scaffold (left; orange = high-control, gray = low-control regions) with the generated asset (right) rendered from the same viewpoint.
Figure S14 : Extended qualitative gallery (Part 2). Additional text-conditioned SpaceFlow results. Each cell pairs the input superquadric scaffold (left; orange = high-control, gray = low-control regions) with the generated asset (right) rendered from the same viewpoint.
Text and image conditioned 3D models now generate convincing assets, but they still offer little direct control over the space an object should occupy or avoid. In authoring, this spatial intent is often known before generation starts. A chair should fit a seating envelope, a prop should leave clearance for motion, or a part should expose a contact surface. Prompts and image views are poor carriers for such constraints, requiring the need for an explicit control interface. We present Arbor, a trainable attachment for text conditioned latent 3D generation. Arbor introduces constraint meshes as a native 3D control interface. The interface uses hull regions where geometry should exist, avoidance regions that should remain empty, and touch regions the object should contact. Unlike completion or whole object scaffold control, these meshes are not target evidence. They are local typed requirements and can include regions where no surface should appear. Arbor keeps this signal as geometry by converting constraint meshes into tokens and learning a routed attachment inside a frozen denoiser. Each latent region can therefore receive the part of the constraint that matters for its spatial location. We evaluate Arbor on automatic and artist curated control benchmarks with hull, avoidance, and touch constraints, and compare the metric trends to a user preference study. Even without dedicated compliance losses, Arbor improves constraint obedience while preserving object quality and variation under fixed constraints.
Jan-Niklas Dihlmann, Andreas Engelhardt, Simon Donne +2
Text-to-3D generation has advanced rapidly, yet state-of-the-art models, encompassing both optimization-based and feed-forward architectures, still face two fundamental limitations. First, they struggle with coarse semantic alignment, often failing to capture fine-grained prompt details. Second, they lack robust 3D spatial understanding, leading to geometric inconsistencies and catastrophic failures in part assembly and spatial relationships. To address these challenges, we propose VLM3D, a general framework that repurposes large vision-language models (VLMs) as powerful, differentiable semantic and spatial critics. Our core contribution is a dual-query critic signal derived from the VLM's Yes or No log-odds, which assesses both semantic fidelity and geometric coherence. We demonstrate the generality of this guidance signal across two distinct paradigms: (1) As a reward objective for optimization-based pipelines, VLM3D significantly outperforms existing methods on standard benchmarks. (2) As a test-time guidance module for feed-forward pipelines, it actively steers the iterative sampling process of SOTA native 3D models to correct severe spatial errors. VLM3D establishes a principled and generalizable path to inject the VLM's rich, language-grounded understanding of both semantics and space into diverse 3D generative pipelines.
We consider the problem of regenerating 3D objects from 2D images and initial 3D shapes. Most 3D generators operate in a one-shot fashion, converting text or images to a 3D object with limited controllability. We introduce instead MeshReGen, a 3D regenerator that is conditioned on an initial 3D shape. This conceptually simple formulation allows us to support numerous useful tasks, including 3D enhancement, reconstruction, and editing. MeshReGen uses a new conditioning mechanism based on VecSet, which allows the regenerator to update or improve the input geometry with consistent fine-grained details. MeshReGen learns a widely applicable regeneration prior from off-the-shelf 3D datasets via self-supervised pretext tasks and augmentations, without additional annotations. We evaluate both the geometric consistency and fine-grained quality of MeshReGen, achieving state-of-the-art performance in controllable 3D generation across several tasks.
Geon Yeong Park, Roman Shapovalov, Rakesh Ranjan +3
KAIST · ⋆Work done during an internship at Meta Reality Labs. · Meta Reality Labs