While Large Language Models (LLMs) advance 3D indoor scene synthesis, current pipelines fail to retain user-specific preferences across sessions, making immersive authoring a repetitive and physically fatiguing process. We present SPHERE, an adaptive VR generation framework that transforms isolated synthesis into continuous human-AI co-creation. SPHERE extracts persistent spatial preferences from natural multimodal interactions (speech and controller edits). To ensure geometric resilience against spatial distortions, it abstracts these raw edits into hierarchical constraints modeling both local functional and global topological contexts. Furthermore, a human-in-the-loop reinforcement learning mechanism dynamically updates retrieval policies based on the user's final edited scenes. A mixed-design user study (N=42) and an offline ablation demonstrate that SPHERE significantly reduces corrective edits and physical demand, preventing bias toward shallow object-level traits to yield geometrically resilient, profile-aligned layouts. Ultimately, SPHERE demonstrates how capturing demonstrated spatial logic enables controlled spatial adaptation, establishing a reliable, governed human-AI collaboration framework for immersive authoring. Project page and source code will be available at: https://github.com/hyeonmin11/SPHERE
Figures & tables
Figure 1 : Given user instruction and the Top-M retrieved preference memory entries as input, SPHERE performs indoor scene generation to produce an initial layout. The generated scene is then refined through scene editing, where users interactively adjust indoor scene in VR. From the edited scene, preference memory construction extracts object-level and layout-level constraints from the edited scene. Finally preference retrieval refinement updates rewards using realization signal, closing the feedback loop and enabling incremental personalization across sessions.
Figure 2 : Users interactively modify objects within the VR environment through natural language utterances and controller-based interactions that specify target objects and spatial positions. The system parses the multimodal input to identify the referenced object and desired placement, converts it into a structured instruction, and dispatches the corresponding edit operation to update the scene.
Figure 3 : From the scene context and user interaction logs, SPHERE constructs structured preference representations using a spatial constraint library. The extracted preferences are organized hierarchically into object-level, intra-area-level, and inter-area-level constraints, and stored in the preference memory.
Figure 4 : Comparison of subjective workload and editing outcomes across Scene 2 and Scene 3. (1) NASA-TLX subscale ratings for the Baseline and Adaptive conditions. (2) Total edit counts, showing fewer corrective edits in the Adaptive condition. (3) Mean attribution scores, indicating stronger perceived personalization in the Adaptive condition.
Figure 5 : Example comparison of inferred preference constraints generated by the Baseline and Adaptive conditions for the same session. The Adaptive condition captures richer spatial organization and profile-consistent layout reasoning than the Baseline.
Figure 6 : SUS and System’s Perceived Intelligence results. The Adaptive condition received higher usability ratings than the Baseline and was also rated significantly higher in adaptability, time saving, and trust.
Scene
Variable Pair
Pearson’s r
p
Scene 2
Mean Spatial Score – P_count
-0.082
.605
Mean Style Score – E_count
-0.424
.005
Mean Functional Score – S_count
-0.334
.031
Scene 3
Mean Spatial Score – P_count
-0.278
.075
Mean Style Score – E_count
-0.644
<.001
Mean Functional Score – S_count
-0.218
.165
Table 1: Correlation between perceived alignment scores and edit counts across Scene 2 and Scene 3. All correlations were computed with df=40 .
Profile Alignment Accuracy
Granularity
Condition
Profile A
Profile B
Profile C
Total
O / total
Condition A
0.600
0.467
0.467
0.511
0.35
Condition B
0.333
0.533
0.400
0.422
0.75
Table 2: Profile alignment accuracy and proportion of object-level constraints (O / total) across conditions. For constraint categorization, evaluation was conducted by two human raters ( κ>0.560 ), and disagreements where only one annotator categorized a constraint as object-level were resolved through adjudication.
Scene synthesis and editing has emerged as a promising direction in computer graphics. Current trained approaches for 3D indoor scene generation either oversimplify object semantics through one-hot class encodings (e.g., 'chair' or 'table'), require masked diffusion for editing, ignore room boundaries, or rely on floor plan renderings that fail to capture complex layouts. LLM-based methods enable richer semantics via natural language, but lack editing functionality, are limited to rectangular layouts, or rely on weak spatial reasoning from implicit world models. We introduce ReSpace, a generative framework for autoregressive text-driven 3D indoor scene synthesis and editing. Our approach features a compact structured scene representation with explicit room boundaries that enables asset-agnostic deployment and frames scene manipulation as a next-token prediction task, supporting object addition, removal, and swapping via natural language. We employ supervised fine-tuning with a preference alignment stage to train a specialized language model for object addition that accounts for user instructions, spatial geometry, object semantics, and scene-level composition. We further introduce a voxelization-based evaluation metric capturing fine-grained geometric violations beyond 3D bounding boxes. Experiments surpass state-of-the-art on object addition and achieve superior human-perceived quality on the application of full scene synthesis, despite not being trained on it.
Automatically generating interactive 3D indoor scenes from natural language is crucial for virtual reality, gaming, and embodied AI. However, existing LLM-based approaches often suffer from spatial errors and collisions, in part because common scene representations-raw coordinates or verbose code-are difficult for models to reason about 3D spatial relationships and physical constraints. We propose SpatialGrammar, a domain-specific language that represents gravity-aligned indoor layouts as BEV grid placements with deterministic compilation to valid 3D geometry, enabling verifiable constraint checking. Building on this representation, we develop (1) SG-Agent, a closed-loop system that uses compiler feedback to iteratively refine scenes and enforce collision constraints, and (2) SG-Mini, a 104M-parameter model trained entirely on compiler-validated synthetic data. Across 159 test scenes spanning five scenarios of different complexity, SG-Agent improves spatial fidelity and physical plausibility over prior methods, while SG-Mini performs competitively against larger LLM-based baselines on single-shot generation scenarios.
Song Tang, Kaiyong Zhao, Yuliang Li +5
The Hong Kong University of Science and Technology (Guangzhou) · XGRIDS · Harbin Institute of Technology (Shenzhen)
Recent advances in large language models (LLMs) have significantly improved language-driven 3D content generation, but most existing approaches still treat scene generation and user interaction as separate processes, limiting the adaptability and immersive potential of interactive multimedia systems. This paper presents a unified framework that closes the loop between language-driven 3D scene generation and immersive user interaction. Given natural language instructions, the system first constructs structured scene representations using LLMs, and then optimizes spatial layouts via reinforcement learning under geometric and semantic constraints. The generated environments are deployed in a virtual reality setting to facilitate HRI-in-the-loop, where user interactions provide continuous feedback to align generated content with human perception and usability. By tightly coupling generation and interaction, the proposed framework enables more responsive, adaptive, and realistic multimedia experiences. Experiments on the ALFRED benchmark demonstrate state-of-the-art performance in task-based scene generation. Furthermore, qualitative results and user studies show consistent improvements in immersion, interaction quality, and task efficiency, highlighting the importance of closed-loop integration of generation and interaction for next-generation multimedia systems. Our project page can be found at https://proj-showcase.github.io/h3ds/.
Anh H. Vo, Sungyo Lee, Phil-Joong Kim +2
Department of Computer Engineering, Sejong University, Seoul, Republic of Korea