cs.CVJul 24, 2025

3D Software Synthesis Driven by Constraint-Expressive Intermediate Representation

Authors: Shuqing Li, Anson Y. Lam, Yun Peng, Wenxuan Wang, Michael R. Lyu

Organizations: The Chinese University of Hong Kong Hong Kong, China · Renmin University of China Beijing, China

Abstract

Graphical user interface (UI) software has undergone a fundamental transformation from traditional two-dimensional (2D) desktop/web/mobile interfaces to spatial three-dimensional (3D) environments. While existing work has made remarkable success in automated 2D software generation, such as HTML/CSS and mobile app interface code synthesis, the generation of 3D software still remains under-explored. Current methods for 3D software generation usually generate the 3D environments as a whole and cannot modify or control specific elements in the software. Furthermore, these methods struggle to handle the complex spatial and semantic constraints inherent in the real world. To address the challenges, we present Scenethesis, a novel requirement-sensitive 3D software synthesis approach that maintains formal traceability between user specifications and generated 3D software. Scenethesis is built upon ScenethesisLang, a domain-specific language that serves as a granular constraint-aware intermediate representation (IR) to bridge natural language requirements and executable 3D software. It serves both as a comprehensive scene description language enabling fine-grained modification of 3D software elements and as a formal constraint-expressive specification language capable of expressing complex spatial constraints. By decomposing 3D software synthesis into stages operating on ScenethesisLang, Scenethesis enables independent verification, targeted modification, and systematic constraint satisfaction. Our evaluation demonstrates that Scenethesis accurately captures over 80% of user requirements and satisfies more than 90% of hard constraints while handling over 100 constraints simultaneously. Furthermore, Scenethesis achieves a 42.8% improvement in BLIP-2 visual evaluation scores compared to the state-of-the-art method.

Figures & tables

Explore similar work

CardsList
  1. MUSE: Agentic 3D Scene Authoring via Memory-Grounded Incremental Requirement Satisfaction

    Jun 12, 2026Ruijie Xu, Xinnan Zhu, Jiayu Ying +33D Scene GenerationScene Understanding

  2. ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning

    Jul 15, 2026Yuan Xiao, Can Wang, Xiangyu Kong +13D Scene Generation3D Scene Understanding

  3. SpatialGrammar: A Domain-Specific Language for LLM-Based 3D Indoor Scene Generation

    Apr 30, 2026Song Tang, Kaiyong Zhao, Yuliang Li +53D Scene Generation3D Layout Generation