Foundation models are powerful generators, but many engineering domains require structured representations that general-purpose systems handle poorly. We introduce FLOORA (Floor Layout Optimization with RL Alignment), a family of small domain-specific language (DSL) models for architectural layout generation. With specialized data and alignment, our 0.6B model outperforms much larger frontier models, achieving VLM judge win rates up to 92.0% on out-of-distribution real-world buildings and 96.0% on synthetic buildings. Human evaluations further corroborate these results, with FLOORA selected as the best model in 89.3% of evaluations. FLOORA combines a token-efficient DSL, custom tokenization, domain-specific pretraining, supervised fine-tuning (SFT), and reinforcement learning (RL) with learned human-preference and verifiable rewards. This pipeline improves architectural and geometric validity, supported by extensive empirical evaluation and ablation studies. Although focused on architecture, our results suggest that similar domain-specific recipes may be useful in other engineering domains with structured, verifiable outputs. Datasets, models, and inference code are available at https://github.com/AutodeskAILab/floora.
Figures & tables
Figure 1: Overview of our DSL pre-training, human feedback collection, and post-training pipeline.
Figure 2: Qualitative comparison of FLOORA-0.6B and frontier models on a real-world test sample. FLOORA outperforms the baselines on geometric and functional checks and architectural quality.
Figure 3: High-level recipe for training and post-training a domain-specific language model.
Figure 4: (a) t-SNE visualization of massing polygons across synthetic, OSM, and SFT datasets, indicating the distribution shift between synthetic and real-world massing geometries. (b) Tokenized sequence length distributions for the pre-training dataset using the DSL tokenizer and the original Qwen3 tokenizer, showing that the DSL tokenizer reduces sequence lengths by 54%, on average.
Figure 5: Structured prompt and completion format used for DSL pre-training and post-training.
Figure 6: Pass@5 for FLOORA models across model sizes and training stages, where success requires all checks to pass. Shading shows 95% CIs across 3 training and 5 evaluation seeds.
Figure 7: Reward model validation accuracy, reported with 95% CIs.
Figure 8: Pairwise VLM judge outcomes for each GRPO (RM+VR) model against its matched GRPO (VR only) counterpart. Bars show the fraction of prompts where RM+VR wins, ties, or loses. As model size increases, the RM+VR variant is increasingly preferred. Error bars indicate 95% Wilson intervals over the evaluation prompts.
Figure 9: (a) Pass@k comparison of FLOORA-0.6B (Qwen3) and frontier baselines, where success requires passing both geometric and functional checks. (b) Pairwise VLM judge outcomes for FLOORA-0.6B (Qwen3) against frontier baselines. Bars show the fraction of prompts for which FLOORA wins, ties, or loses. Our model substantially outperforms all evaluated frontier baselines.
Figure 10: Human evaluation results across 100 test inputs with 95% CIs. FLOORA is selected as the best model in 89.3% of evaluations.
Appendix figures & tables41 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 11: Examples of DSL representation of multifamily residential buildings, illustrating the massing, vertical cores, living units, corridors, and corresponding 2D floor-plan layout.
Stage
Description
Samples
Share of corpus
1,2
TileGPT-derived layouts with scale jitter
3.9M
95%
3
Inference + repair (admitted)
0.2M
5%
Total before rotation augmentation
4.1M
100%
Total after rotation augmentation
82M
Appendix
Table 1: Synthetic corpus composition by generation stage. Counts are before rotation augmentation unless noted.
Parameter
Value
Boundary sample points N
64
Maximum samples per dataset
10,000
t-SNE perplexity
40
t-SNE iterations
1,000
t-SNE initialization
PCA
Appendix
Table 2: Hyperparameters used for the t-SNE analysis of massing footprints.
Figure 12: Samples from the synthetic dataset. Each sample contains the building massing, metadata, and a procedurally generated space layout consisting of cores, corridors, and living units.
Figure 13: Samples from the OSM dataset. Each sample contains only the building massing and metadata, without ground-truth space layouts, and therefore cannot be used for pre-training.
Figure 14: Samples from the SFT dataset. Each sample contains the building massing, metadata, and an architect-edited space layout consisting of living units and, when needed, cores and corridors.
Figure 15: Labelers independently evaluate each model generation according to the criteria.
Figure 16: Model-generated outputs are ranked relative to one another, with ties permitted.
Figure 17: The highest-ranked model-generated output is subsequently edited by the labeler.
Category
Reason
Count
Skipped
Rating −1
61
Skipped
Inconsistent pairs
575
Skipped
Failed geometric checks
287
Errors
Failed DSL encodings
227
Total
1,150
Appendix
Table 3: Number of human feedback datapoints removed during data filtering and processing.
Dataset
Before Aug.
After Aug.
SFT completions
10.4k
135.4k
Pairwise preferences
90.6k
1.2M
Appendix
Table 4: Number of datapoints in the human-feedback dataset after filtering, before and after rotation augmentation.
Tokenizer
Mean
Median
P95
Max
Before Aug.
After Aug.
DSL Tokenizer
235.9
219.0
373.0
785
880M
17.6B
Qwen3 Tokenizer
516.6
475.0
844.0
1,902
1.9B
38.5B
Appendix
Table 5: Tokenization statistics by tokenizer. “Before Aug.” and “After Aug.” denote total token counts before and after augmentation.
Base Model
Effective Params
Learning Rate
Qwen3-0.6B
440M
3e-4
Qwen3-1.7B
1.4B
4e-4
Pythia-70M
20M
2e-4
Pythia-160M
90M
6e-5 1 1 1 A lower learning rate was used because training was unstable at higher learning rates.
Pythia-410M
310M
3e-4
Pythia-1.4B
1.2B
3e-4
Appendix
Table 6: Effective parameter counts after resizing the embedding layers to match the DSL tokenizer vocabulary, and peak learning rates used during pre-training.
Model
SFT
RM
GRPO
Qwen3-0.6B
2e-4
8e-5
6e-5
Qwen3-1.7B
2e-4
3e-4 2 2 2 A higher learning rate was used because the model was under-trained at lower learning rates.
8e-5
Pythia-410M
2e-4
8e-5
4e-5
Pythia-1.4B
2e-4
8e-5
4e-5
Appendix
Table 7: Peak learning rates used during post-training.
Figure 18: Example VLM judge reasoning for representative pairwise outcomes.
Figure 19: Pass@1/3/5 performance of the base models on OSM and synthetic test sets. Pass@k is considered achieved only when both the geometric correctness and functional compliance checks pass. Shaded regions represent 95% CIs across 5 evaluation seeds.
Figure 20: Pre-training validation across model scales. (a) Final validation loss achieved by each model. (b) Validation loss throughout pre-training.
Figure 21: SFT validation loss across model scales. (a) Final validation loss achieved by each model. (b) Validation loss throughout training. Shaded regions represent 95% CIs across 3 training seeds.
Figure 22: Reward model validation accuracy across model scales. (a) Final validation accuracy achieved by each model. (b) Validation accuracy throughout reward model training. Shaded regions represent 95% CIs across 3 training seeds.
Figure 23: Learning curves for GRPO reward components across model scales, including geometric correctness, functional compliance, the normalized reward model score, and the soft overlong penalty. Shaded regions show 95% CIs across 3 training seeds.
Figure 24: Pass@1/3/5 performance of the base and fine-tuned models on OSM and synthetic test sets. Pass@k is considered achieved only when both the geometric correctness and functional compliance checks pass. Shaded regions represent 95% CIs across 3 training and 5 evaluation seeds.
Figure 25: Pairwise VLM judge outcomes for each FLOORA GRPO (RM+VR) model against its matched GRPO (VR only) counterpart. Bars show the fraction of prompts where RM+VR wins, ties, or loses. As model size increases, the RM+VR variant is increasingly preferred. Error bars indicate 95% Wilson intervals over the evaluation prompts.
Functional Compliance
Geometric Correctness
Total Reward
Model
@1
@3
@5
@1
@3
@5
@1
@3
@5
FLOORA-0.6B (Qwen3) – Base
34.2 ± 0.4
51.2 ± 0.8
58.8 ± 1.2
27.8 ± 0.3
38.5 ± 0.6
43.2 ± 0.8
23.8 ± 0.1
33.4 ± 0.4
38.1 ± 0.7
FLOORA-1.7B (Qwen3) – Base
35.6 ± 0.4
53.3 ± 0.5
61.2 ± 0.4
28.7 ± 0.1
40.5 ± 0.4
45.5 ± 0.6
24.9 ± 0.3
35.8 ± 0.6
40.7 ± 0.9
FLOORA-70M (Pythia) – Base
26.6 ± 0.4
47.6 ± 0.6
56.8 ± 0.8
4.3 ± 0.2
8.7 ± 0.5
11.1 ± 0.6
4.2 ± 0.2
8.6 ± 0.4
11.0 ± 0.6
FLOORA-160M (Pythia) – Base
26.3 ± 0.3
43.8 ± 0.5
52.3 ± 0.8
13.1 ± 0.2
21.3 ± 0.2
24.9 ± 0.2
12.4 ± 0.1
20.2 ± 0.3
23.7 ± 0.5
FLOORA-410M (Pythia) – Base
34.6 ± 0.4
52.3 ± 0.6
60.6 ± 0.7
26.7 ± 0.3
36.9 ± 0.5
41.5 ± 0.5
23.0 ± 0.1
32.2 ± 0.4
36.8 ± 0.7
Appendix
Table 8: Pass@1/3/5 performance of the base models on the OSM test set, reported with 95% CIs across 5 evaluation seeds. Best value within each model family is shown in bold.
Functional Compliance
Geometric Correctness
Total Reward
Model
@1
@3
@5
@1
@3
@5
@1
@3
@5
FLOORA-0.6B (Qwen3) – Base
90.9 ± 0.9
94.3 ± 0.9
95.2 ± 0.9
96.9 ± 0.4
98.4 ± 0.3
98.5 ± 0.3
90.2 ± 1.0
94.0 ± 1.0
94.9 ± 1.0
FLOORA-1.7B (Qwen3) – Base
91.3 ± 0.9
94.2 ± 0.8
95.0 ± 0.8
97.2 ± 0.4
98.4 ± 0.3
98.5 ± 0.3
90.7 ± 0.9
93.9 ± 0.8
94.7 ± 0.8
FLOORA-70M (Pythia) – Base
70.2 ± 1.0
90.6 ± 0.7
93.8 ± 0.6
32.0 ± 0.6
54.8 ± 0.8
64.0 ± 0.7
31.0 ± 0.6
53.2 ± 0.9
62.2 ± 0.9
FLOORA-160M (Pythia) – Base
79.2 ± 0.6
91.9 ± 0.6
93.8 ± 0.5
69.6 ± 0.5
88.3 ± 0.6
92.3 ± 0.8
66.1 ± 0.3
84.7 ± 0.5
88.9 ± 0.9
FLOORA-410M (Pythia) – Base
91.1 ± 0.7
94.7 ± 0.7
95.6 ± 0.7
96.0 ± 0.3
98.3 ± 0.3
98.4 ± 0.3
89.6 ± 0.7
94.2 ± 0.7
95.1 ± 0.8
Appendix
Table 9: Pass@1/3/5 performance of the base models on the synthetic test set, reported with 95% CIs across 5 evaluation seeds. Best value within each model family is shown in bold.
Functional Compliance
G eometric Correctness
Total Reward
Δ Total Reward (vs Base)
Model
@1
@3
@5
@1
@3
@5
@1
@3
@5
@1
@3
@5
FLOORA-0.6B (Qwen3) – SFT
45.5 ± 1.0
73.4 ± 0.9
82.7 ± 0.8
45.1 ± 0.9
71.3 ± 1.0
79.6 ± 0.8
37.4 ± 1.0
63.9 ± 1.2
73.9 ± 1.1
+13.7
+30.5
+35.8
FLOORA-0.6B (Qwen3) – GRPO (RM+VR)
85.6 ± 0.3
95.6 ± 0.2
97.5 ± 0.2
80.9 ± 0.7
92.3 ± 0.5
94.8 ± 0.5
79.3 ± 0.7
91.3 ± 0.5
94.0 ± 0.5
+55.5
+57.9
+55.9
FLOORA-0.6B (Qwen3) – GRPO (VR only)
84.0 ± 0.7
95.3 ± 0.4
97.2 ± 0.3
76.3 ± 0.5
90.1 ± 0.3
93.5 ± 0.3
74.7 ± 0.7
89.0 ± 0.5
92.5 ± 0.5
+50.9
+55.5
+54.3
FLOORA-1.7B (Qwen3) – SFT
46.1 ± 0.2
73.5 ± 0.3
82.5 ± 0.3
46.6 ± 0.2
72.1 ± 0.3
79.9 ± 0.3
38.8 ± 0.2
64.8 ± 0.3
74.2 ± 0.3
+13.9
+29.0
+33.5
FLOORA-1.7B (Qwen3) – GRPO (RM+VR)
85.8 ± 0.4
95.7 ± 0.2
97.4 ± 0.2
78.2 ± 0.1
90.5 ± 0.2
93.4 ± 0.2
76.9 ± 0.1
89.7 ± 0.2
92.8 ± 0.3
+52.0
+54.0
+52.0
Appendix
Table 10: Pass@1/3/5 performance of the post-trained models on the OSM test set, reported with 95% CIs across 3 training and 5 evaluation seeds. The final column reports percentage-point changes in Total Reward relative to the base checkpoint. Best absolute value within each model family is shown in bold.
Functional Compliance
Geometric Correctness
Total Reward
Δ Total Reward (vs Base)
Model
@1
@3
@5
@1
@3
@5
@1
@3
@5
@1
@3
@5
FLOORA-0.6B (Qwen3) – SFT
57.0 ± 1.3
86.5 ± 0.9
93.6 ± 0.5
49.3 ± 1.5
79.7 ± 1.4
88.8 ± 0.9
48.7 ± 1.5
79.2 ± 1.4
88.5 ± 0.9
-41.5
-14.8
-6.4
FLOORA-0.6B (Qwen3) – GRPO (RM+VR)
97.4 ± 0.2
98.5 ± 0.1
98.6 ± 0.1
96.8 ± 0.2
98.2 ± 0.2
98.3 ± 0.1
96.8 ± 0.2
98.1 ± 0.2
98.3 ± 0.2
+6.6
+4.1
+3.3
FLOORA-0.6B (Qwen3) – GRPO (VR only)
97.9 ± 0.2
98.4 ± 0.1
98.5 ± 0.1
97.5 ± 0.1
98.2 ± 0.1
98.3 ± 0.1
97.3 ± 0.1
98.1 ± 0.2
98.2 ± 0.2
+7.1
+4.0
+3.2
FLOORA-1.7B (Qwen3) – SFT
57.1 ± 0.3
86.0 ± 0.3
92.9 ± 0.3
50.7 ± 0.3
80.3 ± 0.3
88.9 ± 0.2
50.1 ± 0.3
79.9 ± 0.3
88.7 ± 0.2
-40.6
-14.0
-6.0
FLOORA-1.7B (Qwen3) – GRPO (RM+VR)
97.6 ± 0.2
98.5 ± 0.1
98.6 ± 0.1
97.0 ± 0.2
98.2 ± 0.1
98.3 ± 0.1
97.0 ± 0.2
98.1 ± 0.1
98.2 ± 0.1
+6.3
+4.2
+3.6
Appendix
Table 11: Pass@1/3/5 performance of the post-trained models on the synthetic test set, reported with 95% CIs across 3 training and 5 evaluation seeds. The final column reports percentage-point changes in Total Reward relative to the base checkpoint. Best absolute value within each model family is shown in bold.
Functional Compliance
Geometric Correctness
Total Reward
Model
@1
@3
@5
@1
@3
@5
@1
@3
@5
FLOORA-0.6B (Qwen3) – GRPO (RM+VR)
84.8
95.3
97.2
79.5
91.6
94.0
78.1
90.8
93.5
Claude Opus 4.8
46.1
73.6
83.0
36.4
60.8
72.0
24.7
45.0
55.6
Claude Sonnet 4.6
19.3
38.4
49.4
10.0
20.2
26.8
5.1
10.7
14.5
Gemini 2.5 Flash Image
37.6
62.0
71.7
25.5
44.3
53.8
11.1
20.3
26.0
Gemini 3.5 Flash
52.6
76.6
83.9
42.0
69.1
79.7
33.3
57.4
68.2
Appendix
Table 12: Pass@1/3/5 comparison of the FLOORA-0.6B (Qwen3) – GRPO (RM+VR) and frontier baselines on the OSM test set. Best absolute value is shown in bold. Results are obtained on a single seed.
Functional Compliance
Geometric Correctness
Total Reward
Model
@1
@3
@5
@1
@3
@5
@1
@3
@5
FLOORA-0.6B (Qwen3) – GRPO (RM+VR)
97.4
98.3
98.4
96.5
97.8
98.0
96.5
97.7
97.9
Claude Opus 4.8
66.5
92.2
97.1
38.6
62.9
72.9
30.9
54.2
64.5
Claude Sonnet 4.6
39.0
70.2
82.5
8.1
15.5
20.4
6.9
13.5
17.9
Gemini 2.5 Flash Image
64.1
90.7
96.7
15.6
29.7
38.0
12.9
24.0
30.6
Gemini 3.5 Flash
75.3
96.4
99.2
39.9
66.2
77.0
39.1
65.2
76.0
Appendix
Table 13: Pass@1/3/5 comparison of the FLOORA-0.6B (Qwen3) – GRPO (RM+VR) model and frontier baselines on the synthetic test set. Best absolute value is shown in bold. Results are obtained on a single seed.
Figure 26: Pass@k comparison of the FLOORA-0.6B (Qwen3) GRPO (RM+VR) model and frontier baselines, where success requires passing both geometric and functional checks. Our model substantially outperforms all evaluated frontier baselines.
Figure 27: Pairwise VLM judge outcomes for the FLOORA-0.6B (Qwen3) GRPO (RM+VR) model against frontier baselines. Bars show the fraction of prompts for which the FLOORA model wins, ties, or loses. Our model is preferred over all evaluated frontier baselines. Error bars indicate 95% Wilson intervals over the evaluation prompts.
Figure 28: Qualitative comparison between the FLOORA-0.6B (Qwen3) GRPO (RM+VR) model and frontier models on the OSM test set.
Figure 29: Qualitative comparison between the FLOORA-0.6B (Qwen3) GRPO (RM+VR) model and frontier models on the synthetic test set.
Figure 30: Qualitative comparison between the FLOORA-0.6B (Qwen3) Base, SFT, GRPO (RM+VR), and GRPO (VR only) models on the OSM test set.
Figure 31: Qualitative comparison between the FLOORA-0.6B (Qwen3) Base, SFT, GRPO (RM+VR), and GRPO (VR only) models on the synthetic test set.
Figure 32: Ablation of the RM normalization scale α during GRPO training with the FLOORA-0.6B (Qwen3) model. Moderate scales α∈[1,4] stabilize training and achieve high verifiable rewards. Shaded regions show 95% CIs across 3 training seeds.
Setting
Reward Model
Geometric Correctness
Functional Compliance
Soft Overlong Penalty
No normalization
28.719 ± 5.376
0.383 ± 0.192
0.761 ± 0.224
0.000 ± 0.000
α=1
0.994 ± 0.002
0.853 ± 0.026
0.951 ± 0.016
0.000 ± 0.000
α=2
0.990 ± 0.012
0.846 ± 0.017
0.944 ± 0.002
0.000 ± 0.000
α=3
0.988 ± 0.011
0.851 ± 0.027
0.948 ± 0.007
0.000 ± 0.000
α=4
0.982 ± 0.015
0.850 ± 0.017
0.943 ± 0.012
0.000 ± 0.000
α=10
0.954 ± 0.049
0.771 ± 0.032
0.922 ± 0.026
0.000 ± 0.000
Appendix
Table 14: Final RM and verifiable reward scores for different normalization scales α , reported with 95% CIs across 3 training seeds.
Figure 33: Tokenizer ablation for FLOORA-0.6B (Qwen3) over optimizer steps. Shaded regions show 95% CIs across 3 training seeds, except pre-training which uses a single seed.
Figure 34: Tokenizer ablation during GRPO (RM+VR) training with the FLOORA-0.6B (Qwen3) model. Shaded regions show 95% CIs across 3 training seeds.
Total Reward
Model
@1
@3
@5
FLOORA-0.6B (Qwen3) – Base [DSL tok.]
23.8 ± 0.1
33.4 ± 0.4
38.1 ± 0.7
FLOORA-0.6B (Qwen3) – Base [Orig. tok.]
19.6 ± 0.4
27.3 ± 0.5
31.1 ± 0.5
FLOORA-0.6B (Qwen3) – GRPO (RM+VR) [DSL tok.]
79.3 ± 0.7
91.3 ± 0.5
94.0 ± 0.5
FLOORA-0.6B (Qwen3) – GRPO (RM+VR) [Orig. tok.]
71.6 ± 2.9
83.2 ± 2.0
86.8 ± 1.6
FLOORA-0.6B (Qwen3) – GRPO (VR only) [DSL tok.]
74.7 ± 0.7
89.0 ± 0.5
92.5 ± 0.5
Appendix
Table 15: Pass@1/3/5 performance of FLOORA-0.6B (Qwen3) across training stages using the original and DSL tokenizers on the OSM test set, reported with 95% CIs across 3 training and 5 evaluation seeds. Best absolute value is shown in bold.
Total Reward
Model
@1
@3
@5
FLOORA-0.6B (Qwen3) – Base [DSL tok.]
90.2 ± 1.0
94.0 ± 1.0
94.9 ± 1.0
FLOORA-0.6B (Qwen3) – Base [Orig. tok.]
88.9 ± 0.7
95.1 ± 0.6
96.2 ± 0.6
FLOORA-0.6B (Qwen3) – GRPO (RM+VR) [DSL tok.]
96.8 ± 0.2
98.1 ± 0.2
98.3 ± 0.2
FLOORA-0.6B (Qwen3) – GRPO (RM+VR) [Orig. tok.]
93.5 ± 1.3
97.6 ± 0.5
98.4 ± 0.3
FLOORA-0.6B (Qwen3) – GRPO (VR only) [DSL tok.]
97.3 ± 0.1
98.1 ± 0.2
98.2 ± 0.2
Appendix
Table 16: Pass@1/3/5 performance of FLOORA-0.6B (Qwen3) across training stages using the original and DSL tokenizers on the synthetic test set, reported with 95% CIs across 3 training and 5 evaluation seeds. Best absolute value is shown in bold.
Figure 35: Human feedback data-scale ablation for FLOORA-0.6B (Qwen3) across SFT and GRPO (RM+VR) on the OSM and synthetic test sets. Shaded regions show 95% CIs across 3 training seeds and 5 evaluation seeds.
Indoor scene layout generation is a challenging task in interior design. Existing methods often oversimplify the task by reducing room conditions to coarse 3D bounding boxes and neglecting structural elements such as doors and windows. More fundamentally, many prior approaches formulate spatial reasoning as direct coordinate prediction, thereby casting interior layout design as continuous regression over raw geometric parameters, which hinders the model from learning the underlying reasoning logic of intelligent layout design. We propose \textbf{LayoutDSL}, a novel LLM-based framework for learning an interior layout policy in a domain-specific language (DSL) action space. The DSL provides an explicit symbolic representation of layout information and serves as a structured action space for layout reasoning, where each action corresponds to an interpretable design decision. Under this DSL-based policy learning paradigm, we construct 3D-FrontDSL, a dataset of room-structure annotations paired with synthetic DSL action sequences for supervised fine-tuning. To promote a more generalizable and scalable policy with verifiable feedback, we design rewards grounded in interior design principles and physical plausibility, and optimize the policy via reinforcement learning. Extensive experiments demonstrate that LayoutDSL substantially improves spatial plausibility and design logicality over strong baselines and existing methods.
Furnished floor plans support real-estate visualization, interior design, and architectural workflows, yet automatic furnishing remains challenged by limited real-world data and the need to satisfy interacting geometric and functional constraints. We ask whether professional furnishing knowledge can be learned from real floor plans using a pretrained model, enabling direct constraint-aware layout generation without relying on costly iterative agentic inference. We introduce AntPlan, a curated dataset of 505 real professional architectural floor plans with dense furniture annotations spanning 92 object classes and ten residential room categories, and Architect-Ant, a framework for generating furniture layouts. Architect-Ant represents layouts with an editable coordinate-based DSL and first learns professional furnishing patterns through supervised fine-tuning. It is then optimized with GRPO using a Layout Rule Score (LRS) that aggregates geometric and functional constraints derived from professional plans, providing outcome-level supervision without prescribed reasoning traces. Experiments against diverse state-of-the-art baselines show that Architect-Ant combines low geometric violation rates with high functional completeness, while qualitative results more closely reflect real-world residential furnishing patterns. The resulting layouts remain object-level editable and can be converted into 3D scenes.
Fedor Rodionov, Aleksandar Cvejic, Michael Birsak +2
King Abdullah University of Science and Technology (KAUST), Saudi Arabia · Miami University, United States of America
Interior design is a requirements-to-visual-plan generation process that must simultaneously satisfy verifiable spatial feasibility and comparative aesthetic preferences. While recent multimodal large language models (MLLMs) offer a unified foundation for interpreting user intent and producing design rationales, our empirical analysis reveals a persistent contradiction in real-world deployment: MLLMs often produce layouts that are unbuildable and aesthetically inconsistent. These findings indicate that simply adding in-domain text is insufficient; effective interior design requires an alignment mechanism that separates hard constraints from soft preferences and coordinates them during optimization. To address this, we propose Design-MLLM, a reinforcement alignment framework that optimizes a feasibility-first preference objective via a dual-branch, aesthetic-oriented reward. Specifically, Design-MLLM (i) explicitly evaluates spatial feasibility using programmatic constraint checks, (ii) assesses aesthetic preference only among feasible candidates to avoid visually appealing but unexecutable shortcuts, and (iii) performs group-relative optimization to obtain stable preference signals. Through this process, Design-MLLM learns a controllable policy that consistently selects and generates solutions that are both executable and aesthetically coherent, rather than occasionally producing visually appealing but infeasible designs. Extensive experiments on various benchmark datasets demonstrate the advantages of Design-MLLM.
Yuxuan Yang, Xiaotong Mao, Jingyao Wang
Nanjing Forestry University, Nanjing, China · Université de Lorraine, Nancy, France · Laboratoire Réactions et Génie des Procédés, Nancy, France +2