Design Creativity Bench: Measuring creativity in LLM-Generated UI
Organizations: Kombai Inc.
Abstract
As leading LLMs improve on capability evaluations, their limitations in producing creative outputs on design tasks remain insufficiently characterised. Our work introduces Design Creativity Bench, a benchmark that evaluates diversity and appropriateness in UI designs. It measures distinctiveness among models on the same prompt (originality), how much a model's designs change between two prompts for the same UI goal in different product domains (creative range), and the share of a brief's acceptance criteria each design meets (appropriateness). Originality is 0.592 for same-prompt design pairs from different models (95% CI [0.582, 0.602]), far below the 0.764 for same-prompt human-model pairs (95% CI [0.751, 0.778]). Creative range is 0.581 across models (95% CI [0.567, 0.597]), against 0.902 for human designs (95% CI [0.884, 0.919]). Appropriateness is above 90% for every model, and the best model reaches 99.2%, slightly above the 98.0% for human designs. Our work shows that the default output of LLMs, though generally appropriate, is substantially more repetitive than the human baseline. This calls for strong measures to address the issue.
Figures & tables
| Model (effort) | Provider |
|---|---|
| Qwen3.8 Max (max) | Alibaba |
| Claude Fable 5.1 (max) | Anthropic |
| Claude Opus 5.5 (xhigh) † | Anthropic |
| DeepSeek V4.1 Flash (max) | DeepSeek |
| Gemini 3.8 Flash (high) | |
| Muse Spark 1.3 (xhigh) | Meta |
| Rank | Model | Originality | 95% CI | Cost (USD) | Time (min) |
|---|---|---|---|---|---|
| 1 | GPT-6 Astra | 0.646 | [0.631, 0.661] | 2.564 | 13.8 |
| 2 | Grok 4.7 | 0.643 | [0.631, 0.655] | 0.593 | 18.8 |
| 3 | GPT-6.1 Sol | 0.642 | [0.628, 0.657] | 0.460 | 15.8 |
| 4 | Gemini 3.8 Flash | 0.637 | [0.621, 0.653] | 0.123 | 0.9 |
| 5 | Qwen3.8 Max | 0.623 | [0.609, 0.636] | 0.309 | 18.0 |
| 6 | Grok 4.6 | 0.591 | [0.579, 0.602] | 0.100 | 3.5 |
| Rank | Model | Listing | Dashboard | Form flow | Table | Detail | Settings | Overall | 95% CI |
|---|---|---|---|---|---|---|---|---|---|
| 1 | GLM-5.3 | 0.63 | 0.63 | 0.68 | 0.65 | 0.61 | 0.59 | 0.634 | [0.607, 0.662] |
| 2 | Gemini 3.8 Flash | 0.63 | 0.62 | 0.68 | 0.63 | 0.60 | 0.58 | 0.628 | [0.602, 0.652] |
| 3 | Grok 4.6 | 0.63 | 0.65 | 0.64 | 0.61 | 0.61 | 0.54 | 0.625 | [0.599, 0.650] |
| 4 | Qwen3.8 Max | 0.65 | 0.62 | 0.60 | 0.60 | 0.62 | 0.59 | 0.617 | [0.597, 0.639] |
| 5 | Kimi K3 | 0.60 | 0.59 | 0.62 | 0.63 | 0.67 | 0.62 | 0.615 | [0.593, 0.636] |
| 6 | GLM-5.3-Flash | 0.60 | 0.63 | 0.58 | 0.63 | 0.62 | 0.60 | 0.611 | [0.590, 0.632] |
| Rank | Model | Mean pass rate | Critical criteria met |
|---|---|---|---|
| 1 | Claude Opus 5.5 | 99.2% | 99.9% |
| 2 | GPT-6 Astra | 99.1% | 99.7% |
| 3 | GPT-6.1 Sol | 98.8% | 99.6% |
| 4 | Claude Fable 5.1 | 98.5% | 98.9% |
| 5 | GPT-5.6 Sol | 98.2% | 99.3% |
| 6 | Grok 4.7 | 97.8% | 98.5% |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Space | Pair acc. | AUC | Nuis. | Str. | Unrel. |
|---|---|---|---|---|---|
| Gemini | .645 | .640 | .980 | .966 | .074 |
| Gemini, projected | .732 | .722 | .945 | .879 | |
| UIClip | .661 | .649 | .996 | .990 | .008 |
| UIClip, projected | .730 | .724 | .954 | .862 |
| Key | Page role | Weight |
|---|---|---|
| bg | Page background | 1.0 |
| surface | Panel or card background | 1.0 |
| ink | Primary text | 1.0 |
| ink_muted | Secondary or muted text | 0.5 |
| accent | Primary interactive accent | 1.0 |
| border_color | Hairline or divider colour | 0.5 |
| Participant | Accuracy | AUC | Cases |
|---|---|---|---|
| Grok 4.6 (best single judge) | .885 | — | 78 |
| Six-model judge bench | .868 | — | 76 |
| Hybrid (segments, projection), selected | .825 | .869 | 80 |
| UIClip (whole image, projection) | .787 | .858 | 80 |
| UIClip (segments, projection) | .762 | .855 | 80 |
| UIClip (whole image) | .738 | .765 | 80 |
| Configuration | Passing ladders |
|---|---|
| Selected hybrid + colour | 18/18 |
| Hybrid + colour, geometric decay | 17/18 |
| Hybrid + thirteen theme fields, decay | 16/18 |
| Graded-colour projections | 16/18 |
| Gemini projected segments | 16/18 |
| UIClip projected segments | 17/18 |
| # | Model | Overall | Orig. | Range | Appr. (%) | Cost ($) | Pareto |
|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra | 65.5 | 0.646 | 0.472 | 99.1 | 2.564 | |
| 2 | GPT-6.1 Sol | 63.7 | 0.642 | 0.487 | 98.8 | 0.460 | |
| 3 | Grok 4.7 | 62.2 | 0.643 | 0.603 | 97.8 | 0.593 | |
| 4 | Gemini 3.8 Flash | 60.6 | 0.637 | 0.628 | 97.5 | 0.123 | |
| 5 | Qwen3.8 Max | 58.1 | 0.623 | 0.617 | 97.6 | 0.309 | |
| 6 | GPT-5.6 Sol | 51.3 | 0.587 | 0.505 | 98.2 | 0.598 |
| ID | Category | Critical | Criterion |
|---|---|---|---|
| C01 | Scope | ✓ | The page is a mortgage preapproval application screen where the applicant enters details, not a marketing landing page, rate-comparison page, or sign-in page. |
| C02 | Scope | The page shows that the application is a multi-step flow and marks which step is current. | |
| C03 | Content | The brand name Hearthline appears as the app’s identity. | |
| C04 | Content | ✓ | The borrower is shown as Maya Patel. |
| C05 | Content | ✓ | The property address is shown as 18 Alder Street. |
| C06 | Content | ✓ | The purchase price is shown as exactly $420,000. |