cs.HCOct 8, 2026

Design Creativity Bench: Measuring creativity in LLM-Generated UI

Authors: Aman Rusia, Abhijit Bhole, Prashank Gupta, Dipanjan Dey

Organizations: Kombai Inc.

Abstract

As leading LLMs improve on capability evaluations, their limitations in producing creative outputs on design tasks remain insufficiently characterised. Our work introduces Design Creativity Bench, a benchmark that evaluates diversity and appropriateness in UI designs. It measures distinctiveness among models on the same prompt (originality), how much a model's designs change between two prompts for the same UI goal in different product domains (creative range), and the share of a brief's acceptance criteria each design meets (appropriateness). Originality is 0.592 for same-prompt design pairs from different models (95% CI [0.582, 0.602]), far below the 0.764 for same-prompt human-model pairs (95% CI [0.751, 0.778]). Creative range is 0.581 across models (95% CI [0.567, 0.597]), against 0.902 for human designs (95% CI [0.884, 0.919]). Appropriateness is above 90% for every model, and the best model reaches 99.2%, slightly above the 98.0% for human designs. Our work shows that the default output of LLMs, though generally appropriate, is substantially more repetitive than the human baseline. This calls for strong measures to address the issue.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AGC-Bench: Measuring Artificial General Creativity

    Jul 1, 2026Roger Beaty, Vijeta Deshpande, Clin K. Y. Lai +9Creativity AssessmentLLM Evaluation

  2. How LLMs See Creativity: Zero-Shot Scoring of Visual Creativity with Interpretable Reasoning

    Jun 29, 2026William Orwig, Roger E. BeatyCreativity AssessmentLLM-as-a-Judge

  3. Automated Creativity Evaluation of Language Models Across Open-Ended Tasks

    Jun 10, 2026Min Sen Tan, Zachary Kit Chun Choy, Syed Ali Redha Alsagoff +4Creativity AssessmentLLM Evaluation