cs.CLJul 24, 2026

Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity

Authors: Pengzhao Lyu, Yeun Joon Kim, Hanlin Xiao, Yingyue Luna Luan

Organizations: Judge Business School, University of Cambridge, Cambridge, CB2 1AG, United Kingdom · Institute of Metabolic Science, School of Clinical Medicine, University of Cambridge, CB2 0QQ, United Kingdom · Department of Computer Science, University of Manchester, M13 9PL, United Kingdom · Manchester Institute of Biotechnology, University of Manchester, M1 7DN, United Kingdom · School of Business, University of Queensland, Brisbane, QLD 4067, Australia

Abstract

Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from human judgments. Across three studies and six widely used LLMs, we addressed this gap by identifying the standards underlying LLM creativity evaluation and examining their downstream implications. Study 1 showed that LLMs generally relied on a narrower subset of human creativity evaluation standards. Convergence with human standards was strongest in the novelty dimension, whereas divergence was clearest in the contextual dimension, which captures social, market, and reputational information. Moreover, each LLM exhibited distinct, model-specific standards that varied substantially in breadth. These differences in evaluation standards were reflected in actual creativity judgments. Study 2 (N = 1,103 ideas) showed that LLM evaluations were moderately correlated with human evaluations, and individual LLMs with broader standards better distinguished ideas humans judged as more versus less creative. Study 3 (N = 1,195) showed that LLMs were less sensitive to contextual information: such information significantly altered human creativity ratings but left LLM ratings largely unchanged. Together, our findings help explain the mixed evidence on LLM-human alignment, showing that alignment depends on the evidence a judgment demands and the standards each model applies. LLMs may resemble humans when evaluations emphasize intrinsic qualities such as novelty, yet diverge when judgments require contextual information. Selecting an LLM evaluator is therefore a consequential decision: different models, applying different standards, recognize different ideas as creative.

Explore similar work

CardsList