cs.CLSep 12, 2026

Measuring the Creativity of Frontier LLMs in Automated Research

Authors: Yiheng Zhao, Mengzhuo Chen, Chengming Hu, Pengyi Liao, Yihan Huang, Yiran Pang

Organizations: Concordia University, Montreal, Canada · Independent Researcher · Mcgill University, Montreal, Canada · Florida Atlantic University, Boca Raton, USA

Abstract

Frontier LLMs are increasingly capable of conducting automated research, yet their creativity in this setting has not been systematically evaluated. We propose a set of metrics to evaluate creativity along the two dimensions of valueness and novelty. Valueness assesses whether each proposed idea is useful, while novelty is evaluated from three perspectives: whether the same idea has appeared before (Exact-Match P-Novelty), whether the modified variable or variable combination has been explored before (Variable-level P-Novelty), which reflects the breadth of research-space exploration, and whether the proposed idea is explicitly attributed to external knowledge in the model's reasoning (H-Novelty). Our evaluation shows that the models achieve relatively similar Valueness and Exact-Match P-Novelty scores, while differing substantially in Variable-level P-Novelty. H-Novelty is also consistently high among the models for which it can be evaluated. Notably, further correlation and idea-level performance analyses reveal a strong positive correlation between Variable-level P-Novelty and research performance.

Figures & tables

Explore similar work

CardsList
  1. Automated Creativity Evaluation of Language Models Across Open-Ended Tasks

    Jun 10, 2026Min Sen Tan, Zachary Kit Chun Choy, Syed Ali Redha Alsagoff +4CreativityLarge Language Model Evaluation

  2. Can LLM Agents Discover? Evaluating Creativity on ML Engineering Tasks

    Aug 30, 2026Shitanshu Bhushan, Yunxiang Zhang, Lu WangCreativity

  3. CreativityPrism: A Cross-Domain Evaluation Framework for Large Language Model Creativity

    Oct 23, 2025Zhaoyi Joey Hou, Bowei Alvin Zhang, Yining Lu +9CreativityLarge Language Model Evaluation