AI and the New Science of Measuring Human Creativity

AI and the New Science of Measuring Human Creativity

In 2022 and 2023, large language models and image-generation systems began producing text, art, music, and code that, in blind evaluation studies, judges frequently could not reliably distinguish from human-created work — and in a growing number of specific, narrow domains, rated as more creative or higher quality than human-produced comparison work. This development has done more than generate headlines and anxiety about job displacement; it has created a genuinely novel scientific opportunity and a genuinely novel set of methodological challenges for the psychological study of human creativity itself, a field this series has examined at length through the lens of divergent thinking, neuroscience, and personality research.

AI as a New Instrument for Studying Human Creativity

Perhaps counterintuitively, one of the more scientifically productive applications of AI systems in this space has been using them not to replace human creativity research but to serve as a novel measurement and comparison instrument within it. Several research groups have used large language models to score and evaluate human-generated responses on standard divergent thinking measures, including variants of the Alternative Uses Test discussed in this series’ divergent thinking article, at a speed and consistency that human expert raters — the traditional and considerably more labor-intensive gold standard for scoring these open-ended creativity assessments — cannot practically match at scale. Research validating these AI-based scoring approaches, including work by creativity researchers including Roger Beaty’s research group, has found reasonably strong correlations between AI-generated originality scores and traditional human expert ratings on the same responses, suggesting language models can serve as a genuinely useful, scalable research tool for creativity researchers, potentially allowing far larger-sample studies of divergent thinking than were previously practical given the traditional need for extensive, time-consuming human expert scoring of every individual response.

This is a methodologically significant development in its own right, independent of the broader and more publicly discussed question of AI’s own creative capacities, because sample size limitations have historically constrained the statistical power of much divergent thinking and creativity research — the kind of large-sample, rigorously controlled studies that have proven so valuable in other areas of psychology reviewed throughout this series have been comparatively rarer in creativity research specifically, in part due to the practical cost and difficulty of large-scale expert human scoring of open-ended creative responses.

Do AI Systems Themselves Produce “Creative” Output? A Definitional Problem

The more publicly prominent and more contested question concerns whether AI systems themselves should be considered creative in any meaningful psychological sense, and this question turns out to hinge substantially on which specific definition of creativity, among several competing definitions in the academic literature, one adopts. Applying the standard psychometric framework discussed in this series’ divergent thinking article — Guilford’s components of fluency, flexibility, originality, and elaboration — several published studies administering standard divergent thinking tests directly to large language models have found that these systems can generate responses scoring as high as, and in several specific studies higher than, average human performance on originality and fluency measures specifically, a genuinely striking empirical finding regardless of how one interprets its deeper significance.

However, a substantial number of creativity researchers, including prominent voices such as Margaret Boden, whose extensive work on computational creativity predates the current generation of large language models by decades, have argued that fluency and originality on a standardized test represent an incomplete and potentially misleading operationalization of creativity as a broader human psychological and social phenomenon. Boden’s influential framework distinguishes between “combinatorial creativity” (novel combinations of existing, familiar ideas), “exploratory creativity” (generating novel ideas within the accepted boundaries and rules of an existing conceptual space), and “transformational creativity” (the rarer and more consequential capacity to fundamentally alter or transcend the boundaries of the conceptual space itself, producing ideas that would have been literally impossible or inconceivable within the prior framework) — with Boden and other researchers in this tradition generally arguing that current large language model architectures, which are fundamentally trained to statistically model and recombine patterns present within an enormous existing training corpus of human-generated text, are considerably better characterized as demonstrating combinatorial and exploratory creativity than the rarer, more transformational creativity historically associated with the most celebrated cases of human scientific and artistic genius examined throughout this series, including Einstein’s development of relativity theory or the emergence of entirely new artistic movements.

The Domain-Expert Evaluation Studies

Beyond standardized divergent thinking tests, several more recent studies have directly compared AI-generated creative output against human expert-produced work within specific professional creative domains, using blind evaluation by independent domain experts unaware of which submissions were AI-generated versus human-generated. Results across these studies have varied considerably by domain and specific task: studies of AI-generated short fiction, poetry, and visual art have found blind expert evaluators frequently unable to reliably distinguish AI from human output at rates significantly above chance, and in some specific studies rating AI-generated work as comparably or even more creative than human-generated comparison work on standardized creativity rating scales — a genuinely striking and, for many creativity researchers, unsettling empirical finding that has generated substantial ongoing debate about what exactly these rating scales are capturing and whether they adequately capture the fuller, more socially and historically embedded sense of “creativity” that terms like genius have traditionally implied throughout the rest of this series.

The Missing Ingredients: Intentionality, Stakes, and Historical Situatedness

A recurring theme across the more skeptical scholarly response to these striking AI-creativity findings concerns several dimensions of human creative achievement that standardized creativity assessments, whether applied to humans or AI systems, may simply not capture. Dean Simonton’s historiometric research on genius, discussed extensively elsewhere in this series, has emphasized that historically significant creative achievement is defined substantially by its social and historical reception and influence — a scientific theory or artistic work becomes historically significant not merely by being novel and internally coherent, but by being recognized, adopted, and built upon by a subsequent community of practitioners and by demonstrably reshaping the future trajectory of its field, a criterion that, by definition, cannot be assessed for very recently AI-generated creative output regardless of how favorably it scores on standardized, immediate creativity assessments, since historical influence can only be measured retrospectively, often across a span of years or decades.

A further and more philosophically contested dimension concerns intentionality and genuine stakes: human creative work, including the scientific and artistic genius this series has examined throughout, is typically produced by an agent with genuine, felt personal motivation, genuine risk of failure and its associated psychological and material consequences, and — per this series’ articles on flow states, obsessive passion, and mental illness’s complex relationship to creativity — often substantial psychological investment and cost. Whether these factors are constitutively necessary for something to count as “genuine” creativity, or whether they’re better understood as simply common accompanying features of human creative work that are separable in principle from creativity’s actual cognitive substance, remains a genuinely unresolved philosophical question that current AI-creativity research has brought into much sharper and more practically urgent focus than it previously occupied within a debate that had, until recently, remained largely theoretical.

What This Means for the Science of Genius

Regardless of how the deeper philosophical questions about AI’s own creative status are eventually resolved, the emergence of capable AI creativity-comparison systems has already meaningfully advanced the science of human creativity in at least one clear and less contested way: by providing creativity researchers with a novel comparison baseline and measurement tool that has helped clarify, with unusual empirical precision, exactly which components of standard creativity assessments can apparently be replicated through sophisticated statistical pattern recombination over an enormous training corpus, and — by process of elimination and continued scholarly debate over the specific missing ingredients discussed above — which components of what the popular concept of “genius” has historically referred to may depend on the kind of intentional, historically situated, socially embedded, and personally consequential engagement that current AI systems, however impressive their output, do not straightforwardly possess in the same sense that human creators historically have. This remains among the most actively contested and rapidly evolving areas touched on anywhere in this series, and readers should expect the specific empirical findings and the surrounding scholarly consensus discussed in this article to continue developing rapidly as both AI capabilities and the research methodologies used to study them continue to advance.