Skip to content
AI.info

Research

Stable diffusion models reveal a persisting human and AI gap in visual creativity

Overview Research area: Human versus AI creativity, specifically visual (image-generation) creativity rather than language-based creative tasks. Technical level: Beginner-Friendly. The abstract descri

arXiv
2511.16814
Published
2025-11-20
Authors
Silvia Rondini, Claudia Alvarez-Martin, Paula Angermair-Barkai, Olivier Penacchio, M. Paz, Matthew Pelowski, Dan Dediu, Antoni Rodriguez-Fornells, Xim Cerda-Company

AI summary

Overview

Research area: Human versus AI creativity, specifically visual (image-generation) creativity rather than language-based creative tasks.

Technical level: Beginner-Friendly. The abstract describes a behavioral comparison study — human participants, an image-generation model, and two sets of raters — with no model architecture, training, or computational detail.

Scope: The paper asks how the creative output of an image-generating AI compares to that of human artists and non-artists, and whether giving the AI more human guidance closes the gap.

What This Paper Is About

Recent work has suggested that large language models can match human creative performance on divergent-thinking tasks, but creativity in the visual domain has received far less attention. This study sets out to test that comparison directly by generating images both with humans and with an image-generation AI, then having people and an AI model judge how creative the results are. The goal is to see whether the apparent human-AI parity in language-based creativity carries over to images, or whether visual creativity behaves differently.

Key Contributions

  1. Extends the human-versus-AI creativity comparison from language-centered tasks (such as divergent thinking) into the visual domain, which the abstract identifies as underexplored.
  2. Compares multiple populations at once — Visual Artists, Non Artists, and AI-generated images — instead of a simple human-versus-machine binary.
  3. Uses two prompting conditions that vary how much human input goes into the AI generation (high human input for "Human Inspired," low human input for "Self Guided"), treating the amount of human guidance as a testable variable.
  4. Evaluates the resulting images with both human raters (N=255) and GPT-4o, allowing a direct comparison of how humans and an AI judge creativity.

Main Findings

  • A clear creativity gradient exists: Visual Artists were rated most creative, followed by Non Artists, then Human Inspired generative AI, and finally Self Guided generative AI.
  • Human guidance substantially helps the AI: the Human Inspired condition produced markedly more creative output than the Self Guided condition, bringing the AI's results close to those of Non Artists.
  • The gap does not fully close: even with high human input, the AI did not reach the level of Visual Artists, and its lowest-input condition ranked last.
  • Human and AI raters judge creativity very differently: the abstract reports "vastly different creativity judgment patterns" between the human raters and GPT-4o, meaning the two evaluation sources did not agree on what counted as creative.
  • Visual creativity may be a special case for GenAI: the authors argue that, unlike language-centered tasks, visual domains rely on perceptual nuance and contextual sensitivity — distinctly human capacities that may not transfer readily from language models.

Methodology in Plain English

The researchers had two groups of people — Visual Artists and Non Artists — produce images, and also had an image-generation AI model produce images under two different prompting setups. In the "Human Inspired" condition, people contributed a high level of input to steer the generation; in the "Self Guided" condition, human input was low. All the resulting images were then rated for creativity by a large panel of human raters (255 of them) and separately by GPT-4o. Comparing the ratings across the four sources of images (Visual Artists, Non Artists, Human Inspired AI, Self Guided AI) produced the creativity ranking described above.

Note that the abstract does not report how many images were generated, how the raters were recruited or instructed, what rating scale was used, or any statistical detail — those specifics are not available in the abstract.

Why This Matters

The study pushes back on the idea that AI creativity is simply "solved" by progress in language models. It suggests that visual creativity may not be a straightforward transfer of linguistic capability, and it also raises a methodological warning: human and AI judges of creativity can disagree sharply, which matters for anyone using an AI model to evaluate creative work.

Potential applications and implications (as suggested by the findings):

  • Creative tool design: knowing that more human direction improves AI output argues for co-creation interfaces rather than fully autonomous generation.
  • Evaluation practice: the rater disagreement is a caution for teams that rely on AI models as automatic judges of creative quality.
  • Creative industries and artist workflows: the persistent artist-versus-AI gap is relevant to debates about AI's role in professional visual work.
  • Benchmarking generative models: the study offers a human-comparison framing for assessing image models, not just language models.

Industry relevance: companies building or deploying image-generation tools have a direct stake in how much human input their systems need and in whether automated quality or creativity scoring can be trusted. The result that human guidance materially improves output is relevant to product design, and the rater disagreement is relevant to any pipeline that uses model-based evaluation.

Future Directions

  • Does the gap hold as models change? The abstract does not specify which image model or version was used, leaving open whether newer systems would narrow the distance to Visual Artists.
  • Can perceptual nuance and contextual sensitivity be built in? The authors' interpretation raises the question of whether such capacities can be developed for visual models rather than inherited from language.
  • Why do human and AI raters disagree? Understanding the source of the divergent judgment patterns — and which judgments should be trusted — is an open problem for creativity research and for AI-as-judge evaluation.
  • How should human guidance be structured? Since the degree of human input changed the results substantially, the optimal balance between human direction and machine autonomy is unresolved.

Target Audience

This paper is most useful for creativity and cognitive-science researchers comparing human and machine performance; AI researchers and engineers working on image generation and on automated evaluation of creative output; designers, artists, and creative-industry professionals assessing AI's current capabilities; and anyone interested in whether progress in language models transfers to visual creativity.

Authors’ abstract

While recent research suggests Large Language Models match human creative performance in divergent thinking tasks, visual creativity remains underexplored. This study compared image generation in human participants (Visual Artists and Non Artists) and using an image generation AI model (two prompting conditions with varying human input: high for Human Inspired, low for Self Guided). Human raters (N=255) and GPT4o evaluated the creativity of the resulting images. We found a clear creativity gradient, with Visual Artists being the most creative, followed by Non Artists, then Human Inspired generative AI, and finally Self Guided generative AI. Increased human guidance strongly improved GenAI's creative output, bringing its productions close to those of Non Artists. Notably, human and AI raters also showed vastly different creativity judgment patterns. These results suggest that, in contrast to language centered tasks, GenAI models may face unique challenges in visual domains, where creativity depends on perceptual nuance and contextual sensitivity, distinctly human capacities that may not be readily transferable from language models.

Read the original paper