The gap in the scores
A preprint study found that most automated measures used to assess creativity lined up only weakly with human judgments. Most reported correlations fell between -0.2 and 0.2, a range the paper describes as very weak to negligible.
The researchers compared seven automated metrics with human ratings across 11 dimensions. Each dimension was scored on a five-point scale from 1, Poor, to 5, Excellent. Spearman's rank correlation was used to compare the automated metrics with the ordinal human ratings.
The Creativity Index had almost no relationship with human-rated Creativity: its correlation was 0.07, and it did not exceed 0.11 for any other dimension. CR-POS had the strongest positive correlation reported, at 0.29 with human-rated Elaboration, but that association was still weak.
Perplexity also showed weak negative associations with human ratings. Its strongest links were with Effectiveness, at -0.23, and Elaboration, at -0.21.
When the judge changes
The contrast changed when the comparison focused on who wrote the stories. Human evaluators found human- and AI-generated texts largely indistinguishable in quality, with significant differences only in Authenticity and Elaboration, both favoring LLM-generated stories. The LLM judge, however, produced significant score differences across all 11 dimensions.
The LLM judge also gave every LLM-generated story a perfect score on several dimensions and showed no variation in those dimensions. That is a ceiling effect: scores bunch at the top, leaving no room to distinguish stronger from weaker texts on those dimensions.
For human-written stories, the closest rank agreement between human and LLM ratings was for Elaboration, with a Kendall correlation of 0.31 and p below 0.001. The correlation for Creativity was 0.23, while Surprise was near zero at 0.01, with p = 0.867.
Perplexity had a negative correlation with all 11 dimensions in the LLM judge's ratings. The associations were weak for Surprise and Creativity and moderate for Originality and Novelty.
The pattern of links around the Creativity score also differed between the two rating systems. In human ratings, Creativity was linked to all 10 other dimensions, especially Novelty, Originality, Surprise, Value, Effectiveness and Authenticity. In LLM ratings, Creativity was less broadly aligned, had no correlation with Surprise or Usefulness, and was most strongly related to Elaboration.
How the comparison was built
The comparison used a balanced corpus of 200 stories: 100 human-generated and 100 machine-generated. The named models were each tasked with 20 stories, yielding the 100 machine-generated stories.
Human evaluation comprised 441 individual evaluations from 115 participants, averaging 2.21 independent ratings per text. Human and LLM evaluations were conducted blindly, without identifiers distinguishing human-generated from AI-generated stories.
To compare scores between the two story groups, the researchers used the Mann-Whitney U test. Spearman's rank correlation measured alignment between automated metrics and ordinal human ratings, while Kendall's rank correlation measured agreement between human and LLM-judge rankings.
What the results leave open
The evidence comes from one task and dataset, a 200-text corpus, and one large open-source LLM judge. Proprietary judge biases were not examined.
AI stories used vendor-default generation parameters. Temperature, top-p, top-k, prompt engineering, and differences between base and instruction-tuned models were not explored.
The work is a preprint identified as arXiv:2608.23705v2 and dated 27 Aug 2026. The paper states that generated and collected data, complete source code, and evaluation scripts are publicly available through the listed GitHub repository.
Paper data and sources
Original title: The Limits of Automatic Evaluation of Creativity in Large Language Models
Authors: Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-24
DOI: Not available
Original paper · Full text