Creative Writing AI Rankings Spark Debate: Can Benchmark Tests Really Be Trusted?

AI creative writing rankings spark debate over whether subjective writing quality can ever be reliably benchmarked.
A Reddit discussion about Gemini 3.7 Flash ranking third in creative writing triggered broader debate about benchmark credibility. Users challenged Claude's leaderboard dominance citing its inescapable stylistic fingerprint, criticized Fable's purple prose, and argued that creative writing's inherent subjectivity makes any single ranking unreliable. The community called for blind evaluation by literary professionals as a more trustworthy alternative.
A Ranking That Sparked Controversy
Recently, a discussion about Gemini 3.7 Flash ignited heated debate on Reddit. According to reports, Gemini 3.7 Flash ranked third globally in creative writing ability, only slightly behind the Fable model. However, this seemingly objective ranking provoked strong pushback from the community — not against Gemini itself, but against the credibility of creative writing benchmarks as a whole.
The focal point of the debate was Claude's long-standing dominance on these leaderboards. One user put it bluntly: "As long as Claude models occupy the top spots, I don't trust these benchmarks. I simply can't stand the way Claude writes — no matter how I adjust my prompts, it sounds like a robot." In this user's view, Kimi K3 and even the newly released DeepSeek PRO far surpass Claude in creative writing.
What appeared to be a discussion about a single model actually revealed a deeper challenge in AI creative writing evaluation: Can the quality of creative writing truly be measured objectively?
Subjectivity: The Core Dilemma of Creative Writing Evaluation
In response to the criticism above, other users offered diametrically opposed views: "My experience is exactly the opposite — I think Claude models are the strongest at creative writing." This head-on clash of opinions perfectly illustrates the fundamental issue — creative writing is highly subjective.
As one commenter incisively pointed out: "Creative writing is highly subjective, which is exactly why I think benchmarking creative writing is pointless — everyone has a different opinion on what the best model is."
Current AI creative writing benchmarks primarily follow several paradigms: First, evaluation based on automated metrics like Perplexity, BLEU scores, and ROUGE scores — metrics originally designed for machine translation that have inherent limitations in measuring creativity. Second, arena-style systems based on human preferences (like Chatbot Arena), where users blindly vote between outputs from two anonymous models, generating rankings through an ELO rating system. Third, expert review panels where professionally qualified judges score outputs across multiple dimensions. Each method has its limitations — automated metrics cannot capture the nuances of creativity, crowdsourced voting is susceptible to sample bias and user demographic preferences, and expert reviews face issues of high cost and low inter-rater agreement.
This subjectivity breaks down even further. As one user added, evaluation also depends on what type of creative writing you're doing — there's an enormous difference between fictional story creation and internet content creation. A model might excel at crafting compelling short stories but perform poorly at generating social media copy, or vice versa. A single "creative writing" score can hardly encompass such diverse use cases.
"Claude Flavor": The Inescapable AI Writing Fingerprint
A recurring keyword throughout the discussion was the so-called "Claude flavour." Multiple users reported that no matter how they adjusted their prompts, Claude's output always carried a recognizable fixed style and set of clichés.
One user stated: "I've tried every custom writing instruction I could find online, and it still uses the same phrases and patterns that instantly break my immersion." Another user echoed the sentiment: "I just cannot remove the 'Claude flavour' from the writing no matter what I do, whereas Chinese models seem much more open in terms of writing style."
This phenomenon deserves deeper reflection. It reveals that different companies' models develop unique "stylistic fingerprints" during training. Modern large language models typically go through three stages: pre-training, supervised fine-tuning (SFT), and reinforcement learning from human feedback (RLHF). During the RLHF phase, human annotators rank model outputs by preference, training a Reward Model, which the model then optimizes against using algorithms like PPO to achieve higher reward scores. In this process, if the annotation team has uniform aesthetic tendencies — say, preferring elegant, nuanced expression — the model will systematically learn that style. Claude's distinctive "flavor" likely stems from Anthropic using an annotation team with highly consistent literary aesthetics during alignment training, combined with additional constraints on output style from their Constitutional AI approach.
This fingerprint is a mark of elegance for some users and an immersion-breaking nuisance for others. Interestingly, it's precisely this stylistic "stubbornness" that has driven some users toward Chinese AI models that offer more malleability in writing style. The "openness" in writing style of Chinese AI models frequently mentioned in the discussion (such as Kimi K3 and DeepSeek PRO) has deep technical and cultural reasons. These models are typically pre-trained on large-scale bilingual Chinese-English corpora, and cross-lingual training may endow models with more flexible expression capabilities and less single-style lock-in. Additionally, different companies' alignment strategies during RLHF diverge significantly — compared to Anthropic's heavy emphasis on safety and consistency, some Chinese AI companies may grant models greater stylistic freedom during fine-tuning. Models like DeepSeek that employ Mixture of Experts (MoE) architecture may inherently possess potential for style diversity through their different expert sub-networks.
Divided Opinions on the Fable Model
In the original ranking, Gemini 3.7 Flash was rated as "slightly behind Fable," but this assessment was equally challenged. One user objected: "Worse than Fable? Are you sure? In my experience, Fable is terrible at creative writing. It goes overboard — purple prose, cringeworthy sensory descriptions, and so on."
Fable is a model from a startup focused on narrative AI, designed to generate high-quality story-driven content. The "purple prose" and "cringeworthy sensory descriptions" criticized by users is a well-known term in the AI writing community — referring to an overly ornate writing style that piles on adjectives and metaphors. This phenomenon is particularly common in models fine-tuned on creative writing datasets, because highly-rated texts in training data tend to favor descriptively rich literary works. This causes models to learn an implicit rule of "more ornate is better" while ignoring that concise, powerful writing is equally a hallmark of excellent prose.
This disagreement further reinforces the community's core argument: one model's "eloquent prose" is a bonus for some evaluators and an affectation-laden penalty for others. When evaluation criteria themselves are this polarized, any ranking inevitably carries the personal aesthetic preferences of its judges.
What Kind of AI Writing Benchmark Deserves Trust?
Amid all the skepticism, some users offered constructive thoughts. One commenter suggested that an ideal creative writing benchmark should employ a blind evaluation mechanism: "I would only trust a benchmark where judges with higher education in literature-related fields conduct blind evaluations."
This suggestion touches on the methodological core of AI evaluation. Blind Evaluation is a standard method for controlling assessment bias in academic research, and its application in AI is receiving increasing attention. The core principle is that evaluators make judgments without knowing the source of the output, eliminating brand effects and preconceived biases. In AI writing evaluation, this means judges don't know which text came from which model. Chatbot Arena has partially achieved this, but its evaluator pool consists of ordinary internet users rather than literary professionals. Academic research suggests that the correlation between expert and public evaluation on creative tasks may be lower than expected, further complicating the design of reliable benchmarks.
Many current creative writing leaderboards rely either on automated metrics or on ordinary users' subjective votes — neither of which can truly measure literary quality. An ideal evaluation system might require multi-layered design: both large-scale public preference data and small-scale but high-quality expert deep reviews. Introducing professionally qualified blind reviewers may be a viable direction for improving evaluation credibility — though this would significantly increase evaluation costs and difficulty.
Gemini 3.7 Flash's Actual Writing Performance
Setting aside the ranking controversy, many users expressed positive views about Gemini 3.7 Flash itself. One former heavy user of Gemini PRO said that despite having switched to Chinese models, they still planned to top up their OpenRouter balance to test Flash 3.7, calling their initial experience in Google AI Studio "quite enjoyable."
OpenRouter is a unified AI model API aggregation platform that allows users to access dozens of models from Google, Anthropic, OpenAI, Meta, and other providers through a single interface. Users pay per token consumed without needing separate subscriptions for each model. The rise of such platforms reflects an important trend in the AI application layer: models are moving toward commoditization, where users are no longer locked into a single ecosystem but can flexibly switch models based on specific tasks. Google AI Studio is Google's free model experimentation platform where developers can test various capabilities of the Gemini model family.
Another user, while complaining that "the only thing I hate about Gemini is its clichéd responses," simultaneously acknowledged that "other than that, it's really quite good." Yet another user, referring to a specific capability dimension, commented: "In that regard, the progress is real and tangible."
These evaluations paint a relatively balanced picture: Gemini 3.7 Flash isn't perfect — its formulaic expressions remain a weakness — but it has made substantive progress in overall creative writing ability.
Conclusion: Thinking Beyond Rankings
This discussion around Gemini 3.7 Flash's ranking is far more valuable than the ranking itself. It reminds us that as AI capabilities grow increasingly powerful, not all abilities can be simply quantified. Creative writing, as a task highly dependent on personal aesthetics, cultural background, and use cases, demands that any single ranking be viewed with caution.
For everyday users, perhaps the most practical advice is this: don't blindly trust leaderboards — test candidate models against your actual needs. After all, the model that produces writing you're satisfied with is your "best model."
Related articles

DIY Air Purifier: Building a Silent CR Box with PC Fans and an Aluminum Frame
Learn how to build a quiet Corsi-Rosenthal air purifier using PC case fans and an aluminum frame, covering fan selection, PWM speed control, and cost analysis.

Universality of Gradient Descent Training: Does Neural Network Architecture Choice Really Matter?
Exploring the universal approximation capability of gradient descent training, analyzing the relationship between neural network architecture choice and learnability, from UAT to NTK theory.

From AI to Large Models: Understanding the Conceptual Landscape and Technological Evolution of Artificial Intelligence
Understand how AI, machine learning, deep learning, large models, and generative AI relate to each other. From Deep Blue to ChatGPT, learn how Transformer architecture gave rise to LLMs.