Which AI Model Writes Better Fiction? A Head-to-Head Plot Generation Showdown Across 5 Models

GLM wins on plot logic, Claude and Gemini lead on prose quality in this 5-model AI fiction showdown.
A Bilibili creator tested DeepSeek, Zhipu GLM, Qwen, Gemini, and Claude on power fantasy plot generation using identical prompts. GLM scored highest (89 pts in AI judging) for logical twist design; DeepSeek lost points for having the protagonist spend their own money; Qwen's output was deemed unusable; Claude invented non-existent settings; Gemini likely underperformed due to poor prompting. The verdict: use GLM for plot outlining, Claude or Gemini for final prose.
AI-assisted novel writing has exploded in popularity over the past couple of years, but with so many mainstream models to choose from — DeepSeek, GLM (Zhipu), Qwen, Gemini, and Claude — creators often struggle to figure out which one is best suited for plotting and story development. A Bilibili content creator ran a side-by-side test using identical inputs (the same story opening, skill system, and full world-building setup) across multiple models, focusing specifically on their plot generation capabilities. This article breaks down the findings for anyone interested in using AI for creative writing.
Test Methodology: Same Prompt, Same Stakes — Focus on Plot Generation
The test maintained strict variable control: all models received the same story premise, the same protagonist skill settings, and the same complete world-building document. Each model was tasked with independently generating a story outline and plot progression for the first story arc. Results were scored across three dimensions: AI-ness (how formulaic/robotic the output felt), plot design (especially plot twists and conflict), and instruction-following accuracy.
Models tested included DeepSeek, Zhipu GLM, Qwen, Gemini (Flash version), and Claude (Sonnet 4.5, not the latest version). The tester noted upfront that Claude's performance may have been limited by the lack of access to the newest version.
It's worth emphasizing: this is a subjective evaluation by a single content creator. The scores reflect personal aesthetic preferences and should be treated as a reference point, not a definitive verdict.
Power fantasy (爽文) is a major genre in Chinese web fiction. The core formula involves a protagonist with overpowered abilities — system rewards, skill buffs, etc. — who steamrolls enemies, accumulates resources, and achieves a dramatic rise to power. The pacing is fast, the satisfaction is immediate, and the underlying logic demands that the protagonist always act in their own maximum interest. That's why something like "spending your own money to treat colleagues to dinner" feels instinctively wrong to genre readers — it breaks the core contract of the power fantasy genre. The tester used this as a key evaluation criterion, essentially testing whether models genuinely understood genre conventions rather than just general story coherence.
Full world-building setup (全序设定) refers to providing the model with a comprehensive document in the prompt — covering the complete world rules, system mechanics, and character backgrounds — essentially a detailed story bible. This approach reduces deviation caused by the model improvising, but it also serves as a litmus test: does the model actually follow, remember, and respect these established rules? Or does it ignore, forget, or quietly rewrite them? Claude's "invented setting" issue was a direct failure on this dimension.
Model-by-Model Breakdown
DeepSeek: Solid opening outline planning — the tester gave it roughly 85 points — but it revealed clear signs of its generation process, with internal structures like "AI revision banks" and "material storage banks" leaking into the output. More critically, the plot twist design had a major flaw: the protagonist uses the 100,000 yuan just rewarded by the system to treat the whole team to late-night snacks as a way out of the conflict. The tester flagged this as a serious "landmine."

"Why would a power fantasy protagonist spend their own money on something like this?" the tester asked. The protagonist wasn't established as wealthy, and blowing the freshly-received 100,000 yuan on a dinner outing completely violates power fantasy logic. The other plot device — threatening someone with blackmail material on other artists — also doesn't hold up, as the leverage is too weak. These logic holes were DeepSeek's main demerits.

Zhipu GLM (version 5.3): The tester recalled GLM performing well in the past. This time, its outline showed relatively low AI-ness at the detail level, though the macro-level planning felt more formulaic. Instruction-following was moderate to high, and the score landed around 85 — "not much different from DeepSeek" on the surface. However, GLM's twist and main plot design turned out to be the highlight of the entire test.

Qwen: The lowest-rated model. The tester felt it was "too literary" for the genre, the twist design was poor, and the overall execution was described as "completely off the rails." The verdict: "Qwen doesn't work for this."
Gemini (3.8 Flash): Moderate-to-high AI-ness, had ideas but didn't execute them well, instruction-following was moderate to high, scored 80–85. The tester believed Gemini's actual capabilities were higher than what it showed, suspecting the issue was inadequate prompting and guidance during plot development — and leaned toward re-evaluating after better setup.
Claude (Sonnet 4.5): The least AI-sounding output with the fewest edits needed — the model the tester had highest hopes for in terms of rigor. But the conflict design introduced "splitting of check-in benefits" — content that simply didn't exist in the original world-building — which counts as a "minor landmine." Final score: 80 points, below expectations.

The Decisive Difference: Twist Design
What truly separated the models in this test was how each handled plot twists and conflict design.
After comparing outputs, the tester concluded that GLM (Zhipu 5.3) had the most logically sound twist structure: the first obstacle was the audition scene, the second was online public opinion, and the third involved a frame-up that forced both parties into a "dog-eat-dog" situation. The approach is straightforward, but it follows a coherent cause-and-effect chain that fits the power fantasy genre. By contrast, DeepSeek's "spend your own money" twist, Claude's invented setting, and Qwen's overly literary sensibility all stumbled at this stage.
In other words, plot generation isn't about flowery prose — it's about whether motivations make sense and whether the protagonist's solutions match their actual circumstances. GLM won because its approach was "simple but logically defensible."
Using Gemini as Judge: AI vs. AI Evaluation
The tester also had Gemini (3.8 Flash) perform a multi-dimensional evaluation of all model outputs. The AI-judged results largely aligned with the human assessment:
- GLM (Zhipu 5.3): 89 points — most consistent with power fantasy logic, top score
- DeepSeek: 87 points — inconsistent, with notable logic flaws
- Qwen: Not suitable — execution failed
- Gemini and Claude did not place in the top tier this round
The tester disagreed with DeepSeek's 87-point score, arguing that the core logic flaws in its main plot design shouldn't have earned that rating — but agreed that GLM's 89 points "matched my own perception."
Using one AI model to evaluate the output of other AI models is a common technique in modern prompt engineering, sometimes called "LLM-as-Judge." The advantage is that it quickly generates structured, multi-dimensional scoring that often correlates reasonably well with human judgment. The limitations are equally significant: models tend to score outputs higher when they're fluent and well-formatted, their ability to catch logical flaws varies considerably, and training data overlap can introduce bias toward or against models from related organizations. The tester's skepticism about DeepSeek's 87-point score illustrates this limitation perfectly — the AI judge failed to sufficiently penalize the logic landmines that the human reviewer immediately caught.
Verdict: Which Model Should You Use?
Combining human evaluation with the AI-judged scores, here's where the test landed:
- For plot development (outline, twist logic): GLM (Zhipu 5.3) was the most consistent and most genre-appropriate — the clear winner of this test.
- For final prose quality (writing texture, rigor, low AI-ness): The tester still prefers Claude and Gemini, believing Claude's precision and detail control should be strongest under better conditions, and that Gemini needs better prompting to show its real capability.
The tester's overall recommendation: use GLM for plot development and outline generation; use Gemini or Claude for actual prose writing.
That said, this is just one test run's conclusion. Model versions, prompt design, and genre type all significantly affect results — especially given that Claude wasn't on its latest version, and Gemini may have underperformed due to insufficient guidance. If you want to find your ideal AI writing partner, the best approach is to run a few rounds of your own tests with your specific genre and story type.
Related articles

Gemini 3.5 Transcribe: Transcription Tools Are Becoming Content Understanding Engines
Google's Gemini 3.5 Transcribe supports 85+ languages, timestamps, and up to 3-speaker diarization, plus key point and sentiment recognition — marking a shift from archiving to content understanding.

AI and Data Centers Take Center Stage in U.S. Midterms: Vox Breaks Down Five Key Issues
Vox's midterm election breakdown: Trump's historic low approval, Israel dividing Democrats, AI and data center anxiety crossing party lines, affordability politics, and election security concerns.

Claude Code Recreates Viral Riso Animation: Full Workflow, Prompt Design, and Token Cost Breakdown
A ~10,000-char creative brief drove Claude Code + Opus 5 to recreate a viral risograph animation in pure code. Full prompt design, workflow, and token costs revealed.