A Practical Guide to Removing the AI Flavor from Fiction: Three Techniques to Make Your Writing More Human

Three practical techniques to strip the AI flavor from fiction and make your writing feel genuinely human.
AI-written fiction often has an unnatural "AI flavor" rooted in training-data bias and a lack of genuine creative intent. This article offers three practical techniques—reducing atmospheric padding, driving emotion through events instead of psychological description, and keeping action chains to two steps—to make fiction more natural and human.
Why Does AI-Written Fiction Always Have That "AI Flavor"?
If you frequently use AI to assist with fiction writing, you're surely familiar with this feeling: every sentence flows smoothly and reads coherently, yet something always feels off—the excessive atmospheric rendering, the piled-on rhetoric, the dense psychological descriptions strip the text of its authentic texture. Creators have jokingly dubbed this phenomenon the "AI flavor."
Bilibili content creator "Erming Zawu" cut straight to the heart of the two root causes of the AI flavor. The first is bias in the training data: "For many years, more description, more rhetoric, and expressing emotion through clichéd language and actions—that's what people considered good writing." A massive amount of such text has accumulated throughout history, all fed into AI, so AI learned this "average literary mode of expression."
It's worth understanding the deeper reason why AI generates this "average" style of expression. Large language models (LLMs) are trained through "next-word prediction" on massive amounts of text, essentially learning the statistical distribution patterns of language—that is, which word is most likely to appear given a certain context. This means AI's output naturally gravitates toward the "most common" mode of expression, rather than the most fitting or most distinctive one. Notably, this statistical tendency manifests not only in word choice but also permeates sentence structure, paragraph rhythm, and even the overall narrative framework: AI tends to reproduce the recurring three-part structure found in its training corpus—"emotional rendering → atmospheric setup → interior monologue"—because this structure appears with extremely high statistical frequency.
The formation of this "literary average" is also deeply influenced by "selective digitization": the text that enters AI training corpora is not uniformly sampled from all human writing but is heavily skewed toward content that has already been digitized and publicly published—text from online fiction platforms, writing forums, and e-book libraries vastly outweighs private manuscripts or orally transmitted folk narratives. These channels themselves carry an aesthetic selection bias: text with ornate diction and intense emotion more easily gains traffic boosts from platform recommendation algorithms, thus getting copied and spread widely, further reinforcing its proportion in the training corpus.
In the digital humanities field, this phenomenon is called "Corpus Bias": the composition of training data determines the boundaries of a model's worldview, and those boundaries have themselves been deeply shaped by commercial logic. Take the Chinese online literature ecosystem as an example: platforms like China Literature Group and Qidian have long used "subscription conversion rate" and "reader retention rate" as their core optimization goals, which gives writing styles with high emotional stimulation density, fast pacing, and strong conflict priority in algorithmic recommendations, leading countless authors to imitate and replicate them. These texts constitute one of the highest-proportion samples of creative writing in the training corpora of Chinese LLMs. In other words, the "literary average" that AI learns is actually the "traffic-writing average" filtered a second time through platform algorithms, rather than a fair representation of all excellent writing styles in literary history. From Lu Xun's restrained plainness to Eileen Chang's ornate imagery, from Yu Hua's cold narration to Mo Yan's magical elaboration—these vastly different classic styles of expression, dwarfed by the sheer volume of the training corpus, are all statistically diluted into faint signals.
Even more noteworthy is that this "averaging" tendency has a technical inevitability. Modern LLMs typically introduce a "temperature" parameter during training to adjust the randomness of output: the lower the temperature, the more the model tends to choose the highest-probability word and the more "conservative" the output; the higher the temperature, the more likely it is to sample low-probability words and produce "surprising" output. However, even at high temperatures, the model still selects expressions that are relatively high-frequency in the statistical distribution of the training corpus—truly rare literary expression with strong personal style, being sparsely represented in the training corpus, struggles to become the model's preferred choice at any temperature setting. In the realm of literary creation, the training corpus contains abundant "online literature style" text full of ornate diction and emotional rendering, and these texts statistically constitute AI's benchmark for "quality writing," causing the model to systematically favor this "literary average" when generating creative text, rather than the unique choices a stylistically assertive author would make.

The second reason is more fundamental—AI has no genuine creative intent. It doesn't know what's going to happen next, so it can only "pad" the word count with copious description. This also explains why AI text often feels hollow at crucial plot points yet overexerts itself in atmospheric description. Human authors, before putting pen to paper, have usually already constructed the plot's direction and the characters' arcs of fate in their minds; the weight of every sentence obeys the overall narrative intent. AI, on the other hand, generates each next word based only on the probability calculation of the current context—there is no forward-looking narrative awareness of the kind "I know the protagonist will die three chapters from now."
This absence of intent also has a subtle technical dimension: mainstream LLMs currently generally adopt an "autoregressive" generation mechanism—generating word by word from left to right, where each new word can only see the content preceding it and cannot, like a human author, "reverse-engineer" from a global perspective whether a particular expression serves the overall narrative need. This creates a fundamental asymmetry with the human author's creative psychological process: experienced authors have often already completed the design of the story's overall arc in their minds, and may even have anticipated the emotional tone of key scenes, with the weight of every sentence obeying this forward-looking narrative blueprint.
The autoregressive model is the complete opposite—it cannot make cross-temporal narrative decisions like "I know the protagonist will break down at the end of this chapter, so I'll use a restrained tone at the beginning to build up tension." Cognitive science calls this human capacity "goal-directed reasoning" or "mental time travel": human narrators can "travel" in imagination to future points in the story and then return to the present to decide how much weight to give their words. At the neuroscientific level, this ability depends on the coordinated work of the prefrontal cortex and the hippocampus—the former responsible for goal maintenance and delayed gratification, the latter for retrieving and reorganizing episodic memory. The computational architecture of autoregressive models contains no corresponding functional module, which makes "intent absence" not merely a limitation of current technology but a structural constraint of the existing architectural paradigm. Some researchers are exploring the introduction of a "planning module" to let the model first generate a story outline and then fill in details, but this remains a frontier research direction, still a long way from truly simulating human narrative will. This fundamental absence of intent causes AI, when facing plot junctures that require precise restraint, to instinctively use padding to mask emptiness.
Comparing Two Passages: What Does the AI Flavor Actually Feel Like?
To demonstrate the difference vividly, the creator provided two passages for comparison.
First passage (written by AI): "The night was thick as ink, wind seeping in through the cracks of the window. He stood there, his heart trembling faintly, as if something were about to be lost."
Second passage (written by a human): "His phone buzzed once. He stared at it. On WeChat there was just one line—Let's break up. He froze for a moment, then hurled the phone into the trash can, and then hurled the trash can down the stairwell."

The difference between the two is obvious. The first piles on atmosphere, using "thick as ink" and "heart trembling" to render emotion; the second drives the narrative through concrete character actions. When you strip away all the atmospheric description and replace it with characters and actions, the text's "human flavor" immediately emerges.
From the perspective of the reader's cognitive experience, these two passages activate completely different reading modes. The first requires the reader to passively receive the author's preset emotional labels; the reader plays the role of an emotional "container." The second requires the reader to actively interpret behavior—"Why did he throw the trash can down the stairwell?" That instant of self-questioning is itself the moment emotional resonance arises. Neuroscience research shows that when readers actively interpret the actions of others, the mirror neuron system is activated, providing a physiological basis for "empathizing"; whereas when passively receiving emotional labels, this mechanism remains relatively dormant. It is precisely this process of active interpretation that gives the reader the feeling that "this character is real," thereby forming what we call "human flavor."
This difference can also be understood from the perspective of information theory. Atmospheric descriptions like "the night was thick as ink" provide closed-form information at the semantic level—the emotional meaning has been preset and packaged by the author, requiring almost no semantic inference from the reader. "Hurled the trash can down the stairwell," on the other hand, is open-ended information whose emotional meaning must be actively constructed by the reader drawing on their own social experience and empathic capacity. Information theory has a concept called "information entropy," which refers to the degree of uncertainty in information: the higher the uncertainty, the greater the amount of information, and the more cognitive resources the reader must invest. Excellent narrative writing often maintains an appropriate level of semantic uncertainty at the right moments, inducing the reader to actively fill in the narrative gaps, and this filling-in process is itself the most precious source of the "immersion" in the reading experience. AI-generated text tends to state emotional meaning directly, eliminating this beneficial uncertainty and leaving the reader in a perpetually passive receiving state during reading, naturally making deep emotional resonance difficult to achieve.
This is exactly the first core principle for removing the AI flavor: reduce static atmospheric description and increase dynamic actions and events.
Fill in the Plot with Events, Not Piled-On Psychological Description
Some might ask: doesn't writing this way come out too dry? And what about hitting the word count?
The solution the creator offers is: fill in with events, not with psychological description. Rather than having a character "grieve and brood alone in a room, filled with pain, awash in memories," it's better to let a chain of real events unfold next.
He gives a vivid example: "He threw the trash can down the stairwell; the landlord came to find him; he got into an argument with someone again; then apologized; and finally sat alone in an unlit room, then opened his phone and replied to that WeChat message."

The brilliance of this approach lies in the fact that emotion flows out through events, rather than directly telling the reader "he is in pain." Every one of the character's actions silently conveys his inner state, and as the reader follows the events unfolding, they naturally feel the character's emotions. This reads far more comfortably than lengthy interior monologue, and it better conforms to the classic principle of excellent fiction: "Show, don't tell."
"Show, don't tell" is one of the most central narrative principles in modern fiction writing. The historical roots of this principle can be traced back to the modernist literary movement of the late 19th century—writers like Chekhov and Flaubert were among the first to abandon in practice the omniscient narrator's direct judgments of characters' emotions, turning instead to carefully selected external details to suggest inner psychology. Chekhov summed up this idea with "Don't tell me the moon is shining; show me the glint of light on broken glass." In the early 20th century, Henry James theorized this practice, and Hemingway pushed it to the extreme, developing the famous "iceberg theory": what appears above the water in the text is only a fraction, and the emotional tension beneath the surface is the true force that moves readers. In the latter half of the 20th century, the rise of American creative writing programs (MFA programs) codified this principle into a systematized writing education system, making it a core tenet of writing instruction in the English-speaking world.
It's worth noting that this principle is not universally applicable: the East Asian literary tradition contains a great deal of lyrical expression that "blends scene and emotion," and in classical Chinese poetry, the tradition of "direct expression of feeling" (such as "How I wish for a mansion of thousands of rooms, to shelter all the world's poor scholars in joy") coexists happily with the tradition of "conveying emotion through scenery," each with its own aesthetic legitimacy. In the Japanese aesthetic tradition of "mono no aware" (もののあわれ), the direct presentation of emotion is itself a highly refined art form. AI's problem is not that it uses "telling," but that it cannot judge—based on narrative situation, cultural context, and authorial style—when "showing" versus "telling" is more appropriate. This scene-sensitive meta-judgment capacity is precisely what distinguishes human authors from language models.
This principle has also received empirical support in the field of cognitive narratology. Researchers have found that when readers read "showing-type" narratives, the brain regions responsible for social cognition and mental inference (such as the temporoparietal junction) are significantly more active than when reading "telling-type" narratives. In other words, "showing" is not merely a literary aesthetic preference but a neural mechanism capable of deeply engaging reader cognition. AI-generated text precisely violates this principle en masse, because directly describing emotion is statistically the "safer," higher-frequency choice of expression, conforming to the model's probabilistic prediction of "good writing." The abundance of directly-emotional sentence structures in the training corpus causes the model to default to "he felt sad" rather than "he sat in an unlit room" when it needs to express a character's emotion.
The Key Technique: Keep Action Descriptions to No More Than Two Steps
Beyond the overall approach, the creator also shared a very concrete operational technique: when action descriptions bring out events, don't let the actions exceed two steps.
He uses a passage of original text as an example: "He lifted his head again to look at A by the door, smiling without warmth, as a kind of understated declaration—'Little brother, want to keep going?'" Here the actions are three steps (lifting his head, looking, smiling), which makes the pacing feel sluggish.

Rewritten according to the "no more than two steps" rule: "He tilted his head and looked at A, breaking into an innocent smile—'Little brother, still want to keep playing?'" After the revision, the sense of rhythm noticeably improves, and it reads much more smoothly.
There is cognitive science backing this technique. Cognitive Load Theory, proposed by psychologist John Sweller, treats the capacity of human working memory as a limited resource and distinguishes between "intrinsic cognitive load" (the inherent complexity of the information), "extraneous cognitive load" (the redundancy in how information is presented), and "germane cognitive load" (the effort readers invest in actively constructing meaning). Excellent narrative writing is essentially a fine-grained allocation of cognitive resources: reducing unnecessary extraneous load (such as lengthy action chains) so that readers can devote more cognitive resources to germane load—that is, actively inferring character motivation and experiencing emotional resonance.
Reading research shows that when readers process consecutive action sequences, each additional action node requires extra cognitive resources to maintain an overall understanding of the scene. The so-called "Situation Model" refers to the mental representation of a scene that readers dynamically build in working memory during reading, containing multidimensional information such as character position, posture, and emotional state. Research by cognitive psychologists such as Rolf Zwaan proves that updates to the situation model follow an "event segmentation" pattern: whenever a new action node appears in the narrative, the reader's working memory undergoes a "save-and-reload"-style update, and this process itself produces a tiny but measurable cognitive interruption. When the action chain is too long, the reader's attention shifts from understanding the plot to tracking the actions themselves, causing a break in immersion.
From the perspective of neuroaesthetics, this pattern also has a deeper perceptual-biological basis. The human visual system has evolved a mechanism that prioritizes processing of "key action nodes"—eye-tracking studies show that when readers read action descriptions, their gaze automatically focuses on the verbs with the strongest semantic predictability, rather than scanning all action words evenly. When the action chain is too long, this priority-focusing mechanism is disrupted, and readers are forced into linear, word-by-word tracking rather than jumping to meaning extraction, subjectively producing a "laborious" and "sluggish" reading feeling. This shares common ground with the "Gestalt principles" in visual art: excellent composition always guides the viewer's gaze to flow among a few key nodes rather than distributing attention evenly across dense details. The principle of "actions no more than two steps" essentially applies the same perceptual-aesthetic logic along the narrative timeline.
The field of screenwriting has even stricter norms for this; professional screenwriters typically keep action descriptions within a single shot to the barest minimum, and the logic behind this is highly consistent with the "two-step principle" in fiction writing—both aim to protect the audience's cognitive resources so they can concentrate on emotional experience rather than information tracking. Japanese manga panel theory also has a similar rule: excellent manga artists break complex actions down across multiple panels rather than piling up action lines within a single panel, essentially distributing the "cognitive load" along the narrative timeline rather than compressing it into a single space. AI lacks this dynamic awareness of the reader's cognitive load and often generates lengthy action chains to "fill up" a scene, further intensifying the text's sluggishness.
Of course, this rule is not absolute. The creator also frankly admits that if you have a special expressive purpose, writing four- or five-step action chains is perfectly fine. But you need to be clearly aware that too many consecutive actions will affect the reader's reading flow. Mastering this sense of proportion is an important step in moving from "AI flavor" to "human flavor."
Summary of the Methodology for Removing the AI Flavor
Taken together, the core of removing the AI flavor from fiction is shifting from "piling-on expression" to "narrative-driven expression":
- Reduce atmospheric rendering and use character actions to drive the plot;
- Fill in content with events, letting emotion flow naturally out of action;
- Control the number of action steps to keep the reading rhythm smooth.
Behind these techniques lies a deep understanding of AI's creative limitations—AI lacks genuine creative intent and can only rely on statistical patterns to produce "average" expression. The value of the human creator lies precisely in the command of plot direction and the precise grasp of rhythm. Language models cannot sense the reader's psychological expectations while reading a particular passage, nor can they, like an experienced author, judge by intuition that "this spot should be left blank, that spot should be given force."
From a more macro perspective, the essence of these writing techniques is a set of "anti-averaging" strategies: they all consciously break away from the statistical mean that AI represents, moving toward concrete, contextualized expression carrying personal weight. Every concrete action ("hurled the trash can down the stairwell") is a defection from the statistical average ("he felt pain"), and this defection is precisely the source of literary individuality. This "anti-averaging" narrative choice corresponds, in literary theory, to what Mikhail Bakhtin called "heteroglossia": truly vital literary language always contains deviation from and tension with mainstream discourse norms, whereas the statistical mean AI represents is precisely the most thoroughly "monoglossic" language in Bakhtin's sense—a stagnant linguistic pool in which all differences among individual voices have been erased. The more fundamental difference is this: when a human author writes each word, it obeys a complete narrative will; whereas each generation by AI is essentially a probabilistic sampling—no will, only weights. In today's world of increasingly widespread AI-assisted creation, understanding and compensating for these differences is exactly the discipline every creator must cultivate.
Key Takeaways
Related articles

Qwen3 27B In-Depth Review: A Powerful Reasoner That Overthinks — and How to Fix It
In-depth review of Qwen3 27B's reasoning capabilities and overthinking problem. Analyzes performance advantages, causes of overthinking, and provides practical optimization solutions.

RL for Reasoning Only Changes 1-3% of Tokens? The Truth and Controversy Behind the Claimed 1000x Compute Savings
RL training for LLM reasoning only changes 1-3% of output tokens, with researchers claiming 1000x compute savings. We analyze the deep implications, non-uniform token distribution issues, and the gap between benchmarks and real usability.

AI Algorithm Engineer Self-Study Roadmap: A Complete Plan from Zero to Landing Your First Offer
A detailed AI algorithm engineer self-study roadmap covering foundations, core algorithms, CV/NLP direction selection, and career transition strategies for landing offers.