The Rise of Slow Content: A Tweet Sparks Reflection on the Attention Economy and Content Creation

A single tweet reveals the deep tension between the attention economy and authentic slow content creation.
A minimalist tweet about live-streaming the reading of a book sparks reflection on the attention economy. This article examines the neuroscience behind algorithmic engagement, the revival of slow content like ASMR and companionship streams, the irreplaceability of the human voice in the AI/TTS era, and practical differentiation strategies for content creators.
The Content Creation Dilemma Behind a Single Tweet
Recently, a minimalist tweet sparked widespread discussion on social media: "i want to do a live stream where i just read a book out loud on twitter." This seemingly casual idea actually touches on a deep issue faced by content creators today—in an attention economy driven by algorithms and the pursuit of stimulation, is there still room for the most unadorned forms of content?
It's worth understanding how the "attention economy" operates. This concept was first proposed by economist Herbert Simon in 1971, with the core argument that in an era of overwhelming information, the truly scarce resource is no longer information itself but the limited attention of humans. This theory was later developed further by economist Michael Goldhaber in 1997 into a complete framework of "attention economics." He predicted that attention would replace currency as the new medium of exchange—a prediction borne out by the core business models of tech giants like Meta and Google, which are essentially "attention intermediaries" that package and sell user attention to advertisers. Notably, this evolution has not been unidirectional and linear—changes in the regulatory environment are systematically reshaping the rules. The EU's Digital Services Act (DSA, effective 2023) for the first time required large platforms to offer recommendation options not based on behavioral profiling, effectively acknowledging at the legislative level the structural manipulation of user attention by algorithms. China's mandatory "youth mode" regulations and lawsuits over "dark pattern design" across U.S. states are also forming a global encirclement, signaling that the next major variable in the attention economy will come from the regulatory side rather than the technological side. Social media recommendation algorithms are therefore not neutral technical tools but decision-making systems deeply shaped by commercial goals—their optimization targets have never been "user well-being" or "content quality," but rather engagement metrics like DAU (Daily Active Users) and Session Length. Research by a former Facebook data scientist revealed that while boosting engagement, such algorithms systematically amplify emotionally polarizing content, because anger and anxiety most effectively prolong user dwell time.
The underlying logic of this mechanism has deep neuroscientific support. The dopamine system is key to understanding platform design: social media's "infinite scroll" and random reward mechanisms (the uncertainty of like counts) are highly isomorphic with the reinforcement mechanisms of slot machines, both exploiting the brain's hypersensitivity to uncertain rewards. Research by Stanford neuroscientist Robert Sapolsky shows that the dopamine release triggered by uncertain rewards can even exceed that of certain rewards. At the neural circuit level, this effect occurs primarily in the Nucleus Accumbens, with the Reward Prediction Error (RPE) signal emitted by the ventral tegmental area (VTA) of the midbrain—a signal that is also triggered when "expected rewards fail to materialize." This is precisely the design principle behind the "pull-to-refresh" gesture: every refresh is a bet. Research by neuroscientist Kent Berridge further distinguishes between the "wanting" system and the "liking" system—the former dominated by dopamine, the latter dependent on endogenous opioids. Social media design primarily hijacks the "wanting" system, which explains the contradictory experience of users being unable to stop scrolling yet feeling no pleasure afterward. This research has been consciously converted by Silicon Valley designers into product strategy—the famous concept of "Persuasive Technology" was systematized by Stanford professor B.J. Fogg, whose lab trained practitioners including several early Instagram designers. Former Google design ethicist Tristan Harris later publicly criticized this design tradition and founded the Center for Humane Technology to specifically study the long-term social impact of platform design. Neuroscience research further indicates that chronic overactivation of the dopamine system leads to a decline in baseline satisfaction—requiring ever-stronger stimulation to achieve the same pleasure. Clinically, this closely mirrors addiction mechanisms and partly explains why algorithm iterations on platforms like TikTok and Instagram Reels have compressed content length from 15 seconds to 7 seconds, making the "first 3 seconds hook" an iron law of creation, forcing all creators to compete on the same track for the same type of emotional response from the same audience.
It is precisely against this backdrop that this idea deserves careful consideration—precisely because it runs counter to mainstream content creation logic. While short-video platforms continually compress duration, obsess over the "hook" in the first three seconds, and rely on rapid cuts and visual impact, the form of "reading a book aloud"—slow, linear, and almost entirely devoid of visual gimmicks—instead reveals a rare rebellious quality.
The Revival of Slow Content: From ASMR to Companionship Streams
Humanity's Enduring Need to Be Read To
"Reading aloud" as a content form is not new. From traditional audiobooks and late-night radio to the ASMR reading videos that have risen in recent years, this form has never satisfied information efficiency but rather companionship, focus, and emotional comfort. Placing this form in a broader market context reveals: according to data from the Association of American Publishers (AAP), the U.S. audiobook market has maintained double-digit growth for ten consecutive years since 2013, with the market surpassing $2 billion in 2023; the global podcast audience, according to Edison Research, has exceeded 500 million. These figures reveal a key insight: the demand for "auditory linear narrative" has not been dissolved by the short-video era but has instead formed a complement to it—a large number of users consume audio content while commuting, exercising, or doing housework, a time window that visual content cannot compete for. Audiobooks are professionally produced pre-recorded content, and podcasts are a conversation/commentary format, while "live reading streams" fill precisely the intersection between the two—combining real-time interactivity with textual intimacy.
Understanding the appeal of such content requires a neuroscientific perspective. ASMR (Autonomous Sensory Meridian Response) refers to the low-intensity pleasurable sensation an individual experiences—spreading from the scalp down the spine—upon receiving specific auditory or visual stimuli. Although this phenomenon has been popular among the public for years, its neural mechanism did not receive its first peer-reviewed research support until 2018: a study by Giulia Poerio's team at the University of Sheffield, published in PLOS ONE, first confirmed the physiological reality of the ASMR experience through skin conductance measurements, finding that ASMR content consumers' heart rates dropped by an average of 3.14 beats per minute. A 2022 neuroimaging study further found that the ASMR experience activated brain regions associated with social reward and emotional regulation, including the medial prefrontal cortex and amygdala, and that its neural mechanism may be related to oxytocin secretion and activation of the parasympathetic nervous system.
Meanwhile, "body doubling streams" have deeper clinical roots. Body Doubling was originally a behavioral intervention technique in the clinical treatment of ADHD (Attention Deficit Hyperactivity Disorder), systematized by therapist Judith Kolberg in the 1990s. Its neural mechanism is thought to be related to the mirror neuron system and the social supervision activation of the prefrontal cortex—because ADHD individuals have functional differences in their prefrontal dopamine and norepinephrine systems, their executive function relies significantly more on environmental cues than neurotypical individuals. The presence of others provides external scaffolding that compensates for deficiencies in internal regulation.
At the epidemiological level, the global adult ADHD prevalence under DSM-5 diagnostic criteria is approximately 2.5%–4.4% (Fayyad et al., 2017 WHO cross-national study); if undiagnosed and subclinical populations are included, the actual affected population far exceeds this figure. During the pandemic, remote work removed the natural "social supervision" scaffolding provided by the office environment, and a large number of previously well-functioning ADHD adults recognized their own cognitive characteristics for the first time. Global ADHD diagnoses rose significantly between 2020 and 2022 (increases exceeding 40% in some countries). This context drove the user base of the digital Body Doubling platform Focusmate to grow by more than 600% in a single year during 2020. The increased social acceptance of the Neurodiversity concept has made a broader population willing to use such tools without a clinical label, significantly expanding the potential market boundary for companionship content. Considering that ADHD conservatively affects 5–7% of the global adult population, the scale of this demand far exceeds intuitive perception. With the popularization of remote work during the pandemic, this clinical concept was migrated to the digital space on a massive scale: "Study With Me" channels on YouTube have accumulated billions of views. On platforms like Twitch and YouTube, streamers simply study, work, or read quietly while viewers "virtually accompany" them—the popularity of such content clearly shows that users don't always crave high-intensity stimulation; sometimes what they need is simply a low-pressure sense of presence that satisfies the brain's deep-seated need for social connection.
The Scarcity Value of Anti-Algorithm Content
A purely reading-based stream is essentially an active resistance to "algorithm optimization" logic—it doesn't cater to completion rates, doesn't manufacture controversial topics, and doesn't design interaction hooks. This "purposelessness" instead constitutes a scarce resource in an information-overloaded environment: content that truly allows people to slow down inherently possesses differentiated value.
Content Authenticity in the AI Era: The Irreplaceability of the Human Voice Reading
When TTS Proliferates, the "Human Touch" Becomes Precious
As AI speech synthesis technology matures, the significance of real-person live reading becomes ever more pronounced. Modern TTS systems have gone through three generations of technological evolution: from concatenative synthesis and parametric synthesis to the current neural network synthesis. WaveNet, released by DeepMind in 2016, was a milestone in this field, first replacing traditional vocoders with raw waveform modeling and achieving a qualitative leap in naturalness. Since then, technical progress has exceeded most expectations: Tacotron 2 in 2018 made end-to-end speech synthesis practical, and neural codec architectures—such as Meta's EnCodec and Google's SoundStream—by compressing audio into discrete semantic tokens, enabled deep integration between speech synthesis and large language models, giving rise to systems like VALL-E and VoiceBox that can clone any voice from minimal samples.
The technical significance of this paradigm shift is worth exploring in depth: architectures like VALL-E redefine speech synthesis as an "in-context learning problem for language models"—a 3-second sample serves as a prompt, and the model performs conditional generation directly in discrete audio token space without any gradient updates. This means voice cloning has gone from a "customized service" to "zero-shot inference," with costs dropping from the thousand-dollar range to nearly zero. Microsoft's VALL-E, released in 2023, requires only a 3-second sample to clone a speaker's timbre. A new generation of text-to-speech systems represented by ElevenLabs, Microsoft Azure Neural TTS, and OpenAI TTS can already generate synthetic speech that is highly realistic in prosody, emotion, and breathing rhythm, and AI-dubbed content already accounts for a considerable share of audio content production on various short-video platforms.
However, there remains an as-yet-uncrossed chasm between technical capability and emotional authenticity. TTS is essentially a form of "predictive generation"—it reproduces the statistical regularities of human speech rather than genuine emotional states. The most advanced TTS systems process the "statistical distribution of speech" rather than the "immediate emergence of consciousness." The key point is: the emotional modulation in real-person reading comes from the reader's cognitive processing of the text—understanding metaphors, sensing narrative tension, and generating empathic responses—and these higher-order cognitive processes are projected in real time onto subtle acoustic feature changes. TTS reproduces the distribution patterns of these features but not the generative mechanism. The distinction between the two is especially pronounced in long-form coherent reading, which is precisely why professional audiobook narrators remain irreplaceable. The improvised pauses, occasional slips, emotional fluctuations, and even brief lapses of attention and offhand commentary that naturally occur in real-person reading—slightly choking up because a passage is moving, unable to resist inserting a comment about an interesting detail, speech slowing slightly due to fatigue—are precisely the content dimensions that current TTS systems cannot meaningfully reproduce.
The emotional authenticity of real-person reading has an inseparable holistic quality: it is not merely a collection of audio signal features, but the acoustic projection of a real-time cognitive dialogue between the reader and the text. This signal of "consciousness in presence" cannot, in theory, be fully reproduced by generative models—a distinction that is philosophically highly related to the "Philosophical Zombie Argument": there exists an essential rather than merely gradational difference between a system that behaviorally perfectly simulates a reader and a reader genuinely moved by the words. This poses an unprecedented challenge to the uniqueness of "vocal identity" and gradually makes the "irreproducibility" of the human voice a scarce asset. When synthetic content proliferates uncontrollably, audiences' craving for genuine human voices may well rise rather than fall. The scenario described in this tweet is, to some extent, the most unadorned return to content authenticity.
The Disenchantment and Return of Content Creation
This idea also reflects a stance of "disenchantment" toward content creation—stripping away all commercial packaging and technical means to return to the most primal act of sharing: I'm reading a book, and if you want to come listen, come. This low-barrier, low-cost, high-sincerity mode of creation may be precisely the state that countless creators exhausted by "rat-race competition" truly yearn for deep down.
Practical Insights for Content Creators
Differentiation Strategy: Slower, More Genuine, More Focused
For creators hoping to break through in a crowded content field, this tweet offers a reverse angle of thinking: rather than engaging in close combat on the path of "faster, more explosive, more competitive," it may be better to explore the differentiated path of "slower, more genuine, more focused."
Deep companionship content in vertical niche fields often cultivates a stickier, more loyal community. A sustained reading stream may gradually gather a fixed group of "book friends," forming a unique community culture and creator identity. Media economics researcher Ethan Zuckerman calls the phenomenon of heavy algorithm dependence "Platform Capture"—creators, in the process of optimization, deeply bind their capabilities to platform rules, and once the rules change, they lose the ability to migrate. From this perspective, building audience relationships centered on authentic content is itself a long-term strategy for countering the risk of platform capture.
A Healthy Content Ecosystem Requires Diversity
From the perspective of platform ecosystems, such content experiments also remind us of a deeper systemic problem. The ecological concept of "monoculture" applies equally to content ecosystems: when platform algorithms systematically reward the same type of content form, the creator community undergoes convergent evolution, ultimately leading to increased fragility of the overall ecosystem.
This phenomenon can be quantitatively analyzed using information theory. In Shannon's information theory, entropy H = -Σp(x)log p(x) measures the uncertainty and diversity of a system. When platform algorithms form a powerful "survival of the fittest" selection pressure, the distribution of content forms converges toward a low-entropy state—that is, a few types of high-engagement content templates dominate absolutely. This closely resembles the mechanism of genetic diversity erosion in biology: in the short term, the expansion of "advantageous genotypes" increases the average fitness of the population, but simultaneously weakens its resilience to environmental mutations. Researchers can in fact quantify ecosystem health by calculating the Shannon entropy of the distribution of content forms on a platform: a 2021 study by the MIT Media Lab conducted such an analysis on YouTube recommended content, finding that the form entropy of algorithmically recommended content was significantly lower than that of user-actively-subscribed content, with a gap of about 0.8 bits—meaning the recommendation algorithm compressed content diversity by nearly half. From the perspective of complex systems science, this entropy reduction phenomenon boosts platform engagement metrics in the short term but systematically reduces the platform's "Serendipity Value"—the probability of users encountering unexpected surprises and discovering entirely new interests continues to decline. The contradictory experience of "bored after scrolling but unable to stop" that appears in recent user satisfaction surveys of TikTok and Instagram Reels can be precisely explained within the information-theoretic framework as: the system reinforces the dopamine "wanting" signal while cutting the genuine cognitive satisfaction brought by content diversity—that is, the impoverishment of subjective experience caused by entropy reduction.
Ecologist C.S. Holling's "resilience theory" and "Adaptive Cycle" model can describe this dynamic more completely: a content ecosystem rapidly accumulates dominant species in the "exploitation phase," reaches apparent stability in the "conservation phase" while internal fragility continuously accumulates, and eventually enters the "release phase" of collapse and reorganization due to external disturbances. After YouTube switched its recommendation mechanism from "click-through rate optimization" to "watch-time optimization" in 2012, it spawned a content ecosystem of "quality long videos over ten minutes," but with the impact of TikTok, this ecosystem quickly entered the release phase, and a large number of mid-tier creators were forced to migrate or exit—a real-world case of this very theory. Media researchers have observed related phenomena manifesting across multiple platforms: the "burnout" crisis among the YouTube creator community, the accelerated homogenization of TikTok content, and the overall sluggish user growth of social media are all related to this. When the content distribution entropy of a platform continues to decline, users' marginal satisfaction also decreases accordingly. Those seemingly "out-of-step" attempts are precisely the spontaneous effort to inject diversity at the system level, and their value transcends the commercial gains and losses of individual creators—often serving as the earliest sprout of innovation.
Conclusion
A simple tweet carries deep questions about attention, authenticity, and the essence of content. In an era where AI is intervening in content production on a massive scale, a back-to-basics idea like "reading a book aloud" has instead become a mirror, reflecting our collective yearning for a slow pace, authenticity, and pure connection.
Regardless of whether this stream ultimately comes to fruition, the reflection it has sparked is valuable enough—reminding every content creator: technical tools continue to evolve, but the core of content will forever be the sincere connection between people.
Key Takeaways
Related articles

RisenX Explained: The Coding Agent Officially Recommended by DeepSeek
RisenX is a DeepSeek-native coding agent featured in DeepSeek's official API docs. It supports cache-first loops, tool-call repair, and Flash/Pro smart switching.

ChordViz Review: A Real-Time Visualization Workbench for MIDI and Audio
In-depth review of ChordViz music visualization tool with real-time MIDI and audio input, chord visualization, notation, and audio-reactive visuals, plus OBS, TouchDesigner and Resolume integration.

3D-Printed Robot Desk Lamp: How to Make a Machine Feel Alive Like a Pixar Character
See how an indie developer uses 3D printing, ROS 2, and a custom animation editor to turn Pixar's iconic desk lamp into a real robot with personality, vision, and RL-driven autonomy.