[KongchangAI]
· 1 min read· 898 words

AI vs. Human: Who Writes Better? Reflections on a New York Times Quiz

AI vs. Human: Who Writes Better? Reflections on a New York Times Quiz

A NYT quiz shows detecting AI writing matters less than judging whether content is trustworthy and worth reading.

The New York Times' AI-versus-human writing quiz has sparked wide discussion in the tech community. As LLMs improve, the surface cues once used to spot AI text — overly neat structure, lack of personal detail — are fading, leaving readers unable to tell the difference. But harder detection doesn't mean AI writes better. Good writing carries genuine thought, lived experience, and independent judgment. For creators, the real edge lies in first-hand perspective and a distinct voice; for readers and platforms, learning to judge content as trustworthy and valuable matters far more than identifying its source.

The New York Times recently launched an interactive quiz that has sparked discussion across the tech community: given a few passages of text, can you tell which were written by humans and which were generated by AI? This seemingly simple game touches on an increasingly difficult question at the heart of generative AI development — can machine-written text truly stand alongside human writing?

The Real Question Behind the Quiz

This kind of "Turing test-style" writing identification game isn't new, but its significance is shifting as large language models rapidly improve. A few years ago, AI-generated text often gave itself away through logical gaps, awkward phrasing, or repetitive filler. Today, leading models produce paragraphs that feel remarkably natural in grammar, coherence, and stylistic mimicry — ordinary readers can no longer rely on gut instinct alone to spot the difference.

By designing this quiz, the New York Times is essentially inviting readers to experience this blurry boundary firsthand. When you find yourself repeatedly guessing wrong, the reaction isn't just playful surprise — it's a quiet unsettling of a long-held belief in the uniqueness of human writing. This is likely why the topic landed on Hacker News and sparked extensive comment threads: it hits a nerve for anyone who creates.

The rapid improvement in LLM writing ability is closely tied to how these models work. Trained on vast corpora of text, they learn, at their core, "what sequences of words are most likely to appear in human writing." Combined with instruction fine-tuning and Reinforcement Learning from Human Feedback (RLHF), models are further optimized to produce responses that feel satisfying. In other words, the optimization target is already aligned with producing text that looks like what humans expect to read — which objectively makes the output harder to detect by surface-level criteria, and explains why these identification quizzes grow more difficult with each model generation.

What Does It Mean That Detection Is Getting Harder?

It's worth examining how people actually judge whether a passage "sounds like AI": they tend to rely on surface cues — overly neat structure, a lack of personal detail, bland and noncommittal opinions, and certain high-frequency transition phrases. These traits were genuinely common in earlier models, but through better training and prompt engineering, they can now largely be eliminated or disguised.

Flipping the perspective: growing difficulty in detection doesn't necessarily mean AI is "writing better." The quiz tests whether you can distinguish, not which is superior. A smooth, flawless passage may be precisely what lacks the imperfectly human quality found in genuine writing — a distinctive observational angle, an unexpected associative leap, an expression with emotional warmth. These are dimensions that resist quantification and that a quiz cannot easily capture.

What Is Good Writing, Really?

Reducing "who writes better" to a guessing game sidesteps a more fundamental question: the value of writing lies not just in whether sentences flow smoothly, but in whether the writing carries genuine thought, lived experience, and judgment. The power of a compelling essay or a moving narrative often comes from the author's unique life experience and perspective — not merely from how the words are arranged.

In this sense, AI excels at "generating text that looks plausible," while the core competitive advantage of human writing lies in "expressing something worth expressing." When readers are stumped by the quiz, what's really being challenged may not be human writing ability — it may be the shallow set of standards we've habitually used to evaluate writing in the first place.

Philosopher John Searle's "Chinese Room" thought experiment is instructive here: a person who doesn't understand Chinese manipulates Chinese symbols according to a rulebook, and from the outside appears to respond perfectly to Chinese questions — yet no actual understanding exists inside the room. Large language models are in a strikingly similar position: they can arrange words in grammatically and semantically convincing ways, yet they hold no opinions and have nothing that is "worth expressing." The distinction between "fluent linguistic output" and "genuine expression of thought" is the key entry point for understanding the limits of AI writing — and the philosophical foundation of this article's central argument.

Practical Implications for Creators

For those who write professionally or for pleasure, this kind of quiz offers a pragmatic reminder. Rather than worrying about whether AI can fool readers, it's more productive to ask how to make your writing carry what machines cannot replicate — specific first-hand experience, a clear personal voice, independent judgment on complex questions. These are the genuinely scarce assets in a landscape flooded with AI-generated content.

At the same time, this places new demands on content platforms and readers alike. As the origin of text becomes harder to identify, attribution, credibility, and fact-checking become more important, not less. Distinguishing AI from human writing may ultimately prove a futile endeavor — but the ability to tell trustworthy from untrustworthy, valuable from worthless content is a far more worthwhile skill to develop.

Closing Thoughts

The New York Times quiz offers no definitive answers. It functions more like a mirror, reflecting the collective confusion we face in the age of generative AI. Who writes better? The answer may not matter much. What matters is that now that machines can produce text convincingly enough to pass as human, our understanding of writing, authenticity, and what makes human expression distinctive needs to evolve as well.

Share:

Related articles