Gemini Refuses to OCR Shakespeare: Is AI Copyright Censorship Out of Control?

Gemini refused to OCR Shakespeare on copyright grounds, exposing AI over-alignment and false-positive filtering failures.
Google Gemini recently refused to perform OCR on Shakespeare's Julius Caesar — a work in the public domain for centuries — citing "copyright issues." The incident reveals not a genuine copyright assessment, but a coarse pattern-matching filter that flags any book-like content for blanket rejection. More broadly, it highlights a systemic over-alignment problem: as AI vendors lower safety thresholds to avoid legal risk, models grow increasingly restrictive, eroding usability. Combined with a lack of transparency around policy changes and no clear appeal mechanism for users, this case exposes critical gaps in contextual judgment, public domain exemptions, and communication practices in AI content moderation design.
A Farcical AI Censorship Case
A Reddit user recently reported a baffling incident: Google's Gemini suddenly refused to perform OCR (Optical Character Recognition) on uploaded books, citing "copyright issues." The kicker? The rejected content included Shakespeare's works written in 1599 — over 400 years ago — which have long since entered the public domain.
The user even provided video evidence, noting that the problem appeared not only in the Gemini app but also in Google's AI Studio (the developer-facing experimental platform). With a touch of sarcasm, they wrote: "How am I supposed to study my boy Brutus?" (referring to the character from Julius Caesar)

What might seem like a minor complaint actually reflects an increasingly prominent structural problem in how large language models handle content moderation: overly conservative safety filters are incorrectly flagging legitimate, harmless, and even fully public-domain content.
Why Public Domain Works Get Misidentified as Copyright Infringement by Gemini
The "Better Safe Than Sorry" Logic of Copyright Filters
Shakespeare's Julius Caesar was written in 1599. The author has been dead for over 400 years, and under any nation's copyright law, the original text is unambiguously in the public domain — free for anyone to copy, distribute, and use. So why would Gemini refuse to process it on copyright grounds?
The most plausible explanation is that Gemini's content moderation system doesn't actually evaluate whether a work is truly under copyright. Instead, it relies on a coarse-grained keyword and pattern-matching mechanism. When the system detects that a user has uploaded something resembling a "full book," it likely triggers a generic rule targeting "large-scale text reproduction" — a rule that doesn't bother to verify whether the text is actually protected.
In other words, the model takes an "err on the side of refusal" approach: if the content looks like it might infringe, reject it outright without deeper contextual reasoning. This is a textbook false positive problem.
OCR Tasks Shouldn't Trigger Copyright Blocks in the First Place
What makes this especially puzzling is that the user's request was simply OCR — converting text from an image or scan into editable characters. This is fundamentally a technical format conversion, not "generating pirated content." Equating an OCR request with copyright infringement reveals a glaring lack of contextual awareness in the moderation logic: the system fails to distinguish between "helping a user digitize a book they legally own" and "facilitating large-scale piracy distribution."
A Sudden Change: What Did Gemini's Safety Policy Actually Adjust?
The user specifically highlighted "since yesterday" as a turning point, suggesting this wasn't a long-standing restriction but rather a side effect of a recent model update or policy change.
Possible Causes
This kind of sudden behavioral shift is not uncommon in large language model products. It typically stems from one or more of the following:
- Silent safety policy updates: Vendors adjust content filtering rules in the background without public announcements — users only notice through degraded experiences.
- Responding to copyright litigation pressure: AI companies have faced mounting legal risk around copyright in recent years and may have tightened their handling of book-related content for compliance reasons, without building in any exemption for public domain works.
- Model regression or A/B testing: A newer model version may have been tuned to be more conservative on safety alignment, leading to a higher false positive rate.
The user also mentioned that the upload function "wasn't working either" and complained about instability in Chrome and ClipChamp — suggesting that some issues may have been compounded by client-side technical failures, making the overall experience even worse.
Not an Isolated Incident: The Systemic Problem of AI Over-Alignment
The Cost of Over-Alignment
This incident is a microcosm of the broader AI over-alignment problem. To avoid legal liability and public backlash, vendors keep lowering safety thresholds — and the result is an increasingly "timid" model that routinely refuses reasonable requests and declines to process harmless content.
For everyday users, this directly undermines usability. A tool that won't even recognize text from a 400-year-old classic of the public domain is significantly diminished in value. Studying Shakespeare, researching historical documents, digitizing ancient texts — these are precisely the scenarios where AI should be most useful, yet they get blocked by blunt-instrument filtering rules.
Lack of Transparency Makes the Trust Crisis Worse
The deeper issue is the absence of transparency. Users don't know why the rules changed, when they might be restored, or how to appeal. When AI refuses service with a justification as clearly untenable as "copyright issues," users' first reactions are confusion and frustration — which then erode trust in the product and brand as a whole.
Lessons for AI Content Moderation Design
Small as this incident may seem, it offers several important lessons for designing content moderation systems in AI products:
- Filters need contextual reasoning: Content form alone shouldn't trigger a block. Systems should incorporate multi-dimensional signals — task type (e.g., OCR), content age, copyright status — before making a decision.
- Public domain content should have explicit exemptions: For works definitively in the public domain, systems should implement a whitelist or time-based determination mechanism to avoid penalizing humanity's shared cultural heritage.
- Refusals should come with actionable explanations: When a model declines a request, it should provide a clear, verifiable reason and offer users a feedback or appeal channel.
- Major policy changes require communication: Silently tightening restrictions severely damages user trust; transparent update notes can go a long way toward managing user expectations.
Conclusion
Technically speaking, having an AI recognize a passage of Shakespeare is about as simple a task as it gets. But from a product and compliance standpoint, it sits at the intersection of copyright risk, safety alignment, and user experience — a genuinely complex tradeoff. The absurd situation this Reddit user encountered is a reminder that AI "safety" should not come at the cost of basic usability. When a tool is too afraid to touch classics that humanity has freely shared for centuries, what it needs isn't tighter constraints — it needs smarter judgment.
Related articles

Insufficient Source Material to Generate a Valid Article
The provided source material is a single unrelated tweet with no AI or tech relevance — insufficient to support a complete, valid technical article.

Insufficient Source Material to Generate a Valid AI/Tech Article
This source material is a tweet about the ages of Underworld members — unrelated to AI or tech, and insufficient to support a full article.

Insufficient Material: Unable to Generate a Valid AI/Tech Article
The provided material is a condolence tweet about a San Diego mosque attack — unrelated to AI/tech and too limited to generate a valid technical article.