LLMs Fail the Urdu Test: How Much of Top Models' Multilingual Capabilities Is Hype?

Top LLMs fail at Urdu story generation with grammar, semantic, and cultural errors that few-shot prompting cannot fix.
An empirical study tested GPT-5.1, Qwen-3-Max, and DeepSeek-3.1 on open-ended Urdu story generation using the 93-story Urdu-Stories corpus, annotated with a nine-category error taxonomy across linguistic, semantic, and cultural dimensions. Results show frequent grammar errors, coherence breakdown, unnatural repetition, and pervasive cultural shallowness. Critically, few-shot prompting offered only marginal improvement — cultural and contextual errors persisted, revealing a structural pre-training data deficit that inference-time techniques cannot remedy. The study warns that "supports 100+ languages" claims often mask serious gaps, and calls for systematic corpus development and culturally-aware evaluation benchmarks.
The "Low-Resource" Blind Spot in Multilingual LLMs
Virtually every mainstream large language model (LLM) today markets itself as "multilingual," claiming the ability to handle dozens or even hundreds of languages. Yet how reliable that multilingual capability actually is for low-resource languages has never been rigorously tested at scale. A new study published on arXiv (arXiv:2609.10758) targets exactly this blind spot, selecting Urdu as a representative low-resource language to examine how today's leading LLMs actually perform on open-ended text generation tasks.
The core research question is straightforward: when we ask these models to tell stories in Urdu, how correct and reliable is what they produce? The answer is not encouraging.
Methodology: Three Top Models and a Nine-Category Error Taxonomy
To produce quantifiable evaluations, the research team built a corpus called Urdu-Stories, containing 93 Urdu-language stories generated by three state-of-the-art LLMs:
- GPT-5.1 (OpenAI)
- Qwen-3-Max (Alibaba)
- DeepSeek-3.1 (DeepSeek)
These three models represent the top tier of both Western and Chinese closed-source/open-source AI development, making them highly representative. Researchers then performed manual annotation of errors in these stories and established a taxonomy of nine error labels spanning three core dimensions.
The Three Evaluation Dimensions
- Linguistic: Grammatical, spelling, and morphological correctness
- Semantic: Coherence, logical consistency, and naturalness of expression
- Cultural: Depth of understanding of the cultural context, customs, and social background embedded in Urdu
The value of this layered annotation scheme is that it doesn't just tell us a model "got it wrong" — it pinpoints precisely at which level the model failed. This is critical for diagnosing the deeper deficiencies of multilingual models.
Core Findings: Pervasive Failures Across Grammar, Semantics, and Culture
The results reveal a series of alarming problems. At the most basic level, these so-called "state-of-the-art" models frequently produce fundamental grammatical and semantic errors in Urdu — a stark contrast to their near-flawless performance in high-resource languages like English.
Coherence Collapse and Unnatural Repetition
The generated stories broadly lacked coherence and exhibited unnatural repetition — a classic symptom of a model working in a language where training data is insufficient. When a model has an inadequate grasp of a language's probability distributions, it tends to fall into local loops, repeatedly generating similar sentence structures or phrases.
Pervasive Cultural Shallowness
The most noteworthy finding is pervasive cultural shallowness. Even when models managed to get the grammar roughly right, the stories they generated typically lacked any genuine understanding of the cultural backgrounds, values, and everyday customs of Urdu speakers. The output reads more like the product of a "translated voice" than a narrative rooted in indigenous culture.
Few-Shot Prompting Can't Fix Cultural Deficits
A natural follow-up question: could prompt engineering — specifically providing the model with a few examples — improve these issues? The research team tested few-shot prompting directly.
The results showed that while few-shot prompting helped in some respects, cultural and contextual errors remained largely unresolved. The implications are significant: cultural understanding failures are not simply a matter of "insufficient prompting." They are rooted in a severe lack of linguistic and cultural training data during the pre-training phase. In other words, this is a structural problem at the data and training level — one that cannot be patched by inference-time techniques.
Deeper Implications for the AI Industry
This research sends a wake-up call to the entire AI industry. As major vendors race to advertise that their models support "hundreds of languages," we must be skeptical of the substance behind that "multilingual" label.
Linguistic "Breadth" Is Not the Same as Understanding "Depth"
Supporting a language's input/output capability and truly mastering that language and its culture are two fundamentally different things. For high-resource languages like English and Chinese, massive training data ensures model depth. But for the thousands of low-resource languages around the world, models may have learned only the surface. Urdu, as a major language with over 200 million speakers, is already in this situation — the circumstances for languages with fewer speakers and even less digitized corpus are easy to imagine.
Reliability Risks in Content Generation and Information Retrieval
The researchers explicitly note that these findings highlight the limitations of current LLMs as reliable sources for content generation and information retrieval in low-resource languages. If users rely on these models to access information presented in low-resource languages, they may encounter grammatical errors, logical incoherence, and even cultural misinformation. This is particularly dangerous in applications such as education, journalism, and public services.
Toward Genuinely Equitable Multilingual AI
At its core, this Urdu-focused study is a profound interrogation of AI linguistic fairness. Technological progress should not only benefit the minority who speak dominant languages. Truly multilingual AI must treat all languages equally in terms of both linguistic accuracy and cultural depth.
Achieving this goal requires sustained industry investment in the following areas:
- Building high-quality corpora for low-resource languages
- Systematic injection of indigenous cultural knowledge
- Development of evaluation benchmarks with a cultural dimension
Until then, every claim about an LLM's "multilingual" capability probably deserves a healthy dose of skepticism.
Related articles

Open-Source Python SDK: Measuring AI Agent Reliability with SRE Principles
Agent Reliability is an open-source Python SDK that applies SRE's SLO and error budget concepts to AI Agent evaluation, with PASS/FAIL/UNKNOWN states, CI assertions, and zero forced dependencies.

MiniMax RefMod: A Complete Guide to Training-Free Reusable Identity Workflows
MiniMax RefMod offers training-free reusable identity workflows for image, video, and audio generation. Includes Runpod template and tutorial for quick setup.

Invalid Source Material Notice
The source material provided lacks substantive information and is unrelated to AI/tech topics, making it impossible to produce a complete professional article.