LLMs vs. TCM Physicians: Real Clinical Case Benchmarks Reveal AI's Promise and Risks

Top LLMs rival TCM physicians on diagnostic reasoning but pose safety risks in prescription writing.
A study benchmarking 16 LLMs against 60 licensed TCM physicians across 349 real outpatient cases from 62 hospitals found that top general-purpose models outscored physicians on treatment principles and medical advice, demonstrating strong knowledge integration. However, models showed significant gaps in herb selection, dosing, and treatment strategy at the prescription stage, while also exhibiting hallucination and templated output risks. Researchers recommend positioning LLMs strictly as decision-support tools, requiring physician oversight, safety constraints, and prospective clinical evaluation before clinical use.
Large language models (LLMs) are rapidly making inroads into healthcare, but their real-world performance in traditional Chinese medicine (TCM) — a field heavily reliant on pattern differentiation and accumulated clinical experience — has lacked systematic validation. A large-scale benchmark study based on real outpatient cases now offers the most detailed answer to date: on certain diagnostic dimensions, top general-purpose LLMs scored higher than licensed TCM physicians in expert evaluations. Yet in the critical area of prescription writing, the models revealed significant gaps in drug selection, dosing, and treatment strategy — exposing safety risks that cannot be overlooked.
Study Design: A Rigorous Test Across 349 Real Clinical Cases
The value of this study lies first in the authenticity and scale of its data. The research team assembled a clinical case library comprising 349 de-identified outpatient records from 62 hospitals, then selected 60 representative cases as the evaluation set.
The evaluation covered 16 LLMs alongside a control group of 60 licensed TCM physicians. To ensure fairness, both model outputs and physician reports were anonymized before being submitted to five senior TCM experts, who scored them across nine diagnostic and treatment dimensions. This "double-blind" design effectively prevented raters from being biased by knowing whether a response came from an AI or a human, keeping the scoring focused on the quality of the clinical reasoning itself.

Using real clinical cases rather than standardized exam questions is what distinguishes this study from many existing benchmarks. TCM diagnosis involves integrating observations across multiple sensory modalities, and the complexity and individual variation of real cases far exceeds what question-bank-style evaluations can capture — making this design a more accurate reflection of a model's practical capabilities.
Finding One: Top LLMs Outperform Physicians on Certain Dimensions
The study's most striking conclusion is that frontier general-purpose LLMs scored higher than the physician control group on expert evaluations, particularly excelling in medical advice, treatment principles, and certain diagnostic tasks.
This does not mean AI has comprehensively surpassed TCM practitioners. A more measured interpretation is that LLMs hold a natural advantage in information synthesis, structured articulation, and breadth of knowledge coverage. Given a case, a model can rapidly draw on vast amounts of clinical literature to deliver a well-organized, logically coherent analysis — precisely the kind of output that expert reviewers tend to value.
From another angle, this reflects the fact that general-purpose LLMs, trained on large volumes of TCM classical texts and clinical documents, have achieved a solid grasp of the conceptual framework of pattern differentiation and treatment. When it comes to elaborating on treatment principles and organizing diagnostic reasoning, these models demonstrate real potential as decision-support tools.
Pattern differentiation and treatment (辨证论治) is the core clinical methodology of TCM. It involves collecting symptoms and signs through four diagnostic methods (observation, listening/smelling, inquiry, and palpation), categorizing them into specific "pattern types" (e.g., qi deficiency with blood stasis, liver yang hyperactivity), and then determining treatment principles and prescriptions accordingly. This process relies heavily on the physician's subjective clinical judgment, and differs fundamentally from Western medicine's reliance on objective diagnostic tests. LLMs perform relatively well in this area because TCM literature contains extensive structured knowledge mappings between pattern types and treatment principles — the kind of pattern-matching that models can reproduce reasonably well. However, real outpatient cases often present complex scenarios such as overlapping or shifting pattern types. This kind of dynamic judgment goes well beyond textual pattern matching, and is one of the root reasons why models fall short on prescription specifics.
Finding Two: Prescriptions Expose Critical Weaknesses
What truly separates AI from human physicians is the fine-grained analysis of prescriptions. The study found significant divergence between model outputs and expert judgment on herb selection, dosage calibration, and treatment strategy.
A TCM prescription is the final expression of pattern-differentiated treatment. The hierarchical composition of herbs (sovereign, minister, assistant, and envoy) and the precise calibration of dosages depend on years of clinical experience and a dynamic read of the individual patient's condition. While models can produce prescriptions that appear reasonable on the surface, they lack the precision and caution that real clinical practice demands when it comes to deciding which herbs to include and in what quantities.
Even more concerning are two categories of issues identified in qualitative safety reviews: hallucination, where models generate content that sounds professional but is factually incorrect or fabricated; and templated output, where models default to fixed frameworks rather than tailoring responses to the specific case at hand. In a medical context, both problems can pose direct risks to patient safety.
The TCM prescribing framework of "sovereign, minister, assistant, and envoy" (君臣佐使) is key to understanding this weakness. The sovereign herb targets the primary disease or pattern; the minister herb supports and enhances the sovereign's effect; the assistant herb addresses secondary patterns or counteracts the toxicity of the sovereign and minister herbs; and the envoy herb guides the formula to the target organ or harmonizes the other herbs. This system requires physicians to dynamically balance the role, dosage ratio, and compatibility constraints of each herb within the overall pattern differentiation. When generating prescriptions, LLMs can often reproduce the knowledge mapping of "this pattern calls for this formula," but cannot truly grasp the dynamic balancing logic underlying the composition — for instance, the degree to which aconite (附子) is processed directly affects its toxicity, and the traditional dictum that asarum (细辛) should not exceed one qian (约3克) in dosage reflects a type of experiential clinical constraint that is difficult to reliably internalize from text-based training alone. Hallucination is especially dangerous here, because an incorrect herb combination or dosage is not merely ineffective — it can cause real harm to patients.
Implications for AI Applications in TCM
The study's overall verdict is a cautious one: LLMs hold genuine potential as decision-support tools in TCM, but must not be used independently without physician oversight.
The research team explicitly identified three necessary conditions: physician oversight, safety constraints, and prospective clinical evaluation. This positions LLMs as assistive tools — helping physicians organize diagnostic reasoning and providing reference on treatment principles — while keeping prescription decisions and medication safety firmly in the hands of human experts.
From an industry perspective, this finding offers clear guidance for the practical deployment of TCM AI products. Pursuing "AI as a physician replacement" is neither realistic nor safe. The truly valuable direction is building human-AI collaborative clinical support systems — ones that let models contribute where they excel (knowledge integration and reasoning scaffolding) while applying safety mechanisms to constrain their autonomy in high-stakes areas like prescription writing.
Conclusion
This benchmark, spanning 62 hospitals and 349 real clinical cases, paints a relatively complete picture of LLM capabilities in traditional Chinese medicine: commendable breadth of knowledge and structured expression, but significant gaps in the precision and safety required for actual clinical deployment. For developers and institutions exploring intelligent TCM systems, this is both an encouraging signal and a timely warning — use AI where it performs well, and build guardrails where it doesn't.
Related articles

The Open Source Dilemma: A Non-Autoregressive Architecture Pioneer Overshadowed by Frontier Labs
An indie developer claims a frontier lab repackaged his year-old open-source non-autoregressive RL architecture as a breakthrough. We compare PPO sequence embeddings vs. RLCD parallel sampling and examine open source attribution gaps.

AI Plans an Entire Vineyard: A Real-World Experiment with 100 Grapevines
A Spokane hobbyist let Muse AI plan his entire vineyard — variety, spacing, irrigation, even the logo. He planted 100 Cabernet Franc vines and is documenting everything publicly.

Iceland's Treble Raises $18M to Bet on Voice Simulation Platform
Iceland-based voice simulation company Treble raises $18M. Its platform serves voice AI developers, AI wearables, and robotics firms. A deep dive into the technology and what the funding signals.