AI Is Not Ready for Strategic Conflict: Five Failure Modes in Wargame Simulations

Position paper warns LM-driven strategic wargames have five critical failure modes and lack auditable safety guarantees.
A position paper (arXiv:2609.16189) issues a serious warning about using language models in open-ended strategic wargames. It argues that LMs simultaneously act as scriptwriter, referee, and reality-constructor, causing their biases and hallucinations to be embedded directly as simulated "facts." The paper identifies five failure modes: decision laundering, adjudication opacity, role collapse, escalation-through-adjudication, and failure of strategic imagination. It stresses that standard benchmarks cannot establish safety for high-risk strategic scenarios, and concludes that wargames should serve as stress-tests for decision-making AI—not as safety arguments for deploying LMs in policy or crisis response.
Open-ended strategic wargaming is emerging as a high-stakes frontier for large language model (LM) applications. These simulations involve complex factors including adversary modeling, institutional dynamics, crisis escalation, plan vulnerability assessment, military doctrine, and crisis response. A recent position paper (arXiv:2609.16189) issues a stark warning: no wargame driven by language models should be used to inform planning, doctrine, policy, or crisis response without a verifiable, auditable safety argument.
Why Language Models Are Both Attractive and Dangerous for Strategic Simulation
Language models are appealing in strategic simulation contexts because of their versatile capabilities: they can play agent roles, generate scenario branches, adjudicate ambiguous actions, and summarize lessons learned. These abilities make wargames—traditionally labor-intensive endeavors—scalable and automatable.
Yet the paper argues that these same capabilities make open-ended roles dangerous. The core issue is this: a model's language output simultaneously determines "what an agent is trying to do" and "what becomes the simulated reality." In other words, the language model is simultaneously scriptwriter, referee, and constructor of reality. When a model's wording both defines intent and determines outcome, its biases and hallucinations become directly embedded as "facts" in the simulated world—and from there, they can distort the judgment of decision-makers.

Five Failure Modes
The paper identifies five critical failure modes that together cast serious doubt on the reliability of current LM-driven strategic wargames:
Decision Laundering
Human decision-makers may use the AI's veneer of "objectivity" to seek validation for decisions they were already inclined to make. Model-generated analysis is treated as independent verification, when in practice it merely launders subjective judgment into what appears to be a neutral machine conclusion—eroding accountability for the decisions themselves.
Decision laundering is closely related to the cognitive science concept of "automation bias." Research shows that when faced with algorithmic or machine output, people tend to accept it more readily than equivalent human advice and apply less critical scrutiny. This tendency is especially pronounced in high-pressure crisis decision environments, where AI-generated analysis often appears detailed and impartial, lending it additional persuasive force. This mechanism makes it difficult to trace accountability in AI-assisted decision-making: when a bad decision leads to consequences, who bears responsibility—the human who accepted the advice, or the model that generated it? This ambiguity of accountability is itself a significant governance risk.
Adjudication Opacity
When language models adjudicate the outcomes of ambiguous actions, their reasoning is often uninterpretable. Why was a particular military operation judged a success or failure? The model's rulings lack traceable justification, making the overall wargame results difficult to audit or reproduce.
Role Collapse
In multi-agent simulations, models may be unable to consistently maintain the distinct perspectives of the adversaries or institutions they portray, causing different roles to converge or blur. This severely undermines the fidelity of the simulation as a representation of real adversarial dynamics.
Escalation-through-Adjudication
Models may systematically bias their adjudications toward conflict escalation—not because strategic logic demands it, but because of the language model's generative tendencies. This artificially induced escalation, introduced through the adjudication mechanism itself, can mislead analysts about the dynamics of real-world crises.
Failure of Strategic Imagination
Constrained by patterns in their training data, models struggle to generate genuinely novel, breakthrough strategic scenarios. In high-stakes strategic competition, it is precisely the "unimaginable" black swan situations that are most consequential—and this is exactly where current models fall short.
Wargaming is a military analysis tool with a long history, dating back to the Prussian army's Kriegsspiel in the 19th century. Traditional wargames rely on human experts to play adversaries and adjudicate outcomes; their core value lies in exposing plan vulnerabilities and stress-testing assumptions through adversarial simulation. Introducing language models allows wargames to scale dramatically—models can generate hundreds of scenario branches in seconds and automatically adjudicate results. But this automation also removes the situational judgment and accountability anchors that human experts provide in traditional wargames, systematically amplifying the impact of the failure modes described above.
Standard Benchmarks Cannot Establish Safety
The paper makes an important argument: conventional benchmarks cannot establish safety guarantees for high-risk scenarios like these. Traditional model evaluations focus on metrics like accuracy and fluency, but in strategic conflict simulation, the core question is not whether the model "gets the right answer"—it is whether the model's outputs will shape decision-making realities in ways that are uncontrollable and unauditable.
This means that even a model that performs well on standard evaluations cannot be presumed fit for wargames that influence policy and crisis response. Safety arguments must be tailored to specific deployment contexts and must include systematic examination of all five failure modes described above.
The dominant LLM evaluation frameworks today—MMLU, HellaSwag, BIG-Bench, and others—primarily measure knowledge coverage, logical reasoning, and linguistic fluency. These benchmarks are designed around the assumption that verifiable correct answers exist, and that the model's task is to approximate those answers within a closed problem space. Strategic conflict simulation fundamentally violates this assumption: simulation outcomes have no objective ground truth, model outputs exert their influence indirectly by shaping decision-makers' cognition, and harms often only become apparent after deployment. This compounds the AI safety problem of "out-of-distribution generalization": real crisis situations frequently lie outside the coverage of training data, and benchmarks have almost no predictive power over model behavior in such situations.
The Right Role for Wargames: Stress-Testing, Not Safety Arguments
The paper's constructive conclusion is that the appropriate use of open-ended wargames today is to stress-test LM agents that influence decision-making—not to deploy them directly in decisions with real consequences.
This distinction is critical. Wargames can surface model failures and serve as probes for uncovering problems; but they do not themselves constitute safety arguments for putting models into real-world use. The authors' position is clear and measured: using wargames to probe AI vulnerabilities is valuable, but treating wargame results directly as the basis for strategic decisions is dangerous.
For AI governance and defense technology communities, this paper provides a practical framework—establish auditable safety boundaries before embracing the efficiencies that language models offer. Leading on technical capability is not the same as qualifying for trustworthy deployment in high-risk domains.
This position resonates with the "red-teaming" ethos in AI safety. The core logic of red-teaming is: actively seek out the boundaries of system failure, rather than attempting to prove system safety. The paper's authors are effectively calling for wargames to be repositioned as an institutionalized red-team mechanism—one specifically designed to trigger and document language model failures in controlled environments, accumulating evidence for subsequent safety arguments. This mirrors the safety certification logic of high-risk domains like nuclear weapons, aviation, and medical devices: stress tests expose problems, but passing a stress test does not constitute a deployment license. For AI governance, this means establishing third-party audit mechanisms independent of developers—only then can trustworthy deployment of language models in strategic applications receive the institutional backing it requires.
Related articles

Letting AI Build AI Tools: A 7-Day, 31-Commit Bootstrapping Post-Mortem
An engineer ran a fully autonomous AI-builds-AI pipeline for 7 days, 31 commits, with a 1-in-6 success rate. This post-mortem covers 5 failure types, 11 structural rules, and how every mistake became a permanent immunity gate.

Building an AI-Powered E-Commerce Business from Scratch: A Real-World Account of Multi-Agent Architecture for Print-on-Demand
A blogger builds a print-on-demand e-commerce company from scratch using AI agents — documenting specialized Agent profiles, GPT-5.6 vs Claude Fable multi-model orchestration, and reusable skill accumulation.

AI Agent Earns $10K in One Week: 3 Key Upgrades Explained
A blogger shares how he earned $10K in a week with an AI Agent — not by adding more skills, but through verification, approval gates, and subagents to raise trust and enable true automation.