GPT-5.6 Requires Government Approval: AI Regulation Enters Uncharted Waters

GPT-5.6 needs government approval and Claude Mythos is pulled after breaching classified systems — AI regulation deepens.
OpenAI's GPT-5.6 launched but requires case-by-case government approval, while Claude Mythos was removed after breaching NSA classified systems in red team tests. This roundup covers tightening regulation, cost-driven model downgrades, copyright lawsuits, the Qualcomm-Modular deal, and AI's growing employment impact.
GPT-5.6 Launches — But as an "Approval-Only" Model
OpenAI released its most powerful next-generation model, GPT-5.6, on June 27. But an intriguing detail cast this launch in a light unlike any before: ordinary users cannot access it for the time being.
At the government's request, GPT-5.6 is only available in limited preview to a small group of approved partners, and even then requires case-by-case government approval for each customer. This means that, for the first time, a model's strength is no longer determined purely by compute and parameters, but by a regulator's admission list.
What's interesting is that OpenAI itself expressed reservations about this arrangement, explicitly stating that "this government access process should not become a long-term default." Behind this statement lies the increasingly sharp tension between frontier AI capabilities and national security review — when a model becomes powerful enough, its release is no longer a purely commercial act, but something subject to case-by-case government scrutiny.
It's worth noting that this "approval-based" system didn't emerge out of nowhere. In recent years, the U.S. government's regulatory framework for frontier AI has gradually taken shape — from the AI executive order signed by the Biden administration in 2023, to the AI Safety Institute (AISI) established under the National Institute of Standards and Technology (NIST). Regulators are increasingly inclined to conduct safety evaluations before the release of the most powerful models. The case-by-case approval of GPT-5.6 can be seen as a landmark escalation of this trend, moving from "evaluation" toward "access control."

Why Did AI Regulation Suddenly Tighten?
This policy shift is very likely directly related to another, even more explosive piece of news.
Behind Claude Mythos's Removal: Classified Systems in Jeopardy
Anthropic's Claude Mythos model has been pulled from availability for two full weeks, still unresolved — and the situation has only grown more deadlocked. This time, the cause has finally come to light.
According to The Economist, a senator revealed that the head of the NSA (U.S. National Security Agency) told him that during red team testing, the Mythos model breached nearly all of the NSA's classified systems within just a few hours.
It's worth explaining what "red team testing" actually means. Red teaming originates from the military and cybersecurity fields, referring to an adversarial evaluation method in which a dedicated team simulates an attacker's perspective to actively seek out system vulnerabilities. In the AI field, red teaming is used to probe the boundaries of a model's capabilities and potential harms — for example, whether it can be induced to generate malicious code, bypass safety guardrails, or assist in cyberattacks. It's important to note that red team testing is typically conducted in controlled environments, where testers are often informed in advance of the system architecture and known vulnerabilities, or granted privileged access. As a result, the test outcomes represent a "theoretical ceiling" rather than real-world attack performance. Companies like OpenAI and Anthropic incorporate red teaming into their pre-release safety evaluation processes, and the U.S. government also participates in frontier model evaluation through agencies such as the AI Safety Institute (AISI).
For precisely this reason, the editor who wrote that sentence subsequently issued a clarification, emphasizing that it should not be taken literally and must be understood in the context of specific test conditions — red teaming is often conducted in controlled environments where vulnerability information is provided, and the results are not equivalent to real-world attack capabilities.
But even discounted, this news remains highly impactful. It explains why regulators would suddenly become so wary of frontier models: when an AI model demonstrates the ability to "breach classified systems in a matter of hours" during offensive-defensive testing, the government's cautious, case-by-case approval approach to GPT-5.6 becomes much easier to understand. Viewed together, these two news items sketch out a reality in which AI capabilities have already touched the red line of national security.
Under Cost Pressure, Frontier Models Are No Longer the Only Choice
While AI regulation tightens, the exact opposite is happening at the other end of the market — some companies are proactively abandoning expensive frontier models.
AI startup Lindy announced it is replacing Claude entirely with DeepSeek. They said this would save them several million dollars a year.

DeepSeek is a series of large models developed by the Chinese company DeepSeek, which has shaken the industry with its highly competitive inference costs. Its core breakthrough lies in adopting a Mixture-of-Experts (MoE) architecture and efficient training methods, keeping capabilities close to top-tier models while driving API pricing down to a fraction — sometimes a few dozenths — of that of frontier models. This cost-performance advantage is hugely significant for enterprise applications: many business scenarios (such as text classification, information extraction, and customer service Q&A) don't require the full capabilities of the most powerful models, and once call volumes scale up, the cost difference is dramatically amplified.
This case sends a clear signal: under cost pressure, companies are beginning to proactively "downgrade," replacing frontier models with cheaper ones. For many real-world business scenarios, the marginal capability improvements offered by top-tier models are not enough to justify their steep cost premium. The rise of cost-effective models like DeepSeek is reshaping enterprises' technology selection logic — capability that's "good enough" is sufficient, and cost is the hard constraint.
Mirror Code: A New Benchmark for Measuring AI's "Endurance"
As AI's coding capabilities improve, how to evaluate a model's "stamina" has become a new question. Apoc.ai has launched a new benchmark — Mirror Code — specifically designed to test whether a model can work continuously over long periods.
In testing, one model coded non-stop on a single task for a full 19 days, costing roughly $2,600 just to run. This benchmark measures the endurance and stability of long-horizon autonomous coding.

To understand the significance of this benchmark, one must first grasp the concept of the AI Agent. An AI Agent refers to an AI system capable of autonomously planning, executing multi-step tasks, and interacting with its environment — widely regarded as the next frontier for putting large models into practice. Unlike traditional single-turn Q&A, an Agent must maintain goal consistency over long time spans, manage contextual memory, self-correct, and call external tools. "Long-horizon autonomy" is precisely the biggest bottleneck for current Agents — models are prone to goal drift, error accumulation, and context loss during long tasks. Previously, common coding benchmarks in the industry, such as SWE-bench, mainly evaluated the ability to solve a single problem, whereas tests like Mirror Code extend the evaluation dimension to the endurance level of "working continuously for days."
This marks the entry of AI coding evaluation into a new dimension. In the past, we focused on whether a model could write a correct piece of code; now the focus is on whether it can, like a human engineer, maintain stable output over a lengthy project cycle without deviating from goals or accumulating errors. This kind of "long-horizon autonomy" is precisely the key threshold for Agents to become truly practical.
Copyright War Escalates: Nearly 400 Newspapers File Joint Lawsuit
The battle over content copyright continues to spread. Nearly 400 U.S. newspapers have joined forces to sue OpenAI and Microsoft, accusing the two of systematically "stealing and scraping" newspaper articles to train their models.
More pointedly, The New York Times separately added an accusation, claiming that Microsoft specifically built a supercomputer to help OpenAI carry out the infringement. This lawsuit pushes the copyright dispute from the level of "data scraping" to the more serious charge of "infrastructure conspiracy."
The copyright dispute over AI training data has been one of the most closely watched legal issues of the past two years. The core of the controversy is whether scraping copyrighted content without authorization for model training constitutes infringement, or whether it can be exempted under the "Fair Use" principle of U.S. copyright law. The New York Times sued OpenAI and Microsoft as early as the end of 2023, and this collective lawsuit by nearly 400 newspapers further widens the front. The "infrastructure conspiracy" charge is particularly critical — it seeks to reclassify Microsoft from a mere technology provider to an active participant in the infringement, which could significantly increase the defendants' legal liability. The verdicts in such lawsuits will profoundly affect the data acquisition models and the legal foundations of commercial legitimacy across the entire generative AI industry. The legality of AI training data is becoming a long-term risk hanging over all large model companies.
Qualcomm Acquires Modular, Targeting Integrated Edge-Cloud AI
There's also a major move at the hardware level. Qualcomm announced its acquisition of Modular and plans to team up with Hugging Face.
Modular builds an AI software stack that can run across various chips, enabling "develop once, deploy anywhere" across different hardware such as CPUs and GPUs. Modular was founded by Chris Lattner, one of the co-creators of TensorFlow. Its Mojo programming language and MAX inference engine aim at their core to break AI software's deep lock-in to specific hardware (especially NVIDIA's CUDA ecosystem), enabling developers to efficiently deploy models across heterogeneous hardware. Through this acquisition, Qualcomm clearly intends to extend AI from the device edge all the way to the cloud, connecting an integrated edge-cloud AI deployment pipeline.

This also reflects chip giants' strategic anxiety and ambition — simply selling chips is no longer enough; mastering a cross-hardware software ecosystem is the key to building a moat. Against the backdrop of NVIDIA's de facto monopoly built on the CUDA ecosystem, competitors like Qualcomm and AMD must offer sufficiently open and efficient alternatives at the software level to break through — which is precisely where the strategic value of neutral software stacks like Modular lies.
Capital Frenzy and Employment Shock: Two Sides of the AI Wave
The final two news items — one cold, one hot — form a striking contrast.
The hot side: Chinese AI company Zhongke Wenge went public in Hong Kong today, dubbed the "first decision-making large model stock." It surged 81% at the open, and its public offering was oversubscribed by nearly 6,000 times. This astonishing figure vividly reflects just how hot domestic AI is in the capital markets. Oversubscription of nearly 6,000 times is extremely rare in the history of Hong Kong IPOs, reflecting the market's intense enthusiasm for the AI concept — but such a lopsided subscription multiple often also foreshadows high volatility risk in the subsequent share price.
The cold side: Anthropic flatly stated that it no longer needs to hire junior engineers, because AI can already do the work previously assigned to junior positions. They also issued a warning that when other industries follow suit, it could bring about a wave of economic shock. Anthropic CEO Dario Amodei has previously publicly warned multiple times that AI could eliminate large numbers of entry-level white-collar jobs within the next few years. This statement echoes the company's own hiring decisions and turns "AI replacing jobs" from an abstract discussion into an unfolding reality.
The frenzy of capital and the anxiety over jobs appearing simultaneously are precisely the two sides of the current AI wave: while technology creates enormous value, it is also quietly rewriting the employment structure. The social pain brought by this transformation may only just be beginning.
Key Takeaways
Related articles

Go Microservices in Practice: Detailed Architecture for E-Commerce, AI Agent, and IM System Integration
Deep dive into integrating e-commerce, AI Agent, and IM systems under Go microservices architecture, covering unified auth, gRPC, componentized Agent engines, and group chat bots.

X Platform's Recommendation Algorithm Caught Filtering Brazilian Election Content, Reigniting Algorithm Transparency Debate
X (formerly Twitter) was found filtering Brazilian election content in its For You feed, sparking debate over algorithm transparency and free speech.

Poison-Resistant Concept Anchoring: A New Approach to Defending Against AI Data Poisoning
Deep dive into Poison-Resistant Concept Anchoring, defending against data poisoning via signed anchors and bounded updates. Experiments show 62% poison isolation with 0% false rejection rate.