GPT-5.6 Released: ChatGPT Merges with Codex, Evolving into an AI Workbench

GPT-5.6 merges ChatGPT and Codex into a unified AI workbench with three clear model tiers.
OpenAI's GPT-5.6 officially merges ChatGPT with Codex, transforming it from a chat box into a complete AI workbench. With three new tiers—Soul, Terra, and Nuna—it supports coding, task automation, and web deployment, going head-to-head with Anthropic's Claude Fable 5 in the Agent era.
OpenAI's Strategic Shift: From Chat Box to Workbench
OpenAI has just released GPT-5.6, but what truly deserves attention this time isn't how much the model's capabilities have improved—it's a key product strategy decision: officially merging ChatGPT with the programming tool Codex. Behind this move is OpenAI's desire to catch up with Anthropic in the commercialization market, where the latter has established a certain advantage.
This merger means ChatGPT is no longer just a simple chat box, but is beginning to transform into a true "AI workbench." Users can freely switch between the workspace on the left side of the interface and Codex, forming a complete productivity workflow and finally moving beyond the old single-purpose Q&A experience.

Not Just Answering, But Getting Work Done
The new version's Chat mode is positioned well beyond the traditional scope of "answering questions." It can directly write code, edit documents, create spreadsheets, generate slide decks, and even let an AI Agent take over the computer and execute a series of tasks end-to-end. With the newly added Site feature, users can also deploy their output directly as an accessible web page.
The leap from "generating content" to "completing tasks" is precisely the core focus of the current AI product competition. Models are no longer merely expected to provide information—they must be able to deliver genuinely usable results.
Three-Tier Model Naming: The Logic Is Finally Clear
GPT-5.6 has also made significant improvements in model tiering, with the three versions named as follows:
- Soul: The flagship version, aimed at the most complex professional workflows
- Terra: The everyday balanced version, offering both performance and efficiency
- Nuna: The economical version focused on cost-effectiveness

Compared to the past confusing naming system—with the O series, 4O, Mini, and High all jumbled together, making it hard for users to tell them apart—this three-tier logic is far clearer. Users can quickly match the appropriate version based on task complexity and budget. This maturity in product thinking is itself an important step for OpenAI on the path to commercialization.
Performance: Going Head-to-Head with Claude Fable 5
On the performance front, GPT-5.6 Soul focuses on long-range professional workflow capabilities. In the LexExam Agent capability test, GPT-5.6 Soul scored 13.1 points higher than Anthropic's Claude Fable 5. On programming Agent-related metrics, OpenAI also emphasized a lower error rollback rate, shorter task completion times, and better cost control.

Beyond Benchmarks, Usability Is the Real Threshold
However, let's be frank: the reference value of these benchmark scores is quite limited. Current mainstream evaluation benchmarks already struggle to objectively measure the real gaps between top-tier models. For first-tier models like Claude Fable 5 and GPT-5.6, what truly matters isn't the numbers, but real-world stability and usability.

This precisely highlights the most fundamental shift in the AI industry right now: entering the Agent era, a model can't just be "smarter"—it must also be able to run stably over long periods, keep costs under control, and genuinely integrate into actual workflows. A model that is occasionally brilliant but frequently crashes is far less valuable than one that can reliably deliver results.
The Agent Era: A Comprehensive Upgrade in Competition Dimensions
From the release of GPT-5.6, it's clear that competition among large AI models is entering a new phase. In the past, everyone competed on parameter scale and benchmark scores, but the key words have now become: stability, long-range task handling, cost efficiency, and end-to-end task delivery.
By merging ChatGPT with Codex, OpenAI is essentially building a complete AI productivity platform, attempting to reclaim the initiative in the commercialization market from Anthropic. The clear division of the three model tiers shows that OpenAI is becoming more pragmatic and more user-centric in its product strategy.
Who's Stronger? Real-World Practice Is the Ultimate Judge
Benchmark data is, after all, just a reference on paper. Whether GPT-5.6 Soul or Claude Fable 5 truly performs better in real-world scenarios still needs to be tested through hands-on coding. The upcoming real task comparison tests may be the answer most worth looking forward to.
For developers and enterprise users, the ultimate beneficiaries of this competition are none other than ourselves—as two top vendors fiercely compete on stability, usability, and cost, the practical value of the AI workbench will continue to be pushed to new heights.
Related articles

The Truth Behind Codex 'Build a Website in 5 Minutes': AI Isn't Creating Sites—It's Helping You Copy Them
Exposing the truth behind viral Codex 5-minute website videos: creators aren't building original sites with AI—they're copying shared prompts or scraping others' work. Learn AI coding tools' real limits.

Getting Started with AI Agent Development: A Complete Guide from Concept to Practice
A comprehensive guide to AI Agent architecture and development, covering automated marketing, intelligent customer service, and investment analysis scenarios with single and multi-agent collaboration.

The Truth Behind Codex 'Build a Website in 5 Minutes': AI Isn't Creating Sites — It's Helping You Copy Them
Exposing the truth behind viral Codex 5-minute website videos: creators aren't building original sites with AI — they're copying shared prompts or scraping others' work.