GPT-5.6 Three-Model Hands-On Test: A Complete Comparison of Sol, Terra, and Luna

In-depth testing of GPT-5.6's Sol, Terra, and Luna: Sol wins overall at nearly half the price.
A product manager put OpenAI's GPT-5.6 series (Sol, Terra, Luna) through weeks of hands-on testing against Fable 5 and Sonnet 5. Using a 70% human + 30% LLM-judge weighted framework across PRD writing, prototyping, debugging, and browser use, Sol won overall—at nearly half Fable's price—thanks to its practical, get-things-done approach.
Three New Models Debut: Sol, Terra, and Luna
OpenAI has released three new versions of the GPT-5.6 series, forming a clear capability tier. According to a product manager blogger who spent several weeks putting them through their paces, each model has its own positioning:
- Sol: The next-generation frontier flagship model, aimed at the most complex tasks;
- Terra: A balanced model, aimed at efficient day-to-day work;
- Luna: A lightweight option that's inexpensive and suited for high-concurrency, high-volume tasks.
The author's testing focused on the head-to-head between Sol and its competitor Fable—the model tier she actually uses in her daily work.
Sol Is Nearly Half the Price of Fable: A Look at the Cost Advantage
Pricing is a major highlight of this test. The author provides clear comparison data:
- Sol: $5 per million input tokens, $30 per million output tokens;
- Fable: $10 per million input tokens, $50 per million output tokens.
By API pricing, Sol costs almost half as much as Fable.
Token Pricing and the Economics of Large Models: LLM API pricing is measured in units of "per million tokens." A token is the basic unit a model uses to process text, roughly corresponding to about 3/4 of an English word or 1-2 Chinese characters. Input tokens (prompt) and output tokens (completion) are billed separately, because generating text consumes more compute than understanding it—the model must generate autoregressively token by token rather than processing in parallel. As a result, output pricing is typically 4-10x the input price. Sol's 6x spread ($5→$30) and Fable's 5x spread ($10→$50) both align with industry norms. For product teams making high-frequency API calls, the input-side half-price advantage alone can save thousands of dollars at a scale of millions of calls—which is precisely why pricing has become a core competitive edge.
The author also notes that OpenAI subscriptions come with a built-in usage allowance, and the combination of low price and high performance makes up Sol's first line of competitiveness.
Evaluation Methodology: The "How I AI Vibe Review" Framework
To move beyond pure gut judgment, the author built a semi-automated evaluation framework covering real product scenarios.
Five Testing Dimensions
- PRD Writing: The quality of generating product requirement documents;
- Prototype Design: Complete designs and wireframes for various app ideas;
- Code Debugging: Accuracy of multi-step agentic debugging;
- Agentic Voice: The model's human-like conversational ability;
- Overall Aesthetics.
Why PRD and Prototype Design Matter as AI Capability Test Dimensions: Product requirement documents (PRDs) and prototype designs are core deliverables in a product manager's daily work, and they're excellent test dimensions for gauging whether an AI "truly understands product thinking." A PRD requires the model to structurally decompose user needs, define feature boundaries, balance technical feasibility against business goals, and organize the document with clear logic—this reflects not just language ability but business comprehension. Prototype design involves visual hierarchy, user journeys, and interaction logic, exposing differences in a model's spatial imagination and design taste. It's a dimension harder to quantify than code generation, yet closer to real product value.
The participating models included Fable 5, Sonnet 5, and the three versions of GPT-5.6. The evaluation used a hybrid approach of LLM judging (with GPT-5.5 as the "strictest judge") plus human taste scoring, ultimately combined into a Clairvaux weighted index at a 70% human + 30% machine weighting.
The LLM-as-Judge Evaluation Methodology: Using large language models as judges is an important method that has emerged in AI evaluation in recent years, first systematized by Zheng et al. in the 2023 MT-Bench paper. Traditional human evaluation is time-consuming and highly subjective, whereas LLM judges can perform standardized scoring at scale and low cost. But this method has known biases: models tend to give higher scores to themselves or to models in the same family (self-enhancement bias), and they score well-formatted long outputs higher (verbosity bias). The author's use of a 70% human + 30% machine hybrid weighting is precisely intended to harness the LLM's efficiency while suppressing its biases—and using GPT-5.5 rather than one of the contestant models as the judge is a key design choice to avoid self-evaluation bias.

Test Conclusion: Sol Wins Overall
Under this weighted index, GPT-5.6 Sol took the highest taste score by a significant margin. The author emphasizes this was a blind test—she only learned which output corresponded to which model after scoring was complete, making the conclusion more convincing.
The winners by task are as follows:
- Prototype Design: Sol wins, with the most complete features and the most thoughtful design;
- PRD Writing: Terra was preferred, for being concise and direct;
- Code Bug Fixing: Sonnet 5 earned the highest score from the LLM judge;
- Agentic Voice: Sonnet 5 was described as "aside from the em dashes, you're just like a real person."

In the full-fidelity prototype comparison, Sol won three or four out of five times. Worth noting: Sol has an intense fondness for a "woodland" forest-green color scheme—this stylistic consistency suggests the GPT-5.6 series may have reinforced a particular aesthetic tendency during training data curation or the RLHF stage. For product teams that need consistent brand tone, this is both an advantage and a potential constraint to keep in mind.
Sol's Core Strength: Practical Rather Than Flashy
The author sums up the essential difference between Sol and Fable in one line: "Fable is extremely smart in theory; Sol actually works in practice."
Writing Like a Normal Person
Fable's expression is extremely technical and pedantic, "like an engineer who just arrived on Earth for the first time." Sol communicates clearly and directly, with less filler—whether for executive summaries or long documents, it's more readable and easier to collaborate with.
Breaking Through Self-Imposed Limits to Deliver User Value
The author shares a real example: her prototyping tool had a tool-calling loop that Fable had "over-hardened," with Fable insisting "this is a problem with the model itself." After switching to Sol, she told it to "do it the way you think is right," and Sol broke out of the fixed mindset and fixed the problem in one shot.
Agentic AI and Tool-Calling Loops: "Agentic" AI refers to a mode of operation where a model doesn't just answer questions, but proactively plans steps, calls external tools (code executors, browsers, APIs), and iterates its actions based on results. Fable getting stuck in an "over-hardened tool-calling loop" means the model, constrained by overly strict rules or self-verification logic, repeatedly called tools without being able to advance the task, forming a dead loop. This is a known pain point of current frontier models: over-alignment or overly cautious safety mechanisms can sometimes prevent a model from being as flexible as a human collaborator in actual tasks. Sol's ability to "break out of the fixed mindset" and complete the task reflects its superior task completion rate in agentic scenarios—consistent with OpenAI's recent strategy of weighting "helpfulness" more heavily in model training.

Similarly, in an insight-ingestion product, Fable obsessed over making every result "reproducible, verifiable, and citable," deviating from the goal of building a good product. Sol, on the other hand, generated a genuinely useful wiki page.

The author comments: "I find it hard to work with colleagues who are theoretically smart but can't get anything done." Sol understands user goals, is willing to relax constraints appropriately, and gets things done—and that's exactly where its value lies.
Sol's Two Killer Use Cases
Video Editing
Drag a long video into Codex and say "cut this into 5 horizontal social short clips, with a faster, tighter pace," and Sol automatically picks the highlights to generate a hype video. Import it into CapCut, add music, and it's ready to publish—saving a huge amount of manual editing time.
Browser Use
The author calls Sol "a beast" at browser operations. By hooking Chrome into Codex, she had Sol handle around 500 LinkedIn messages, automatically filtering and replying based on the criterion of "only reply to executives at top-tier companies." She also used it to test web apps and fill out tedious forms.
Browser Use and Computer-Use Agents: Browser control (Browser Use) is a frontier direction in the expansion of AI Agent capabilities in recent years. It refers to a model directly operating a browser via programmatic interfaces (such as the Chrome DevTools Protocol, Playwright, or Puppeteer) to complete tasks like clicking, filling forms, scraping, and navigating—essentially giving the AI a pair of "digital hands." After Anthropic launched Computer Use and OpenAI launched Operator in 2024, this capability began moving from research to practical use. Unlike traditional RPA (robotic process automation) that relies on preset rules, LLM-based browser agents can understand natural language instructions and dynamically respond to page changes. This kind of scenario demands a high degree of "real-world awareness" from the model: it needs to understand the unstructured layout of real web pages, handle abnormal states, and maintain the task goal across multiple steps—making it one of the gold standards for testing a model's practicality. Sol's outstanding performance here confirms its overall advantage in agentic tasks.
Summary: A Division of Labor Where Each Model Has Its Strengths
After this hands-on test, the author offers pragmatic recommendations for dividing labor among the models:
- GPT-5.6 Sol: The daily workhorse, excelling at zero-to-one prototyping, clear writing, complex technical problems, video editing, and browser operations;
- Terra: Concise, direct business writing and PRDs;
- Fable: When conversation isn't needed, its web app code quality is still excellent;
- Sonnet 5: Human-like agentic voice and some code-debugging scenarios.
For product teams pursuing cost-effectiveness and real delivery efficiency, GPT-5.6 Sol offers capabilities closer to real-world workflows at a lower price, making it worth prioritizing. Sol's value proposition can be captured in one formula: lower token cost × higher task completion rate × stronger agentic adaptability = the optimal practical solution for product teams.
Key Takeaways
Related articles

Mecanum Wheel Motion Simulation Platform: A Detailed Guide to Low-Cost VR Haptic Solutions
A detailed look at a Mecanum wheel-based omnidirectional motion simulation platform using VR trackers for 3-DOF motion simulation and recentering correction — a viable low-cost VR immersion solution.

LangChain Managed DeepAgents: Hosted Agent Infrastructure So You Can Focus on Core Logic
LangChain launches Managed DeepAgents public beta, hosting evals, memory, OAuth, Slack integration, and sandbox infrastructure so developers can focus on Agent core logic.

Stripe's In-House AI Platform Architecture Explained: A Practical Guide to Enterprise AI Implementation
Deep dive into how Stripe built its internal AI platform, covering unified model access layers, RAG knowledge integration, security governance frameworks, and lessons for enterprise AI implementation.