GPT-6 Sol and Luna Cut Prices in Half — So Why Didn't the Leaderboard Move?

GPT-6's Sol and Luna are half the price, but leaderboard scores held flat — value improved more than raw capability.
OpenAI slashed prices for Sol and Luna in the GPT-6 family — Sol at $2/$10 and Luna at $0.1/$0.5 per million input/output tokens — but their Artificial Analysis composite scores of 48 and 37 trail far behind Opus 5.5's 58. This "price halved, scores flat" contrast doesn't mean the models stagnated; the composite index's averaging effect masks task-level gains such as Sol's improvement on coding agent tasks and Luna's progress on automation workflows. The three models serve distinct roles: Astra for peak capability, Sol for complex reasoning and agents, Luna for cheap high-volume tasks. Developers are best served by testing real workloads — tracking success rate, revision rounds, and total cost — rather than relying on a single leaderboard number to make model selection decisions.
Half the Price, Almost No Score Bump
OpenAI has significantly cut prices for Sol and Luna, two models in the GPT-6 family, slashing them to roughly half the discounted rate of the previous generation. At standard API pricing for short contexts, Sol comes in at around $2 per million input tokens and $10 per million output tokens, while Luna goes as low as $0.1 input and $0.5 output. The price drop is dramatic — but the comprehensive leaderboard scores? Barely moved.
According to analysis from the Bilibili podcast "Cahua's Bedtime AI," the most notable change this time is simply that these models got cheaper. Capability improvements need to be evaluated task by task. In other words, the real headline about Sol and Luna isn't that they topped the overall leaderboard — it's that the cost of deploying them at scale has come down substantially.

From a product positioning standpoint, the three models each serve distinct roles: Astra remains OpenAI's top-tier offering; Sol targets more complex coding and agentic tasks; Luna is aimed at cheap, fast, high-volume repetitive workloads. This tiered structure means that measuring all models with a single "composite index" ruler is inherently prone to misleading conclusions.
Behind the Leaderboard Numbers: Progress Hides in Task-Level Details
In Artificial Analysis's composite index, the contrast from models released the same day is striking: the newly released Opus 5.5 scored 58, Sol scored 48, and Luna scored 37. Opus's high score stands out — next to it, Sol and Luna look like they've barely progressed.
But interpreting 48 and 37 as "basically no improvement" isn't accurate. The composite index blends ten different tests together — useful for a bird's-eye view, but not for deciding which model handles your specific workload best. More granular observations from evaluators reveal the nuance: Sol gained two points on the coding agent index while Luna actually dropped two; both showed improvement on certain automation and terminal tasks, but both are more prone to slipping up on details in knowledge work deliverables.

This isn't a simple "upgrade" or "downgrade" — different tasks yield different answers. Some Zhihu users have noted that OpenAI and Anthropic seem to be diverging onto two separate paths: one keeps pushing the ceiling of raw capability, while the other focuses on making AI accessible to more users. The contrast in this release does lend some credibility to that observation.
Artificial Analysis's composite index is a widely used tool for cross-model comparison in the industry. It combines weighted scores from roughly ten benchmarks — including coding, mathematical reasoning, knowledge QA, and instruction following — into a single number, enabling quick comparisons across models and vendors. The advantage is intuitive readability; the drawback is just as obvious: if a model improves dramatically on a few tasks while staying flat on others, those gains get "averaged out," masking real task-level capability differences. On top of that, the weighting scheme itself affects rankings — a system that heavily weights math reasoning may make a model that excels at code generation look unremarkable. This is exactly why the same set of benchmark data can lead a casual writing user and a full-time developer to completely different model selection conclusions.
Don't Just Look at One Overall Score
One point is worth emphasizing repeatedly: don't judge cutting-edge models solely by a single composite ranking. If you're using a model every day to refactor code, what actually matters is whether the changes it proposes can be merged directly. If you're using it for research, what matters is whether the report misses any key requirements.
As for speculation that OpenAI is "intentionally ignoring certain leaderboards" — there's no evidence to draw that conclusion on the company's behalf. The one clear fact is this: Astra is positioned at the highest capability tier, while Sol and Luna are explicitly aimed at everyday tasks that more people can afford. Cutting prices in half is fundamentally about selling cost-efficiency at scale — not about putting these two models in contention for the top overall rank.
What Does Half-Price Mean for Developers?
For anyone calling APIs regularly, the most immediate benefit of a price cut is simple: the same budget gets you more runs. A practical approach to model routing is to hand off simple classification and data organization tasks to the cheaper Luna, then pass anything requiring sustained reasoning to Sol. The combined cost could be significantly lower than running everything through a large model.

That said, this is just a directional guide — not a guaranteed money-saver. A full accounting of any task needs to include failed retries, human review time, and latency. You can't just look at the listed price per million tokens. A cheaper model that requires constant rework may not be the better deal when you tally everything up.
User feedback on Zhihu reflects this task-dependence: some praise Sol for being direct and fast; others report that it tends to "drift" on complex projects, losing track of the original workflow in long tasks and requiring human nudges to stay on course. These seemingly contradictory reviews are actually consistent — they're describing different outcomes for different tasks. Early anecdotes are worth hearing, but no single success or failure should be treated as a verdict on the model.
"Per million tokens" is the standard billing unit for large language models today. Input and output tokens are priced separately, because generating text is computationally more expensive than processing input — which is why output prices are consistently higher. With Luna as an example, at $0.1 per million input tokens and $0.5 per million output tokens, a single conversation with roughly 750 Chinese characters of input and 3,750 characters of output costs less than $0.001. That's negligible for individual calls — but when an automated pipeline runs millions of tasks per day, the input-to-output ratio, retry frequency, and prompt length can all shift the monthly bill by an order of magnitude. So when estimating costs, beyond comparing list prices, you also need to benchmark your own typical input/output ratio to arrive at a real per-task cost.
How to Choose Without Being Led by the Leaderboard
With Sol, Luna, and the newly released Opus 5.5 all on the table, how do you actually choose? Here's a practical approach that keeps you from being driven by rankings alone: pick three tasks you actually do — one simple and repetitive, one moderately complex coding task, one piece of writing that requires a complete deliverable — then test all models against the same requirements, tracking whether they succeed, how many revision rounds it takes, how long it takes, and what it costs in total.

Luna, Sol, Opus, and even Astra are all competing in that kind of real-world ledger. The highest score and the best fit for day-to-day use are often not the same answer.
So the real shift here isn't just another model name change. The competition is expanding — from "who can push the leaderboard the highest" to "who can get more work done at lower cost." Sol and Luna cutting prices in half is genuinely impactful, but the flat leaderboard scores are a reminder: cheaper doesn't mean every capability went up. Don't rush to declare a winner — run your own tasks first, then decide whether it's worth making the switch.
Related articles

Receipt Forgery Detection Near Random? Real-World Struggles and Solutions in Document Image Forensics
A receipt forgery detection project with ROC-AUC near random reveals the pitfalls of small-sample document forensics. Explores anomaly detection, self-supervised pre-training, and numerical consistency as viable alternatives.

From Workflows to Eval-Driven Development: A Paradigm Shift in How We Solve Problems with AI
AI problem-solving is shifting from deterministic workflows to "define evals + hillclimb." This piece explores how eval-driven development reshapes tasks, data vendors, human roles, and Agent UX.

Tesla Powerwall + Electric Vehicle: A Dual Backup Power Solution for Outages
Tesla Powerwall combined with EV bidirectional charging can provide multi-layer home backup power during outages. We break down runtime, V2H realities, and Supercharger loop feasibility.