GPT-6 Astra Lands on Devin: Near Top-Tier Performance at 64% Lower Cost

GPT-6 Astra joins Devin platform, ranking 2nd on FrontierCode 1.1 at 64% lower cost.
OpenAI's GPT-6 Astra is now live across all three Devin product lines — desktop, CLI, and cloud. The model ranks second on the FrontierCode 1.1 coding benchmark, surpassing Fable 5.1 while costing 64% less than comparable models. As cost efficiency becomes a key competitive dimension in AI coding tools, Devin Cloud's model mixture mechanism lets GPT-6 Astra quietly absorb a large share of routine coding workloads. Real-world stability and compatibility remain to be validated through broader use.
GPT-6 Astra Now Available Across the Entire Devin Platform
OpenAI's GPT-6 Astra has gone live across the Devin ecosystem. According to official announcements, the model is now directly available in Devin Desktop and Devin CLI, and has also been added to Devin Cloud's model mixture. This means developers can access the new model immediately, regardless of whether they're using the desktop client, command-line tool, or cloud service.

For teams that rely on AI coding assistants, model availability and integration method often determine real-world productivity. Devin's decision to deploy GPT-6 Astra simultaneously across all three product lines signals strong confidence in the model's capabilities, and lets users experience the latest code generation features without any additional configuration.
Benchmark Performance: Near the Top, at a Fraction of the Cost
What's truly noteworthy is GPT-6 Astra's performance on the FrontierCode 1.1 benchmark. According to disclosed data, the model outperforms Fable 5.1 on this coding evaluation, placing second only to Fable 5.
The ranking alone may not sound spectacular, but when cost is factored in, the value proposition becomes clear — GPT-6 Astra achieves this performance level at a cost 64% lower than comparable models. In other words, it delivers near-top-tier coding capability at a dramatically lower price than competing options.
Why Cost Efficiency Matters
In today's AI coding tool landscape, raw capability comparisons have become intensely competitive, and cost efficiency is emerging as the new battleground. For engineering teams that need to call models at scale, a 64% cost reduction can directly reshape a project's economics. A model that performs close to the best but costs significantly less is especially attractive for high-frequency, high-volume code generation workloads.
FrontierCode 1.1 is a benchmark specifically designed to evaluate AI coding capabilities, built to simulate real-world software engineering tasks — including code completion, bug fixing, multi-file refactoring, and unit test generation. Compared to earlier benchmarks like HumanEval, which focused primarily on code fill-in-the-blank tasks, FrontierCode 1.1 is substantially harder and broader in scope. Its test cases are typically drawn from real open-source projects, requiring models to demonstrate cross-file context understanding and long-sequence reasoning — which is why the industry treats it as a reliable indicator of an AI coding assistant's practical engineering ability. Fable 5 and Fable 5.1, mentioned above, currently sit at the top of the leaderboard and represent the approximate upper bound of performance on this benchmark. GPT-6 Astra placing second signals that it has reached a top-tier level for complex engineering coding tasks.
What This Means for Developers
Adding GPT-6 Astra to Devin Cloud's model mixture means Devin can dynamically select the most appropriate model for different tasks. Given its cost advantage, Astra is likely to handle a significant portion of everyday coding workloads. Direct support in the desktop client and CLI gives developers who care about balancing performance and cost a more flexible set of options.
It's worth noting that the information currently available is largely limited to availability and benchmark data. GPT-6 Astra's stability, context-handling capability, and compatibility with other toolchains in real-world projects still await validation through broader usage.
The "model mixture" mechanism used by Devin Cloud refers to the platform automatically routing and distributing tasks across multiple models in the background, based on task type, complexity, and cost constraints — rather than using a single fixed model for all requests. The core advantage of this approach is that it can reserve high-cost flagship models for tasks that genuinely require deep reasoning, while assigning routine, repetitive code generation work to more cost-effective models. With its 64% cost advantage, GPT-6 Astra is a natural fit for the latter category, maintaining high output quality while keeping overall costs in check. For development teams, this means platform-level cost optimization is already quietly happening — without any need to actively adjust their workflows.
Takeaway
GPT-6 Astra's arrival on Devin comes down to one core proposition: exceptional value. It ranks second on FrontierCode 1.1 behind only Fable 5, while achieving that at 64% lower cost. For cost-conscious developers and teams who don't want to compromise on code quality, this combination is worth evaluating. As the model rolls out across the full Devin platform, its real-world performance will be put to the test in the weeks ahead.
Related articles

rag-eval: A Zero-Dependency, No-API-Key RAG Evaluation Tool
rag-eval is a zero-dependency, framework-agnostic open-source RAG pipeline evaluation tool. It supports free local lexical and retrieval metrics with no API keys required, and offers optional LLM Judge for semantic validation. Compatible with Haystack, LangChain, and LlamaIndex.

Vercel AI SDK Releases workflow-harness 1.0.115 Patch Update
Vercel AI SDK releases @ai-sdk/workflow-harness 1.0.115 patch update, syncing the @ai-sdk/harness dependency. Learn about the update, release mechanism, and what it means for developers.

GLM 5.3 Now Available on Serverless Training API — No Sales Process Required
GLM 5.3 is now available on Serverless Training API alongside Kimi K3 and Qwen 3.8 27b. No sales process needed — start fine-tuning directly via docs or pre-made recipes.