Cognition Launches SWE-2 Coding Model: Near Top-Tier Performance at 64% Lower Cost

Cognition's SWE-2 delivers near top-tier coding performance at one-quarter the cost of leading rivals.
Cognition has launched SWE-2, a next-generation coding model built on Kimi K3 post-training with reinforcement learning optimizing both capability and cost. On the FrontierCode 1.1 benchmark, it scores 50.0% — within one point of Fable 5.1 but at 64% lower cost — and costs just one-quarter of GPT-6 Astra. Compared to its predecessor SWE-1.7, it cuts interaction turns by 58% and cost by 81% while scoring higher. Now integrated into Devin Desktop and CLI, SWE-2 marks a shift in AI coding from benchmark racing toward efficiency and value.
The competition among AI coding models has shifted from a pure performance race to a dual battle of capability and cost. Cognition's newly released SWE-2 coding model is a prime example of this trend — rather than chasing benchmark leaderboard supremacy, it maintains near top-tier performance while cutting costs to a fraction of its competitors.

What Is SWE-2
SWE-2 is Cognition's next-generation coding model, and its technical approach is worth noting: it is post-trained on Kimi K3 and uses reinforcement learning (RL) to optimize across two dimensions simultaneously — cost and capability.
This "dual-objective optimization" philosophy sets it apart from models that solely chase benchmark scores. In real-world software engineering scenarios, developers care not just about whether a model can solve a problem, but also how many tokens each call consumes and how many interaction rounds it takes to complete a task. SWE-2 incorporates these economic metrics into its training objectives from the ground up — that's the core design principle that distinguishes it from pure "performance monsters."
SWE-2 is already integrated into Devin Desktop and CLI, available for users right now. As the company behind the AI software engineer Devin, Cognition has tightly coupled its in-house model with its own product, forming a complete loop from the underlying model to the end-user application.
Performance vs. Cost: Breaking Down the Key Numbers
The official data paints a clear picture of SWE-2's positioning — trading a modest performance dip for a significant cost advantage.
On the FrontierCode 1.1 Main benchmark, SWE-2 scores 50.0%. This is less than one percentage point behind competitor Fable 5.1, but at 64% lower cost. In other words, you get nearly equivalent coding capability for just over a third of the price.
Compared to the stronger GPT-6 Astra, SWE-2 trails by a few percentage points, but costs only one-quarter as much. For most everyday coding tasks, that small performance gap rarely justifies a fourfold cost difference — especially in engineering workflows that require large-scale, high-frequency model calls.
Improvements Over the Previous Generation
The most telling comparison for SWE-2 is actually against its predecessor, SWE-1.7.
According to official data, SWE-2 requires 58% fewer interaction turns to complete the same tasks, reduces cost by 81%, and actually scores higher. These numbers reveal a key insight: model progress is no longer just about raising the capability ceiling — it's increasingly about improving efficiency.
Fewer interaction rounds mean the model understands requirements faster and delivers usable results in fewer attempts, cutting down on repeated clarification and trial-and-error. In agent-based coding workflows, reduced turn counts translate directly into faster response times and lower costs. This explains why Cognition optimized for cost during training — in the agentic era, a single task often involves multiple rounds of model calls, so any per-turn efficiency gain gets amplified significantly.
Why the "Value-for-Money" Approach Deserves Attention
SWE-2 earned 119 upvotes on Product Hunt, ranking third for the day in the Artificial Intelligence category. That level of attention reflects genuine market demand for high-value coding models.
For a while, competition in AI coding tools devolved into a "benchmark arms race," with everyone fighting for the top spot. But for real commercial deployment, being a few points ahead on benchmarks matters less than whether costs can be kept at a sustainable level — and that's what ultimately determines whether these tools get adopted at scale.
SWE-2 takes a more pragmatic path: rather than chasing absolute performance supremacy, it maximizes cost advantages within a "good enough" performance range. When a model can deliver near top-tier results at one-quarter the cost, its appeal to businesses and individual developers may far exceed that of more expensive top-ranked alternatives.
Closing Thoughts
The launch of SWE-2 signals a shift in the AI coding space from "competing on capability" to "competing on efficiency." Its technical approach — post-training on Kimi K3 with reinforcement learning to jointly optimize cost and capability — combined with deep integration into the Devin product ecosystem, gives it a differentiated position in a crowded market. For developers and teams focused on return on investment, high-value models like this may offer more practical value than chasing leaderboard leaders. That said, the official benchmark numbers still need to be validated in real development environments, and the model's actual performance remains to be confirmed through broader user feedback.
Related articles

Perplexity Hybrid Compute: Cloud Research, Local Mac Privacy Protection
Perplexity's Hybrid Compute splits AI tasks by data sensitivity: cloud handles research, local Mac handles private files. Requires Apple Silicon, 24GB RAM, Pro/Max/Enterprise plan.

Resurf: The Personal Context Manager That Feeds Your Data to AI
Resurf is a personal context app for Apple devices that saves notes, links, images, and PDFs, then feeds them to AI via MCP and CLI. Local storage with iCloud sync keeps your data private.

ScreenCursor: A Screen Recorder That Automatically Generates Zoom Effects
ScreenCursor is a Chrome extension that auto-generates zoom effects during screen recording, turning every click and drag into camera motion — no editing, fully local.