Fable 5.1 and Mythos 5.1 Launch Together: Topping Eight Benchmarks with 45% Cost Reduction

Fable 5.1 and Mythos 5.1 top eight benchmarks while cutting Agent task costs by up to 45%.
Fable 5.1 and Mythos 5.1 launched simultaneously, with Fable 5.1 ranking first across all eight public benchmarks and doubling its predecessor's score on the Terminal Bench Science 0.1 Agent benchmark. Mythos 5.1 achieved protein design success rates 3–5x the industry average. Cached read prices were slashed by 75%, reducing overall Agent task costs by up to 45%, signaling a strong push to accelerate complex Agent application development.
Dual Model Launch: Differentiated Positioning for Different Scenarios
The new flagship models Fable 5.1 and Mythos 5.1 have arrived simultaneously, continuing the strong performance of their predecessors while making clear adjustments in both capability boundaries and pricing strategy. Though the two models share common roots, they are specifically tailored for entirely different use cases.
Fable 5.1 is positioned as a general-purpose model, fully available to everyday users, developers, and enterprise customers, covering a broad range of needs from daily Q&A to complex Agent tasks. Mythos 5.1, on the other hand, is placed within a more restrictive security framework, primarily targeting vetted partners in high-risk domains. This tiered access approach is fundamentally about finding the balance between capability release and safety control — the more powerful the model, the stricter the access requirements need to be.
This dual-track product logic reflects how frontier model providers are shifting from a one-size-fits-all release model to a new phase of fine-grained management based on risk levels and user qualifications.
Performance Breakthrough: Topping All Eight Benchmarks
Fable 5.1 has achieved a significant leap in performance over its predecessor, Fable 5. According to data released by the developer, the model secured first place in all 8 public benchmark tests, demonstrating comprehensive capability leadership.

The most noteworthy result is the score on Terminal Bench Science 0.1, a scientific research Agent benchmark. Fable 5.1 scored 52.6% on this test — more than double the previous generation's score. This number matters because scientific research tasks typically require the model to combine long-chain reasoning, tool use, and result verification. A failure in any single step can cause the entire task to fail. Doubling the score signifies a qualitative improvement in the model's stability and accuracy on complex tasks.
Protein Design: Success Rate 3–5x the Industry Average
In hands-on testing, Mythos 5.1's performance was particularly impressive. When using tools for protein design, roughly half of its generated designs were actually viable.

A "half viable" ratio might not sound high, but in the field of protein design, this success rate is already 3 to 5 times the industry norm. Computational design in life sciences has long been a recognized challenge, and multiplying the effective output rate several times over holds considerable value for real-world applications like drug development and enzyme engineering.
Additionally, Fable 5.1 demonstrated cross-domain processing capabilities — it processed radar images of Venus captured by NASA over thirty years ago into a much clearer topographic map. This kind of image enhancement and scientific data reconstruction further validates the model's practical utility in specialized scenarios.
Creative Generation: Game Scenes That Look Like the Real Thing
After the model launch, users quickly jumped in with all kinds of creative tests, and the results were equally impressive.

One user generated Minecraft-style scenes in a single pass — the mountains, clouds, and iconic pigs were virtually indistinguishable from the original game. Others used Blender MCP to generate racing games — with the same prompts, Fable 5.1's output was noticeably superior to the previous Fable 5.
These real-world test cases demonstrate that the model's improvements go beyond benchmark scores. In actual creative generation and tool collaboration scenarios, users can intuitively feel the quality improvement. The integration with tool chains like Blender MCP, in particular, showcases the model's maturity in agentic workflows.
Key Highlight: More Powerful Yet Cheaper
The most surprising aspect of this release is that prices actually dropped while performance improved.

Specifically, Fable 5.1 maintains the same base input/output pricing as Fable 5 — no price increase. However, cached read prices have been slashed by 75%. For Agent tasks involving long contexts and frequent tool calls, the overall cost can be reduced by up to 45%.
The intent behind this pricing strategy is crystal clear: encourage developers to build more complex Agent applications. Long contexts and high-frequency tool calls are hallmarks of Agent tasks, and these workloads have historically been difficult to scale due to high costs. The dramatic reduction in cached read pricing directly lowers the economic barrier for long-running, repeated model invocations.
The Industry Signal Behind the Pricing
"Stronger models at lower prices" has become the dominant theme in today's LLM competition. This is driven both by cost reductions from inference efficiency optimizations and by strategic considerations as providers compete for developer ecosystems. As the cost of calling high-performance models continues to drop, the real beneficiaries are the developers and enterprises looking to deeply integrate AI into their products.
It's foreseeable that as cost barriers for Agent applications are progressively dismantled, more real-world products built on long contexts and multi-tool collaboration will emerge at an accelerating pace. For developers on the fence, this price cut may be the perfect moment to jump in.
Summary: A Dual Leap in Capability and Cost-Effectiveness
The simultaneous launch of Fable 5.1 and Mythos 5.1 is more than a routine version update. From topping all eight benchmarks to achieving protein design success rates several times the industry average, to a pricing strategy that's up to 45% cheaper — this update sends positive signals on both the capability and commercialization fronts. For developers and enterprises tracking cutting-edge AI applications, it's well worth trying out firsthand.
Related articles

curl Project Exposes AI Code Auditing Shortcomings: 6 CVEs Found by Humans After AI Detected Zero
After OpenAI and Anthropic AI audits found zero issues in curl, human reviewers uncovered 6 CVEs. Explore AI code auditing limitations and human-AI collaboration best practices.

Perplexity Pro Service Downgrade? Long-Time Users Complain About Model Downgrades and Tighter Censorship
A Perplexity Pro long-time user complains on Reddit about weakened Deep Research, ignored system prompts, and silent model downgrades. We analyze the structural causes behind declining user experience.

Perplexity Pro Renewal Date Suddenly Changed: Shrinking Free Perks Spark Contract Dispute
A Perplexity Pro user's free subscription was cut short with a surprise $236 charge. The incident raises questions about AI subscription contracts and billing transparency.