Unverified50% confidenceFactExact time
OpenAI发布材料中未公布Sol在SWE-bench Pro的得分,该项测试目前由Fable 5领跑
1
Sources
50%
Confidence
Medium-term (~90 days)
Relevance
7/12/2026
First Seen
Valid until: 10/10/2026
Sources
GPT-5.6 Sol实测:多智能体并行如何重塑AI编程工作流
bilibili地层世界7/11/2026
Related Claims
UnverifiedOpenAI did not publish official scores for the standard SWBench for GPT-5.2 Codex, suggesting they may not have surpassed Opus 4.5's performance on that benchmark.73% similarUnverifiedFable 5在SWE Bench Pro编程测试上得分80.3分,而OpenAI的GPT 5.5仅为58.6分69% similarUnverified基准测试结果显示,Sol和Sol Ultra在大多数领域优于竞争对手Misos 5,Tara优于Fable 567% similarUnverified根据Artificial Analysis Arena的排名数据,xAI的Grok 4.6在综合能力上被认为与Sol 5.6处于同一水平63% similarUnverifiedGrok 4.6在UI设计测试中表现平庸,逊于Fable和5.6 Sol63% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/492458API
curl https://kongchang.com/api/v1/knowledge/claims/492458MCP
get_claim(id=492458)