待验证50% 置信事实精确时间
OpenAI发布材料中未公布Sol在SWE-bench Pro的得分,该项测试目前由Fable 5领跑
1
来源数
50%
置信度
中期 (~90 天)
时效性
2026/7/12
首次发现
有效期至:2026/10/10
来源
GPT-5.6 Sol实测:多智能体并行如何重塑AI编程工作流
bilibili地层世界2026/7/11
相关事实
待验证OpenAI did not publish official scores for the standard SWBench for GPT-5.2 Codex, suggesting they may not have surpassed Opus 4.5's performance on that benchmark.73% 相似待验证Fable 5在SWE Bench Pro编程测试上得分80.3分,而OpenAI的GPT 5.5仅为58.6分69% 相似待验证基准测试结果显示,Sol和Sol Ultra在大多数领域优于竞争对手Misos 5,Tara优于Fable 567% 相似待验证根据Artificial Analysis Arena的排名数据,xAI的Grok 4.6在综合能力上被认为与Sol 5.6处于同一水平63% 相似待验证Grok 4.6在UI设计测试中表现平庸,逊于Fable和5.6 Sol63% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/492458API
curl https://kongchang.com/api/v1/knowledge/claims/492458MCP
get_claim(id=492458)