待验证60% 置信观点精确时间
OpenAI did not publish official scores for the standard SWBench for GPT-5.2 Codex, suggesting they may not have surpassed Opus 4.5's performance on that benchmark.
2
来源数
60%
置信度
中期 (~90 天)
时效性
2026/7/2
首次发现
有效期至:2026/9/30
来源
相关事实
待验证OpenAI发布材料中未公布Sol在SWE-bench Pro的得分,该项测试目前由Fable 5领跑73% 相似待验证Fable 5在SWE Bench Pro编程测试上得分80.3分,而OpenAI的GPT 5.5仅为58.6分64% 相似部分验证Claude Opus 4.8在SWE-bench基准测试中得分69.2%,GPT 5.5为58.6%,Google Gemini为54.2%63% 相似待验证Claude Opus 4.8在AI Coding基准测试中表现超越GPT-5.562% 相似待验证在Terminal Bench 2.0基准测试中,OpenAI 4.6得分为58%,GPT 5.4得分为75.1%62% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/51094API
curl https://kongchang.com/api/v1/knowledge/claims/51094MCP
get_claim(id=51094)