待验证50% 置信事实时间未知
Codex One在SWE-bench上的表现略优于O3 High
1
来源数
50%
置信度
中期 (~90 天)
时效性
2026/7/2
首次发现
有效期至:2026/9/30
来源
OpenAI Codex Deep Dive: How Does the Async AI Coding Agent Actually Perform?
bilibiliKrillinAI小林
相关事实
待验证Codex One在SWE-bench上的表现略优于O3 High,在OpenAI内部软件工程任务上准确率达到75%,而O3 High为70%78% 相似待验证在编程基准测试中,Codex-1的表现超越了Claude 3.7和O3 High76% 相似待验证在SWE-Bench评分中,o1的得分远高于o3-mini medium,但略低于o3-mini-high64% 相似待验证SW1E1的能力接近Claude 3.7和Claude 3.5的水平,明显优于DeepSeek V361% 相似待验证Codex Max在SWE-Bench验证集上表现优于原始GPT 5.1 Codex,同时减少了30%的思考token60% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/54738API
curl https://kongchang.com/api/v1/knowledge/claims/54738MCP
get_claim(id=54738)