待验证60% 置信事实时间未知
在编程基准测试中,Codex-1的表现超越了Claude 3.7和O3 High
2
来源数
60%
置信度
中期 (~90 天)
时效性
2026/6/2
首次发现
有效期至:2026/8/31
来源
OpenAI Codex深度解析:AI编程从代码补全迈向全仓库自主开发
bilibiliAI领航员-园园
涉及实体
相关事实
待验证Codex One在SWE-bench上的表现略优于O3 High76% 相似待验证多个独立基准测试一致显示Claude Code与Codex的token消耗比在3.2倍到4.2倍之间73% 相似待验证Codex One在SWE-bench上的表现略优于O3 High,在OpenAI内部软件工程任务上准确率达到75%,而O3 High为70%72% 相似待验证In recent hands-on tests, Codex has surpassed Claude Code in multiple evaluation dimensions.71% 相似待验证SW1E1的能力接近Claude 3.7和Claude 3.5的水平,明显优于DeepSeek V368% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/36275API
curl https://kongchang.com/api/v1/knowledge/claims/36275MCP
get_claim(id=36275)