Unverified50% confidenceBenchmarkExact time
早期Codex在HumanEval基准测试中以37%的准确率独立解决编程题,达到当时最高水平
1
Sources
50%
Confidence
Long-term
Relevance
7/11/2026
First Seen
Sources
GPT-5.6登陆Codex、语音模型迭代,AI产品密集更新
bilibiliinfinite灵感港7/7/2026
Related Claims
VerifiedCodex在HumanEval基准测试上达到了28.8%的pass@1准确率81% similarUnverified测试者评估Codex大概能完成看板上30%的任务73% similarUnverified初版 Codex 采用约120亿参数规模,在 HumanEval 基准测试中达到 28.8% 的 pass@1 通过率72% similarUnverified向 Codex 发出'无论消耗多少额度,修复完要做真实测试,真实发请求验证API能否返回结果,反复测试直到成功'的指令后,Codex 运行了36分钟,最终自主完成了调试70% similarUnverifiedIn recent hands-on tests, Codex has surpassed Claude Code in multiple evaluation dimensions.67% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/488600API
curl https://kongchang.com/api/v1/knowledge/claims/488600MCP
get_claim(id=488600)