Unverified50% confidenceFactExact time
AI 编程领域最受认可的评测基准包括 HumanEval(测试从函数文档字符串生成正确实现)和 SWE-bench(在真实 GitHub 仓库中定位并修复实际 issue)
1
Sources
50%
Confidence
Long-term
Relevance
7/9/2026
First Seen
Sources
GPT-5.6 Sol Ultra传闻解析:Codex编程能力将如何升级
hackernewshackernews7/6/2026
Related Claims
UnverifiedHumanEval、MBPP、SWE-bench是AI编程评测领域的行业标准基准测试78% similarUnverified评估智能体编程能力的主流基准包括SWE-bench、SWE-bench Verified和HumanEval77% similarUnverifiedHumanEval、SWE-bench等代码评测基准的排名是评估执行型Agent底层模型选型的重要参考维度74% similarVerifiedAI领域模型能力的可信评估依赖标准化基准测试,常见评测集包括MMLU、HumanEval、MATH、GPQA73% similarUnverifiedSWE-bench Verified由人工标注者筛除描述模糊或测试不稳定的样本,是业界权威的AI编程能力排行榜之一73% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/361068API
curl https://kongchang.com/api/v1/knowledge/claims/361068MCP
get_claim(id=361068)