Unverified50% confidenceFactExact time
评估智能体编程能力的主流基准包括SWE-bench、SWE-bench Verified和HumanEval
1
Sources
50%
Confidence
Long-term
Relevance
8/17/2026
First Seen
Sources
Related Claims
Unverified代码生成领域的模型评估如HumanEval、SWE-bench格外强调功能正确性而非流畅度79% similarUnverifiedAI 编程领域最受认可的评测基准包括 HumanEval(测试从函数文档字符串生成正确实现)和 SWE-bench(在真实 GitHub 仓库中定位并修复实际 issue)77% similarUnverified评估模型智能体能力的主流基准测试包括BFCL(Berkeley Function Calling Leaderboard)、AgentBench、SWE-Bench;编码能力评测集包括HumanEval、MBPP、LiveCodeBench76% similarUnverifiedSWE-bench 涉及跨文件依赖分析、版本兼容性判断、回归测试通过等高阶能力,难度比 HumanEval 等写函数类基准更贴近真实工程场景75% similarVerified评估编程大模型的主流基准包括HumanEval、MBPP、SWE-bench和LiveCodeBench74% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/763820API
curl https://kongchang.com/api/v1/knowledge/claims/763820MCP
get_claim(id=763820)