Verified80% confidenceFactExact time
SWE-bench 由普林斯顿大学研究团队创建,从 GitHub 真实开源项目中抽取 issue,要求模型生成能通过测试用例的补丁
6
Sources
80%
Confidence
Long-term
Relevance
7/11/2026
First Seen
Sources
Grok 4.5 实测:更快更便宜的编码黑马,能否取代 Opus 与 GPT?
youtubeForrestKnight7/10/2026
Related Claims
VerifiedSWE-bench由普林斯顿大学团队开发,Pro版本要求模型在真实GitHub代码仓库中定位Bug并生成修复补丁84% similarPartially VerifiedSWE-Bench was developed by a Princeton University team and extracts tasks from real GitHub issues, requiring models to locate problems within complete codebases and generate fix patches83% similarUnverifiedSWE-bench从主流Python开源项目提取真实GitHub Issue,要求模型生成能通过官方测试套件的代码补丁82% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/487943API
curl https://kongchang.com/api/v1/knowledge/claims/487943MCP
get_claim(id=487943)