Partially Verified65% confidenceFactExact time
SweetBench Pro (SWE-bench) is a coding benchmark developed by a Princeton University research team that uses real GitHub Issues and Pull Requests
3
Sources
65%
Confidence
Medium-term (~90 days)
Relevance
7/2/2026
First Seen
Valid until: 9/30/2026
Sources
DeepSeek V4 Pro In-Depth Review: Performance Rivaling GPT-5.5 at 1/12 the Cost
bilibili63号炼金工坊6/18/2026
Related Claims
UnverifiedSWBench is an AI coding evaluation benchmark developed by a Princeton University team that extracts issues and pull requests from real GitHub open-source projects.75% similarPartially VerifiedSWE-Bench was developed by a Princeton University team and extracts tasks from real GitHub issues, requiring models to locate problems within complete codebases and generate fix patches69% similarUnverifiedSWE-bench(Software Engineering Benchmark)是目前衡量智能体编码能力的核心基准,从真实的GitHub Issue出发要求模型自主定位代码库中的问题并提交修复补丁68% similarVerifiedSWE-bench由普林斯顿大学团队开发,Pro版本要求模型在真实GitHub代码仓库中定位Bug并生成修复补丁66% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/44337API
curl https://kongchang.com/api/v1/knowledge/claims/44337MCP
get_claim(id=44337)