部分验证65% 置信事实精确时间
SweetBench Pro (SWE-bench) is a coding benchmark developed by a Princeton University research team that uses real GitHub Issues and Pull Requests
3
来源数
65%
置信度
中期 (~90 天)
时效性
2026/7/2
首次发现
有效期至:2026/9/30
来源
DeepSeek V4 Pro In-Depth Review: Performance Rivaling GPT-5.5 at 1/12 the Cost
bilibili63号炼金工坊2026/6/18
相关事实
待验证SWBench is an AI coding evaluation benchmark developed by a Princeton University team that extracts issues and pull requests from real GitHub open-source projects.75% 相似部分验证SWE-Bench was developed by a Princeton University team and extracts tasks from real GitHub issues, requiring models to locate problems within complete codebases and generate fix patches69% 相似待验证SWE-bench(Software Engineering Benchmark)是目前衡量智能体编码能力的核心基准,从真实的GitHub Issue出发要求模型自主定位代码库中的问题并提交修复补丁68% 相似已验证SWE-bench由普林斯顿大学团队开发,Pro版本要求模型在真实GitHub代码仓库中定位Bug并生成修复补丁66% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/44337API
curl https://kongchang.com/api/v1/knowledge/claims/44337MCP
get_claim(id=44337)