Unverified50% confidenceFactExact time
SWE-Bench Pro has a false negative rate of 24%, meaning 1 out of every 4 correct answers is incorrectly marked as a failure.
1
Sources
50%
Confidence
Medium-term (~90 days)
Relevance
7/2/2026
First Seen
Valid until: 9/30/2026
Sources
DeepSWE Benchmark Reveals the Truth: GPT 5.5 Leads Opus 4.7 by a Wide Margin
bilibili7k的每日搬运5/29/2026
Related Claims
UnverifiedSWE-Bench Pro的假阳性率为8.5%,假阴性率高达24%;而DeepSWE的假阳性率仅为0.3%,假阴性率仅为1.1%78% similarUnverifiedSWE-Bench Pro has a false positive rate of 8.5%, meaning the verifier accepts incorrect implementations at that rate.75% similarUnverifiedFrontier Code的误报率为6.9%,相比SWE-Bench Pro的36.0%降低了约81%66% similarUnverifiedGPT-5.4相对于GPT-5.2,单独声明出错的概率降低了33%,整个回复包含任何错误的概率降低了18%66% similarVerified在长链路任务中,若每一步约5%的错误率,经过20步后整体成功率可能跌至不足36%63% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/46949API
curl https://kongchang.com/api/v1/knowledge/claims/46949MCP
get_claim(id=46949)