待验证50% 置信事实精确时间
SWE-Bench Pro has a false positive rate of 8.5%, meaning the verifier accepts incorrect implementations at that rate.
1
来源数
50%
置信度
中期 (~90 天)
时效性
2026/7/2
首次发现
有效期至:2026/9/30
来源
DeepSWE Benchmark Reveals the Truth: GPT 5.5 Leads Opus 4.7 by a Wide Margin
bilibili7k的每日搬运2026/5/29
相关事实
待验证SWE-Bench Pro的假阳性率为8.5%,假阴性率高达24%;而DeepSWE的假阳性率仅为0.3%,假阴性率仅为1.1%76% 相似待验证SWE-Bench Pro has a false negative rate of 24%, meaning 1 out of every 4 correct answers is incorrectly marked as a failure.75% 相似待验证Frontier Code的误报率为6.9%,相比SWE-Bench Pro的36.0%降低了约81%64% 相似待验证接受率随位置逐级衰减:GSM8K从位置0的0.92下滑至约0.53,MTBench从0.78急跌至0.13,因误差在连续猜测中累积62% 相似待验证GPT-5.4相对于GPT-5.2,单独声明出错的概率降低了33%,整个回复包含任何错误的概率降低了18%58% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/46947API
curl https://kongchang.com/api/v1/knowledge/claims/46947MCP
get_claim(id=46947)