Unverified70% confidenceFactTime unknown
2024 年初最先进的模型在 OS World Verified 测试上通常连 20% 都拿不到
1
Sources
70%
Confidence
Long-term
Relevance
6/1/2026
First Seen
Sources
Related Entities
Related Claims
Verified2024年初时,最强AI模型在SWE-Bench上的得分仅在10%-20%区间66% similarUnverified2024年多项研究发现,部分模型在去污染版本的评测集上得分下降幅度可达10-20个百分点65% similarUnverified在PlanBench的神秘方块世界测试中,2023年的模型几乎完全失败,但O1等推理模型的出现带来了显著改善65% similarUnverified斯坦福大学2023年的研究表明,主流AI检测工具误判率可高达20%以上65% similarUnverified在某些测试条件下,主流AI编程工具推荐不存在包名的概率可达5%~20%64% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/35322API
curl https://kongchang.com/api/v1/knowledge/claims/35322MCP
get_claim(id=35322)