Unverified50% confidenceBenchmarkExact time
o1在国际数学奥林匹克级别问题上正确率超过80%,远高于GPT-4的约13%
1
Sources
50%
Confidence
Long-term
Relevance
8/8/2026
First Seen
Sources
Related Claims
UnverifiedO4-mini在GPQA博士级问答评测中准确率超过80%72% similarUnverifiedo1模型在国际数学奥林匹克预选题上的表现接近金牌水平,在Codeforces编程竞赛中达到了89百分位70% similarUnverifiedARC-AGI-2由AI安全研究员François Chollet设计,早期GPT-4在该测试上得分不足10%,人类平均得分约85%65% similarUnverifiedGPT-4等大模型在Spider等基准测试上的零样本或少样本SQL生成准确率已超过80%64% similarUnverified一个校准良好的模型,其输出的0.8分应当在统计意义上真的对应80%的正确率63% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/705748API
curl https://kongchang.com/api/v1/knowledge/claims/705748MCP
get_claim(id=705748)