Verified65% confidenceBenchmarkExact time
GPT-4、Claude 3.5 Sonnet、Gemini 1.5 Pro等新一代大语言模型在HumanEval、SWE-bench等代码能力基准测试上的得分已大幅超越此前版本
3
Sources
65%
Confidence
Medium-term (~90 days)
Relevance
7/16/2026
First Seen
Valid until: 10/14/2026
Sources
Related Claims
Cite This Claim
Stable URI
https://kongchang.com/claim/533888API
curl https://kongchang.com/api/v1/knowledge/claims/533888MCP
get_claim(id=533888)