[KongchangAI]
Verified65% confidenceBenchmarkExact time

GPT-4、Claude 3.5 Sonnet、Gemini 1.5 Pro等新一代大语言模型在HumanEval、SWE-bench等代码能力基准测试上的得分已大幅超越此前版本

3
Sources
65%
Confidence
Medium-term (~90 days)
Relevance
7/16/2026
First Seen
Valid until: 10/14/2026

Sources

Related Claims

Cite This Claim

Stable URI
https://kongchang.com/claim/533888
API
curl https://kongchang.com/api/v1/knowledge/claims/533888
MCP
get_claim(id=533888)