[KongchangAI]
Verified65% confidenceBenchmarkExact time

AI领域模型能力的可信评估依赖标准化基准测试,常见评测集包括MMLU、HumanEval、MATH、GPQA

3
Sources
65%
Confidence
Long-term
Relevance
7/10/2026
First Seen

Sources

Related Claims

Cite This Claim

Stable URI
https://kongchang.com/claim/457627
API
curl https://kongchang.com/api/v1/knowledge/claims/457627
MCP
get_claim(id=457627)