Unverified50% confidenceBenchmarkExact time
研究表明GPT-4作为评估者与人类专家评分的一致性超过80%,显著高于BLEU、ROUGE等基于词汇匹配的自动化指标
1
Sources
50%
Confidence
Long-term
Relevance
7/25/2026
First Seen
Sources
Related Claims
UnverifiedGPT-4级别的模型作为Judge时,其评判结果与人类专家的一致性超过80%84% similarUnverified在标准化测试中,GPT-4的表现已超过90%的人类考生78% similarUnverified经过优化的提示词可以将GPT-4在特定任务上的准确率从60%提升至95%以上76% similarVerifiedGPT-4、Claude 3.5、DeepSeek等模型在HumanEval编程基准测试上的通过率已超过80%74% similarUnverifiedGPT-4、Claude等大模型在Spider等基准测试上的Text-to-SQL准确率已突破85%72% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/613444API
curl https://kongchang.com/api/v1/knowledge/claims/613444MCP
get_claim(id=613444)