Unverified50% confidenceFactExact time
Scale AI在2024年的分析显示,部分模型在MMLU上的分数膨胀可能高达5-10个百分点,因测试集已泄露到训练数据中
1
Sources
50%
Confidence
Long-term
Relevance
8/5/2026
First Seen
Sources
Related Claims
Unverified2024年研究表明即使AI生成内容仅占训练数据的10-20%,在多代训练后也足以引发可测量的质量退化78% similarUnverified自2024年下半年以来,主流大模型在MMLU和HumanEval等标准基准测试上的分数提升明显收窄76% similarUnverified到2024年中,顶级模型在MMLU上的分数已普遍超过85-90%,该基准的区分度急剧下降73% similarUnverifiedAI benchmark suites like MMLU, HumanEval, and MATH may produce inflated scores due to benchmark contamination, where models have indirectly encountered test data during training.72% similarUnverified2022年Hoffman等人的Chinchilla论文重新校准了算力最优的参数-数据比例,指出此前许多大模型训练不充分70% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/688242API
curl https://kongchang.com/api/v1/knowledge/claims/688242MCP
get_claim(id=688242)