Unverified80% confidenceFactExact time
Multiple studies in 2023 found that abnormally high scores from certain models on specific benchmarks were highly correlated with data leakage
1
Sources
80%
Confidence
Long-term
Relevance
8/3/2026
First Seen
Sources
Related Entities
Related Claims
Unverified基准测试面临饱和(Saturation)和数据污染(Data Contamination)两个系统性挑战,当多个主流模型得分超过95%时测试集失去区分能力67% similarUnverified2024年下半年以来,业界普遍观察到主流大模型在标准评测集上的得分提升幅度明显收窄,这一现象被称为'Scaling Law放缓'63% similarUnverified2023年多项独立研究表明部分顶级模型在标准数学推理测试中分数极高,但在等价的重新表述版本上成绩大幅下滑,表明模型学习的是测试模式而非真正的推理能力63% similarUnverifiedAI benchmark suites like MMLU, HumanEval, and MATH may produce inflated scores due to benchmark contamination, where models have indirectly encountered test data during training.62% similarUnverified在样本严重不均衡时,精度(Accuracy)会产生误导,模型把所有人都预测为健康精度照样能达到99%以上62% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/680095API
curl https://kongchang.com/api/v1/knowledge/claims/680095MCP
get_claim(id=680095)