待验证80% 置信事实精确时间
Multiple studies in 2023 found that abnormally high scores from certain models on specific benchmarks were highly correlated with data leakage
1
来源数
80%
置信度
长期有效
时效性
2026/8/3
首次发现
来源
涉及实体
相关事实
待验证基准测试面临饱和(Saturation)和数据污染(Data Contamination)两个系统性挑战,当多个主流模型得分超过95%时测试集失去区分能力67% 相似待验证2024年下半年以来,业界普遍观察到主流大模型在标准评测集上的得分提升幅度明显收窄,这一现象被称为'Scaling Law放缓'63% 相似待验证2023年多项独立研究表明部分顶级模型在标准数学推理测试中分数极高,但在等价的重新表述版本上成绩大幅下滑,表明模型学习的是测试模式而非真正的推理能力63% 相似待验证AI benchmark suites like MMLU, HumanEval, and MATH may produce inflated scores due to benchmark contamination, where models have indirectly encountered test data during training.62% 相似待验证在样本严重不均衡时,精度(Accuracy)会产生误导,模型把所有人都预测为健康精度照样能达到99%以上62% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/680095API
curl https://kongchang.com/api/v1/knowledge/claims/680095MCP
get_claim(id=680095)