Verified75% confidenceFactTime unknown
基准污染(Benchmark Contamination)是当前大模型评估领域面临的重要挑战,模型训练数据中可能已包含公开的测试题目导致评估结果失真
3
Sources
75%
Confidence
Long-term
Relevance
6/1/2026
First Seen
Sources
Percy Liang确认出席CAIS 2026:AI安全与大模型评估的前沿对话
twitterJeffDean
Related Entities
Related Claims
Unverified数据泄漏分为目标泄漏和训练-测试污染两类,会导致模型性能被系统性高估74% similarUnverifiedBenchmark overfitting is a systemic limitation where models score high through targeted training on test sets without corresponding improvements in generalization.73% similarUnverified部分模型存在针对测试集过度优化的风险,导致榜单分数与真实场景表现出现偏差(Benchmark饱和)72% similarVerified若用含测试集的全量数据计算均值和标准差,会导致预处理泄漏,使评估指标虚高,无法真实反映模型泛化能力72% similarVerified基准数据污染是指测试集公开后其题目可能出现在后续模型的预训练语料中,导致得分反映记忆而非真实推理能力71% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/22964API
curl https://kongchang.com/api/v1/knowledge/claims/22964MCP
get_claim(id=22964)