Unverified50% confidenceFactExact time
当前AI领域广泛使用的benchmark包括代码能力的HumanEval、MBPP,数学推理的GSM8K、MATH,以及综合能力的MMLU等
1
Sources
50%
Confidence
Long-term
Relevance
9/9/2026
First Seen
Sources
Related Entities
Related Claims
UnverifiedAI模型能力评估长期依赖MMLU、HumanEval、GSM8K等标准化基准集73% similarUnverified业界常用的AI评测基准包括MMLU、HumanEval、GSM8K等,每个基准只能衡量特定维度的能力73% similarUnverified本地优先架构在AI领域的可行性得益于模型量化技术(GGUF格式、INT4/INT8)和端侧推理框架(llama.cpp、ONNX Runtime、Core ML)的成熟71% similarUnverifiedMMLU、HumanEval、GSM8K等评测榜单曾是AI模型竞争与公众注意力的焦点69% similarUnverifiedAI Prep内置了330+个已讲解的概念,覆盖ML基础、深度学习、NLP/LLMs、GenAI、MLOps、AI Agents等领域69% similar
Cite This Claim
Stable URI
https://kongchang.com/claim/881627API
curl https://kongchang.com/api/v1/knowledge/claims/881627MCP
get_claim(id=881627)