待验证50% 置信事实精确时间
Community benchmarks showed DeepSeek V3-0324 achieved significant improvements on mainstream code evaluation benchmarks including HumanEval and SWE-bench
1
来源数
50%
置信度
中期 (~90 天)
时效性
2026/7/2
首次发现
有效期至:2026/9/30
来源
相关事实
待验证DeepSeek-Coder performs excellently on mainstream code evaluation benchmarks like HumanEval and MBPP.77% 相似已验证DeepSeek在多项代码基准测试(如HumanEval、MBPP、LiveCodeBench)中表现优异,部分指标接近甚至超越GPT-4级别模型71% 相似待验证DeepSeek-Coder-V2 在 HumanEval 等代码评测榜上长期位居前列70% 相似待验证DeepSeek的代码能力在多项基准测试中达到国际领先水平69% 相似待验证DeepSpec配套九个评测基准,涵盖数学推理(GSM8K、MATH-500)、代码生成(HumanEval、MBPP、LiveCodeBench)和通用对话(MT-Bench、Alpaca、Arena-Hard v2)69% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/58320API
curl https://kongchang.com/api/v1/knowledge/claims/58320MCP
get_claim(id=58320)