待验证50% 置信基准精确时间
Llama 3.1 405B在MMLU上的得分已接近GPT-4o的水平,Qwen 2.5 72B在数学和代码生成基准上某些指标超越了GPT-4 Turbo
1
来源数
50%
置信度
中期 (~90 天)
时效性
2026/9/7
首次发现
有效期至:2026/12/6
来源
涉及实体
相关事实
待验证Qwen3-4B在多项推理类评测指标上已与DeepSeek V3或GPT-4o相差无几76% 相似待验证Qwen2.5-72B 在多项基准测试中接近 Llama-3.1-405B 的表现72% 相似待验证Models like GPT-4 and Claude 3.5 surpassed 90% accuracy on MMLU, significantly diminishing its discriminative power71% 相似待验证Glimmer智能指数35分,超越Gemma 2 27B(30分),几乎追平Qwen 2.5(36分),Qwen 2.5 72B以38分居榜首70% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/871702API
curl https://kongchang.com/api/v1/knowledge/claims/871702MCP
get_claim(id=871702)