待验证90% 置信事实时间未知
Mistral Medium 3.5在HumanEval基准测试中通过率为77.6%
1
来源数
90%
置信度
长期有效
时效性
2026/6/1
首次发现
来源
Mistral Vibe完全指南:免费API驱动的终端AI编程助手
bilibili薛定猫AI
涉及实体
相关事实
待验证在自研基准PinBash 3测试中,SOL约得78.6%,TERRA约62.9%,LUNA约44.3%65% 相似待验证Mistral 3.8B 指令模型使用 Q4 量化和 llama.cpp 原生调用时,综合 Score 从 27.3% 提升到 78.3%,区别仅在于是否有 Forge 护栏63% 相似待验证在综合基准测试中,So得55分(约78.6%),Terra得44分(约62.9%),Luna得31分(约44.3%)62% 相似待验证在HumanEval、MBPP、GSM8K等基准测试上,Gemma 7B的表现显著优于同参数量级的LLaMA 2 7B和Mistral 7B62% 相似待验证在Artificial Analysis的Omniscience幻觉指标上,GPT-5.6 Sol Max高达89%,高于Claude Fable的55%和GLM 5.2的28%62% 相似
引用此条事实
Stable URI
https://kongchang.com/claim/25125API
curl https://kongchang.com/api/v1/knowledge/claims/25125MCP
get_claim(id=25125)